Agent performance is often attributed primarily to the underlying model.
Our research suggests the harness recovers what a model can already do, but it can't create capability the model doesn't have.
@dhruvrnaik and I saw this with Qwen3.6-27B on the
@harvey Legal Agent Benchmark. Harness optimization alone moved pass rate from 67% to 85%. Adding post-training pushed it to 88%.