ByteDance dropped a banger paper on self-evolving agent harnesses!
HarnessDev evaluates whether AI agents can build a runnable harness from scratch and iteratively improve it using execution feedback.
Most agent benchmarks keep the harness fixed and only evaluate the model inside it. HarnessDev changes the target of evaluation itself.
The agent starts from a minimal seed, builds the harness around the task, runs it, observes what worked or failed, and then modifies that harness across multiple iterations.
That means the agent is not only solving the task. It is also changing the planning, memory, tool use, state management, and execution logic around itself.
The paper evaluates this in two stages:
• Creation: can the model build a complete runnable harness from a minimal starting point?
• Evolution: can it improve that harness using feedback from previous runs?
The interesting part is that runnable does not automatically mean better.
Some generated memory and state mechanisms existed in the code but were barely used during execution, and improvements on visible feedback did not always transfer to held-out tasks.
Only 34 of 64 harness changes moved in the same direction on both visible feedback and held-out evaluation, and only 2 of 9 final harness versions were actually the best-performing version on the held-out set.
So the paper is really exposing a new challenge:
Agents can already start modifying the infrastructure they run on.
The harder part is making sure those changes actually generalize.
I've shared the paper in the comments!
显示更多