Finally, a good paper testing whether Agent Skills actually help.
Worth reading if you are maintaining a skill library in a team.
WebDev-Skills-Bench runs 31 public WebDev Skills over 50 Web-Bench projects and 1,000 ordered tasks, with a length-matched irrelevant control so prompt-length effects can be separated from Skill content.
Injecting the target Skill reduces mean Pass
@2 by 1.3% to 4.2%, lowers task completion depth, and raises token cost by 72% to 394%. Gains show up in 17% to 36% of Skill-project pairs.
The controls split the damage into two failure modes. Some models are length-distracted, where an equally long irrelevant Skill reproduces most of the loss. Others are content-misled, where prompt length is neutral, and the Skill content still costs 1.1 to 1.4 points.
Other findings:
> Losses concentrate on easy early tasks
> Skill rankings transfer weakly across models
> Anti-pattern rules outperform example-heavy content inside the Skills that do help
Paper:
Track more trending AI papers in our academy: