With frontier models getting more capable by the month 🚀, everyone is talking about Recursive Self-Improvement (RSI). But how do we actually measure this progress? 🤔
Excited to present RSI-Exam — a benchmark testing whether AI agents can improve an existing method through autonomous, long-horizon experimentation, and, importantly, whether those improvements generalize to hidden data.
Across 88 executable research tasks spanning 6 domains, the trend is clear: Claude and GPT form the first tier, with a clear gap over the rest of the field. Yet there is still huge room for improvement.
More importantly, the trajectories tell both stories: agents that discover fundamentally better methods — and agents that spend hours rigorously optimizing the wrong idea.
Check it out:
显示更多