注册并分享邀请链接,可获得视频播放与邀请奖励。

Cihang Xie (@cihangxie) “With frontier models getting more capable by the month 🚀, everyone is talking a” — TopicDigg

Cihang Xie 的个人资料封面
Cihang Xie 的头像
Cihang Xie
@cihangxie
加入 July 2014
0 正在关注    0 粉丝
With frontier models getting more capable by the month 🚀, everyone is talking about Recursive Self-Improvement (RSI). But how do we actually measure this progress? 🤔 Excited to present RSI-Exam — a benchmark testing whether AI agents can improve an existing method through autonomous, long-horizon experimentation, and, importantly, whether those improvements generalize to hidden data. Across 88 executable research tasks spanning 6 domains, the trend is clear: Claude and GPT form the first tier, with a clear gap over the rest of the field. Yet there is still huge room for improvement. More importantly, the trajectories tell both stories: agents that discover fundamentally better methods — and agents that spend hours rigorously optimizing the wrong idea. Check it out:
显示更多
0
2
68
17
转发到社区