注册并分享邀请链接,可获得视频播放与邀请奖励。

OpenAI (@OpenAI) “We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and” — TopicDigg

OpenAI 的个人资料封面
OpenAI 的头像
OpenAI
@OpenAI
加入 December 2015
0 正在关注    0 粉丝
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval.
显示更多
0
297
7.5K
528
转发到社区