注册并分享邀请链接,可获得视频播放与邀请奖励。

Arena.ai (@arena) “GPT-5.6 Sol by @OpenAI is #2 on the Agent Arena leaderboard, based on 7.8K real-” — TopicDigg

Arena.ai 的个人资料封面
Arena.ai 的头像
Arena.ai
@arena
加入 March 2023
0 正在关注    0 粉丝
GPT-5.6 Sol by @OpenAI is #2# on the Agent Arena leaderboard, based on 7.8K real-world agentic sessions! It is a notable uplift from GPT-5.5 (xHigh) of +1.6% Net Improvement, narrowing the gap with the frontier Claude Fable 5. The biggest difference comes from ‘Praise vs Complaint’, a signal that captures implicit user satisfaction with an agent’s responses and artifacts. Claude Fable 5 scores +17.3%, compared with +10.9% for GPT-5.6 Sol. See detailed signal-level comparison below. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. Congrats again to the @OpenAI team!
显示更多
0
34
608
57
转发到社区