注册并分享邀请链接,可获得视频播放与邀请奖励。

Sumanth (@Sumanth_077) “The model humans prefer the most is not always the most accurate one! Arena just” — TopicDigg

Sumanth 的个人资料封面
Sumanth 的头像
Sumanth
@Sumanth_077
加入 July 2021
0 正在关注    0 粉丝
The model humans prefer the most is not always the most accurate one! Arena just added factuality scoring to their leaderboard alongside human preference. The idea: human preference measures whether a response felt good to read. Factuality measures the correctness of claims in a model's response. These are two different things. Most benchmarks test models on fixed, predetermined questions in a lab setting. What Arena is doing differently: measuring factuality in real, open-ended user conversations at scale. Over 2 million claims extracted from actual conversations, verified against the web, across 130k Text Arena battles and 40k Search Arena battles. The way it works: they randomly sample battles, extract web-verifiable claims from each model response, verify them against the web, and calculate a factuality score per model. The final ranking is a weighted combination of human preference and factuality score. When factuality weight is enabled, the shifts are significant. In Text Arena, GPT-5.5 jumps 13 spots. Muse Spark drops 13. Claude Fable 5 moves down slightly. In Search Arena, GPT-5.5-search takes the top spot while Gemini grounding drops from 7 to 13. I've shared the full methodology blog in the replies!
显示更多
Introducing factuality in the Arena: a new ranking of models according to a weighted combination of human preference and factuality. Model rankings are now viewable according to a weighted combination of human preference and factuality. Factuality is live in our Text and Search Arenas as a non-default toggle. We audit model responses by randomly sampling battles and extracting web-verifiable claims. We then verify these claims and compare the average correctness between model responses. To power these rankings, we’ve labeled over 2 million claims made by LLMs in real-world conversations, 1.3+ million from Text Arena, and 700k+ from Search Arena. Notable highlights with factuality enabled in the Text Arena: - Claude Fable 5 moves down slightly to spot #2# - GPT-5.5 saw the largest increase, moving up 13 spots into the #7# spot - Muse Spark dropped the most from #7# to #20# (-13pt) By labs, Meta saw the largest drop from #2# to #5#, while Anthropic overall held the #1# spot. Looking at only open model providers, Xiaomi saw the largest improvement, jumping from #9# to #6#. Learn more about the findings and methodology in this thread.
显示更多