注册并分享邀请链接,可获得视频播放与邀请奖励。

Arena.ai (@arena) “Introducing factuality in the Arena: a new ranking of models according to a weig” — TopicDigg

Arena.ai 的个人资料封面
Arena.ai 的头像
Arena.ai
@arena
加入 March 2023
0 正在关注    0 粉丝
Introducing factuality in the Arena: a new ranking of models according to a weighted combination of human preference and factuality. Model rankings are now viewable according to a weighted combination of human preference and factuality. Factuality is live in our Text and Search Arenas as a non-default toggle. We audit model responses by randomly sampling battles and extracting web-verifiable claims. We then verify these claims and compare the average correctness between model responses. To power these rankings, we’ve labeled over 2 million claims made by LLMs in real-world conversations, 1.3+ million from Text Arena, and 700k+ from Search Arena. Notable highlights with factuality enabled in the Text Arena: - Claude Fable 5 moves down slightly to spot #2# - GPT-5.5 saw the largest increase, moving up 13 spots into the #7# spot - Muse Spark dropped the most from #7# to #20# (-13pt) By labs, Meta saw the largest drop from #2# to #5#, while Anthropic overall held the #1# spot. Looking at only open model providers, Xiaomi saw the largest improvement, jumping from #9# to #6#. Learn more about the findings and methodology in this thread.
显示更多
0
12
397
44
转发到社区