注册并分享邀请链接,可获得视频播放与邀请奖励。

Arena.ai 的个人资料封面
Arena.ai 的头像

Arena.ai (@arena)

@arena
0 正在关注    0 粉丝
DeepSeek-V4-Flash-20260731 (High) by @deepseek_ai landed in Agent Arena at #21# overall (+1.98% net-improvement), and among open source #3# based on +12.5K real-world agentic sessions. This release is ranked 6 higher than DeepSeek-V4-Pro (-0.47%) and 13 higher than the previous DeepSeek-V4-Flash model (-2.81%). Across key signals, DeepSeek-V4-Flash-20260731 (High) is strong in Confirmed Success (+6.99%), but weaker in Steerability (-2.31%). Below we break down how it scored across all 5 signals, drawn from tasks submitted by a global community of users. Here’s an overview on the signals: - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist
显示更多
🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇 🔷 The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex! Check out the configuration details in our official API docs:
显示更多
0
49
590
37
转发到社区
Big news: MiniMax-H3 by @MiniMax_AI is now the #1# open model in Video Arena: across both Text-to-Video and Image-to-Video. This is +280pts over the next best open model, hunyuan-video-1.5, and a huge improvement from Hailuo-2.3 at #27# (1199 pts) and Hailuo-02-pro at #28# (1197 pts). In Image-to-Video specifically, it scored 1476 pts, just 2 pts away from Dreamina Seedance-2.0 (1478), making it tied for#1# overall! Learn more about how MiniMax-H3 performed in Text-to-Video Arena in thread.
显示更多
0
33
848
98
转发到社区
Qwen3.8-Max ranks #2# in Vision Arena scoring 1,305. Second only to Claude Fable 5 (High) which has only a 13pt lead.
0
17
365
35
转发到社区
Big news: Qwen3.8-Max by @Alibaba_Qwen just landed at #4# on the Frontend Code Arena leaderboard with a score of 1,668! With 1,668 points, Qwen3.8-Max is trailing only Claude Opus 5 (Max) with 1,705 pts and Kimi K3 (Max) with 1,676 pts, on par with Claude Opus 5 (High) with 1669 pts. It also ranks high across all domains: #2# in Consumer Product #3# in Brand & Marketing, Reference-based design, Gaming, and Content Creation Tools #4# in Data & Analytics #5# Simulations Dig into the thread for more details and to see how it also performs on real-world tasks in the Text Arena. Congrats to @Alibaba_Qwen on this huge release!
显示更多
0
139
3.3K
307
转发到社区
DeepSeek-V4-Flash-High by @deepseek_ai is #7# overall in the Frontend Code Arena with 1,586 pts! It’s #3# among open. In categories it’s #4# in Consumer Product, #6# in Reference-based Design, Data & Analytics, Gaming, and #7# in Brand & Marketing. This is an impressive improvement from DeepSeek-V4-Flash-High-Preview, a +154 pt jump. Even more impressive is the jump from DeepSeek-V4-Pro-Preview, a +121 point improvement.
显示更多
0
120
2.6K
254
转发到社区
Big news from @Kimi_Moonshot this week. Check out Kimi K3 head-to-head with Fable 5 on identical prompts in Frontend Code Arena.
Big news: Kimi-K3 by @Kimi_Moonshot is now #1# in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18# -> #1#). In Frontend, Kimi-K3 ranked #1# in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2# only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!
显示更多
0
20
704
70
转发到社区
For the first time, China has taken the lead over the US in Frontend Code Arena with the launch of Kimi-K3 by @Kimi_Moonshot. The last time a Chinese model came close was in early 2025, with DeepSeek-R1.
显示更多
0
98
3.1K
512
转发到社区
Factuality and human preference are complementary signals. 00:00 Why we're adding factuality as a signal 00:14 Human preference vs. factuality: complementary, not the same 01:06 How this differs from style control 02:32 The math: reviewing the existing Bradley-Terry objective 04:27 Extending the objective to factuality 06:23 The composite objective: weighting human preference + factuality 07:14 Why rankings shift between human preference and factuality 07:34 Defining the factuality label 09:48 Why models aren't penalized for abstaining from claims 13:44 Confidence intervals via M-estimation and asymptotic normality 17:27 Wrap-up: what's next for factuality at Arena
显示更多
Introducing factuality in the Arena: a new ranking of models according to a weighted combination of human preference and factuality. Model rankings are now viewable according to a weighted combination of human preference and factuality. Factuality is live in our Text and Search Arenas as a non-default toggle. We audit model responses by randomly sampling battles and extracting web-verifiable claims. We then verify these claims and compare the average correctness between model responses. To power these rankings, we’ve labeled over 2 million claims made by LLMs in real-world conversations, 1.3+ million from Text Arena, and 700k+ from Search Arena. Notable highlights with factuality enabled in the Text Arena: - Claude Fable 5 moves down slightly to spot #2# - GPT-5.5 saw the largest increase, moving up 13 spots into the #7# spot - Muse Spark dropped the most from #7# to #20# (-13pt) By labs, Meta saw the largest drop from #2# to #5#, while Anthropic overall held the #1# spot. Looking at only open model providers, Xiaomi saw the largest improvement, jumping from #9# to #6#. Learn more about the findings and methodology in this thread.
显示更多
0
9
93
10
转发到社区
Kimi-K3 just topped the Frontend Code Arena with a 76% pairwise win rate. When its output was compared head-to-head against other models on the same task, it was picked as the better output 76% of the time on average. For reference: Claude Fable 5 (63%), GPT-5.6 Sol (58%). 50% is baseline, a model winning and losing equally often.
显示更多
Big news: Kimi-K3 by @Kimi_Moonshot is now #1# in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18# -> #1#). In Frontend, Kimi-K3 ranked #1# in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2# only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!
显示更多
0
84
2.6K
297
转发到社区
Big news: Kimi-K3 by @Kimi_Moonshot is now #1# in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18# -> #1#). In Frontend, Kimi-K3 ranked #1# in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2# only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!
显示更多
0
1K
25.6K
3.8K
转发到社区
Introducing factuality in the Arena: a new ranking of models according to a weighted combination of human preference and factuality. Model rankings are now viewable according to a weighted combination of human preference and factuality. Factuality is live in our Text and Search Arenas as a non-default toggle. We audit model responses by randomly sampling battles and extracting web-verifiable claims. We then verify these claims and compare the average correctness between model responses. To power these rankings, we’ve labeled over 2 million claims made by LLMs in real-world conversations, 1.3+ million from Text Arena, and 700k+ from Search Arena. Notable highlights with factuality enabled in the Text Arena: - Claude Fable 5 moves down slightly to spot #2# - GPT-5.5 saw the largest increase, moving up 13 spots into the #7# spot - Muse Spark dropped the most from #7# to #20# (-13pt) By labs, Meta saw the largest drop from #2# to #5#, while Anthropic overall held the #1# spot. Looking at only open model providers, Xiaomi saw the largest improvement, jumping from #9# to #6#. Learn more about the findings and methodology in this thread.
显示更多
0
12
397
44
转发到社区
GPT-5.6 Sol by @OpenAI is #2# on the Agent Arena leaderboard, based on 7.8K real-world agentic sessions! It is a notable uplift from GPT-5.5 (xHigh) of +1.6% Net Improvement, narrowing the gap with the frontier Claude Fable 5. The biggest difference comes from ‘Praise vs Complaint’, a signal that captures implicit user satisfaction with an agent’s responses and artifacts. Claude Fable 5 scores +17.3%, compared with +10.9% for GPT-5.6 Sol. See detailed signal-level comparison below. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. Congrats again to the @OpenAI team!
显示更多
0
34
608
57
转发到社区
Exciting news: @OpenAI’s GPT-5.6-sol is now joint #1# in the Code Arena: Frontend, matching Claude Fable 5! This marks the first time an OpenAI model has reached the top spot in Code Arena, demonstrating major gains in agentic coding, frontend and web app development. Highlights: - Significant improvement from GPT-5.5-xhigh (#18# -> #1#) - #1# in Data & Analytics, Brand Marketing, Consumer product, and Gaming - Priced at $5/$30 per million input/output tokens - roughly 2× cheaper than Claude Fable 5 Huge congrats to the @OpenAI team for this incredible milestone!
显示更多
0
101
1.5K
141
转发到社区
Exciting news: GLM-5.2 (Max) ranks #2# in Code Arena: Frontend, with +29pt over Claude Opus 4.7 (Thinking) and only behind Fable 5! GLM-5.2 is the best open model vs Kimi-K2.6 and Minimax-M3 by a large margin. - #2# React and #4# HTML sub-leaderboards - Ranks as the top model in nearly all sub categories: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Gaming, and Simulations. Congrats @Zai_org for the incredible milestone!
显示更多
0
160
4.3K
492
转发到社区
GLM-5.2 (Max) by @Zai_org ranks #10# on the new Agent Arena leaderboard, closely matching Claude-Opus-4.8 (non-thinking) and is the #1# open model by a wide margin! In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. Compared to 5.1, GLM-5.2 (Max) climbs from #13# to #10#. Its clearest gains are confirmed task success, and user praise vs. complaint. Bash capabilities and tool hallucination remain stable. There is a tradeoff in steerability compared to the previous model (-6.0% vs. +1.2%). GLM-5.2 remains the same price as GLM-5.1, $1.4/$4.4 per input/output MTokens. 1M context window. Huge congrats @Zai_org for the incredible release! See thread for details on how GLM-5.2 (Max) performs across 5 different signals.
显示更多
0
17
527
50
转发到社区
Qwen3.7 Max (20250517) debuts at #4# in Code Arena: Frontend - the top-ranked Chinese lab on the board, surpassing GLM-5.1 and is now on par with Claude Opus 4.6 on agentic web development tasks. Huge congrats to @Alibaba_Qwen on this achievement!
显示更多
0
50
942
91
转发到社区
Qwen3.7 Preview By @Alibaba_Qwen lands on Arena for Text and Vision. In Text Arena, Qwen3.7 Max Preview ranks #13# overall. Alibaba is now the #6# lab in this arena. - #7# Math - #9# Expert - #9# Software & IT - #10# Coding In Vision Arena: Qwen3.7 Plus Preview ranks #16# overall, making Alibaba the #5# lab. Congrats to the @Alibaba_Qwen team on the latest progress!
显示更多
0
32
403
39
转发到社区
US vs China update. Stanford's AI Index put the US–China gap at 2.7%. Here's what two years of real-world use from the Text Arena shows. Gap three years ago: +278. Today: +29. @AnthropicAI's Claude Opus 4.6 Thinking vs. Baidu's @ErnieforDevs Ernie 5.1 at the top. The US has never lost #1#, but the race keeps closing.
显示更多
0
22
292
33
转发到社区