注册并分享邀请链接,可获得视频播放与邀请奖励。

X Freeze (@XFreeze) “Grok 4.6 just took the #1 spot on MedAgentBench....one of the most interesting b” — TopicDigg

X Freeze 的个人资料封面
X Freeze 的头像
X Freeze
@XFreeze
加入 July 2024
0 正在关注    0 粉丝
Grok 4.6 just took the #1# spot on MedAgentBench....one of the most interesting benchmarks for real-world agentic healthcare tasks and Grok now has two generations sitting in the top 3 • #1# Grok 4.6 — ~95.9% • #2# GPT-5.6 Sol — ~94.7% • #3# Grok 4.5 — ~93.4% MedAgentBench goes far beyond answering medical questions It tests AI agents on 300 clinically derived tasks across 10 categories inside a realistic electronic health-record environment, requiring the model to reason, use tools and actually execute multi-step clinical workflows Grok 4.5 was already one of the strongest medical agents tested. Grok 4.6 just pushed it to another level Grok is rapidly becoming one of the strongest AI model families for agentic healthcare work
显示更多
0
70
393
64
转发到社区