From Mecha Hitler to SOTA rare-disease diagnosis in children?
@SpaceXAI's
@grok 4.6 has taken the 👑 on RareBench, edging out
@AnthropicAI Claude Opus 5 for about 1/3 the cost. This was not on my 2026 bingo card!
@deepseek_ai's new v4-pro-0813 model underperformed my expectations, v4-flash, and seemingly the entire internet's. We accessed using DeepSeek's 1P API on the day of release and I almost wonder if they didn't switch over their model endpoint correctly. We will re-benchmark and report back.
@Zai_org has attracted a following with GLM5.2, but they, too, underperformed. This doesn't surprise me because when I compared GLM and
@Kimi_Moonshot K3 for coding use-cases, I found Kimi substantially stronger, but the internet seems to love this model.