注册并分享邀请链接,可获得视频播放与邀请奖励。

与「gpt」相关的搜索结果

gpt 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 gpt 的内容
Grok 4.6用最低的成本,最优秀的表现,在真实编程测试里拿下了第一。 它的表现和第二名 Fable 5只差0.3个百分点,但每个任务的成本却只有对方的6分之1。 这套测试是Cursor用真实编程任务做的,主要看AI能不能理解项目,修改多个文件,并完成开发工作。 实际表现中,Grok拿到70.8分,Claude Fable 5是70.5分,Claude Opus 5是70分,GPT-5.6 Sol是67.2分。 最离谱的是成本差距。 Grok完成一个任务平均只花2.81美元,Claude Fable 5高达17.32美元。 Claude Opus 5要8.23美元,GPT-5.6 Sol要5.69美元。
显示更多
0
10
13
1
转发到社区
Grok 4.6 just ranked #1# on CursorBench 3.2 Outperforming Claude Fable 5, Opus 5 and GPT-5.6 Sol on real-world coding performance And what makes this even crazier is the efficiency....the chart gives CursorBench performance against average cost per task, and Grok 4.6 is sitting right at the top of the frontier Grok 4.6 is insanely capable at coding....delivering frontier coding performance at ridiculous efficiency
显示更多
看到 DeepSeek Harness 开源发布,满怀期待的安装好 这是用「DeepSeek V4 Flash + 1390 万 Token + 30 多分钟」做出来的网站 呃。。怎么说呢,一定是我不会用,是我的问题 🤦‍♂️ btw... 图1 是 DeepSeek Harness 图 2 是 Codex + GPT-5.6 Sol
显示更多
0
52
15
17
转发到社区
Grok 4.6 from @SpaceXAI on ARC-AGI (Verified): - ARC-AGI-1: 87.5%, $0.30/task - ARC-AGI-2: 67.1%, $0.76/task - ARC-AGI-3: 2.11%, $5.6K On ARC-AGI-3, Grok 4.6 with xhigh reasoning scored comparably to GPT-5.6 Sol with high reasoning, but cost $5.6K versus Sol's $15.2K.
显示更多
0
12
275
19
转发到社区
睡醒看了很多人测评的deepseek的agent “DSH”,很多人聊的最多的就是token消耗量,kv键缓存命中率,同时DSH的同任务时间消耗是gpt或者clude的2倍。 这点透露的信息强化了我对英特尔陈立武老爷子说的CPU和dram存储堆叠封装的路线理解。 agent代理的普及很快会迎来cpu/gpu的1:1时刻,那时的cpu也非常需要像英伟达的gpu和hbm堆叠封装一样,需要把cpu和hbm堆叠封装在一起,极大的提高读取效率和降低计算时间。
显示更多
0
11
96
22
转发到社区
Ultrafast mode for GPT-5.6 Sol is now in limited preview, powered by Cerebras. We gave @OpenAI's GPT-5.6 Sol the same prompt on Ultrafast and Standard: build a financial terminal-style dashboard for analysts. Ultrafast: 1 min 50 seconds Standard: 12 min 20 seconds Same result, nearly 7x faster.
显示更多
0
20
371
27
转发到社区
Grok 4.6 wins again. 👑 Grok 4.6 takes the #1# spot on GPQA Diamond with a score of 94.9%, beating GPT-5.6, Gemini 3.1 Pro, Claude Opus 5, and every other model tested by Artificial Analysis.
显示更多
0
32
169
29
转发到社区
Grok 4.6 ranks #1# on the GPQA Diamond leaderboard 🧠 Grok 4.6 (high) scores 95% - the highest score on the chart for graduate-level scientific reasoning It outperforms Claude Fable 5, Opus 5, GPT-5.6 Sol and Kimi K3
显示更多
0
20
123
10
转发到社区
Previewing Ultrafast mode for @OpenAI's GPT 5.6 Sol, powered by Cerebras. GPT-5.6 Sol Ultrafast generates responses at up to 750 tokens per second. That's the full, GPT-5.6-Sol model -- up to 14× faster than the same model on Standard processing. It speedran Humanity’s Last Exam in 11h 11m, nearly 7× faster than Claude Fable 5 with comparable accuracy.
显示更多
0
42
893
97
转发到社区
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows.
显示更多
0
197
2.5K
191
转发到社区