注册并分享邀请链接,可获得视频播放与邀请奖励。

与「AgenticCoding」相关的搜索结果

AgenticCoding 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 AgenticCoding 的内容
GLM-5.2 刚刚正式发布! 给大家带来实测! 直接说结论本次测试中, 提升最大的是Agent能力, 而且是有质的变化! 测试中GLM-5.2 完全不用搜索附近的位置, 就能直接去想要到达的地方. 这一切竟然是它在一开始把地图背下来了! 这在我测试的20多个模型中之前是没有一个模型能做到的, 比如之前的模型想去换电站, 那么都要搜一下附近有哪些换电站(这就会浪费一次tool_call), 而GLM-5.2直接就知道换电站的位置! 从来没用过搜索函数. 这种一开始就把需要的数据内化到上下文中, 并且能够贯穿整个1M上下文进行推理的能力真的是叹为观止. 除此之外, 本次测试后端代码的 Agentic Coding 能力也有提升, 来到了总榜的第二名. 而本次测试暴露出最大的短板则是空间理解. 其实成也萧何败也萧何, 它虽然把换电站的位置都背下来了, 但是去的换电站却不是最近的, 所以虽然记住了, 但是记住了之后在用之前再根据自己当前所在位置推理一下, 他还是没有做到的, 这也是最大的短板了, 强烈建议官方优化一波. #GLM52# #智谱# #智谱AI# #AgenticCoding# #长上下文能力#
显示更多
0
10
80
5
转发到社区
Sonnet 5.5 发布了,这个模型很神奇:它在 Agentic coding 测试上的得分比 Opus 5.5 还要高 4 分。 定价上也非常便宜,每百万 token 输入输出的价格分别是 $2 和 $10,与 GPT-6 Sol 一样。而且它的运行速度相对于 Sonnet 5 快了 30%,Claude 官方表示成本也降低了 30%。 至于为什么它在 Agentic coding 上的表现跟 Opus 5.5 差不多甚至更高,可以看一下价格测试。 Sonnet 5.5 在 max 努力程度下,同样任务的花费金额比 Opus 5.5 还要高得多,快要持平甚至反超了,这可能就是它得分高的原因。 所以大家在用 Opus 5.5 或者 Sonnet 5.5 的时候千万别开 max:在一些复杂任务上,它有可能跑出非常贵的价格,几乎赶上 Claude Fable 5.1 了
显示更多
Claude Opus 5.5 在 Agentic coding: FrontierCode 下不同思考程度下的得分 low 是确实不行,med 和 max 居然非常接近,高于 high 和 xhigh 😂
📢 Claude Opus 5.5 Is Now Live on As the first model in the new Claude 5.5 family from @AnthropicAI, Claude Opus 5.5 is built for long-running agentic coding and complex knowledge work. It performs at the level of Claude Fable 5.1 on most tasks while costing ~40% less to run than Opus 5. Supporting a 1M-token context window and up to 128K output tokens, it delivers industry-leading capabilities across autonomous workflows and Computer Use. Now available on both API and Web Chat! 👉 Try now: 🔗 Learn more:
显示更多
0
11
31
4
转发到社区
what if ai agentic coding be like:
0
182
7.4K
751
转发到社区
BREAKING: Grok 4.6 just took the #1# spot on CursorBench 3.2 — while delivering a massive efficiency advantage. ⚡💻 • Grok 4.6 Extra High — 70.8% | $2.81/task • Fable 5 Max — 70.5% | $17.32/task • Opus 5 Max — 70.0% | $8.23/task • GPT-5.6 Sol Max — 67.2% | $5.69/task Grok achieved the highest score while costing roughly 6× less than Fable 5 Max and nearly 3× less than Opus 5 Max per task. For AI agents, raw intelligence is only part of the equation. The ability to maintain high performance across long coding tasks without burning massive amounts of compute could be a major advantage. Grok’s agentic coding efficiency is becoming seriously impressive. 🚀 Source: CursorBench 3.2
显示更多
Grok 4.6 just took the #1# spot on CursorBench 3.2.....and the efficiency is insane Here's the cost comparison: • Grok 4.6 Extra High — 70.8% | $2.81/task • Fable 5 Max — 70.5% | $17.32/task • Opus 5 Max — 70.0% | $8.23/task • GPT-5.6 Sol Max — 67.2% | $5.69/task Grok achieved the highest score while costing roughly 6X less than Fable 5 Max and nearly 3X less than Opus 5 Max per task That’s what makes Grok so powerful for agents Top-tier intelligence is great.....but top-tier intelligence that can keep working across long coding tasks without burning ridiculous amounts of compute is even better Grok’s agentic coding efficiency is insane
显示更多
(1/2) Just in time for back to school: the power of agentic coding has landed in Search. Anyone can now create custom tools and simulations to visualize different topics. One cool example: I asked AI Mode “make an interactive 3D fractal Mandelbulb visual I can zoom into” and it coded and rendered this visualization, right in Search & on the fly. Available globally in English (free of charge). Try it by asking AI Mode to create a simulation for whatever’s on your mind. Plus, we’ve started to roll out these interactive visuals in AI Overviews. lmk what you create!
显示更多
📣 Kimi K3 is now available in GitHub Copilot for @code! Try Kimi Moonshot's latest open-weight model for agentic coding, now hosted by @FireworksAI_HQ. 📖 Learn more:
显示更多
0
18
474
57
转发到社区
📣 @Kimi_Moonshot's Kimi K3, an open-weight model, is now generally available and rolling out in GitHub Copilot. The model shows frontier-level abilities on agentic coding with highly cost-effective pricing. It is hosted by @FireworksAI_HQ. Learn more. 👇
显示更多
0
56
732
68
转发到社区