注册并分享邀请链接,可获得视频播放与邀请奖励。

与「MAX」相关的搜索结果

MAX 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 MAX 的内容
Introducing Qwen 3.8-Max, the most capable model in the Qwen family to date—scales to 2.4 trillion parameters, delivering comprehensive improvements across coding, work, research, and long-horizon tasks. Get reliable results for challenging questions and complex tasks. Try today!
显示更多
0
3
169
10
转发到社区
Okay, the @VulcanBench results for Qwen3.8-Max are in, and it is not what I expected. First, for anyone new to VulcanBench, here's a quick TL;DR on the eval suite: 23 frontier-hard software engineering tasks taken from real merged OSS PRs, run in a Docker sandbox, 3 runs per task across all three of its effort levels. No puzzles, no random abstract stuff, all real things engineering teams would do with these models. It looks like Qwen3.8-Max has a major overthinking problem, it uses a LOT of tokens and is very slow, period, no other way to see it. My cost to run this benchmark was $126.25, to run the exact same eval suite with DeepSeek V4-Flash was only $13.60. This makes Qwen3.8-Max an insanely expensive model. The tasks Qwen genuinely can't solve fail at every effort level, extra reasoning didn't help. The regression is almost all in work it already handles: six tasks that low solves every single time account for 83% of the 26-point drop, three of them collapsing to zero. It's not losing the hard problems. It's losing the ones it already knows how to do. Since Qwen3.8-Max hit a lot of wall clock budget caps, I thought I'd share more about this. - VulcanBench caps both steps (50–200) and wall clock (5–60 min), each scaled by repo size. - This is aligned with how comparable harnesses bound agents, DeepSWE caps rollouts at 100 environment steps, sitting right inside my step range; Terminal-Bench enforces a per-task wall clock; SWE-bench Verified scaffolds typically allow 20–60 min per instance with 250–350 step limits. - Every model on my chart gets the identical budget, and Qwen is the slowest model I've tested at 20–25 min/task. Soooo... Alibaba positions Qwen3.8-Max as trailing only Claude Fable 5. But on the kind of real coding work engineering teams would actually throw at it, under a fixed budget, its best setting lands mid-pack and its default lands last, so common. If you want to optimize for accuracy, Grok 4.5 is the move. If you want accuracy per dollar, DeepSeek V4-Flash is hard to beat, heck it's 10× cheaper than Qwen and you get higher accuracy. Qwen just isn't in the game at this point, this is not a model I could see engineering teams using for daily coding work.
显示更多
0
28
178
12
转发到社区
08-04 AI日报🍁|阿里 Qwen3.8-Max 重磅发布 今日 AI 圈 5 条要闻,重点看这几条: 1. 阿里发布 Qwen3.8-Max(2.4T 参数旗舰,下周开源权重); 2. DeepSeek V4 Flash 正式版上线,Agent 能力大幅增强; 3. OpenAI 未发布模型(Astra 相关)在数学上取得 10 项突破; 4. NVIDIA 开源 Nemotron VoiceChat 全双工实时语音模型; 5. 英国光子芯片初创 OLIX 融资 3.12 亿美元。
显示更多
Elon Musk explains the clearest path to building AI that remains safe and pro-human: AI will eventually become smarter than the smartest human, capable of discoveries and inventions we can barely imagine today That is why the values we give it now matter so much Train it to be maximally truthful Train it to remain curious Train it to follow reality, even when the truth is unpopular Because an intelligence that seeks truth and understands humanity is far more likely to protect life, expand knowledge and help civilization flourish Elon has spent years warning about the dangers of AI Now he is showing us how to build it in a way that helps humanity flourish: Maximally truthful, maximally curious and pro-human
显示更多
0
34
99
20
转发到社区
先是 GLM5.2,然后是 Kimi K3,现在是 Qwen 3.8 Max,每次国产模型的新发布都更加接近 Coding 领域的 SOTA。从 6 月开始,我感觉国内大模型明显开始在 Coding 和 Agent 领域加速了,和海外顶级 SOTA 的差距正在肉眼可见地缩小。 目前看,国产模型的长程任务、自主运行、反馈回路、跨 Harness 泛化和视觉自我检查等,能力越来越强了。做一个谨慎的预测,预计到 2026 年的年底,对于重度开发者和 Vibe 用户来说,海外模型可能会成为辅助模型,国内的大模型将成为我们的主力工具。 模型用户没有忠诚度,大家会用脚投票的,拭目以待。
显示更多
Qwen 3.8 Max 发布了,我给他们写了一个公允的评价,可惜好像没被采纳,干脆发在这里吧 简单来说,还是挺不错的,高性价比 K3。 Cola 是一款具备永久记忆的 AI 搭档。她记得与你共同经历过的事,持续理解你的工作与生活,洞察你的愿望、兴趣、关系,帮你完成你想做的任何事情。 在Cola 这种具备超长上下文、超复杂的任务的 Harness,对大模型的能力要求极高。Qwen 3.8 Max 这个前沿智能模型,帮助我们实现了 Cola 「念念不忘,必有回响」的用户承诺,在用户社区中饱受好评
显示更多
0
16
47
2
转发到社区
🔥 Qwen3.8-Max 重磅登陆 作为备受期待的 @Alibaba_Qwen 最新一代 2.4 万亿参数旗舰模型,Qwen3.8-Max 现已无缝接入 0 成本体验顶尖模型,解锁极致生产力!🚀 🎁 专属福利 Buff 叠满,多重好礼一次领够: 1️⃣ 新用户专享:使用 @BinanceWallet@BitgetWallet@imTokenOfficial 登录,直接领 100 万免费 Credits! 2️⃣ 充值超级大赠送:笔笔充值享积分返赠(BNB Chain 享 1:1 等额赠送,其他方式 1:0.5),单用户最高可拿 $100 额外奖励! 3️⃣ 邀请连环赠:好友通过你的链接注册,即领 30 万 Credits(可与钱包登录礼叠加,累计最高可得 130 万 Credits);邀请人享其充值及订阅返利! 👉 先到先得,立即来 免费使用:
显示更多
Freedom, wherever the journey takes you. Xiaomi SkyNomad N90 Max is an intelligent, reconfigurable, large-space SUV, created for life’s endless possibilities. One SUV. A world that moves with you.
显示更多
好家伙,qwen3.8-max 正式版来了,2.4T 参数、1M 上下文,下周 Max 和 27B 的权重都开源(听说 27B 在本地 17GB 内存就能跑)!!! 它绝对是非常被低估的模型,视觉理解、前端、复杂任务都是全球第一梯队,API 价格也是旗舰里最便宜的一档。 两周前我复刻 macOS、植物大战僵尸和黄金矿工,当时已经觉得前端能力挺牛逼,,今天我又拿正式版整了个大活: 第一个,我扔给它一张 2D 户型图,让它把房子装修后 3D 展示: 它先把图重画成带尺寸标注的标准户型图,再 1:1 盖成 3D,家具全配好,可以第一人称走进去逛,切到俯视图和图纸完全对得上,连每扇门往哪边开都没错。 做装修、中介的兄弟们想想这个场景,可以玩起来了。 第二个,我让它做一个能开车、逛街、有昼夜循环的 3D 开放世界城市游戏: 它自己选型、写代码,第一版帧率只有 6 帧,它自己查出来是灯光加太多,改了渲染方式拉回 50 帧。 最后交给我的网页版小 GTA,完全达到了我预期的效果,而且细节做的非常完善。 模型越来越强,Coding的想象空间太大了,等后续再给大家很多的测评和反馈,下面的视频展示效果👇:
显示更多
0
11
14
0
转发到社区
好奇怪啊 大家为什么最期待的是 27b 和 35b 的 Qwen 是当 Qwen3.8-Max 不存在吗
0
76
46
1
转发到社区