注册并分享邀请链接,可获得视频播放与邀请奖励。

与「MiMO」相关的搜索结果

MiMO 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 MiMO 的内容
荣耀发明了WiFi8,所以到底能不能4*4mimo?
刚看完中关村学院的何纪言老师带领 7 位博士生,用 3 个月时间从头训练出了 7B 模型,从预训练到中训练、后训练完全都是自己搞的,在 7B 尺寸模型中达到 SOTA 水平。 首先,虽然很多人在吹 RSI,但技术报告指出,现在模型的能力远未到全自主的程度,在模型架构设计、学习算法设计、数据清洗等方面仍然只能达到 L2 的辅助水平。我自己也有一样的感觉,不管 Astra 还是 Fable,都不能替代我的架构设计,我也得不断提醒自己不要外包思考。 其次,模型训练就是个会者不难、难者不会的问题,对会的人来说,用不了多少算力资源,但对不会的人来说,用再多的算力资源、堆再多的人,都是做不出来的。比如 MiMo V2.6 RL 只花了 300 多万美金,对基模公司来说很便宜了,整个 MiMo core team 只有几十位正式员工,更没有走蒸馏之类捷径。但烧了上亿美金、投入上千人的模型未必就能搞好。中关村学院一个老师加上 7 个学生只用 3 个月就完成造数据、预训练、中训练、后训练,是非常 impressive 的。 最后,数据是新的代码,数据质量就像代码质量一样非常重要。过去几年,我总是想用代码的方式实现 Agent 的自我进化,但很快就达到上限了。最近一年我才发现,这些我用代码方式做的东西更应该用训练数据的方式表达,训到模型里面。我们都知道脏代码的危害,脏数据其实也是一样的。中关村学院这个 7B 模型做了非常深入的数据工作,预训练数据清洗和课程学习,中训练按照上下文长度扩增,后训练利用开源和蒸馏来的 trace 建立指令遵循、长上下文等基础能力,进而构建长思维链推理、工具调用等高阶能力;在 RL 中不断移除已经高概率解决的问题,保持 GRPO学习效率。 技术报告:
显示更多
0
9
358
62
转发到社区
🎉 Tiered Model Perks Are Officially Active Starting today at 15:00 (SGT), updated model perk structure is officially live! The newly integrated MiMo-V2.6-Flash spearheads the lineup at 90% OFF, flagship MiMo-V2.6-Pro debuts at 50% OFF, and popular Flash models enter the 70% OFF tier! remains dedicated to democratizing AI compute, making top-tier big models accessible, affordable, and stress-free for every developer! 👉 Start building now:
显示更多
📢 Responses API Major Upgrade: Universal Upstream Protocol Support & Plug-and-Play for New Models! Existing models on that support the Responses API upstream can now be called via the unified /v1/responses endpoint. When adds support for new models from currently integrated vendors in the future, developers can automatically access them in the Codex client with zero extra configuration! 🔸 Official API Group: Added support for MiMo, Qwen, and Hunyuan model series. 🔸 Mix & OL Station Provider 40% OFF Groups: Added support for OpenAI, Gemini, and Grok model series. 📊 Check out the poster below for the full list of supported models 👇 Simply configure your API Key and Base URL to bring the powerful reasoning, coding, and text processing capabilities of leading models seamlessly into your daily workflow, flexibly meeting a wide range of needs. 👉 Try it now: 📚 More details:
显示更多
MiMo-V3 is getting a new architecture. The core of it, HySparse2, is out today. Less prefill, a smaller KV cache, better long-context retrieval—and we got all three at once. Compared with MiMo-V2.6's Hybrid SWA architecture: • 5.02× lower prefill FLOPs at 1M tokens • 4.5× smaller KV cache at 1M tokens • Better MRCRv2 and RULER-v2 scores, plus lower AgentPPL and LongPPL Why build a new architecture? Agentic inference is a very different workload. Each round, a short action can return a long observation that needs to be prefilled, while the context keeps growing. That puts prefill cost, KV-cache size, and retrieval accuracy on the critical path at the same time. HySparse2 tackles all three with two levels of KV sharing: • KV Bridging: Following YOCO, full-attention layers in the cross-decoder build their K/V from self-decoder hidden states. • KV Reuse: Within each hybrid block, sparse layers reuse the preceding full-attention layer's KV cache and selection indices. Two more changes: token-level selection replaces block-level selection, and a forced window of recent tokens replaces the separate SWA branch, so local and global tokens share one KV cache. Since all cross-decoder KV caches now come from the self-decoder, prefill can stop once the self-decoder finishes. Paper:
显示更多
0
162
3.6K
304
转发到社区
📢 Xiaomi MiMo-V2.6 Series Models Are Now Live on Developed by Xiaomi MiMo, the newly released V2.6 series features open-weights native multimodal reasoning models with a 1M context window and full text/image/video/audio understanding: 🔸 MiMo-V2.6-Pro: Flagship 1.02T total / 42B active parameter sparse MoE architecture, engineered for repository-level software engineering, long-horizon agent workflows, and complex multimodal reasoning. 🔸 MiMo-V2.6-Flash: High-efficiency 309B total / 15B active parameter architecture, optimized for high-frequency office automation, fast agent execution, and cost-sensitive workloads. Now available on both API and Web Chat! 👉 Try now:
显示更多
0
27
58
10
转发到社区
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:
显示更多
0
352
8.7K
857
转发到社区
Anthropic发布《检测与应对AI滥用》报告,指控阿里巴巴、月之暗面Moonshot、DeepSeek、小米Mimo、MiniMax、商汤等对Claude进行蒸馏。 报告声称发现了6起常规武器案例,其中有3起在中国。其中一例称,中国境内行为体用Claude起草反鱼雷火控系统中文规格书、200多页技术建议书和汇报材料,面向国内国防制造商审批与试验。另有无人机蜂群软件、电子战/压制防空瞄准软件等方向。 Anthropic还声称发现,Claude Code被当作软件工程师替代品,写制导、导航与控制(GNC)软件,参与也门手机级飞控制导火箭、射程超2000公里的多级弹道导弹的开发。
显示更多
0
15
46
2
转发到社区
📢 MiMo 系列现已登陆 Web Chat! 继 API 端上线后,MiMo 系列现已进一步开放 Web Chat 使用: 🔸MiMo V2.5 已同步开启限时免费,Web Chat 与 API 双端均可 0 成本体验 🔸MiMo V2.5 Pro 也已上线 Web Chat,欢迎来 体验小米 MiMo 系列模型 🔗
显示更多