注册并分享邀请链接,可获得视频播放与邀请奖励。

与「SOTA」相关的搜索结果

SOTA 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 SOTA 的内容
5/ 当某项能力刚达到SOTA时,成本通常下降得最快。 综合5项基准测试,SOTA性能的成本平均每季度下降66%。 两年后,降价速度会放缓至每季度32%。
刚看完中关村学院的何纪言老师带领 7 位博士生,用 3 个月时间从头训练出了 7B 模型,从预训练到中训练、后训练完全都是自己搞的,在 7B 尺寸模型中达到 SOTA 水平。 首先,虽然很多人在吹 RSI,但技术报告指出,现在模型的能力远未到全自主的程度,在模型架构设计、学习算法设计、数据清洗等方面仍然只能达到 L2 的辅助水平。我自己也有一样的感觉,不管 Astra 还是 Fable,都不能替代我的架构设计,我也得不断提醒自己不要外包思考。 其次,模型训练就是个会者不难、难者不会的问题,对会的人来说,用不了多少算力资源,但对不会的人来说,用再多的算力资源、堆再多的人,都是做不出来的。比如 MiMo V2.6 RL 只花了 300 多万美金,对基模公司来说很便宜了,整个 MiMo core team 只有几十位正式员工,更没有走蒸馏之类捷径。但烧了上亿美金、投入上千人的模型未必就能搞好。中关村学院一个老师加上 7 个学生只用 3 个月就完成造数据、预训练、中训练、后训练,是非常 impressive 的。 最后,数据是新的代码,数据质量就像代码质量一样非常重要。过去几年,我总是想用代码的方式实现 Agent 的自我进化,但很快就达到上限了。最近一年我才发现,这些我用代码方式做的东西更应该用训练数据的方式表达,训到模型里面。我们都知道脏代码的危害,脏数据其实也是一样的。中关村学院这个 7B 模型做了非常深入的数据工作,预训练数据清洗和课程学习,中训练按照上下文长度扩增,后训练利用开源和蒸馏来的 trace 建立指令遵循、长上下文等基础能力,进而构建长思维链推理、工具调用等高阶能力;在 RL 中不断移除已经高概率解决的问题,保持 GRPO学习效率。 技术报告:
显示更多
0
9
358
62
转发到社区
Introducing Qwen Intelligence, bringing personal intelligence within everyone's reach. 📱✨ It launches with three SOTA agents: 🥳 - Mobile Planner Agent: plans, decomposes & orchestrates complex tasks. #1# on MobilePA-Bench, MobilePA-Bench Business & Memory. - Mobile-Use Agent: gets things done, API-first with GUI fallback. MobileWorld 82.1, MobileWorld-Real 92.2, AndroidDaily 97.2, 90% end-to-end success rate. - Mobile Creative Agent: turns one sentence into ready-to-use creations. Image generated in 3s, about 2x faster than leading peers. We're also opening up our benchmark suite: MobilePA-Bench, MobileWorld, MobileWorld-Real, and MobileWorld-Safety, covering planning, cross-app execution, real-device performance and safety. 🔗 Learn more about the agents: - Qwen Intelligence official website: - Mobile Planner Agent: - Mobile-Use Agent: - Mobile Creative Agent: 🔗 Explore our open benchmark suite: - MobilePA-Bench: - MobileWorld (GitHub): - Leaderboard:
显示更多
0
107
2.5K
252
转发到社区
Speech recognition is easy—until you ask it to listen forever. Today we’re open-sourcing Audio8 ASR Infinite: Ultra-low latency, unlimited audio, 24/7 transcription, no drift. Built-in semantic turn detection keeps it listening like a human ear. New SOTA for streaming ASR.
显示更多
0
36
1.5K
126
转发到社区
[RAG 论文分享] VikingRAG:匹配 SOTA 准确率、Token 成本降到 5%–32% 现有问题:RAG 高准确率依赖结构上下文与多轮交互,是 token 开销的主要来源 企业问答、法律、财报等场景的语料都是章、节、段落组成的结构化文档,结构本身就是检索线索,它指示事实归属与局部和全局的关联。 现有 RAG 方法陷入两难:不用结构(朴素向量 RAG、图 RAG、SQL-RAG)丢失导航线索,准确率低;用结构(MoDora、BookRAG、DeepRead)准确率高,但 DeepRead 要把候选文档的完整目录塞进 prompt,开销随目录规模线性增长,加上多轮交互历史不断累积,token 成本巨大。 论文地址 # 三个核心设计 1. 层次化语义存储:把结构从 prompt 搬进可查询的外部状态。 文档分块后保留所属结构节点,自底向上生成节点摘要,目录、块、摘要全部物化为 URI 可寻址对象(如 viking://Pasta/Carbonara/),祖先-后代关系编码为 URI 前缀。系统向代理暴露 Search / List / Grep / Read 四个工具,语义与结构路径共享同一 URI 空间。效果:结构 token 与实际访问的目录片段成正比,而非与完整目录成正比——直接消解 DeepRead 的线性开销。 2. 证据缺口驱动的多轮检索。 Agent 每轮判断证据是否充分,不充分则继续调用工具(轮数预算 B=15),充分即作答。设计哲学是粗定位与细验证分离:Search 锚点 → Read 查看后发现缺口 → List 相邻块 → Grep 精确命中 → Read 验证,每步把搜索空间收窄到相关子树内。 3. 经验边 + 自适应升级——让相似查询不必重复探索。 经验边从历史检索轨迹中把“Search 命中的 URI”连向“真正支撑答案的 URI”,边上存历史问题嵌入做查询时过滤;新查询沿边做条件化多跳扩展,复用路径而非重新探索(VikingRAG-E)。经验积累足够后,多数查询一轮检索即可作答——用约束感知的充分性判断器先验证证据是否支撑答案的关键约束,验证不过关才升级为完整多轮代理检索(VikingRAG-E+)。 # 实验结果:token 降至 SOTA 的零头 6 个真实结构化文档数据集(从 0.24M 词元的课程大纲到 8.78M 词元的财报),8 个基线,骨干 LLM 为 DeepSeek-V4-Pro,并在 GPT-5.5、Seed-2.0、GLM-4.7 上验证稳健性。主要结论: · 准确率与所有基线持平或更高(DeepRead 通常是最强基线); · 基础版 VikingRAG 仅消耗 SOTA 方法的 11.6%–51.9% token,完整版 VikingRAG-E+ 降至 5.1%–32.5%,延迟同样显著更低; · 逐层消融:经验边再省 12%–33% token,自适应升级再省 19%–50%; · 可扩展性:LightRAG、HippoRAG-2 在最大数据集上 24 小时内无法完成摄入,BookRAG 在多数数据集上超时;而 VikingRAG 在文档数从“仅相关文档”增至 991 篇时性能基本稳定——因为它是按需定位,不随语料规模膨胀; · 存储方面:插入延迟与 DeepRead/MoDora 相当、远快于图方法;代价是摄入期 token 更高(为每个索引对象生成预览);文档删除零 LLM 成本。
显示更多
Grok 4.6 multimodal is a step change from Grok 4.5. It’s one of the under-discussed improvements, and I’ve been very impressed by it. My daily work includes reviewing lots of videos and understanding the context; Grok 4.6 improves the productivity of such workloads by at least 10x if not 100x. Such workflows may not be captured by common VLM benchmarks, but in my use cases it outperforms Gemini 3.5 and Gemma 4, which is considered the SOTA of VLMs in my opinion. Hats off to the multimodal teams—you did a great job.
显示更多
0
43
507
27
转发到社区
From Mecha Hitler to SOTA rare-disease diagnosis in children? @SpaceXAI's @grok 4.6 has taken the 👑 on RareBench, edging out @AnthropicAI Claude Opus 5 for about 1/3 the cost. This was not on my 2026 bingo card! @deepseek_ai's new v4-pro-0813 model underperformed my expectations, v4-flash, and seemingly the entire internet's. We accessed using DeepSeek's 1P API on the day of release and I almost wonder if they didn't switch over their model endpoint correctly. We will re-benchmark and report back. @Zai_org has attracted a following with GLM5.2, but they, too, underperformed. This doesn't surprise me because when I compared GLM and @Kimi_Moonshot K3 for coding use-cases, I found Kimi substantially stronger, but the internet seems to love this model.
显示更多
BREAKING: MiniMax H3 by @MiniMax_AI is 1st across 3 Video categories (Multi-Image to Video, Image to Video, and Video Editing) on Design Arena. MiniMax H3’s performance in Video Arena puts it ahead of other top performing models like Seedance 2.0 by @BytePlusGlobal, Grok Imagine Video 1.5 Preview by @SpaceXAI, and Gemini Omni Flash by @GoogleDeepMind. This marks another category to be led by open-weight models, following Kimi K3’s first place in our coding categories. Congratulations to the @MiniMax_AI team for establishing a new SOTA in video generation!
显示更多
0
15
1.3K
34
转发到社区
7 days of unlimited Seedance 2.5 on Higgsfield. The most realistic and production-ready video model. Now we offer generations at ZERO credit cost for 7 DAYS. Up to 50 references, 30 seconds of perfect continuity, realism in physics,  advanced VFX/SFX, and SOTA video-to-video editing. 7-day unlimited access starts on August 7.
显示更多
0
169
997
362
转发到社区
先是 GLM5.2,然后是 Kimi K3,现在是 Qwen 3.8 Max,每次国产模型的新发布都更加接近 Coding 领域的 SOTA。从 6 月开始,我感觉国内大模型明显开始在 Coding 和 Agent 领域加速了,和海外顶级 SOTA 的差距正在肉眼可见地缩小。 目前看,国产模型的长程任务、自主运行、反馈回路、跨 Harness 泛化和视觉自我检查等,能力越来越强了。做一个谨慎的预测,预计到 2026 年的年底,对于重度开发者和 Vibe 用户来说,海外模型可能会成为辅助模型,国内的大模型将成为我们的主力工具。 模型用户没有忠诚度,大家会用脚投票的,拭目以待。
显示更多