注册并分享邀请链接,可获得视频播放与邀请奖励。

与「ml」相关的搜索结果

ml 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 ml 的内容
才发现 firecrawl 新出的这个开源的 anydoc 文件转换器非常棒:支持 Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF 等等转换为 GitHub-Flavored Markdown. Rust 写的,还有 Node, Python, browser, and CLI bindings,还有skills 支持,这是文件转换的大一统了吧。 还有在线的 demo 可以用一下: 简单试了下,效果很不错。
显示更多
从目前AI炒作逻辑来看,再次出现了,缺光,缺HBM,缺MLCC的市场局面!🧐 今天太阳诱电将公司全年盈利预期上调了 50%,大幅超出市场预期。我们知道太阳诱电是村田 $MRAAY 的直接竞争对手,是MLCC市场的龙二,龙二都这么牛的业绩,龙一村田可想而知! 👇这个配置策略,依旧很能打,有「存储」有「云」有「光」有「MLCC」! 本文由 @binancezh 赞助,「币安买美股:全球资产,零秒时差,一键即达」!
显示更多
0
24
18
2
转发到社区
加仓了330ml无糖可口可乐现货,我的逻辑很简单,存储再有价值也得放进手机电脑里,而玩手机玩电脑,都得喝可乐。 短期缺薯条,长期缺炸鸡,永远缺可乐。
显示更多
0
444
1.6K
64
转发到社区
来杯好茶摇一摇 唉?鼠鼠也要摇吗
0
155
7.5K
255
转发到社区
The scale of what Elon is building in Bastrop is not well understood.
0
23
839
28
转发到社区
转一下苏剑林老师对 K3 架构的复盘。 一句话概括,K3 = KDA + MLA + Stable LatentMoE + AttnRes。整套设计没什么炫技,核心就是在模型效果、计算效率和训练稳定性之间做取舍。 这里稍微解释一下: KDA,一种线性 Attention MLA,一种 KV Cache 较小的 Full Attention 变体 Stable LatentMoE,对 LatentMoE 做过稳定性改造的 MoE 方案 AttnRes,用可学习的跨层注意力,替代固定等权的残差累加 几个比较有意思的点: ① K3 同时使用了 KDA 和 MLA。另外,训练仍然用 Moonlight 版本的 Muon 优化器,Attention 权重改成 Per-Head Muon,每个 Head 独立优化。苏神说这不会直接提高效果,主要是数学和结构上更合理:各个 Head 本来就相对独立,不应该在优化时耦合在一起。 ② 解决MoE“容易炸”的问题。 LatentMoE 会先降维,再使用更多专家,最后升维,在训推成本大致相同的情况下效果略好。不过连续的矩阵投影也让训练更容易出现数值不稳定。 K3 把 SwiGLU 换成了 SiTU-GLU,用 softcap 压住异常激活;又在 LatentMoE 升维前加了一层 RMS Norm。这个 Norm 不只是让训练更稳定,在 Valid Loss 差不多的情况下,不加它,某些 Benchmark 会稳定变差。 (这里有超级多的技术细节,感兴趣的可以看苏神原文) ③ 关于MLA DSV4 看上去都换设计了,K3 怎么还在用 MLA?苏神的答案是,目前 MLA 仍然很难被全面击败。它在训练阶段是 MHA 形态,推理时 KV Cache 较小;在固定训练成本和 KV Cache 大小时,MLA 依然近乎最优。 它的问题是对 MTP(推测解码)不够友好。MTP 的思路是“用计算换速度”,而 MLA 在 Decoding 阶段本身就比较吃计算,再叠一个 MTP,两边就容易抢算力。 换别的方案也有代价。比如换成 128+128 的 GQA8,效果很难打赢 MLA,KV Cache 还是 MLA 的三倍多;换成 256+256 的 MFA(本质上是 MQA),训练和 Prefill 成本又会上去。 目前还没有一个简单的 Attention 设计,能同时占住效果、训练成本、Prefill、KV Cache 和 Decoding 计算量。在 KDA+MLA 的混合架构下,MLA 的部分问题得到缓解,所以 K3 最后还是选了 MLA。 ④ 苏神觉得,DSV4 也不算真正“抛弃 MLA”。 DSV4 看上去重新设计了 Attention,但底层仍然有 MLA 的影子。它采用的是 head_dims=512、K=V 的 MQA,这正是 MLA 在 Decoding 阶段的形态;再加上 Sparse + Compress,前者减少计算,后者进一步压缩 KV Cache,同时减少计算。 方向很激进,不过 Infra 也更复杂。 ⑤ K3 的 MLA 可以去掉 RoPE,是因为 KDA 已经隐含提供了一种广义的位置信息。 这只适用于 KDA+MLA 的混合结构,像 K2 那种全 MLA 模型,直接去掉 RoPE 还是会明显掉效果。 感觉苏神真正的观点是:大模型架构里很少有免费的升级。K3 的设计哲学不是找一个碾压所有方案的新架构,而是在效果、训练成本、Prefill、Decoding、KV Cache 和稳定性之间,找一个当前更合适的组合。 苏神原文里还有很多技术内容,非常值得一看。 原文地址:
显示更多
0
28
326
66
转发到社区
🚨 BREAKING: Apple filed for a PRELIMINARY INJUNCTION against OpenAI AND asked a federal judge to put them under forensic supervision "Apple respectfully moves the Court for a preliminary injunction to stop THE THEFT OF ITS TRADE SECRETS" Apple filed NINE sworn declarations, a 28-page memorandum and a concurrent motion for expedited discovery What Apple now says, under oath: Chang Liu: 8 years at Apple, now OpenAI "Member of Technical Staff" exploited an authentication bug to steal Apple trade secrets "on AT LEAST FIVE SEPARATE OCCASIONS" from February to April 2026, WHILE working for OpenAI Liu downloaded "THOUSANDS OF PAGES of Apple's most sensitive trade secrets" The stolen files, NAMED: >DisplayNotes.key — "several hundred pages" on Apple's custom display power development program >Architecture analyses. Fabrication decisions. Testing results >Engineering data for an UNANNOUNCED Apple product: 'touch, display, and power systems" >Final.key + V2.key — compilations of two undisclosed Apple R&D projects >and those are "only four of the dozens of proprietary documents Mr. Liu stole" Liu fed OpenAI "a steady stream of Apple proprietary information that he actively concealed" Liu also "coached Yu-Ting "Alyssa" Peng, then still INSIDE Apple, how to access and copy files from Apple workstations "to avoid trouble with the security team" and directed her to communicate with him on the encrypted LINE app "to avoid detection" Tang Yew Tan: 24-year Apple VP, now OpenAI's Chief Hardware Officer, "used an Apple internal project codename for an unannounced product to elicit still more trade secrets from job candidates." Tan's own messages, quoted in the motion: >"Just like last time, bring some parts you worked on" >"mlb, battery, shields type of stuff is interesting" OpenAI recruiter, quoted: "No, you won't sign anything at the exit interview. If they do ask you to sign anything, let me know asap." APPLE TOLD FEDERAL JUDGE: >"OpenAI knows its misappropriation is wrong and has tried to conceal it." >"This is not a case of 'mere hiring'... it is a case of repeated instances of deliberate theft." Apple says OpenAI went after its SUPPLIERS: >OpenAI "directed a trusted Apple partner [name redacted] to perform [Apple's proprietary metal finishing] process for them, knowing it was proprietary to Apple... because they were involved in this partnership while at Apple." Apple put its own Surface Finishing Manager, Jackie Hughes, under oath to prove it. Apple named ELEVEN MORE former Apple employees at OpenAI — beyond Liu, Tan, and Peng — Fourteen people total. Apple also filed a concurrent motion for EXPEDITED DISCOVERY demanding depositions: - Liu. Tan. Peng. - A fourth unnamed OpenAI employee - Plus OpenAI itself, under oath, through Rule 30(b)(6) Apple has asked a federal judge to put OpenAI under forensic supervision RIGHT NOW: >Forensic inspection of ALL OpenAI devices >ALL cloud storage, Slack, email >Including anything that "previously contained" Apple data — deleted included Demanding the "first available hearing date," citing "imminent threat" to its trade secrets. APPLE: > "The harm is happening now — every day that passes without an injunction allows OpenAI to embed their knowledge of Apple's stolen information into its hardware development efforts." Hearing: October 1, 2026. Judge Edward J. Davila. ITS HAPPENING
显示更多
0
41
190
30
转发到社区
"Almost every team in baseball, with good management and a little bit of luck, can contend [for a title], at least for a period of time." @realbobcostas breaks down what separates MLB's best franchises ⚾️
显示更多
0
67
249
19
转发到社区
The 🐐 turns 49 today. Happy Birthday @tombrady!