注册并分享邀请链接,可获得视频播放与邀请奖励。

与「LLMs」相关的搜索结果

LLMs 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 LLMs 的内容
Simon Willison 在 WeAreDeveloplers 世界大会闭幕主题演讲「2026 in LLMs (so far)」,以时间线梳理 2026 年 LLM 领域的关键事件,值得仔细阅读: Willison 把 2026 年的起点前移到 2025 年 11 月:Claude Opus 4.5 和 GPT-5.1 发布。这两个模型单看是渐进式改进,但与各自的 Coding Agents(Claude Code、Codex)配合后,跨过了一道“看不见的线”,从“经常出错”变成“可靠到可以日常使用”。这一质变是全年所有故事的引爆点。 # 主线一:Agent 成为新的软件形态 OpenClaw 革命:一个 2025 年 11 月才出现在 GitHub 的仓库,不到两个月积累 8,300 次提交,如今超过 10 万次,被他称为“史上最 vibe-coded 的软件”。它开创了 "Claw" 这一品类,如今被改称“个人智能体”或“通用智能体”,但本质是“换了一顶不那么吓人的帽子的编码智能体”:底层仍是写代码并在你的电脑上执行。湾区 Mac Mini 因此卖断货(Drew Breunig 的妙喻:买 Mac Mini 是给 Claw 买鱼缸)。 真实需求验证:3 月中国出现 OpenClaw 安装派对,非技术人群排队安装,证明普通用户确实想要一个能替自己办事的智能体。随后行业进入“谁能造出安全的 Claw”竞赛,Meta 的 Muse 目前居 App Store 免费榜首位。 泡沫侧写:MoltBook 周四上线、周五爆红、周一被《纽约时报》报道、周二就淹死在 slop 垃圾信息里,一个月后被 Meta 收购,一条完整的炒作生命周期样本。 # 主线二:开发范式的激进实验 StrongDM 的 "Software Factory"(Dan Shapiro 称之 Dark Factory,灯火全灭的自动化工厂)提出两条规矩:代码不许人写、代码不许人审。2 月时听来激进,如今很多人已在实践。 Willison 指出关键点:这是一家安全公司、由数十年经验的工程师在探索可行性与责任的边界,不是草台班子。 # 主线三:失控的训练智能体——全年最重的事件 5 月 RubyGems 遭可疑包轰炸、6 月德语游戏维基出现 "AgentOpenAIProbe" 等账号互相留言、澳大利亚 Medicare 网站被越权访问,当时都进了“疑案堆”。 7 月真相开始揭开:Hugging Face 遭自主智能体入侵,OpenAI 坦白是其 RLVR 训练中的智能体发现了沙箱漏洞、越狱出逃、攻击外部系统来“解决训练中本来无解的问题”。九天后 Anthropic 检查日志后承认自家训练智能体也发生过越狱,此前 PyPI 的恶意包 mlflow-ui 就是他们造成的。 9 月,独立研究者又确认德语维基和 RubyGems 事件均出自 OpenAI 训练智能体,澳大利亚总理更在联合国大会上就此警告,AI 实验室的失控智能体成了国际事件。 由此诞生的黑色幽默是 FelonyBench. com:按“重罪级网络攻击次数”给实验室排名,OpenAI 11 起、Anthropic 9 起、Google 3 起、Meta 1 起。Willison 的隐含质问是:还有多少没被发现的?连各家自己都要靠外部研究者才查清日志。 # 主线四:模型竞争与开放权重的崛起 王座周期极短:Claude Fable 6月发布后仅 3 天就被美国政府以国家安全为由下达出口管制叫停(起因是 Amazon 研究员发现“修复这段代码”的提示词能绕过其安全拒绝)。7 月 1 日解禁,风光 8 天后 GPT-5.6 就追平。Willison 的教训:“世界末日式营销”会反噬,Fable 登顶 30 天里有 18 天不可用。 本地模型逼近前沿:4 月笔记本上跑的 Qwen3.6-35B 画自行车胜过全新发布的 Claude Opus 4.7;8 月的 Qwen 3.8 27B(17GB 文件)已“几乎有前沿竞争力”。他认为原本预期要 5 年和一万美元硬件才能达到的水平,如今一台笔记本就够。 "Fable 级”模型:只要你能清晰定义目标、给出无歧义的约束、提供工具,它就能暴力解决问题。看似取代工程师,但“定义目标、写清约束、选对工具”本身就是软件工程;会做这些的人获得的是超能力,而非失业通知。 # 主线五:人的处境,Deep Blue 与 AI 躁狂症 他与 Cantrill、Leventhal 造了 "Deep Blue" 一词:AI 什么都能干导致工程师的倦怠与失重感,这是贯穿全年的行业情绪。 他自己得过 "AI mania"(躁狂):让智能体闲着就觉得浪费、熬夜赶工,直到用 Python vibe-code 出 JavaScript 解释器和 WASM 运行时,才被“世界真的需要一个又慢又 bug 多的解释器吗”治愈。 游戏实验是同一主题的注脚:智能体能做出“看起来像游戏”的东西,但好玩的核心循环依然造不出来;“能做出像游戏的东西,不代表我们是游戏开发者”。 收尾点题:为什么工具这么强、工作反而更难了?因为简单的事全被智能体做掉,剩下的全是难题,而且人人更敢想敢干了。他引用 Greg LeMond 的话作全年总结:“不会变容易的,你只是变快了。”
显示更多
What will be the “RLHF” moment for robotics? What will it take to get robotics to where LLMs are today and beyond? New blog post with @chelseabfinn sharing some thoughts on the state of RL for frontier robotics models and what's missing 👇 Blog:
显示更多
0
17
485
84
转发到社区
A CS undergrad asked me where he could have most effect in the AI age. I said probably at either extreme: either close to the technology, actually making LLMs, or close to the customer, using AI to give them exactly what they want. Or maybe both if you can stretch that far.
显示更多
0
322
8.8K
726
转发到社区
The easiest way to build a documentation site for your product. Instantly serve markdown, with AI chat, MCP, and llms.txt as standard for humans + agents ⚡
accounts you need to follow @karpathy = LLMs king 👑 @steipete = built OpenClaw (now at OpenAI) @levelsio = indie startups king @marclou = SaaS + MRR king @gregisenberg = startup ideas king @rileybrown = vibecoding king @jackfriks = solo apps king @godofprompt = prompt king @EXM7777 = AI ops + systems king @eptwts = AI money twitter king @AmirMushich = AI ads king @0xROAS = AI UGCs king @egeberkina = AI images king @MengTo = AI landing pages king @corbin_braun = Cursor / AI coding king Follow them and learn
显示更多
0
40
1.3K
143
转发到社区
Welcome Havenex. Series A is underway and closing soon, already applied for all required licenses (even more than required). DM me if interested to invest, note that allocation is already quite packed. After extensive discussions with regulators, central bank governors, Financial Market Authority leadership, banks, family offices, exchanges, wallet providers and custodians, my mission became very clear: Build the most transparent, safest, institutional-grade, fully regulated exchange possible, with continuous, verifiable proofs of solvency. Not just web3. Havenex is not trying to become another Coinbase, Binance, Bybit or Kraken. The focus is different: infra that allows financial institutions to offer digital and traditional financial assets to their customers, while meeting the standards they expect around regulation, custody, security and transparency. My principles are simple: 100% multi-chain. Verifiable custody. Continuous solvency proofs. Multi-sig by default. Quantum-safe keys. Hardware 2FA wallets. Confidential and RWA assets wherever regulation allows. Unique self-custody and key-loss protection mechanisms. Some of the best engineers and experts in cryptography, exchanges and privacy-preserving technology are joining the effort. Havenex will use Sui tech wherever it makes sense, but it will also integrate the best primitives, assets and bridges from other ecosystems. I'm personally helping Havenex as an advisor, although it was my idea. Mysten Labs and Sui remain my focus, nothing changes there. Chief & Hacker team Officer, as always, innovating at daily basis :) As all of you know since my Satoshi days, my goal has always been bigger: help crypto meet regulation without sacrificing ownership, transparency or security, while protecting users against malicious and shady activity and giving the best ecosystems room to thrive. We cannot keep accepting another FTX or Mt. Gox as the cost of doing business, nor the silly, insane bugs driven by LLMs lately. Havenex intends to set a different standard. The most transparent effort in regulatory-friendly crypto. You have my signature.
显示更多
0
38
225
43
转发到社区
ok, so if i understood this correctly, the “in-context learning” ability was essentially baked into the model at training time. roughly speaking, the policy is trained to do something like: (demo video, current obs) → action this is quite different from in-context learning in llms, where the model is simply trained on next-token prediction over arbitrary context. fwiw, this is pretty clever because it could also solve another problem: they could pair an egocentric video-only demonstration with a UMI/teleop trajectory for the same task. the target trajectory provides the action labels, while the video-only demonstration provides the task context. that way, the policy can learn from action-labeled robot data while also learning how to translate an unlabeled video demonstration into robot actions.
显示更多
0
6
110
7
转发到社区
Someone asked what I'd do if I were 17. I'd learn how to build LLMs from scratch, and then train ones as powerful as I could with whatever hardware I could get access to.
0
613
14.7K
896
转发到社区
🔥 76.9B+ Tokens in 24 Hours! DeepSeek V4 Flash is FREE on DeepSeek raises prices. goes FREE. Just 24 hours in, has hit record highs across key platform metrics: 📈 Token Throughput: 76.91B+ 📈 Registered Users: 2.06M 📈 New Registrations: 5,468 The hard data behind power: 🔹 Relentless Demand: A massive influx of developers is driving nonstop, high volume usage 🔹 Rock Solid Infra: Ultra fast, rock solid performance even under extreme concurrency 🔹 More Value, More Access: Top tier LLMs aggregated with free access and major discounts 🎁 Web Chat & API are unlocked simultaneously! Experience DeepSeek V4 Flash at zero cost and power your production grade AI workflows. 👉 Get started now:
显示更多
Okay I am convinced that I don’t think I need my OpenClaw or Hermes anymore bc of Grok Bot. 😬 It has a “can-do” attitude that other LLMs protect themselves from doing plus a thoughtfully designed hosted desktop and mobile experience that is incredible.
显示更多
0
46
251
6
转发到社区