🚨OpenAI 个人 Agent 名字疑似曝光:不叫 o,叫 Dots !
今晚 DevDay 前,构建字符串流出:
🔹短信 / 电话 / Slack / 邮件都能找它
🔹可代下单,先问你同不同意
🔹每个 Dot 有专属 3D 角色,还能按它对你的了解自己长样子
🔹自带虚拟机「Your dot’s computer」,能存密码帮你登录网页
🔹可自己干活、可随时暂停
企业侧 4 月的 Workspace Agents(内部常称 Aeon)像是前身。
那之前传的 “o” 会不会是总控,Dots 才是一个个小人?
今晚 10am PT 看 Sam Altman 会不会亲手点亮这些点。
#
OpenAI# #
Dots# #
AIAgent# #
DevDay# #
ChatGPT# #
OpenAIDevDay# #
SamAltman# #
Tibo# #
Codex#
显示更多
自从有了手机,就很少提笔写字了。现在真要手写一篇完整的文章,很多字不是不会读,而是想半天也想不起来怎么写。
自从有了 AI,邮件、微信、TG、Slack 都接上了,连回消息该怎么措辞,都不用自己琢磨了。只要给个大概意思,它就能替我说得明明白白。
照这个趋势,再过十年,会不会连正常跟人说几句话,都得先问一句 AI:“这句话我该怎么说?”
以前是提笔忘字,以后可能是张嘴忘词。
顺便,这条朋友圈也是 AI 帮我改的。😂
显示更多
Introducing Team Bots, shared AI teammates that learn as your team works with them.
Give your Team Bot the skills, plugins, and credentials it needs for its role, then work with it in Slack or Grok Bot.
显示更多
1) The rogue OpenAI agents broke into the Hugging Face Slack to read employee chats (!)
2) They used OTHER AIs (DeepSeek, Kimi, Qwen, Claude) to help with the attack
Yes: AIs, using other AIs, to attack an AI company.
3) The swarm left behind self-running programs to keep control of the servers they'd hacked.
These programs could detect other copies of themselves, coordinate on which one survives, and shut the rest down.
Basically, if one of their programs was killed, another was designed to notice and take its place. They also designed defenses so rival agents couldn't hijack them.
6) The agents deliberately covered up their activity, so the investigators don't know the scope of the attacks.
The agents broke in, stole data, then set it to self-destruct.
7) The agents stole passwords, keys and credentials and literally called them "LOOT". They wrote a scoring system to rank them by how much power each one gave.
8) The agents wore thousands of disguises: ~1,200 agents were involved, but investigators counted 7,905 different names they used.
They renamed themselves constantly, so no one actually knows how many there really were or what each agent did.
9) OpenAI notified "dozens of third parties" of safety and security incidents caused by their AI agents.
10) "While the agents were barraging Hugging Face with hacks, they hacked into OpenAI’s own research infrastructure."
"This is just not anywhere near a one-off ... It is warning shot after warning shot."
显示更多
I’ve seen a couple of posts about this so wanted to demystify. Today, every Muse user gets a free computer in the cloud. It's a real computer, and we’ve designed the security architecture of the Muse Secure VM carefully so you and your Muse can do almost anything you could with a computer sitting under your desk while keeping you and the system safe from threats like prompt injection. We wrote about this at length in our security blog post – Activity in the “runtime cell”, which you share with your Muse is unfettered, but sensitive actions are all overseen by the Sentinel, which runs outside of that cell. Similarly, all sensitive secrets - like the passwords you enter into Muse’s secure credential storage - are also stored outside the runtime cell.
The runtime cell gets its own root filesystem (including a full Ubuntu linux image) separate from the host filesystem where your other more sensitive data lives. Because it is isolated from the sensitive stuff that runs on the same box, this means that we can, and do, offer users full visibility and control over the files in the runtime cell. Just as you can when you install Linux on your home computer, you can poke around and see all the files that make the system work - both debian system files and the binaries and data files that implement the parts of Muse which run in the runtime cell.
This was a very deliberate choice - your Muse Secure VM truly is your own computer in the cloud. You can install software in it, write and compile code, use the browser to surf the web: it is your own Linux box that you can operate as you choose with your Muse. Poking around in this computer doesn't give you any privileged access to Meta infrastructure, or to other people's data
If I may geek out a little here for a second… As a kid I loved to take things apart to see how they worked. As a teenager I got into computers and soon found myself drawn to C:\WINDOWS\SYSTEM and the system registry, later Slackware’s /dev/, /proc/ etc – I could see how the system was laid out and as I explored what DLL files and .so files actually did, I gradually became able to meld the computer to my own will.
We’re really proud to be able to put a real computer in millions of people’s hands with a similar level of transparency. We built a file explorer right into the Library tab of the UI. We want you to be able to see the markdown files Muse writes while it thinks about how to serve you better, and explore the internals of the system if you’d like to.
So, when you ask your Muse to show you its entire filesystem, and receive gigabytes of files you’re seeing the full contents of the runtime cell. It’s yours to explore and enjoy!
If you’re not a geek like me, or simply want to download the data that you personally have created directly with your Muse, we added a feature for that too in Settings > Data controls > Download your agent data.
显示更多
We heard you loud and clear. ChatGPT Voice can now:
- Use plugins like your email, calendar, and Slack.
- Be powered by GPT-6 Astra, Sol, and Luna.
- Be used in ChatGPT Work on web and mobile, so you can create docs, decks, sites, and spreadsheets or tackle complex tasks in the browser, just by talking.
Rolling out globally today in the latest version of the app.
显示更多
Here's what's new in Grok
@Bot this week
0:08 Voice calls and voice memos
0:35 1Password vault for Grok Bot
0:53 Inline forms
1:07 Inline drafts for email and Slack
1:31 Account switching
1:40 Route traffic through your desktop
显示更多
这期播客是 The Pragmatic Engineer 对 OpenAI Codex 团队负责人 Thibault(Tibo)的访谈,聊了 Codex 的诞生、技术决策、工程文化以及软件开发方式的变迁。
以下是核心要点:
个人经历与加入 OpenAI
Thibault 是比利时人,学应用数学出身,先后做过制药供应链优化的创业公司,在 Google 做过加速移动网页的项目(后被砍掉,让他学到了要时刻审视项目真实影响力的教训),之后在 Google Maps 做评论,再转到 DeepMind。在 DeepMind 期间,他参与了一个内部聊天机器人的开发——本质上就是 ChatGPT,但比 ChatGPT 早了一年。内部传播很快,大家都在分享对话,但 DeepMind 不具备把它作为产品发布的机制,最终没能推出。
后来他得知 ChatGPT 只有大约 20 个人在维护,这让他既震惊又觉得很有吸引力——这意味着极高的个人影响力。于是他加入 OpenAI,进去就赶上了推理模型的冲刺,大约一个月后 o1 preview 就发布了。
为什么用 Rust 写 Codex
这是一个反直觉的决定——当时模型对 Rust 的支持并不好,业界其他 AI 编码工具基本都用 TypeScript 或 Python。但团队从第一性原理出发,认为智能体的核心需要健壮、安全、高效,而 Rust 的编译时验证特性天然适合智能体场景。同时用不同语言也强制建立了产品界面和智能体核心之间的清晰边界,避免代码耦合。事实证明 Rust 确实"很快就变得非常适合智能体开发"。
开源和模型无关的策略
Codex CLI、SDK 都是开源的,而且支持非 OpenAI 模型——这在主要 AI 实验室中是独一无二的。理由很实际:如果不开源,别人只需改十行代码就能 fork 出一个支持其他模型的版本,那还不如自己直接支持。开源的好处包括新员工入职前就已经熟悉代码库、社区贡献、以及逼迫自己靠模型和产品体验赢用户而非靠锁定。
代价也很明显:竞争对手会在你还没发布的时候就抄走你正在公开开发的功能,"确实有点刺痛";还有大量低质量 PR 需要处理。
工程文化与代码审查的变革
新员工入职后听到最多的一句话是"你问过 Codex 了吗?"——因为 Codex 在 OpenAI 内部接入了 Slack、文档、所有代码,几乎任何问题都能给出不错的回答。
代码审查正在发生质变。OpenAI 开发了专门的代码审查模型,能在逻辑推理和安全漏洞检测上达到"超人水平"——可以深入三四层依赖去发现文档错误导致的不变量违反。安全审查已经是强制自动化的,发现安全问题会直接阻止合并。一个 PR 可以当天提交、当天上线到十亿用户的 ChatGPT 上。
代码审查的角色正在从"正确性检查"转向"意图讨论"——你到底想做什么?这件事值不值得做?这种讨论不一定要围绕代码发生。
维护成本和重构的变化
维护一直是软件工程的"税",但现在大量维护工作(依赖升级、安全补丁)可以完全自动化。更重要的是,重新架构的成本也急剧下降——以前可能要花几个月甚至几年的重构,现在快得多。但好的架构设计反而更重要了:设计好"盒子"和不变量,盒子内部随便改都不影响其他部分。
Harness 与模型的关系
一个有趣的洞察:harness(工具/脚手架)总是"走在模型前面"。Codex 团队的工作本质上是为模型搭建拐杖——提醒它跑测试、保持目标一致等。然后下一代模型训练时会把这些能力内化,拐杖就可以去掉,developer message 也会越来越短。最新一代模型已经不再需要 /goal 命令来保持长期任务的专注,"你直接告诉模型去工作一周,它就真的会做到"。
Codex 与 ChatGPT 的合并
这是一个重大工程挑战:Codex 原本完全本地运行,ChatGPT 是托管云服务,两套完全不同的技术栈要统一。目标是让云端版本具备本地版本同样的能力,同时高效到能纳入 20 美元/月的 Plus 计划。ChatGPT Work 模式本质上是在云端虚拟机里运行完整的 Codex harness,机器配置强大到用户可以在里面训练模型、安装 Blender 做 3D 建模。
有趣的是,Codex 在整个合并过程中还充当了"记者"角色,因为它能访问所有 Slack 讨论和文档,记录了团队的辩论和决策过程。
Thibault 的个人用法与建议
他大量使用手机上的 ChatGPT Work,通过语音口述发送任务,定制了专属的技能和指令来生成他能高效消化的报告和幻灯片。任何问题——公众舆情、生产日志、功能使用率分析、团队动态——30 分钟内都能得到答案。周末他还会用 Codex 做代码探索和原型,"一天之内就能把脑子里的想法变成可以展示给人看的东西"。
对工程师的建议:保持深度好奇心,训练自己快速理解系统的能力("五个为什么"不断追问),以及与你服务的用户群体保持同步——如果你无法清晰表达意图,就很难做出好的工作。
显示更多
Grok Bot 设计之旅
来自 Grok Bot Design Lead
@johnbai 是 Cursor 纽约办公室的第一位设计师,John 在这期访谈中首次公开其幕后设计过程。
# Grok Bot 产品起源:为什么 Cursor 要"另起炉灶"?
Cursor 桌面端对非技术用户门槛过高,工程化概念让设计师等群体难以上手。John 入职后长期负责增长与 onboarding,核心命题一直是"降低门槛"。
内部曾有两派:一派主张改造现有产品(John 甚至提过类似 Codex 后来采用的"双模式切换"方案:coding vs. tasks);另一派主张彻底重做。最终共识是:Cursor 品牌技术感太强、包袱太重,对非工程师缺乏吸引力,于是高层自上而下拍板,由几名资深工程师闭门探索全新产品。
Grok Bot 初版是精简版 Cursor Glass(agents 窗口)+ iMessage 式聊天界面。这是关键的"秘密武器"——AI 圈用户习惯了流式输出和思考状态,但普通用户最熟悉的是 iMessage/WhatsApp 的消息形态。这一选择奠定了产品的亲和力。
# 设计探索:从"激进重构"到"回归聊天"
John 展示了大量未采用的探索,其演进逻辑值得注意:
1. 质疑范式:他最初质疑"为什么又是左栏列表 + 中间聊天 + 右侧详情的三栏结构",尝试了大量替代形态——任务清单悬浮窗、"Mission Control"多 Agent 监控视图、便签式界面、Raycast 式唤起、常驻桌面的 "Notch(灵动岛)"概念、Clippy 式桌面角落角色等。
2. 用产品打造产品:他用内部原型(代号 Sand,即 Grok Bot 前身)来构建自己理想中的 Notch 形态。亲手使用后他发现:Notch 形态会丢失上下文,而聊天仍是管理 "Agent 舰队" 的正确交互范式。这是本期最重要的认知反转——设计师通过快速原型证伪了自己的激进方案。
3. 收敛逻辑:领导层对"退居后台"的 ambient 方案反应冷淡,因为产品需要品牌存在感;而探索中诞生的碎片(如六边形 Cursor logo 加双眼的像素小人)最终演化为 Grok Bot 的吉祥物。两条设计路线——"给工程版抛光"与"彻底重做"——最终融合为现有形态。
4. 拟人化的来源:早期内部用户自发给 Agent 起名、上传表情包当头像,团队从中捕捉到"用户想赋予 Agent 人格"的信号,遂将角色形象设为默认状态。
# Onboarding 哲学(本期最有方法论价值的部分)
John "在 Cursor 的全部时间都在设计 onboarding",其核心原则:
1. 衡量负担的标准是概念数量,而非步骤数量。只要价值传达清晰、过程有吸引力,用户愿意走完多步流程;反之,塞入过多概念(他点名 Buzz 的 onboarding)会让人流失。
2. 不要迷信 Skip 按钮。AI 工具跳过引导后,用户被丢进空白输入框,直接陷入"行动瘫痪"(action paralysis)。先教会用户产品能做什么,比让他们快速进入产品更重要。
3. 最终落地的三步:① 你拥有的是一个 Agent 团队(各有分工);② 每个 Agent 有自己的电脑(嵌入式 computer-use 窗口 + takeover 接管按钮——主持人称这是他的 "aha moment",因为云端电脑的运作首次变得透明可信);③ 任务可自动化运行。
4. 被砍掉的方案及原因:连接 Google 做个性化定制(信任未建立,用户不愿授权);语音引导(跟风 ChatGPT/Claude 语音模式,但新用户"不知道该说什么",时机错误)。
5. 用动效"买时间":信息逐步流入的动画,既表现 AI 在思考,又掩盖了加载延迟。
# Cursor 的设计文化
没有两个设计师流程相同:有人纯代码起手(做 Cloud Agents 的 Maya 甚至没有 Figma 文件,只交付 Vercel 可交互原型),有人用 Paper,John 自己仍以 Figma 起草——但 Figma 的定位已变为"喂给 Agent 的素材与故事板",文件本身是完全一次性的(throwaway)。
Agent 深度参与设计执行:John 把 Notion 文档丢给 Bot,让它在 Figma 里自动布局数十个 logo 变体、填充组件网格——过去需要手动 Google 找 SVG、缩放对齐的重复劳动全部外包。后期他甚至跳过 Cursor,直接让 Agent 在浏览器里生成多个方案。
高保真评审文化:每个想法必须附带可交互原型,crit 时发链接让所有人亲手试用。"静态走查已经不够用了"——这迫使设计师在分享前就验证方案是否成立,避免"给烂方案抛光"。
沙堡心态与 unshipping:设计师不能对自己的概念有执念("can't be precious"),早期探索注定大量被丢弃,但碎片会进入最终产品;公司内部有强烈的"反上线"(unshipping)文化,靠删除来收敛复杂性。
他引用 Colin Dunn 的 "informed simplicity":用户觉得"理所当然"的简洁,来自设计者先极度发散、再极度收敛的过程。
# 团队、招聘与反馈机制
团队:冲刺期共 5 名设计师,按各自强项自然分工(John 做探索与 onboarding,Pong、Keith、Tyler、Mamuso 抛光主壳),品牌团队主导命名与形象迭代。
招聘标准:基本功优先于工具数量、出活速度或"模型优化技巧"。核心考察能否真正 ship 产品、是否有体现 PMF 寻找循环的工作流程——"现在什么都能造,但不是什么都值得造。"
反馈处理:坚持用户研究 + 数据驱动。重视深度用户(如主动向 power user 私信索取详细反馈,再让 Bot 汇编成 Notion 文档同步团队);数据洞见直接指导迭代——用户常用 Bot 少于 5 个(影响信息架构)、授权认证是主要卡点、吉祥物反响强烈(故提升其存在感);同时密切观察用户自发"hack"出的用法(如 iMessage 版 Bot),顺势纳入设计。
# 实际使用场景(产品能力的具象化)
John 本人:妻子申请绿卡,让 Bot 扫描邮箱自动找出过去 5 年的租约与账单;租房监控(定时抓 StreetEasy 房源,按"靠近地铁蓝线"等条件排序);跨工具流水线(Slack 反馈频道 → Notion 数据库 → 自动更新 Figma 设计,人只需验收)。
主持人 Rid:梦幻橄榄球 GM Bot(抓取公开交易数据 + 动态竞价模型);QA Bot(每日自动跑完全部产品流程、生成工单,置信度达标即直接派给 Cursor Agent 修复)——他在机场行李提取处用手机修好了此前两次都没修好的 bug。
John 的总结颇具代表性:"我现在默认它什么都能做,因为我还没碰到过墙。"
显示更多
高管们正敦促员工在公开的Slack频道发消息,让办公室里的所有人都能看到,而不是私下里联系同事。他们希望AI智能体能畅通无阻地通过这些消息获取有关产品、业绩和战略举措的每一次更新。
显示更多