阿里云开源「企业级 Agent 白皮书」
2026 年最新发布,是 2025 年 9 月「AI 原生应用架构白皮书」的升级续作。全书按 架构 → 构建 → 运行 → 治理 → 调优 的全生命周期组织,共 7 篇 30 章,由阿里云数十位一线工程师分工撰写,并纳入吉利、塔斯汀、MiniMax、哔哩哔哩、信永中和等外部企业案例。
它的写作动机很明确:过去一年市场重心已经从 “如何快速搭出一个 Agent” 转移到三个新挑战,工程化(从概率智能到可靠生产力)、规模化(从单点试验到智能基础设施)、组织化(从 Agent 孤岛到进入核心业务流程)。现有的框架文档和教程基本不回答这些问题,这本白皮书填补的正是这个空白。
开源地址
# 各篇核心内容
架构篇(1–2 章) 建立认知框架。给出 Agentic Application 的六个判定特征(以任务结果为中心、运行时决定部分执行路径、能作用于环境、维持跨请求状态、受确定性机制约束、可观测可评估)和成熟度四级模型(L1 辅助生成 → L2 受控自动化 → L3 Agentic Execution → L4 规模运营)。两个重要的解耦判断:用哪种形态取决于任务结构,处于哪级成熟度取决于治理完备程度;单 Agent / Long-Horizon / 多 Agent 是沿时间跨度和协作结构两个正交维度的扩展,不存在“必须升级到多 Agent”的路径。贯穿的原则是“最低充分架构”,为任务选择成本与风险可接受的最低复杂度。
构建篇(3–6 章) 是方法浓度最高的部分,按“范式—任务—信息—行动”还原构建过程:
· 任务:Agent Loop 五阶段(Prepare→Model→Act→Observe→Verify)+ 十态任务状态机,要害是“消息历史不应是任务状态的唯一来源”;完成判定的核心原则是“模型只能申请完成,Harness 依据环境证据提交完成”,验证分五级并与风险匹配。
· 信息:Context 是动态“编译”而非静态字符串。本章的独创设计是 Context Manifest,每次调用记录上下文每个片段的来源、作用域、版本、信任级别、选中理由和内容哈希,使“模型看见了什么”变得可解释、可回放、可审计。信息被五分为 Context/State/Memory/Knowledge/Skill,其中 Memory(个人经验)与 Knowledge(组织内容)必须分列,因为治理责任不同,“放进同一个向量库会同时失去两类治理能力”。
· 行动:统一 Action Plane(意图→Schema 校验→身份绑定→策略决策→执行→观测),关键三分:“模型看见工具 ≠ Harness 注册了工具 ≠ 获得执行授权”。协议定位清晰:Function Calling 是模型-Harness 意图接口,MCP 是 Harness-能力提供方连接协议,A2A 面向拥有独立任务循环的远程 Agent,“协议选择由能力是否拥有独立任务循环决定,而非新旧或流行度”。
运行篇(7–12 章) 处理规模化后的工程问题,大量内容达到了分布式系统的专业深度:沙箱后端选型判据(容器/gVisor/MicroVM 按代码可信度与租户边界取舍);状态外置后 Event Log / Checkpoint / 工作区快照三者不可互相替代,且“Durable Execution ≠ 外部动作恰好执行一次”;AI 网关对 LLM/MCP/Agent 三类流量按不同粒度治理,其中“严格预算需要原子预留而非阈值检查”的数学化分析(余额 100、两笔 80 的并发请求都会通过)是真实的并发工程细节;多 Agent 编排强调“最小充分共享”,共享的是上下文来源而非同一个 Context 窗口。
治理篇(13–16 章) 让自主运行的系统变得可信。可观测性的判据是“请求成功 ≠ 任务成功”;安全章同时把 Agent 当被攻击对象和行为主体来防护(身份是全章最扎实的部分:数字工牌、Token Exchange 权限收敛、On-Behalf-Of 且 Agent 权限 ≤ 用户权限);资产管理把 Prompt/Skill/MCP/Agent 当作运行时依赖做注册与版本治理。第 16 章 Agent Simulation 是全书原创性最强的一章:Agent 行为之所以不可验证,是缺制度前提(角色无外部标准、失败无自然代价、身份不连续),模拟是当下唯一可做的事,本质是“用可靠 Harness 约束不可靠内核”。它甚至给出诚实的统计学提醒:n 次零违规的 95% 置信上界约为 3/n 而非零。
调优篇(17–24 章) 的组织原则是“归因决定方法”:先排除环境故障、再修 Harness、最后才动模型,“把本应由上下文或工具协议解决的问题当成模型不行,是代价最高的一类误判”。主线是数据飞轮:Trace→Trajectory→黄金数据集(输入/轨迹/结果/判据四要素)→Badcase 闭环→受控自进化(模型生成的改进一律是候选变更,须回流构建、过门禁、可回滚)。模型调优章对 SFT/Agentic RL/蒸馏的适用边界、奖励投机的三套机制分离(训练奖励、独立评测、系统硬约束)论述相当严谨,广引 DeepSeek-R1、Tulu 3、FrugalGPT 等外部工作。
总结篇(第 30 章) 是全书思想密度最高的总结。当企业同时运行多 Agent、多框架、多租户时,同样的工程要求在每个应用里被重复且不一致地实现,这本质上是缺一个共享的系统层。Agentic OS 被给出“窄定义 + 三条否定”:为 Agent 任务提供公共运行对象、能力接入、可强制边界与统一证据的系统层,它不持有任务语义、不是又一个框架、不必然改内核。能力下沉有三条判据(复用性 + 强制性或可验证性),九类管理对象(其中 Budget Lease 预算租约最易被忽略),并提出“自治上限由可撤销范围与可证明范围决定,而非模型能力”。
调研报告 的 1906 份问卷给出一个关键发现:已开发或开发中 Agent 的企业占 46%,但真正上生产的仅 18%;有评估体系的企业任务成功率是无评估者的约两倍,卡点不是模型能力,是 Harness 层的工程配套。这与全书立意互为印证。
显示更多
I have to say - for me it started 2 years ago. I remember the moment so well.
I was on vacation in Montenegro. Early night. Everyone is asleep and I am on the balcony watching the cruise ships in Kotor bay as the horizon faded to dark. And I think it was Sonet 3.5 (don't quote me).
I had just setup a plugin in NeoVim called Avante by
@yetone . And it was magic.
Coding with AI at the time wasn't the same - I was going function by function. I was accepting diffs one hunk at a time. But I remember that night - from 10 to like 2 in the morning - I built what I thought would have taken me a week (I was rusty and learning the stack). It was an internal tool - a partial JSON parser / repair tool - so you can take half finished JSON and safely turn it into something that would parse. It is more complicated than it sounds.
Anyway. I was hooked that day.
At the time I was an executive at a Fortune 50, so coding was not in my job description as such. I had hundreds of engineers... but I have to say - it relit my passions, and now as I am (hopefully) building another startup - I still love it. And I still yell at it. And constantly hunt where it is messing up the architecture and adding bloat and.... well you know how it is.
显示更多
Locked in 🔒. Big 12 on the horizon.
📢 Xiaomi MiMo-V2.6 Series Models Are Now Live on
Developed by Xiaomi MiMo, the newly released V2.6 series features open-weights native multimodal reasoning models with a 1M context window and full text/image/video/audio understanding:
🔸 MiMo-V2.6-Pro: Flagship 1.02T total / 42B active parameter sparse MoE architecture, engineered for repository-level software engineering, long-horizon agent workflows, and complex multimodal reasoning.
🔸 MiMo-V2.6-Flash: High-efficiency 309B total / 15B active parameter architecture, optimized for high-frequency office automation, fast agent execution, and cost-sensitive workloads.
Now available on both API and Web Chat!
👉 Try now:
显示更多
📢 DeepSeek-V4.1-Flash Is Now Live on
DeepSeek’s newly released native multimodal MoE model, DeepSeek-V4.1-Flash, features a 552B parameter backbone and a pioneering Causal Encoder-Decoder architecture. Supporting a 1M context window and up to 384K output, it is engineered for long-horizon coding agents, multimodal software workflows, and high-concurrency agentic systems.
Web Chat and API access are now live. Stay tuned, even bigger perks are coming soon!
👉 Try now:
🔗 Learn more:
显示更多
Comparison between GPT-Images 2.5 vs GPT-Images 2.0
Here's the prompt:
"Animate the supplied pixel-art heroine in one continuous 8-second side-view shot. Preserve her exact design, colors, proportions, crisp pixels, and outlines. Lock the camera and pale blue-gray background. Keep her full body near center, facing right, with no more than 25% horizontal travel. From 0–2.5s, run mostly in place with alternating strides and arm pumps. From 2.5–4.5s, perform one two-foot jump and land crouched. From 4.5–6.5s, perform one full forward ground somersault: hands down, tuck, roll over the shoulders, feet over head, and land on both feet. From 6.5–8s, rise into a balanced ready pose. Animate at exactly 6 FPS with 167ms frame holds, normal speed, and no interpolation, blending, optical flow, or blur. On every frame, shift the whole sprite slightly and unpredictably on both axes to create strong constant jitter, especially in the final pose. Only the heroine moves; the camera and background remain still. No cuts, cartwheels, aerial flips, scene changes, text, particles, weapons, duplicates, audio, 3D, color flicker, morphing, or design drift."
显示更多
everyone's got that one rule they swear by. "just stay above the 200 day moving average and you dodge every crash, easy money, sleep like a baby." never seen anyone actually back it with numbers though, so i had horizon run it instead of trusting vibes.
dead simple rule on QQQ going back to 2018. long when price closes above the 200 day sma, flat the second it closes below, back in when it reclaims. no extra filters, no discretion, just the line doing the talking.
9 trades total in over 8 years. thats the whole strategy. 7 winners, 2 losers, 77.78% win rate, +268% total return, ~17% CAGR, profit factor of 7.05. horizon scored it 64/100, "viable." ngl thats a solid card for something this basic.
heres the part that never makes it into the "just follow the trend" tweets though. even with all that green, the equity curve was sitting underwater from its last high 78.77% of the time. one single trade in april 2022 ate a straight -21% before the rule even caught the whipsaw. it's not a smooth line up and to the right, its long stretches of nothing followed by a handful of trades doing all the work.
the rule isnt fake. it's just way less comfortable to actually hold than it sounds when someone posts it as a golden rule with no context.
built and backtested the whole thing on horizon in plain english, no spreadsheet, no coding it myself. you can go try it yourself and test your own strategy
显示更多
吴说获悉,Midas 在 Aave 治理论坛提交 ARFC,提议将 mWIN 作为抵押品纳入 Aave Horizon。mWIN 由 Midas 发行、Wellington Management 管理,是代币化多部门主动管理固定收益投资组合,已在以太坊主网上线。提案称,mWIN 面向白名单机构及合格投资者,底层资产包括 CLO、CMBS、RMBS、ABS 及投资级公司债等;LlamaRisk 正进行独立风险审查,风险参数将在 Snapshot 投票前公布。
显示更多
Ever seen the Summer Palace glow like this at sunset? As the sun dips below the horizon, fiery colors fill the sky and bathe the red walls and historic architecture in a warm golden light. Paired with the beauty of a traditional Chinese garden, it's a picture-perfect summer
显示更多