注册并分享邀请链接,可获得视频播放与邀请奖励。

meng shao (@shao__meng) “EverMind 开源了一个多智能体 Host Agent 框架「Raven」,定位是 "the harness of har” — TopicDigg

meng shao 的个人资料封面
meng shao 的头像
meng shao
@shao__meng
加入 November 2023
0 正在关注    0 粉丝
EverMind 开源了一个多智能体 Host Agent 框架「Raven」,定位是 "the harness of harnesses" @evermind Raven 站在 Claude Code、Codex 等 Coding Agent 之上:把任务拆解为 DAG,调度内置和第三方的专业 Agent 协同完成,并能迭代改进自己的 harness 本身,也就是递归自我改进(Recursive Self-Improvement, RSI)。 开源地址: 为什么是 "Harness of Harnesses"? 当前 Agent 生态的碎片化问题很明显:Claude Code、Codex 擅长编码、各类 Agent 各有所长,但没有一个统一层去编排它们。Raven 的切入点正是这个空档,它本身是一个“宿主 Agent”,接收复杂任务后: 1. 生成任务 DAG,明确依赖关系,可并行执行的并行执行; 2. 为每个节点分派最合适的 Agent(内置的四个之一,或通过 ACP / CLI / OpenAI 兼容 API 接入的第三方 Agent,官方内置了 Claude Code、Codex、Qwen Code、Kimi Code 等 13 个预设); 3. 沉淀为可复用的工作流,下次类似任务直接复用。 架构的核心模块 四个内置专业 Agent,各自有明确的职责边界和对应的 benchmark: · Raven-Research:自主深度研究、文献综述、带可溯源引用的结构化报告; · Raven-Code:Agentic 软件开发(SWE-bench Pro/Verified、DataAgentBench 等,README 称 DataAgentBench 以 0.8762 Pass@1 居首); · Raven-Design:视觉交付物,PPT、海报、图表、Web 界面(PresentBench、ArtifactsBench); · Raven-Oncall:无人值守的长时工作流自动化,跑数小时的实验、过夜监控、持续优化,官方称在 AI4AI 和 AI4S 基准上于质量、成功率、成本三项均优于 Claude Code。 这是理解 Raven 的关键:Oncall 这类“跑一整夜自己迭代实验”的场景,才是它区别于交互式编码工具的差异化所在。 Evolver 是 RSI 叙事的落地:它把 Agent 主循环拆为四个解耦的策略模块,Memory(本轮看到什么上下文)、Planning(用什么方法)、Capability(每轮可用哪些工具)、Action(下一步做什么及执行前判断)。一个实验性的 "Curator" 会重写这四块的实际代码(不只是 prompt,还包括工具、技能和判断逻辑),候选改动必须先在 benchmark 上打败基线、通过验证门禁才会被安装,且反馈中会剥离参考答案以防“作弊”。首次使用时 Curator 还会根据用户的任务描述组装一个定制化的 "Persona" 团队(一个主理角色 + 若干专家 Agent)。 EverOS Memory 提供跨会话的三层记忆:用户上下文、Agent 自身经验、世界知识。SkillForge 则统一检索本地技能库、记忆和 SkillHub 上约 11.4 万个 Skills。Proactivity 提供事件监听和定时调度(提醒、跟进)。 实证 Showcase · THRESHOLD:约 4 天、42 轮规划/开发/验证循环,全自主产出一款 Godot 4 FPS 游戏,外加海报、演示文稿和官网,这是“长时程多 Agent 协作”的展示; · RSI 实验:7 轮共 172 次 nanochat 训练全程无崩溃,在固定 20 分钟单 GPU 预算下将 val_bpb 降低 5.8%;另有 CFD(溃坝问题超调量降低三个数量级)和 FEA 极限载荷搜索的优化案例; · 发布物料(浏览器物理小游戏、16 页 deck、README 本身)均由 Raven 自己生成,这个“自证”传统与 Claude Code 一脉相承。
显示更多
Raven 0.2.0 — The Harness of Harnesses, built for RSI. 🐦‍⬛ One harness can't be best at everything. Raven combines its own specialist harnesses (Research, Code, Design, Oncall) with the agents you already use (Claude Code, Codex and more) into one team. And it's built for RSI, and not just at the skill level. The whole harness can be rewritten by AI: prompts, policies, strategy code, playbooks. Every sub-harness, including the orchestration layer itself, is its own instance that can be improved. With Raven you can: 1. Orchestrate many agents as one team. Raven's sub-harnesses and external agents work in one task graph with shared memory across sub-agents, powered by leading orchestration (0.963 Node F1 on the Multi-Agent Orchestration Benchmark). 2. Run long, complex tasks. Oncall and proactive execution keep work going for days, from scientific research loops to shipping a full Godot game. 3. Build vertical agents with RSI. Use Raven's RSI to develop and refine an agent for your domain, and we'll optimize it with you. Experimental for now; reach out to the Raven team(Discord: More in the video and slides below. Open source, Apache-2.0. (lots of work made with Raven lives there, and much of this launch's material was made with Raven too)
显示更多