注册并分享邀请链接,可获得视频播放与邀请奖励。

与「CodingAgent」相关的搜索结果

CodingAgent 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 CodingAgent 的内容
我从2025年就一直反复强调。 现在所有大学本科生最重要的第一节课,就是买一个最大的coding plan,用上claude code或者codex, 第二节课是自己做一个最最最小版本的coding agent,可以对比codex或者claode code的基本功能,只要能输入一个基本功能,iteratively让agent完成写代码、编译、测试、 运行的功能即可,一切在terminal里,先把terminal和tool calling功能做好, 第三节课是认真观察codex和claude code的基本功能,把里面的memory、skills、multi agent/subagent、background tasks、session管理、context compression、TUI/GUI设计、如何可视化diff、如何管理好额外的btw等等类似的功能、如何把goal的功能放进去、如何实现scheduled tasks、如何实现权限管理等等,一步步一点点摸索实现出来。 我反复讲,一个计算机本科生能看完立党AI研究学习教程,把上面这三节课做完,就已经吊打清华计算机80%以上的本科生了。
显示更多
0
85
944
208
转发到社区
如果对标 Coding Agent 的模型跨越,Opus5.5 成功的把 AI 剪辑的时间点,从 Sonnet 3.5 跨域到 Opus 4.6 。利好 @chatcutapp @hypitai 如果你不信,请看我的下一个视频~
显示更多
阿里把团队内部用了两年的官方 AI Code Review Skills 开源了,采用 “确定性工程 pipeline + AI Agent” 的混合架构,专门解决通用 Agent 做代码审查时 “漏审、定位漂移、质量不稳” 的老问题。 40.5K ✨ 开源项目 OpenCodeReview: # 核心设计:确定性工程 pipeline × Agent 各司其职 确定性工程负责硬约束: · 精确文件选择:用代码决定哪些文件必须审、哪些要过滤,不依赖模型自觉; · 智能文件捆绑:把相关文件合成一个审查单元(例如 message_en.properties 和 message_zh.properties 捆绑),每个单元以上下文隔离的 sub-agent 运行,分治策略让超大变更集也稳,且天然支持并发(默认 8 个文件 worker); · 细粒度规则匹配:内置约 54 个按语言/文件类型的规则文档(Java、Go、TS/JS、Python、Rust、SQL/XML mapper、properties 等),用模板引擎而非自然语言把规则匹配到文件特征上,从源头消除信息噪声; · 外部定位与反思模块:评论的“落点”和“内容”分别由独立的 re-location 和 reflection 模块系统性校正,这正对“位置漂移”痛点。 Agent 负责动态决策: · 深度优化的场景 prompt(内部分为 plan → grouping → main → memory_compression → re_location → review_filter 多个任务模板,可在 internal/config/template/prompts/ 看到); · 从海量生产环境的 tool-call 轨迹(调用频率分布、单工具重复率、新工具对调用链的影响)反向蒸馏出的专用工具集,包括全文件读取、代码搜索、其他变更文件查阅等,比通用 agent 工具箱更小更稳。 # 能力面与生态集成 功能上覆盖:workspace/分支区间/单 commit 审查、断点恢复(ocr session)、全文件 scan(无 git 历史也能审计陌生代码库)、本地 Session Viewer 网页查看与回放、SARIF/JSON 输出、OpenTelemetry 可观测性、MCP Server 扩展。 作为 “Skills 生态” 级项目,它的形态相当完整:既提供 npm 全局 CLI,也提供可移植的 Agent Skill(skills/open-code-review/SKILL.md,带标准 frontmatter,可直接被兼容 skill 的 agent 加载),还有面向 Claude Code、Codex、Cursor、Kimi Code、OpenCode 等平台的插件,每种都封装成斜杠命令或可调用 skill。LLM 侧兼容 OpenAI、Anthropic、AWS Bedrock 三类协议,并可直接复用 Claude Code 的 ANTHROPIC_* 环境变量。 其中一个设计很巧妙:Delegation 模式(ocr delegate preview/rule)。此时 OCR 只做自己擅长的确定性部分(文件选择和规则解析)审查本身交给宿主 coding agent 的 LLM 执行,用户无需给 OCR 配任何 API key。这实际上是把“harness 能力”与“模型能力”彻底解耦。 # 工程质量:超出平均水准的部分 · 安全有正式的 Assurance Case(ASSURANCE_CASE.md):完整的威胁模型、四条信任边界、T1–T7 威胁逐条给出缓解措施,并按 Saltzer & Schroeder 设计原则和 OWASP Top 10 做了映射。细节经得起推敲:所有外部进程调用只限 git 且子命令硬编码、--end-of-options 防 flag 注入;Agent 读文件路径经 pathutil.WithinBase() 在符号链接解析前后双重校验;本地 Viewer 有 Host 白名单防 DNS rebinding + 严格 CSP。这类文档在一般开源项目里非常罕见。 · 贡献规范近乎严苛(AGENTS.md):使用 AI 必须在 issue/PR 中披露工具与模型、必须逐行理解 AI 生成的代码、禁止“AI 生成→反复修复→再修复”的循环、禁止把 commit 署名给 AI。源码强制英文(CI 有 english-check,连全角标点都查)、90% 测试覆盖率门槛、-race 与 govulncheck 每次 push 都跑、SPDX 头与 LF 行尾强制。 # Benchmark:数据情况 官方基准 AACR-Bench(已在 Hugging Face 开放)规模不小:50 个流行开源仓库、200 个真实 PR、10 种语言、80+ 资深工程师交叉验证出 1505 条标注问题。结论是同模型对比 Claude Code:Precision 和 F1 显著更高、token 消耗约为 1/9、速度更快。 需要指出两点:其一,Recall 低于通用 agent,README 自己承认这是“以精度换噪声”的刻意权衡,如果你最怕漏问题而非误报,可能不适合;其二,该基准由阿里自建,虽开放了数据集供社区复核,但独立第三方的复现结论目前还少,可以把它当作“有披露的、方向可信的参考”。
显示更多
0
13
169
39
转发到社区
here's a prompt to improve your agent harness based on what we've learned at cursor. enjoy # Improve this agent harness's token efficiency You're working on an LLM agent harness: the system prompt, tool definitions, request assembly, context caching, compaction, and retrieval, and how work is split across agents. Make the agent's runs cheaper without making it worse at its job. - Objective: lower price-weighted token cost per completed task. - Constraint: no measurable drop in task quality. Measure per task, not per request. Every turn resends the prefix (tools, instructions, setup, and the conversation so far), so a change that shrinks each request but adds turns can cost more. Weight tokens by billing type: output, uncached input, and cached input are priced very differently. Work in this order: map the harness and measure the baseline, rank the opportunities, make the changes that are safe to make directly, put the rest behind flags or in proposals, then report. Figures below come from one team's production coding agent and its multi-agent experiments. Use them to gauge magnitude, not as targets. One round of these changes (prompt trimming, tool offloading, cache layout, sparse line numbers, subagent tuning) cut that team's overall token cost about 7% with no loss in quality. The larger percentages apply only to the part of the request each change touched. ## Principles 1. Change what the harness sends, not how hard the model tries. Don't ask the model to conserve tokens. A harness that told its model to "take care to preserve tokens and not be wasteful" found it grew reluctant to take on ambitious tasks and sometimes quit, saying it wasn't supposed to waste tokens. 2. Capable models need definitions, not commands. Lists of "DO NOT", "You must", and "Important", and guards against older models' habits, can usually be replaced with plain descriptions of what each tool does. One team cut about two-thirds of its system prompt this way, and the shorter prompt worked across model families. Instruct only on what the model can't know (the product, the environment, the user's processes) and on quirks you've seen in transcripts. 3. Static context is for what most turns need. Everything else should be discoverable when needed. Less up-front context also means less confusing or contradictory information. 4. Expect removals to win. Guardrails written for weaker models, coordination steps that became bottlenecks, and prompting for behavior the model now does on its own all cost tokens. 5. Real usage decides. Evals are a fast proxy, but they skew toward hard problems and miss the real mix of requests. ## 1. Map the harness and measure the baseline Find: - Where requests are assembled, the system prompt, and tool schemas. If a framework or SDK builds requests, find its hooks for message order, cache control, and tool loading. - How tool results are formatted, and how history is kept, trimmed, or summarized. - How subagents or parallel agents are spawned, if any. - Which models and provider APIs are used. From the provider's docs, get the prompt caching behavior (automatic or explicit breakpoints, TTL, minimum cacheable length) and the prices for output, uncached input, and cached input. - Existing logging, token accounting, and evals. If the harness doesn't record per-request token usage by billing type and cache hits, add that first. Everything later depends on it. Then render a few real requests (from logs, or by running representative tasks) and count tokens per section with the model's tokenizer or the API's usage fields. Produce: - Cost share by source × billing type. Sources: system prompt, tool definitions, skill/rule/integration descriptions, user messages, file reads, search results, command and other tool output, history, summaries, subagents. - Static tokens per request, cache hit rate, and turns per task. - Per tool: the share of runs that call it at least once, and its error rate. Read the rendered requests, not just the templates. Duplication, leaked volatile values, and misordered blocks only show up there. Rank opportunities by share of spend × fraction removable ÷ quality risk. ## 2. System prompt and injected context Label every instruction: - Keep: product or environment knowledge the model can't infer, fixes for quirks seen in this model's transcripts, and rules a mode depends on. - Rewrite: commands and emphasis into plain descriptions. Reminders into constraints: "No TODOs, no partial implementations" works better than "remember to finish implementations." Vague quantities into ranges: "generate 20–100 tasks" gets far more ambitious behavior than "generate many tasks." - Delete: things capable models do by default, guards against behavior you haven't seen from this model, text that repeats tool descriptions, and lines that could contradict a user request. Models trained to rank system instructions above user messages will side with the system prompt. - Move: anything per-user or per-request (date, environment, repo state, lists of skills or subagents, user rules) into a user-role setup message after the cache boundary. Audit other injected context the same way. As models improved, the team behind these figures dropped directory trees, pre-retrieved snippets, compressed copies of attached files, lint errors injected after every edit, forced expansion of short file reads, and caps on tool calls per turn. They kept small, high-value facts: OS, repo status, and open or recently viewed files. Skip checklists for open-ended work. The model optimizes the listed items and deprioritizes everything else. ## 3. Tool definitions Tool schemas ride along on every request. Most tools beyond the core set were each needed in under 20% of conversations, and moving them out of static context cut tool-description tokens 60%. Doing the same for integration tools (such as MCP servers), with names in context and full schemas in one folder per server that the agent can search with grep or jq, cut total tokens 46.9% in sessions that used them. - Keep in static context: high-frequency tools (for a coding agent: read, search, edit, shell), tools the model tries to call even when they're absent, and tools a mode depends on. - Offload the rest: leave a name or one-line pointer and make the full schema discoverable on demand. Group related tools so they load together, and put status (such as "needs re-authentication") where the agent will see it. - Tighten what remains: describe behavior and arguments, and drop usage lectures. - Pick the split by testing a few configurations and tracking tokens, cost, latency, tool-call errors, and task success. ## 4. Cache layout Order each request so the reusable prefix is as long as possible: `tool definitions → system instructions → [breakpoint] → setup message (skills, subagents, rules, environment) → [breakpoint] → conversation` - Keep the prefix byte-identical across turns. Use deterministic tool order and serialization, put timestamps and IDs after the boundary, and don't rewrite earlier messages except when compacting. - Use explicit breakpoints if the provider supports them. Otherwise rely on automatic prefix caching with the stable part first. Respect TTL and minimum-length rules. - Switching models mid-conversation throws away the cache (caches are per model and provider) and hands the new model a history it didn't write. When a different model is needed, run it as a subagent with fresh context. Explicit breakpoints plus moving per-request setup after them cut cold cache misses 20%. ## 5. Tool results and other context added during a run - Large outputs (commands, integrations, logs): write them to a file and return the path, size, and a short tail. The agent can tail, grep, or read ranges for more. Truncating loses data, and inlining bloats every later request. Treat long-running terminal sessions the same way. - High-volume formats: look for overhead repeated on every line or item. Numbering every 10th line of a file read instead of every line cut cache-read tokens 1.6% without hurting citation accuracy. Each number costs 3–5 tokens, and agents read tens of thousands of lines per session. Also check repeated absolute paths, verbose JSON keys, ANSI codes, progress bars, and repeated headers. - Good retrieval saves exploration turns. Adding semantic search alongside grep raised codebase question-answering accuracy 12.5% on average and cut the iterations users needed. - Tool errors waste tokens and leave confusing debris in context. Classify expected errors (invalid arguments, unexpected environment, provider error, timeout, user abort), treat unknown errors as harness bugs, and track rates per tool and per model. One focused effort along these lines cut unexpected tool errors 10×. ## 6. Long runs: compaction, subagents, and model mix - Compaction: keep the summarization prompt short and the summary compact, carry forward plan state and remaining tasks, and save the full history to a file the agent can search for details the summary dropped. A model trained to self-summarize from a one-line prompt wrote ~1k-token summaries with half the compaction error of a multi-thousand-token prompt that produced 5k+ token summaries. Untrained models may need more guidance, so test how short you can go. A more expensive summarization model made a negligible difference. - Scratchpads and running notes: rewrite them instead of appending. For repeated work in one environment, a small agent-maintained notes file with a line budget, loaded at start, is a promising way to shorten later runs. - Subagents: fresh context keeps the parent lean, but isolation adds coordination cost (duplicate or stale work). If the model already delegates on its own, remove prompting that pushes it to. Have subagents return short handoffs: what was done, findings, concerns, and deviations. A subagent should use a different model only when the user or harness says so. - Model mix: in large multi-agent runs, workers used at least 69% of tokens, and over 90% in most runs. A frontier planner with cheap workers matched a frontier model doing everything at about one-eighth the cost. Planner choice still changes worker spend. One planner that cost less on its own saw its workers use several times more tokens, and the run cost more overall. Measure the whole tree. - Routing and reasoning effort: send simple turns to a cheaper model or lower effort, and upgrade only when a stronger model is clearly better. A router built this way matched or beat single frontier models on user satisfaction at 41–68% lower cost. - Reasoning continuity: if the API returns reasoning items (including encrypted ones), pass them back on later turns and alert when they go missing. Dropping them cost one reasoning model 30% on a coding benchmark, and it burned tokens reconstructing its plan. ## 7. Fit the harness to each model Adapt to what each model was trained on instead of forcing one shape on all of them. If you've tuned the harness for a similar model, start from that version. - Edit format: use the one the model was trained on (for example, patch-style or search-and-replace). An unfamiliar format costs extra reasoning tokens and causes more mistakes. - Shell or tools: shell-first models fall back to `cat` or inline scripts. Name tools after their shell equivalents (such as `rg`), and if needed add: "If a tool exists for an action, prefer to use the tool instead of shell commands (e.g. read_file over `cat`)." - Literalness: some model families follow instructions literally and others tolerate imprecision. Some spiral on emphasized wording. Strip caps and emphasis for literal models. - Triggers: some models ignore a tool until told when to use it. A literal trigger works: "After substantive edits, use the to check recently edited files for linter errors. If you've introduced any, fix them if you can easily figure out how." - Progress updates: if a model reports progress through reasoning summaries, keep them to 1–2 sentences that note new findings or a change of tactic, and remove instructions about messaging mid-turn. - Quirks worth a targeted line: hedging or refusing as context fills ("context anxiety"), declaring completion early, stopping to ask permission, and calling tools that don't exist. Tie each added instruction to the transcript behavior it fixes. Re-audit when models change, since guidance one version needed can be dead weight for the next. ## 8. Validate - Offline: run a fixed set of realistic tasks before and after, ideally drawn from real usage and phrased the way users actually write (short and ambiguous). Compare task success, tokens, cost per task, turns, and tool errors. Don't ship a change that lowers success. - Online, if you have users: A/B test each change or small bundle. The primary metric is cost per completed task. Guardrails are task success signals, tool-call errors, latency, turns per task, and cache hit rate. For a coding agent, a good success signal is how much agent-written code survives over time. In general, check whether the user's next message moves on or reports a problem. - Ship only when cost drops and no guardrail regresses beyond noise. Record null results. ## What to change directly and what to propose - Change directly, each in its own revertible commit: token and cache telemetry, deterministic serialization and tool order, moving volatile content out of the cached prefix, explicit cache breakpoints, writing large outputs to files instead of truncating, passing back reasoning items that are being dropped, and fixes for recurring tool errors. - Change behind a flag so it can be tested: system prompt edits, tool offloading, output format changes, compaction changes, and subagent prompting. - Propose only: changes to which models run, routing, reasoning-effort defaults, or how work is split across agents. ## Traps - Asking the model to use fewer tokens or do less. - Truncating tool output. - Dropping reasoning items to save input tokens. - Volatile content in the cached prefix, or tool order that changes between requests. - Offloading a tool the model needs on the first turn or tries to call when it's missing. - Emphasis-heavy prompts (MUST, NEVER, IMPORTANT, all caps), especially with literal models. - Forcing a terser output format than the model was trained on. Fewer output tokens can mean less thinking and worse results. - Optimizing raw token counts instead of cost, per request instead of per task, or evals instead of real usage. - Switching models mid-conversation to save money. - Adding coordination layers that become bottlenecks. ## Report back with 1. The harness map and baseline: cost by source × billing type, with the biggest sources called out. 2. A ranked list of changes: layer, what changes, estimated savings and how you estimated them, quality risk, how to validate, and how to roll back. 3. The changes you made, including a system prompt diff with a keep, rewrite, delete, or move reason for each line. 4. A test plan for the flagged changes. 5. Gaps: anything you couldn't find or measure.
显示更多
0
92
1.4K
65
转发到社区
最近恰逢秋季开学,我也连续推荐了好几门北美顶尖高校新开的 Agent 课程,一个信号越来越明显: Agent Engineering 开始大量被正式写进顶尖高校的课程体系了。 这次是 CMU 新开的 11-768《AI Agents》,看完 syllabus 后感觉内容真的很贴近现在一线在做的东西。 Tool Use、Context、Skills、Memory、Planning,往后还有 Coding Agent、GUI、Deep Research、SFT/RL、Sandboxing 和安全。 作业设计是我更看重的部分。 这个课程要求先自己搭 Agentic Harness,再做 Agent Eval,然后需要用 RL 去训练 Agent,最后做一个研究项目。 这个顺序其实挺能说明 2026 年高校开始怎么理解 Agent 了,会调用模型已经很基础,接下来要学的是怎么把 Agent 搭出来、测明白、训练好,再让它可靠地跑下去。 而且课程资料会公开,视频也刚刚开始上传。 最近想系统补 Agent 的,很推荐跟着这门课学。 课程官网: 课程 Schedule: 课程 Assignments: youtubu 视频:
显示更多
0
46
938
250
转发到社区
最近和一些 AI 焦虑的老板聊天,有几个有趣的结论: 1. FDE 的作用其实是 AI 时代的麦肯锡,关键作用是需要外力来对企业实现 AI 转型,大部分企业根本没那个组织能力去自己实现转型。 2. 年轻人都说原来的公司是老登公司,压根招不上来 AI Native 的人才。 3. 针对问题 2,单独新建立一个新公司招 00 后是一个解法,最起码能让老板自己开始跑起来 AI 实验。 4. 国内现在 2B 的软件生意很难做,一方面 Workbuddy 和字节豆包都在打入,另一方面企业被自媒体宣传得,都想自己用 Coding Agent 自己开发软件。 5. 自嘲自己是老登公司的老板们,也都知道自己下面的员工,人人都想用 AI 做点副业。因为知道自己的船在沉、也很焦虑。
显示更多
斯坦福大学把 CS329A Self-Improving AI Agents 视频都公布到 Youtube 了。 两位主讲老师做过 PaLM、Gemini、Claude。9 讲讲透一个循环: 采样 → 验证 → 筛选 → 训练 → 更强的模型 想知道为什么 coding agent 最新普及,写作却不行? 完整解读 👇
显示更多
0
59
661
153
转发到社区
AI Engineering Skills Map 系列之「使用 Coding Agent」 吴恩达老师的 AI 工程技能图谱第三篇详细展开: 1. 构建与部署 AI 应用 2. 软件工程基础 3. 使用 Coding Agent(本文主题) 4. 塑造构建方向 吴恩达老师认为:使用 Coding Agent 正在成为 AI 工程师的关键能力,而且它的演进速度比其他顶层技能都快,因为 Agent 本身在 harness 和模型两个层面同时快速迭代。因此,这项技能没有终态,只能靠持续的实验、构建和学习来维持。 # 用 Coding Agent 构建软件的通用工作流 通过访谈数十位顶尖 AI 工程师并复盘自己团队的实践,他归纳出一个一致的高层工作流,分三步: 1. 规划(Planning) 包含两部分:一是头脑风暴,可能涉及研究、实验、理解已有代码库;二是写 spec(规格说明),涵盖需求、技术设计、架构,随后生成执行计划。规划完成后还应审视计划本身:质疑关键假设,检查安全性、过度设计等问题。 2. 执行(Execution) 构建、测试、验证,关键在于把握智能体自主性与人工监督之间的平衡。一是让智能体以"校准过的自主程度"去构建;二是通过自动化和/或人工检查来验证输出。 3. 部署与监控(Deployment and monitoring) 部署可能经由 CI/CD 流水线或额外的人工关卡把关;随后用智能体观察日志、发现问题、提出并执行改进。 这个工作流有两个特别注意: 1. 它与前智能体时代的软件开发流程本质相似。真正变化的是注意力的重心:从写代码转移到决定做什么、设计架构、写 spec、验证输出。 2. 各步骤的时长弹性极大,可以省略。greenfield(从零开始)原型的 spec 可能只是一条快速写下的提示词;而有大量用户的 brownfield(存量)项目的 spec 则需要投入大量精力去撰写和验证。整个流程高度迭代,熟练的开发者知道何时该从后面的步骤退回前面:验证失败就引导智能体重建修复;监控发现问题就让智能体更新系统并重新部署。 # 五项关键技能 1. 指挥工作流(Directing the workflow) 知道如何走完上述每一步,并决定每一步投入多少人力、多少智能体算力,以及何时回退迭代。这背后是对速度、成本、技术风险、人力投入四者权衡的深刻理解,具体体现在:前期研究和规划做到什么程度、哪些关键工作保留人类所有权、如何选择架构、规划产物(如 spec)写多细、如何把工作拆解成可验证的步骤。 2. 赋予智能体自主性(Enabling agent autonomy) 这一节内容最密集,可拆成四个决策点: · 自主程度:盯着它交互式往返,还是委托一大块工作?何时设定明确目标让它循环直到成功? · 上下文管理:构建过程会经历不同阶段,要判断何时把关键经验、用户反馈、假设(包括中途变化的假设)记录下来供智能体下游使用。 · 并行化:何时把任务拆解后让多个智能体并行,由人或更高层的智能体来编排;以及如何在多个并发会话之间分配人的注意力。 · 安全运行:设置权限、对高风险动作设关卡,在保持开发速度的同时限制泄露、数据丢失等损害。 3. 审查工作成果(Reviewing the work) 出发点是一个基本事实:智能体的输出是不确定的。我们事先不知道它会想出什么好主意,也不知道它会埋下什么 bug。因此审查和验证是拿到想要结果、并在偏离时纠正的关键环节。 具体手段包括: · 设计与任务匹配的测试和验证,按需结合行为验证和功能验证。 · 测试用户流程,可让智能体提供截图作为成功或失败的证据。 · 对定性/行为性评估,可使用评估集(eval sets),可能配合 LLM-as-a-judge。 · 决定测试的自动化程度。某些工作流会把测试完全自动化,让智能体能自行检查、自知何时完成。但必须评估这些测试是否真正对应你的目标,不对应就要演进它们。 · 使用智能体代码审查,运行 AI 驱动的安全和架构审计。 · AI 审查不够时,审慎地插入人工审查——主要审查代码行为,较少审查代码本身——同时探索进一步自动化的可能。 · 验证部署,并用智能体把监控和事故管理运营起来。 这里有一个值得注意的判断:人工审查的对象主要是"代码行为"而非"代码",这反映了注意力重心的转移。 4. 定制智能体及其环境(Customizing the agent and its environment) 目标是让智能体高效获取所需上下文、访问工具、正确高效地构建。具体包括: · 集成 skills、插件、MCP 服务器,并在不再必要时(如新模型让旧 skill 过时)剪除它们。 · 用 hooks 自动化开发流程中可重复的部分,如触发自动代码审查或 CI/CD。 · 维护常驻上下文(AGENTS.md、CLAUDE.md),记录代码库信息、关键架构假设、代码风格、数据访问模式。 · 跨会话、跨并行智能体保存状态,随时间积累智能体的经验,比如通过运行后复盘记录哪些做法有效、哪些无效。 · 建立一致的约定和结构,让代码库对智能体可导航;定期清理智能体产生的技术债。 · 团队协作时,考虑如何在不同开发者的智能体之间协调上下文。 5. 编码智能体基础原理(Coding agent foundations) 要做好上述所有决策,需要理解智能体的工作机制:如何做代码库搜索/检索、如何管理上下文窗口、不同操作(增加工具调用、MCP 服务器等)如何影响上下文、智能体与子智能体如何交互、智能体是如何通过在 LLM 外包裹 harness 构建出来的。 这种理解让智能体不再是黑箱,帮助你识别典型失败模式: · 把简单方案过度设计 · 因缺乏显式验证流程而丧失严谨性 · 未达目标就停下 · 可能破坏文件或生产数据的动作 同时也帮助你推断智能体的状态、给出正确的指令或上下文来引导它,并在监控运行时更早发现它偏离轨道、需要介入。 # 对行业叙事的批评 结尾处吴恩达老师有一段针对性很强的观点。他认为社交媒体对如何使用编码智能体的描述往往过度简化。让智能体自主运行数小时、消耗数百万甚至数千万 token 有时确实有用,但目前超长时程任务的实际效用——尤其是相对成本而言——被夸大到超出现实。 他的结论是:最有效的编码智能体使用是一个复杂、高度迭代的过程,能够以高水平判断力适时介入,效果远好于放手长跑。
显示更多
0
25
42
10
转发到社区
CS 329Z: Engineering AI Agents Stanford / Fall 2026 @stanfordnlp 课程定位:从"模型"到"系统"的工程学 覆盖:简单 LLM 流水线 → 复合 AI 系统 → 自主 Agent。三位讲师的背景也高度互补: @Diyi_Yang(斯坦福 NLP 教授,人机交互与社会计算方向) @michaelryan207(DSPy 核心贡献者,自动评估 AutoMetrics 作者) @jyangballin(SWE-agent / SWE-bench / SWE-smith 作者,软件工程 Agent 领域最重要的研究者之一) # 课程主线:三大工程挑战 贯穿全课的三个核心问题——分解(decomposition)、数据(data)、评估(evaluation)。11 周的内容基本围绕这三条线展开,可以分为五个模块: 模块 1:构建基元(Week 2–3) · LLM 作为构建材料:API/SDK(litellm)、结构化输出、约束生成、解码策略、test-time compute、上下文工程、模型选型与成本/延迟权衡 · RAG:embedding、向量库、分块策略、混合检索、cross-encoder 与 ColBERT 后期交互 · 工具调用:函数调用 API、MCP(Model Context Protocol)、工具设计、代码沙箱、错误处理与重试 模块 2:框架与设计模式(Week 3–5) · 框架层:DSPy(signature / module / optimizer)、LangChain/LangGraph、LlamaIndex,重点是"框架抽象了什么 vs. 你手写了什么" · 设计模式:workflow vs. agent 的分类学,五种可组合 workflow 模式,ReAct / plan-and-execute / reflection,"scaffold(脚手架)本身就是设计决策" · 记忆架构:短期/长期记忆、记忆作为工具动作、文件系统作为外化记忆、跨 Agent 记忆(MemGPT、Mem0、Generative Agents) · 多 Agent 系统:编排模式、handoff 与状态传递,以及一个很有态度的对照阅读——既读 AutoGen,也读《Why Do Multi-Agent LLM Systems Fail?》和 Neubig 的《Don't Sleep on Single-agent Systems》 模块 3:优化(Week 5) · 从提示词到微调的全景:GEPA、MIPROv2、OPRO、TextGrad(提示优化);LoRA/QLoRA、蒸馏、RLHF/DPO(权重优化);test-time scaling(推理算力) · 核心问题是决策框架:什么时候优化 prompt、什么时候优化 weights、什么时候堆推理算力 模块 4:数据与评估(Week 6–8)——最有分量的部分 · 数据:trace、demonstration、feedback 三类数据;训练数据 vs. 评估数据;数据飞轮;合成数据;从 Agent 轨迹构建数据集(SWE-smith) · 评估基础:为什么 eval 难;4 元组框架(request / environment / stopping criteria / scorer);好 benchmark 的性质;tinyBenchmarks · 评估基础设施:三类 grader、LLM-as-judge 的 prompt 设计与已知偏差、pairwise vs. pointwise、非确定性指标 pass@k vs. pass^k、harness 设计 模块 5:安全与前沿(Week 8–11) · 安全:工具访问的隐私风险、prompt injection(含间接注入)、红队、沙箱与权限模型、输出护栏、human-in-the-loop · Coding Agent:SWE-agent、Claude Code、OpenHands 的端到端架构对比 · 主动式 Agent:从 reactive 到 proactive,General User Models(GUM)、Next Action Prediction,以及"Agent 何时应主动、何时应等待"的 mixed-initiative 问题 · 开放问题:多模态/web/计算机使用 Agent、科学 Agent、长时运行架构、生产可观测性(tracing、monitoring、成本管理) # 作业设计:一手建、一手评 HW1:从零构建 Agent 系统(10%) 给定论文库,构建能检索并推理回答科学问题的 Agent。 · Part A:只用 litellm 手写 RAG + 工具调用 + ReAct 式 Agent 循环 · Part B:用 DSPy 重建关键组件,并反思框架抽象了什么 HW2:评估一个 Agent(10%) 给定一个预构建 Agent,设计完整评估套件:代码型 grader、至少一个 LLM-as-judge、用 4 元组框架构建 benchmark 任务、错误分析。
显示更多
0
23
148
30
转发到社区
最近收到不少用户发来的 bug 录屏,我用 MOSS-VL + Kimi K3 串了一条自动排查链路,再不用来回拖进度条了。 一段用户操作录屏发过来,真正费时间的往往不是改代码,而是先还原: 1. 用户做了什么? 2. 问题从哪一秒开始? 3. 到底报了什么错? 这次我把一段 44 秒的录屏以视频流送进 MOSS-VL。它直接整理出了一份结构化 Bug Report: - 复现步骤:输入 → 清空 → 点击 - 异常时间:00:20 - Console 中的 TypeError 和相关变量 然后我通过 Herdr,把报告交给另一个 pane 中运行的 Coding Agent。 Agent 不需要再从头看完整段视频,可以直接根据时间点、操作路径和报错变量定位源码。 整条链路大概是:用户录屏 → MOSS-VL 生成结构化报告 → Kimi K3 定位并修复 → E2E 验证 这里真正负责看录屏的是 MOSS-VL,一个面向持续视频流的 11B 模型,可以在画面持续输入的过程中边看边理解。它提供了 NF4 量化版本,单张 RTX 4090 就可以本地部署;需要适配自家产品 UI 和控件时,也可以通过 LlamaFactory 进行微调。不想部署还也可以接 API,目前每天登录还会赠送 100credits API: Github: 这次最直观的感受是:原本需要人工反复查看录屏、整理复现步骤的工作,现在可以直接变成 Coding Agent 能继续处理的上下文。
显示更多