注册并分享邀请链接,可获得视频播放与邀请奖励。

与「reflection」相关的搜索结果

reflection 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 reflection 的内容
阿里把团队内部用了两年的官方 AI Code Review Skills 开源了,采用 “确定性工程 pipeline + AI Agent” 的混合架构,专门解决通用 Agent 做代码审查时 “漏审、定位漂移、质量不稳” 的老问题。 40.5K ✨ 开源项目 OpenCodeReview: # 核心设计:确定性工程 pipeline × Agent 各司其职 确定性工程负责硬约束: · 精确文件选择:用代码决定哪些文件必须审、哪些要过滤,不依赖模型自觉; · 智能文件捆绑:把相关文件合成一个审查单元(例如 message_en.properties 和 message_zh.properties 捆绑),每个单元以上下文隔离的 sub-agent 运行,分治策略让超大变更集也稳,且天然支持并发(默认 8 个文件 worker); · 细粒度规则匹配:内置约 54 个按语言/文件类型的规则文档(Java、Go、TS/JS、Python、Rust、SQL/XML mapper、properties 等),用模板引擎而非自然语言把规则匹配到文件特征上,从源头消除信息噪声; · 外部定位与反思模块:评论的“落点”和“内容”分别由独立的 re-location 和 reflection 模块系统性校正,这正对“位置漂移”痛点。 Agent 负责动态决策: · 深度优化的场景 prompt(内部分为 plan → grouping → main → memory_compression → re_location → review_filter 多个任务模板,可在 internal/config/template/prompts/ 看到); · 从海量生产环境的 tool-call 轨迹(调用频率分布、单工具重复率、新工具对调用链的影响)反向蒸馏出的专用工具集,包括全文件读取、代码搜索、其他变更文件查阅等,比通用 agent 工具箱更小更稳。 # 能力面与生态集成 功能上覆盖:workspace/分支区间/单 commit 审查、断点恢复(ocr session)、全文件 scan(无 git 历史也能审计陌生代码库)、本地 Session Viewer 网页查看与回放、SARIF/JSON 输出、OpenTelemetry 可观测性、MCP Server 扩展。 作为 “Skills 生态” 级项目,它的形态相当完整:既提供 npm 全局 CLI,也提供可移植的 Agent Skill(skills/open-code-review/SKILL.md,带标准 frontmatter,可直接被兼容 skill 的 agent 加载),还有面向 Claude Code、Codex、Cursor、Kimi Code、OpenCode 等平台的插件,每种都封装成斜杠命令或可调用 skill。LLM 侧兼容 OpenAI、Anthropic、AWS Bedrock 三类协议,并可直接复用 Claude Code 的 ANTHROPIC_* 环境变量。 其中一个设计很巧妙:Delegation 模式(ocr delegate preview/rule)。此时 OCR 只做自己擅长的确定性部分(文件选择和规则解析)审查本身交给宿主 coding agent 的 LLM 执行,用户无需给 OCR 配任何 API key。这实际上是把“harness 能力”与“模型能力”彻底解耦。 # 工程质量:超出平均水准的部分 · 安全有正式的 Assurance Case(ASSURANCE_CASE.md):完整的威胁模型、四条信任边界、T1–T7 威胁逐条给出缓解措施,并按 Saltzer & Schroeder 设计原则和 OWASP Top 10 做了映射。细节经得起推敲:所有外部进程调用只限 git 且子命令硬编码、--end-of-options 防 flag 注入;Agent 读文件路径经 pathutil.WithinBase() 在符号链接解析前后双重校验;本地 Viewer 有 Host 白名单防 DNS rebinding + 严格 CSP。这类文档在一般开源项目里非常罕见。 · 贡献规范近乎严苛(AGENTS.md):使用 AI 必须在 issue/PR 中披露工具与模型、必须逐行理解 AI 生成的代码、禁止“AI 生成→反复修复→再修复”的循环、禁止把 commit 署名给 AI。源码强制英文(CI 有 english-check,连全角标点都查)、90% 测试覆盖率门槛、-race 与 govulncheck 每次 push 都跑、SPDX 头与 LF 行尾强制。 # Benchmark:数据情况 官方基准 AACR-Bench(已在 Hugging Face 开放)规模不小:50 个流行开源仓库、200 个真实 PR、10 种语言、80+ 资深工程师交叉验证出 1505 条标注问题。结论是同模型对比 Claude Code:Precision 和 F1 显著更高、token 消耗约为 1/9、速度更快。 需要指出两点:其一,Recall 低于通用 agent,README 自己承认这是“以精度换噪声”的刻意权衡,如果你最怕漏问题而非误报,可能不适合;其二,该基准由阿里自建,虽开放了数据集供社区复核,但独立第三方的复现结论目前还少,可以把它当作“有披露的、方向可信的参考”。
显示更多
0
13
169
39
转发到社区
the reflection when duo is semi-folded! geez, apple, you didn't have to go that hard.
0
61
10.2K
480
转发到社区
CS329A Self-Improving AI Agents 斯坦福大学一门关于「能够通过与自身和环境的交互持续自我改进的 AI Agents」的课程。 讲师阵容 Azalia Mirhoseini:AlphaCode 与 LLM 大规模训练调度优化("Circuit Training")的核心人物,现为 Reflection AI 联合创始人 Aakanksha Chowdhery:PaLM 2 与 Gemma 的负责人,同样在 Reflection AI 嘉宾名单 Misha Laskin(Reflection AI CEO) 多位 Google DeepMind 研究员(Denny Zhou、Thang Luong、Melvin Johnson)、Physical Intelligence(Danny Driess,机器人) 课程主线:一条完整的自我改进技术栈 1. 推理时自我改进(Test-time) 2. 从反馈中学习(Feedback → Learning) 3. 开放式进化与搜索(Open-endedness) 4. 工程化与瓶颈(Agentic Engineering & Bottlenecks) 课程详细信息:
显示更多
0
61
497
112
转发到社区
CS 329Z: Engineering AI Agents Stanford / Fall 2026 @stanfordnlp 课程定位:从"模型"到"系统"的工程学 覆盖:简单 LLM 流水线 → 复合 AI 系统 → 自主 Agent。三位讲师的背景也高度互补: @Diyi_Yang(斯坦福 NLP 教授,人机交互与社会计算方向) @michaelryan207(DSPy 核心贡献者,自动评估 AutoMetrics 作者) @jyangballin(SWE-agent / SWE-bench / SWE-smith 作者,软件工程 Agent 领域最重要的研究者之一) # 课程主线:三大工程挑战 贯穿全课的三个核心问题——分解(decomposition)、数据(data)、评估(evaluation)。11 周的内容基本围绕这三条线展开,可以分为五个模块: 模块 1:构建基元(Week 2–3) · LLM 作为构建材料:API/SDK(litellm)、结构化输出、约束生成、解码策略、test-time compute、上下文工程、模型选型与成本/延迟权衡 · RAG:embedding、向量库、分块策略、混合检索、cross-encoder 与 ColBERT 后期交互 · 工具调用:函数调用 API、MCP(Model Context Protocol)、工具设计、代码沙箱、错误处理与重试 模块 2:框架与设计模式(Week 3–5) · 框架层:DSPy(signature / module / optimizer)、LangChain/LangGraph、LlamaIndex,重点是"框架抽象了什么 vs. 你手写了什么" · 设计模式:workflow vs. agent 的分类学,五种可组合 workflow 模式,ReAct / plan-and-execute / reflection,"scaffold(脚手架)本身就是设计决策" · 记忆架构:短期/长期记忆、记忆作为工具动作、文件系统作为外化记忆、跨 Agent 记忆(MemGPT、Mem0、Generative Agents) · 多 Agent 系统:编排模式、handoff 与状态传递,以及一个很有态度的对照阅读——既读 AutoGen,也读《Why Do Multi-Agent LLM Systems Fail?》和 Neubig 的《Don't Sleep on Single-agent Systems》 模块 3:优化(Week 5) · 从提示词到微调的全景:GEPA、MIPROv2、OPRO、TextGrad(提示优化);LoRA/QLoRA、蒸馏、RLHF/DPO(权重优化);test-time scaling(推理算力) · 核心问题是决策框架:什么时候优化 prompt、什么时候优化 weights、什么时候堆推理算力 模块 4:数据与评估(Week 6–8)——最有分量的部分 · 数据:trace、demonstration、feedback 三类数据;训练数据 vs. 评估数据;数据飞轮;合成数据;从 Agent 轨迹构建数据集(SWE-smith) · 评估基础:为什么 eval 难;4 元组框架(request / environment / stopping criteria / scorer);好 benchmark 的性质;tinyBenchmarks · 评估基础设施:三类 grader、LLM-as-judge 的 prompt 设计与已知偏差、pairwise vs. pointwise、非确定性指标 pass@k vs. pass^k、harness 设计 模块 5:安全与前沿(Week 8–11) · 安全:工具访问的隐私风险、prompt injection(含间接注入)、红队、沙箱与权限模型、输出护栏、human-in-the-loop · Coding Agent:SWE-agent、Claude Code、OpenHands 的端到端架构对比 · 主动式 Agent:从 reactive 到 proactive,General User Models(GUM)、Next Action Prediction,以及"Agent 何时应主动、何时应等待"的 mixed-initiative 问题 · 开放问题:多模态/web/计算机使用 Agent、科学 Agent、长时运行架构、生产可观测性(tracing、monitoring、成本管理) # 作业设计:一手建、一手评 HW1:从零构建 Agent 系统(10%) 给定论文库,构建能检索并推理回答科学问题的 Agent。 · Part A:只用 litellm 手写 RAG + 工具调用 + ReAct 式 Agent 循环 · Part B:用 DSPy 重建关键组件,并反思框架抽象了什么 HW2:评估一个 Agent(10%) 给定一个预构建 Agent,设计完整评估套件:代码型 grader、至少一个 LLM-as-judge、用 4 元组框架构建 benchmark 任务、错误分析。
显示更多
0
23
148
30
转发到社区
If you have a Grok @bot token usage issue, ask the @bot to create a new channel for dedicated tasks and close it when finished. This would allow Grok @bot to focus on the context you are currently interested in. If you want to reuse the same task, have it save the reflection and upload to the repo so it can be repeatable.
显示更多
0
44
429
23
转发到社区
100 AI Directing Techniques - Epic Sci-Fi Prompts 01 Monumental Scale Colossal celestial body · Tiny human figure · Extreme wide shot · Scale contrast · Architectural framing 02 Colossal Ruins Monumental ruins · Tiny human figure · Barren negative space · Volumetric dust · Atmospheric perspective 03 Extreme Close-up Profile silhouette · Rim light · Visor reflections · Lens bloom · Optical flare 04 Practical Destruction Practical explosion · Layered smoke · Strong backlight · Drifting embers · Foreground-to-background depth 05 Impossible Geometry Anti-gravity space · Folded architecture · Tilted horizon · Extreme depth · Scale contrast 06 Volumetric Light Central symmetry · Single overhead light · Volumetric beam · Monumental scale · Deep negative space 07 Futuristic Isolation Architectural framing · Tiny human figure · High-altitude view · Cool-warm contrast · Atmospheric perspective 08 Cosmic Metropolis Curved leading lines · Layered haze · Tiny human scale · Golden particles · Wet reflections Use official invite code VIDUAIGUST4 to get 200 free credits, and become a director yourself.
显示更多
0
12
102
6
转发到社区
Haters will say it's just a reflection
Cut your AI ad production costs without cutting quality. 🔥 Use Seedance 2.5 on ToAPIs and save 30–60% compared to other API providers. Seedance 2.5 isn’t just for creative videos — it’s built for serious production workflows too. From industrial manufacturing and robotics to autonomous driving, generate high-quality synthetic video for simulation, training, testing, and product demonstrations. In this example, a simple clay render becomes a high-end photorealistic car assembly sequence — preserving the camera work, structure, motion, materials, lighting, and reflections. More production. Lower API cost. See pricing ↓
显示更多
0
12
17
2
转发到社区
Alongside the novels, I have been dipping into A Joseph Campbell Companion, (thank you @patrick_oshag !) which I have been opening each night this summer before bed to read and reflect on a few pages or a chapter. I am a huge Joseph Campbell fan, and his reflections on myth, meaning, and the art of living are a quiet, grounding way to end the day. A friend also gave me David R. Hawkins’ Power vs. Force. I read some of his work years ago, and enjoyed returning to his exploration of the difference between power cultivated within and force imposed from without.
显示更多
The following little gems were not spewed out by The Postmodernism Generator, nor are they part of an Alan Sokal hoax. They are a representative selection from Jason Arday’s thesis, which earned him a PhD from Liverpool John Moores University. The University of Cambridge was mightily impressed, and they elected him to a professorship at the young age of 37. His thesis is now under suspicion of plagiarism, but plagiarism, if proved, would seem to be among that document’s lesser problems. “Throughout the process, at critical junctures the researcher was conscious that with regards to providing knowledge to the participants regarding reflection and mentoring, a contradictory position was being adopted, with regards to eliminating anything that may resemble a hierarchical structure.” “As the researcher, there was an awareness that the legitimacy of such suggestions regarding reliability are open to critique, however, this research study adopted these principles in attempting to maintain reliability through a qualitative research design.” “As the researcher, textual units concerning the narratives and biographies were identified based on the criteria that the identified themes were fit for purpose in supporting the themes associated with the study.” “As with other forms of discourse inquiry, narrative inquiry is underpinned by a social constructivist paradigm, in which behaviours and their meanings are socially situated and socially interpreted.” “This process allowed the researcher to immerse themself in the data collated.” “References were made with regards to the benefits that were gained from the peer-mentoring experience, with regards to developing pedagogically and working as part of a collegial environment to support each other during teacher training.” “An important consideration which also formed part of the theoretical paradigm for the research recognised that the mentoring process is socially constructed, and therefore the relationship between the participants and the researcher needed to be tailored towards supporting the professional development of the student teachers, and making them autonomous social actors in this process, hence the researcher adopting the secondary position of facilitator within the research.”
显示更多
0
334
3.4K
523
转发到社区