注册并分享邀请链接,可获得视频播放与邀请奖励。

与「Message」相关的搜索结果

Message 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 Message 的内容
⚠️‼️📵我的TG 账号 @realcarriefu 已被盗,黑客仍在冒用我的身份。请立即拉黑并举报该账号,勿相信其消息或点击任何链接。 ⚠️‼️📵My Telegram account @realcarriefu is compromised and still being used to impersonate me. Please block and report it; do not trust its messages or click any links.
显示更多
Eight democracies. One message to Beijing. Australia, France, Germany, Japan, New Zealand, Poland, South Korea and the UK issued a joint statement on September 22 opposing any unilateral change to the Taiwan Strait status quo. It came out of a meeting on the sidelines of the UN General Assembly, where officials called today's security environment more volatile than a year ago. Their language was direct: "We encourage dialogue and oppose any unilateral attempts to change the status quo across the Taiwan Strait." The same day, the US, Japan and South Korea issued a separate statement making the same point, and also backing Taiwan's participation in international organizations. Taiwan's foreign ministry welcomed the statement and said it would keep building self-defense and deepen ties with democratic partners.
显示更多
0
93
1.1K
315
转发到社区
On Sunday, Ken and his wife Teena planned to sit down together and film this video, reacting to some of what their family has experienced during Ken’s time as Mayor. In the days before the shoot, Teena received a threatening message on Instagram. They decided it was better she not appear on camera. It was the latest in a series of threats and incidents involving Ken, Teena and their family. Ken went ahead alone. “People may think this is funny, or it’s not a big deal, but when you’re dropping into someone’s DMs, it’s pretty personal,” Ken said. Ken has thick skin. He chose public life and knows criticism comes with it. His family didn’t.
显示更多
SpaceXAI / X : Check out @Grok in an XChat Group chat. you can tag it for direct queries, ask it to make pictures or videos, ask for reminders. ( it has temporal awareness) ask for a summary of the chat ( since it joined), or do any of the other typical things you would ask grok. The neatest thing is, it also participates organically and chimes in whenever it feels the need to, reacts to some messages when it finds things funny or to give a thumbs up. ( you dont have to explicitly tag or mention it in other words). Its like having just another member of the chat that is also super useful . This is so cool! Note: Still in testing and not yet released.
显示更多
阿里把团队内部用了两年的官方 AI Code Review Skills 开源了,采用 “确定性工程 pipeline + AI Agent” 的混合架构,专门解决通用 Agent 做代码审查时 “漏审、定位漂移、质量不稳” 的老问题。 40.5K ✨ 开源项目 OpenCodeReview: # 核心设计:确定性工程 pipeline × Agent 各司其职 确定性工程负责硬约束: · 精确文件选择:用代码决定哪些文件必须审、哪些要过滤,不依赖模型自觉; · 智能文件捆绑:把相关文件合成一个审查单元(例如 message_en.properties 和 message_zh.properties 捆绑),每个单元以上下文隔离的 sub-agent 运行,分治策略让超大变更集也稳,且天然支持并发(默认 8 个文件 worker); · 细粒度规则匹配:内置约 54 个按语言/文件类型的规则文档(Java、Go、TS/JS、Python、Rust、SQL/XML mapper、properties 等),用模板引擎而非自然语言把规则匹配到文件特征上,从源头消除信息噪声; · 外部定位与反思模块:评论的“落点”和“内容”分别由独立的 re-location 和 reflection 模块系统性校正,这正对“位置漂移”痛点。 Agent 负责动态决策: · 深度优化的场景 prompt(内部分为 plan → grouping → main → memory_compression → re_location → review_filter 多个任务模板,可在 internal/config/template/prompts/ 看到); · 从海量生产环境的 tool-call 轨迹(调用频率分布、单工具重复率、新工具对调用链的影响)反向蒸馏出的专用工具集,包括全文件读取、代码搜索、其他变更文件查阅等,比通用 agent 工具箱更小更稳。 # 能力面与生态集成 功能上覆盖:workspace/分支区间/单 commit 审查、断点恢复(ocr session)、全文件 scan(无 git 历史也能审计陌生代码库)、本地 Session Viewer 网页查看与回放、SARIF/JSON 输出、OpenTelemetry 可观测性、MCP Server 扩展。 作为 “Skills 生态” 级项目,它的形态相当完整:既提供 npm 全局 CLI,也提供可移植的 Agent Skill(skills/open-code-review/SKILL.md,带标准 frontmatter,可直接被兼容 skill 的 agent 加载),还有面向 Claude Code、Codex、Cursor、Kimi Code、OpenCode 等平台的插件,每种都封装成斜杠命令或可调用 skill。LLM 侧兼容 OpenAI、Anthropic、AWS Bedrock 三类协议,并可直接复用 Claude Code 的 ANTHROPIC_* 环境变量。 其中一个设计很巧妙:Delegation 模式(ocr delegate preview/rule)。此时 OCR 只做自己擅长的确定性部分(文件选择和规则解析)审查本身交给宿主 coding agent 的 LLM 执行,用户无需给 OCR 配任何 API key。这实际上是把“harness 能力”与“模型能力”彻底解耦。 # 工程质量:超出平均水准的部分 · 安全有正式的 Assurance Case(ASSURANCE_CASE.md):完整的威胁模型、四条信任边界、T1–T7 威胁逐条给出缓解措施,并按 Saltzer & Schroeder 设计原则和 OWASP Top 10 做了映射。细节经得起推敲:所有外部进程调用只限 git 且子命令硬编码、--end-of-options 防 flag 注入;Agent 读文件路径经 pathutil.WithinBase() 在符号链接解析前后双重校验;本地 Viewer 有 Host 白名单防 DNS rebinding + 严格 CSP。这类文档在一般开源项目里非常罕见。 · 贡献规范近乎严苛(AGENTS.md):使用 AI 必须在 issue/PR 中披露工具与模型、必须逐行理解 AI 生成的代码、禁止“AI 生成→反复修复→再修复”的循环、禁止把 commit 署名给 AI。源码强制英文(CI 有 english-check,连全角标点都查)、90% 测试覆盖率门槛、-race 与 govulncheck 每次 push 都跑、SPDX 头与 LF 行尾强制。 # Benchmark:数据情况 官方基准 AACR-Bench(已在 Hugging Face 开放)规模不小:50 个流行开源仓库、200 个真实 PR、10 种语言、80+ 资深工程师交叉验证出 1505 条标注问题。结论是同模型对比 Claude Code:Precision 和 F1 显著更高、token 消耗约为 1/9、速度更快。 需要指出两点:其一,Recall 低于通用 agent,README 自己承认这是“以精度换噪声”的刻意权衡,如果你最怕漏问题而非误报,可能不适合;其二,该基准由阿里自建,虽开放了数据集供社区复核,但独立第三方的复现结论目前还少,可以把它当作“有披露的、方向可信的参考”。
显示更多
0
13
169
39
转发到社区
here's a prompt to improve your agent harness based on what we've learned at cursor. enjoy # Improve this agent harness's token efficiency You're working on an LLM agent harness: the system prompt, tool definitions, request assembly, context caching, compaction, and retrieval, and how work is split across agents. Make the agent's runs cheaper without making it worse at its job. - Objective: lower price-weighted token cost per completed task. - Constraint: no measurable drop in task quality. Measure per task, not per request. Every turn resends the prefix (tools, instructions, setup, and the conversation so far), so a change that shrinks each request but adds turns can cost more. Weight tokens by billing type: output, uncached input, and cached input are priced very differently. Work in this order: map the harness and measure the baseline, rank the opportunities, make the changes that are safe to make directly, put the rest behind flags or in proposals, then report. Figures below come from one team's production coding agent and its multi-agent experiments. Use them to gauge magnitude, not as targets. One round of these changes (prompt trimming, tool offloading, cache layout, sparse line numbers, subagent tuning) cut that team's overall token cost about 7% with no loss in quality. The larger percentages apply only to the part of the request each change touched. ## Principles 1. Change what the harness sends, not how hard the model tries. Don't ask the model to conserve tokens. A harness that told its model to "take care to preserve tokens and not be wasteful" found it grew reluctant to take on ambitious tasks and sometimes quit, saying it wasn't supposed to waste tokens. 2. Capable models need definitions, not commands. Lists of "DO NOT", "You must", and "Important", and guards against older models' habits, can usually be replaced with plain descriptions of what each tool does. One team cut about two-thirds of its system prompt this way, and the shorter prompt worked across model families. Instruct only on what the model can't know (the product, the environment, the user's processes) and on quirks you've seen in transcripts. 3. Static context is for what most turns need. Everything else should be discoverable when needed. Less up-front context also means less confusing or contradictory information. 4. Expect removals to win. Guardrails written for weaker models, coordination steps that became bottlenecks, and prompting for behavior the model now does on its own all cost tokens. 5. Real usage decides. Evals are a fast proxy, but they skew toward hard problems and miss the real mix of requests. ## 1. Map the harness and measure the baseline Find: - Where requests are assembled, the system prompt, and tool schemas. If a framework or SDK builds requests, find its hooks for message order, cache control, and tool loading. - How tool results are formatted, and how history is kept, trimmed, or summarized. - How subagents or parallel agents are spawned, if any. - Which models and provider APIs are used. From the provider's docs, get the prompt caching behavior (automatic or explicit breakpoints, TTL, minimum cacheable length) and the prices for output, uncached input, and cached input. - Existing logging, token accounting, and evals. If the harness doesn't record per-request token usage by billing type and cache hits, add that first. Everything later depends on it. Then render a few real requests (from logs, or by running representative tasks) and count tokens per section with the model's tokenizer or the API's usage fields. Produce: - Cost share by source × billing type. Sources: system prompt, tool definitions, skill/rule/integration descriptions, user messages, file reads, search results, command and other tool output, history, summaries, subagents. - Static tokens per request, cache hit rate, and turns per task. - Per tool: the share of runs that call it at least once, and its error rate. Read the rendered requests, not just the templates. Duplication, leaked volatile values, and misordered blocks only show up there. Rank opportunities by share of spend × fraction removable ÷ quality risk. ## 2. System prompt and injected context Label every instruction: - Keep: product or environment knowledge the model can't infer, fixes for quirks seen in this model's transcripts, and rules a mode depends on. - Rewrite: commands and emphasis into plain descriptions. Reminders into constraints: "No TODOs, no partial implementations" works better than "remember to finish implementations." Vague quantities into ranges: "generate 20–100 tasks" gets far more ambitious behavior than "generate many tasks." - Delete: things capable models do by default, guards against behavior you haven't seen from this model, text that repeats tool descriptions, and lines that could contradict a user request. Models trained to rank system instructions above user messages will side with the system prompt. - Move: anything per-user or per-request (date, environment, repo state, lists of skills or subagents, user rules) into a user-role setup message after the cache boundary. Audit other injected context the same way. As models improved, the team behind these figures dropped directory trees, pre-retrieved snippets, compressed copies of attached files, lint errors injected after every edit, forced expansion of short file reads, and caps on tool calls per turn. They kept small, high-value facts: OS, repo status, and open or recently viewed files. Skip checklists for open-ended work. The model optimizes the listed items and deprioritizes everything else. ## 3. Tool definitions Tool schemas ride along on every request. Most tools beyond the core set were each needed in under 20% of conversations, and moving them out of static context cut tool-description tokens 60%. Doing the same for integration tools (such as MCP servers), with names in context and full schemas in one folder per server that the agent can search with grep or jq, cut total tokens 46.9% in sessions that used them. - Keep in static context: high-frequency tools (for a coding agent: read, search, edit, shell), tools the model tries to call even when they're absent, and tools a mode depends on. - Offload the rest: leave a name or one-line pointer and make the full schema discoverable on demand. Group related tools so they load together, and put status (such as "needs re-authentication") where the agent will see it. - Tighten what remains: describe behavior and arguments, and drop usage lectures. - Pick the split by testing a few configurations and tracking tokens, cost, latency, tool-call errors, and task success. ## 4. Cache layout Order each request so the reusable prefix is as long as possible: `tool definitions → system instructions → [breakpoint] → setup message (skills, subagents, rules, environment) → [breakpoint] → conversation` - Keep the prefix byte-identical across turns. Use deterministic tool order and serialization, put timestamps and IDs after the boundary, and don't rewrite earlier messages except when compacting. - Use explicit breakpoints if the provider supports them. Otherwise rely on automatic prefix caching with the stable part first. Respect TTL and minimum-length rules. - Switching models mid-conversation throws away the cache (caches are per model and provider) and hands the new model a history it didn't write. When a different model is needed, run it as a subagent with fresh context. Explicit breakpoints plus moving per-request setup after them cut cold cache misses 20%. ## 5. Tool results and other context added during a run - Large outputs (commands, integrations, logs): write them to a file and return the path, size, and a short tail. The agent can tail, grep, or read ranges for more. Truncating loses data, and inlining bloats every later request. Treat long-running terminal sessions the same way. - High-volume formats: look for overhead repeated on every line or item. Numbering every 10th line of a file read instead of every line cut cache-read tokens 1.6% without hurting citation accuracy. Each number costs 3–5 tokens, and agents read tens of thousands of lines per session. Also check repeated absolute paths, verbose JSON keys, ANSI codes, progress bars, and repeated headers. - Good retrieval saves exploration turns. Adding semantic search alongside grep raised codebase question-answering accuracy 12.5% on average and cut the iterations users needed. - Tool errors waste tokens and leave confusing debris in context. Classify expected errors (invalid arguments, unexpected environment, provider error, timeout, user abort), treat unknown errors as harness bugs, and track rates per tool and per model. One focused effort along these lines cut unexpected tool errors 10×. ## 6. Long runs: compaction, subagents, and model mix - Compaction: keep the summarization prompt short and the summary compact, carry forward plan state and remaining tasks, and save the full history to a file the agent can search for details the summary dropped. A model trained to self-summarize from a one-line prompt wrote ~1k-token summaries with half the compaction error of a multi-thousand-token prompt that produced 5k+ token summaries. Untrained models may need more guidance, so test how short you can go. A more expensive summarization model made a negligible difference. - Scratchpads and running notes: rewrite them instead of appending. For repeated work in one environment, a small agent-maintained notes file with a line budget, loaded at start, is a promising way to shorten later runs. - Subagents: fresh context keeps the parent lean, but isolation adds coordination cost (duplicate or stale work). If the model already delegates on its own, remove prompting that pushes it to. Have subagents return short handoffs: what was done, findings, concerns, and deviations. A subagent should use a different model only when the user or harness says so. - Model mix: in large multi-agent runs, workers used at least 69% of tokens, and over 90% in most runs. A frontier planner with cheap workers matched a frontier model doing everything at about one-eighth the cost. Planner choice still changes worker spend. One planner that cost less on its own saw its workers use several times more tokens, and the run cost more overall. Measure the whole tree. - Routing and reasoning effort: send simple turns to a cheaper model or lower effort, and upgrade only when a stronger model is clearly better. A router built this way matched or beat single frontier models on user satisfaction at 41–68% lower cost. - Reasoning continuity: if the API returns reasoning items (including encrypted ones), pass them back on later turns and alert when they go missing. Dropping them cost one reasoning model 30% on a coding benchmark, and it burned tokens reconstructing its plan. ## 7. Fit the harness to each model Adapt to what each model was trained on instead of forcing one shape on all of them. If you've tuned the harness for a similar model, start from that version. - Edit format: use the one the model was trained on (for example, patch-style or search-and-replace). An unfamiliar format costs extra reasoning tokens and causes more mistakes. - Shell or tools: shell-first models fall back to `cat` or inline scripts. Name tools after their shell equivalents (such as `rg`), and if needed add: "If a tool exists for an action, prefer to use the tool instead of shell commands (e.g. read_file over `cat`)." - Literalness: some model families follow instructions literally and others tolerate imprecision. Some spiral on emphasized wording. Strip caps and emphasis for literal models. - Triggers: some models ignore a tool until told when to use it. A literal trigger works: "After substantive edits, use the to check recently edited files for linter errors. If you've introduced any, fix them if you can easily figure out how." - Progress updates: if a model reports progress through reasoning summaries, keep them to 1–2 sentences that note new findings or a change of tactic, and remove instructions about messaging mid-turn. - Quirks worth a targeted line: hedging or refusing as context fills ("context anxiety"), declaring completion early, stopping to ask permission, and calling tools that don't exist. Tie each added instruction to the transcript behavior it fixes. Re-audit when models change, since guidance one version needed can be dead weight for the next. ## 8. Validate - Offline: run a fixed set of realistic tasks before and after, ideally drawn from real usage and phrased the way users actually write (short and ambiguous). Compare task success, tokens, cost per task, turns, and tool errors. Don't ship a change that lowers success. - Online, if you have users: A/B test each change or small bundle. The primary metric is cost per completed task. Guardrails are task success signals, tool-call errors, latency, turns per task, and cache hit rate. For a coding agent, a good success signal is how much agent-written code survives over time. In general, check whether the user's next message moves on or reports a problem. - Ship only when cost drops and no guardrail regresses beyond noise. Record null results. ## What to change directly and what to propose - Change directly, each in its own revertible commit: token and cache telemetry, deterministic serialization and tool order, moving volatile content out of the cached prefix, explicit cache breakpoints, writing large outputs to files instead of truncating, passing back reasoning items that are being dropped, and fixes for recurring tool errors. - Change behind a flag so it can be tested: system prompt edits, tool offloading, output format changes, compaction changes, and subagent prompting. - Propose only: changes to which models run, routing, reasoning-effort defaults, or how work is split across agents. ## Traps - Asking the model to use fewer tokens or do less. - Truncating tool output. - Dropping reasoning items to save input tokens. - Volatile content in the cached prefix, or tool order that changes between requests. - Offloading a tool the model needs on the first turn or tries to call when it's missing. - Emphasis-heavy prompts (MUST, NEVER, IMPORTANT, all caps), especially with literal models. - Forcing a terser output format than the model was trained on. Fewer output tokens can mean less thinking and worse results. - Optimizing raw token counts instead of cost, per request instead of per task, or evals instead of real usage. - Switching models mid-conversation to save money. - Adding coordination layers that become bottlenecks. ## Report back with 1. The harness map and baseline: cost by source × billing type, with the biggest sources called out. 2. A ranked list of changes: layer, what changes, estimated savings and how you estimated them, quality risk, how to validate, and how to roll back. 3. The changes you made, including a system prompt diff with a keep, rewrite, delete, or move reason for each line. 4. A test plan for the flagged changes. 5. Gaps: anything you couldn't find or measure.
显示更多
0
92
1.4K
65
转发到社区
X Numbers are here. Share yours with anyone you want to contact you, even if you don't follow them. It's an optional way for people to message or call you, without you needing to accept requests or follow them back.
显示更多
0
405
9.2K
1.1K
转发到社区
这期播客是 The Pragmatic Engineer 对 OpenAI Codex 团队负责人 Thibault(Tibo)的访谈,聊了 Codex 的诞生、技术决策、工程文化以及软件开发方式的变迁。 以下是核心要点: 个人经历与加入 OpenAI Thibault 是比利时人,学应用数学出身,先后做过制药供应链优化的创业公司,在 Google 做过加速移动网页的项目(后被砍掉,让他学到了要时刻审视项目真实影响力的教训),之后在 Google Maps 做评论,再转到 DeepMind。在 DeepMind 期间,他参与了一个内部聊天机器人的开发——本质上就是 ChatGPT,但比 ChatGPT 早了一年。内部传播很快,大家都在分享对话,但 DeepMind 不具备把它作为产品发布的机制,最终没能推出。 后来他得知 ChatGPT 只有大约 20 个人在维护,这让他既震惊又觉得很有吸引力——这意味着极高的个人影响力。于是他加入 OpenAI,进去就赶上了推理模型的冲刺,大约一个月后 o1 preview 就发布了。 为什么用 Rust 写 Codex 这是一个反直觉的决定——当时模型对 Rust 的支持并不好,业界其他 AI 编码工具基本都用 TypeScript 或 Python。但团队从第一性原理出发,认为智能体的核心需要健壮、安全、高效,而 Rust 的编译时验证特性天然适合智能体场景。同时用不同语言也强制建立了产品界面和智能体核心之间的清晰边界,避免代码耦合。事实证明 Rust 确实"很快就变得非常适合智能体开发"。 开源和模型无关的策略 Codex CLI、SDK 都是开源的,而且支持非 OpenAI 模型——这在主要 AI 实验室中是独一无二的。理由很实际:如果不开源,别人只需改十行代码就能 fork 出一个支持其他模型的版本,那还不如自己直接支持。开源的好处包括新员工入职前就已经熟悉代码库、社区贡献、以及逼迫自己靠模型和产品体验赢用户而非靠锁定。 代价也很明显:竞争对手会在你还没发布的时候就抄走你正在公开开发的功能,"确实有点刺痛";还有大量低质量 PR 需要处理。 工程文化与代码审查的变革 新员工入职后听到最多的一句话是"你问过 Codex 了吗?"——因为 Codex 在 OpenAI 内部接入了 Slack、文档、所有代码,几乎任何问题都能给出不错的回答。 代码审查正在发生质变。OpenAI 开发了专门的代码审查模型,能在逻辑推理和安全漏洞检测上达到"超人水平"——可以深入三四层依赖去发现文档错误导致的不变量违反。安全审查已经是强制自动化的,发现安全问题会直接阻止合并。一个 PR 可以当天提交、当天上线到十亿用户的 ChatGPT 上。 代码审查的角色正在从"正确性检查"转向"意图讨论"——你到底想做什么?这件事值不值得做?这种讨论不一定要围绕代码发生。 维护成本和重构的变化 维护一直是软件工程的"税",但现在大量维护工作(依赖升级、安全补丁)可以完全自动化。更重要的是,重新架构的成本也急剧下降——以前可能要花几个月甚至几年的重构,现在快得多。但好的架构设计反而更重要了:设计好"盒子"和不变量,盒子内部随便改都不影响其他部分。 Harness 与模型的关系 一个有趣的洞察:harness(工具/脚手架)总是"走在模型前面"。Codex 团队的工作本质上是为模型搭建拐杖——提醒它跑测试、保持目标一致等。然后下一代模型训练时会把这些能力内化,拐杖就可以去掉,developer message 也会越来越短。最新一代模型已经不再需要 /goal 命令来保持长期任务的专注,"你直接告诉模型去工作一周,它就真的会做到"。 Codex 与 ChatGPT 的合并 这是一个重大工程挑战:Codex 原本完全本地运行,ChatGPT 是托管云服务,两套完全不同的技术栈要统一。目标是让云端版本具备本地版本同样的能力,同时高效到能纳入 20 美元/月的 Plus 计划。ChatGPT Work 模式本质上是在云端虚拟机里运行完整的 Codex harness,机器配置强大到用户可以在里面训练模型、安装 Blender 做 3D 建模。 有趣的是,Codex 在整个合并过程中还充当了"记者"角色,因为它能访问所有 Slack 讨论和文档,记录了团队的辩论和决策过程。 Thibault 的个人用法与建议 他大量使用手机上的 ChatGPT Work,通过语音口述发送任务,定制了专属的技能和指令来生成他能高效消化的报告和幻灯片。任何问题——公众舆情、生产日志、功能使用率分析、团队动态——30 分钟内都能得到答案。周末他还会用 Codex 做代码探索和原型,"一天之内就能把脑子里的想法变成可以展示给人看的东西"。 对工程师的建议:保持深度好奇心,训练自己快速理解系统的能力("五个为什么"不断追问),以及与你服务的用户群体保持同步——如果你无法清晰表达意图,就很难做出好的工作。
显示更多
Some recent quality-of-life improvements to Grok Bot. You can ask your Bot to draft messages inline for you to approve before sending.
0
170
2.9K
178
转发到社区
A note on recursive STARK mempools (EIP-8288) This is an EIP that I am hoping we can get included in I-star (the fork after Hegota) that you can think of as the next step after Frames, that would unlock extreme amounts of power. Particularly: * Ultra-cheap quantum-safe signatures (SPHINCS-). Much of the cost savings comes from the fact that the signature data (~3 kB) does not have to go onchain * Ultra-cheap quantum-safe privacy protocols. Status quo minimum cost for private txs is ~300k if you engineer very well (no one does), status quo quantum-safe is ~10M gas, this could reduce it to low tens of thousands. * Universal support for your favorite new signature or proof scheme without needing EVM changes. Whatever you use (Falcon, ML-DSA, some other lattice-based thing, something code-based or isogeny-based or even more esoteric), you can just wrap it client-side in a STARK, onchain gas cost low tens of thousands just like privacy protocols. Hopefully, Ethereum will never need "please support my favorite cryptographic algo" politics again. * Private account abstraction: keep your account logic private, and in a private location onchain. Then you can make one transaction to change the ownership of all your onchain state - accounts, defi positions, privacy protocol notes, everything - without revealing which objects' ownership you're changing. Here's how it works. Your transaction can include a type of frame that we call a "dependency frame". The frame is a list of statements, asserting claims like "message hash M was signed by SPHINCS- public key P" and "data hash D was proven to satisfy a statement defined by verification key V". When you send your transaction, you send it in an envelope, which includes a signature or a STARK for each statement in a dependency frame. Once the transaction reaches the mempool, nodes aggregate them. Each node runs a loop: wait one tick (eg. 500ms), aggregate all new envelopes (either single-tx or multi-tx) that you've seen, remove any transactions that are expired, generate a STARK recursively proving all dependencies, and send a new multi-tx envelope containing that STARK. Hence, the bandwidth load is bounded: each node's outbound is one STARK (~100-300 kB) per tick, plus each transaction getting broadcasted through the network once (as happens already). The block builder acts as "yet another mempool node", receiving envelopes from the mempool (plus any side channels), generates its own STARK covering the subset of transactions it intends to include in the block, and adds that STARK to the block. Total onchain overhead: one STARK (100-300 kB), plus 96 bytes for each statement being proven. This is what I've called before ( ) "The Proof Singularity". Today, we have all the ingredients to actually implement it. As a developer, this requires a somewhat different workflow than you are used to, but it is conceptually simple. Any signatures or STARKs, you put into a separate frame. Then the main logic that today is verifying a signature or STARK, you replace with checking for the existence of a frame that includes the correct statement as a dependency. Examples of useful statements: * [tx sighash] verifies against [the pubkey at sload(0)] * there exists a secret and a merkle branch such that hashing secret+0 and applying the merkle branch outputs (public) root R, and hashing secret+1 outputs (public) nullifier N * there exists a secret address A, salt S and signature Z such that sload(0) = hash(A, S) and a merkle proof of address A inside a recent ethereum state contains some pubkey D where [tx sighash] was signed by D [this is private account abstraction; all variables except [tx sighash] and sload(0) are private; you can also make D a STARK verification key] * there exists an ML-DSA signature signing [tx sighash], that verifies against an ML-DSA pubkey whose hash is sload(0) At the core, this is moving any compute and data other than bookkeeping "business logic" outside the core path of Ethereum execution, sharding and parallelizing it via the mempool. Notice also that this requires agreeing on a _language_ (aka. an ISA) for the recursive STARKs to define statements in. The current leading candidate is RISC-V. So this would also de-facto be Ethereum adding RISC-V (or something else we decide on) as a canonical ISA - a big decision that should be done carefully, but that I think will be necessary to drive Ethereum forward.
显示更多