注册并分享邀请链接,可获得视频播放与邀请奖励。

与「mix」相关的搜索结果

mix 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 mix 的内容
BREAKING: Chinese illicit actors laundering funds from the $387M Bitget exploit on behalf of the alleged DPRK attackers are openly asking for support with orders in public Discord servers and Telegram channels of services they use. Notably, Alias 4 (below) was also seen laundering funds from the Kelp DAO $292M exploit earlier this year. I've observed the same pattern after multiple TraderTraitor attributed exploits, and I've closely tracked these groups. I plan to share more of my data on them in coming weeks. Currently, funds are being chain-hopped via bridges and being deposited into mixing services such as Wasabi. Alias 1 - Cc Discord: cc02006 Discord ID: 1351486674386948148 Txn: F08657EFAEAE7B58217CD17A22BF4779582E5D080C2E92E173BC239A5D828363 Alias 2 - jack Discord: jack_34808 Discord ID: 1553705721768714377 Txn: 68583D313A0CCC99F2702D61D05A242A69C86CAE09ED252B65F34B677C377F69 Alias 3 - Melon Discord: under0346 Discord ID: 1394240539108573215 Txn 1: ABE2AEF8B10057E60F8259CA2CA2CD5D71D38D254960FEEE89CB3C95A964873F Txn 2: 7BE290865901DEB1680A6D692A11392D1FDDACA0D31C48F26DCB249F2444BD6B Alias 4 - lolo / Marin TG: pvpcz TGID: 6223514198 Discord: losern Discord ID: 1024415186527985704 Txn: 8E935C19D00F40639B78BF1FD094FE48C92118F3E586AB52DF93851080CCCE15 Alias 5 - HELP ME Discord: helpme031897 Discord ID: 1554035533817188384 Txn: 7C58CAD760EBBED54F2CA4D2146D056910B7B6688B4F524396DC8807353C2379
显示更多
0
556
5.7K
750
转发到社区
📢 Responses API Major Upgrade: Universal Upstream Protocol Support & Plug-and-Play for New Models! Existing models on that support the Responses API upstream can now be called via the unified /v1/responses endpoint. When adds support for new models from currently integrated vendors in the future, developers can automatically access them in the Codex client with zero extra configuration! 🔸 Official API Group: Added support for MiMo, Qwen, and Hunyuan model series. 🔸 Mix & OL Station Provider 40% OFF Groups: Added support for OpenAI, Gemini, and Grok model series. 📊 Check out the poster below for the full list of supported models 👇 Simply configure your API Key and Base URL to bring the powerful reasoning, coding, and text processing capabilities of leading models seamlessly into your daily workflow, flexibly meeting a wide range of needs. 👉 Try it now: 📚 More details:
显示更多
here's a prompt to improve your agent harness based on what we've learned at cursor. enjoy # Improve this agent harness's token efficiency You're working on an LLM agent harness: the system prompt, tool definitions, request assembly, context caching, compaction, and retrieval, and how work is split across agents. Make the agent's runs cheaper without making it worse at its job. - Objective: lower price-weighted token cost per completed task. - Constraint: no measurable drop in task quality. Measure per task, not per request. Every turn resends the prefix (tools, instructions, setup, and the conversation so far), so a change that shrinks each request but adds turns can cost more. Weight tokens by billing type: output, uncached input, and cached input are priced very differently. Work in this order: map the harness and measure the baseline, rank the opportunities, make the changes that are safe to make directly, put the rest behind flags or in proposals, then report. Figures below come from one team's production coding agent and its multi-agent experiments. Use them to gauge magnitude, not as targets. One round of these changes (prompt trimming, tool offloading, cache layout, sparse line numbers, subagent tuning) cut that team's overall token cost about 7% with no loss in quality. The larger percentages apply only to the part of the request each change touched. ## Principles 1. Change what the harness sends, not how hard the model tries. Don't ask the model to conserve tokens. A harness that told its model to "take care to preserve tokens and not be wasteful" found it grew reluctant to take on ambitious tasks and sometimes quit, saying it wasn't supposed to waste tokens. 2. Capable models need definitions, not commands. Lists of "DO NOT", "You must", and "Important", and guards against older models' habits, can usually be replaced with plain descriptions of what each tool does. One team cut about two-thirds of its system prompt this way, and the shorter prompt worked across model families. Instruct only on what the model can't know (the product, the environment, the user's processes) and on quirks you've seen in transcripts. 3. Static context is for what most turns need. Everything else should be discoverable when needed. Less up-front context also means less confusing or contradictory information. 4. Expect removals to win. Guardrails written for weaker models, coordination steps that became bottlenecks, and prompting for behavior the model now does on its own all cost tokens. 5. Real usage decides. Evals are a fast proxy, but they skew toward hard problems and miss the real mix of requests. ## 1. Map the harness and measure the baseline Find: - Where requests are assembled, the system prompt, and tool schemas. If a framework or SDK builds requests, find its hooks for message order, cache control, and tool loading. - How tool results are formatted, and how history is kept, trimmed, or summarized. - How subagents or parallel agents are spawned, if any. - Which models and provider APIs are used. From the provider's docs, get the prompt caching behavior (automatic or explicit breakpoints, TTL, minimum cacheable length) and the prices for output, uncached input, and cached input. - Existing logging, token accounting, and evals. If the harness doesn't record per-request token usage by billing type and cache hits, add that first. Everything later depends on it. Then render a few real requests (from logs, or by running representative tasks) and count tokens per section with the model's tokenizer or the API's usage fields. Produce: - Cost share by source × billing type. Sources: system prompt, tool definitions, skill/rule/integration descriptions, user messages, file reads, search results, command and other tool output, history, summaries, subagents. - Static tokens per request, cache hit rate, and turns per task. - Per tool: the share of runs that call it at least once, and its error rate. Read the rendered requests, not just the templates. Duplication, leaked volatile values, and misordered blocks only show up there. Rank opportunities by share of spend × fraction removable ÷ quality risk. ## 2. System prompt and injected context Label every instruction: - Keep: product or environment knowledge the model can't infer, fixes for quirks seen in this model's transcripts, and rules a mode depends on. - Rewrite: commands and emphasis into plain descriptions. Reminders into constraints: "No TODOs, no partial implementations" works better than "remember to finish implementations." Vague quantities into ranges: "generate 20–100 tasks" gets far more ambitious behavior than "generate many tasks." - Delete: things capable models do by default, guards against behavior you haven't seen from this model, text that repeats tool descriptions, and lines that could contradict a user request. Models trained to rank system instructions above user messages will side with the system prompt. - Move: anything per-user or per-request (date, environment, repo state, lists of skills or subagents, user rules) into a user-role setup message after the cache boundary. Audit other injected context the same way. As models improved, the team behind these figures dropped directory trees, pre-retrieved snippets, compressed copies of attached files, lint errors injected after every edit, forced expansion of short file reads, and caps on tool calls per turn. They kept small, high-value facts: OS, repo status, and open or recently viewed files. Skip checklists for open-ended work. The model optimizes the listed items and deprioritizes everything else. ## 3. Tool definitions Tool schemas ride along on every request. Most tools beyond the core set were each needed in under 20% of conversations, and moving them out of static context cut tool-description tokens 60%. Doing the same for integration tools (such as MCP servers), with names in context and full schemas in one folder per server that the agent can search with grep or jq, cut total tokens 46.9% in sessions that used them. - Keep in static context: high-frequency tools (for a coding agent: read, search, edit, shell), tools the model tries to call even when they're absent, and tools a mode depends on. - Offload the rest: leave a name or one-line pointer and make the full schema discoverable on demand. Group related tools so they load together, and put status (such as "needs re-authentication") where the agent will see it. - Tighten what remains: describe behavior and arguments, and drop usage lectures. - Pick the split by testing a few configurations and tracking tokens, cost, latency, tool-call errors, and task success. ## 4. Cache layout Order each request so the reusable prefix is as long as possible: `tool definitions → system instructions → [breakpoint] → setup message (skills, subagents, rules, environment) → [breakpoint] → conversation` - Keep the prefix byte-identical across turns. Use deterministic tool order and serialization, put timestamps and IDs after the boundary, and don't rewrite earlier messages except when compacting. - Use explicit breakpoints if the provider supports them. Otherwise rely on automatic prefix caching with the stable part first. Respect TTL and minimum-length rules. - Switching models mid-conversation throws away the cache (caches are per model and provider) and hands the new model a history it didn't write. When a different model is needed, run it as a subagent with fresh context. Explicit breakpoints plus moving per-request setup after them cut cold cache misses 20%. ## 5. Tool results and other context added during a run - Large outputs (commands, integrations, logs): write them to a file and return the path, size, and a short tail. The agent can tail, grep, or read ranges for more. Truncating loses data, and inlining bloats every later request. Treat long-running terminal sessions the same way. - High-volume formats: look for overhead repeated on every line or item. Numbering every 10th line of a file read instead of every line cut cache-read tokens 1.6% without hurting citation accuracy. Each number costs 3–5 tokens, and agents read tens of thousands of lines per session. Also check repeated absolute paths, verbose JSON keys, ANSI codes, progress bars, and repeated headers. - Good retrieval saves exploration turns. Adding semantic search alongside grep raised codebase question-answering accuracy 12.5% on average and cut the iterations users needed. - Tool errors waste tokens and leave confusing debris in context. Classify expected errors (invalid arguments, unexpected environment, provider error, timeout, user abort), treat unknown errors as harness bugs, and track rates per tool and per model. One focused effort along these lines cut unexpected tool errors 10×. ## 6. Long runs: compaction, subagents, and model mix - Compaction: keep the summarization prompt short and the summary compact, carry forward plan state and remaining tasks, and save the full history to a file the agent can search for details the summary dropped. A model trained to self-summarize from a one-line prompt wrote ~1k-token summaries with half the compaction error of a multi-thousand-token prompt that produced 5k+ token summaries. Untrained models may need more guidance, so test how short you can go. A more expensive summarization model made a negligible difference. - Scratchpads and running notes: rewrite them instead of appending. For repeated work in one environment, a small agent-maintained notes file with a line budget, loaded at start, is a promising way to shorten later runs. - Subagents: fresh context keeps the parent lean, but isolation adds coordination cost (duplicate or stale work). If the model already delegates on its own, remove prompting that pushes it to. Have subagents return short handoffs: what was done, findings, concerns, and deviations. A subagent should use a different model only when the user or harness says so. - Model mix: in large multi-agent runs, workers used at least 69% of tokens, and over 90% in most runs. A frontier planner with cheap workers matched a frontier model doing everything at about one-eighth the cost. Planner choice still changes worker spend. One planner that cost less on its own saw its workers use several times more tokens, and the run cost more overall. Measure the whole tree. - Routing and reasoning effort: send simple turns to a cheaper model or lower effort, and upgrade only when a stronger model is clearly better. A router built this way matched or beat single frontier models on user satisfaction at 41–68% lower cost. - Reasoning continuity: if the API returns reasoning items (including encrypted ones), pass them back on later turns and alert when they go missing. Dropping them cost one reasoning model 30% on a coding benchmark, and it burned tokens reconstructing its plan. ## 7. Fit the harness to each model Adapt to what each model was trained on instead of forcing one shape on all of them. If you've tuned the harness for a similar model, start from that version. - Edit format: use the one the model was trained on (for example, patch-style or search-and-replace). An unfamiliar format costs extra reasoning tokens and causes more mistakes. - Shell or tools: shell-first models fall back to `cat` or inline scripts. Name tools after their shell equivalents (such as `rg`), and if needed add: "If a tool exists for an action, prefer to use the tool instead of shell commands (e.g. read_file over `cat`)." - Literalness: some model families follow instructions literally and others tolerate imprecision. Some spiral on emphasized wording. Strip caps and emphasis for literal models. - Triggers: some models ignore a tool until told when to use it. A literal trigger works: "After substantive edits, use the to check recently edited files for linter errors. If you've introduced any, fix them if you can easily figure out how." - Progress updates: if a model reports progress through reasoning summaries, keep them to 1–2 sentences that note new findings or a change of tactic, and remove instructions about messaging mid-turn. - Quirks worth a targeted line: hedging or refusing as context fills ("context anxiety"), declaring completion early, stopping to ask permission, and calling tools that don't exist. Tie each added instruction to the transcript behavior it fixes. Re-audit when models change, since guidance one version needed can be dead weight for the next. ## 8. Validate - Offline: run a fixed set of realistic tasks before and after, ideally drawn from real usage and phrased the way users actually write (short and ambiguous). Compare task success, tokens, cost per task, turns, and tool errors. Don't ship a change that lowers success. - Online, if you have users: A/B test each change or small bundle. The primary metric is cost per completed task. Guardrails are task success signals, tool-call errors, latency, turns per task, and cache hit rate. For a coding agent, a good success signal is how much agent-written code survives over time. In general, check whether the user's next message moves on or reports a problem. - Ship only when cost drops and no guardrail regresses beyond noise. Record null results. ## What to change directly and what to propose - Change directly, each in its own revertible commit: token and cache telemetry, deterministic serialization and tool order, moving volatile content out of the cached prefix, explicit cache breakpoints, writing large outputs to files instead of truncating, passing back reasoning items that are being dropped, and fixes for recurring tool errors. - Change behind a flag so it can be tested: system prompt edits, tool offloading, output format changes, compaction changes, and subagent prompting. - Propose only: changes to which models run, routing, reasoning-effort defaults, or how work is split across agents. ## Traps - Asking the model to use fewer tokens or do less. - Truncating tool output. - Dropping reasoning items to save input tokens. - Volatile content in the cached prefix, or tool order that changes between requests. - Offloading a tool the model needs on the first turn or tries to call when it's missing. - Emphasis-heavy prompts (MUST, NEVER, IMPORTANT, all caps), especially with literal models. - Forcing a terser output format than the model was trained on. Fewer output tokens can mean less thinking and worse results. - Optimizing raw token counts instead of cost, per request instead of per task, or evals instead of real usage. - Switching models mid-conversation to save money. - Adding coordination layers that become bottlenecks. ## Report back with 1. The harness map and baseline: cost by source × billing type, with the biggest sources called out. 2. A ranked list of changes: layer, what changes, estimated savings and how you estimated them, quality risk, how to validate, and how to roll back. 3. The changes you made, including a system prompt diff with a keep, rewrite, delete, or move reason for each line. 4. A test plan for the flagged changes. 5. Gaps: anything you couldn't find or measure.
显示更多
0
92
1.4K
65
转发到社区
Sam Harper has gone from a player whose state career was on the brink to being in the mix for an Australian debut. @DanielCherny examines the Victorian gloveman’s stunning turnaround. READ ▶️
显示更多
Today, we are announcing the second generation of Agentic Document Extraction (ADE). It is faster, more accurate, and more affordable than ever before. ADE Gen2 delivers higher performance while optimizing cost for every document. Pricing has been completely overhauled. Customers running mixed workloads should see 25% to 80% cost reductions. Groundings and citations now reach the word level following a page > block> line > word hierarchy. This allows every extracted value to point to the exact location on the page it came from. The outputs from the v2 Parse and Extract APIs are fully agent-ready and agent-friendly. Your agents get a clear hierarchy, stable IDs, and clean, standardized Markdown designed just for them. ADE Gen2 is a step change in document intelligence not an incremental update. Read the full announcement on the blog (link in comments). Create an account and get started for free at
显示更多
0
2
31
16
转发到社区
CS 329Z: Engineering AI Agents Stanford / Fall 2026 @stanfordnlp 课程定位:从"模型"到"系统"的工程学 覆盖:简单 LLM 流水线 → 复合 AI 系统 → 自主 Agent。三位讲师的背景也高度互补: @Diyi_Yang(斯坦福 NLP 教授,人机交互与社会计算方向) @michaelryan207(DSPy 核心贡献者,自动评估 AutoMetrics 作者) @jyangballin(SWE-agent / SWE-bench / SWE-smith 作者,软件工程 Agent 领域最重要的研究者之一) # 课程主线:三大工程挑战 贯穿全课的三个核心问题——分解(decomposition)、数据(data)、评估(evaluation)。11 周的内容基本围绕这三条线展开,可以分为五个模块: 模块 1:构建基元(Week 2–3) · LLM 作为构建材料:API/SDK(litellm)、结构化输出、约束生成、解码策略、test-time compute、上下文工程、模型选型与成本/延迟权衡 · RAG:embedding、向量库、分块策略、混合检索、cross-encoder 与 ColBERT 后期交互 · 工具调用:函数调用 API、MCP(Model Context Protocol)、工具设计、代码沙箱、错误处理与重试 模块 2:框架与设计模式(Week 3–5) · 框架层:DSPy(signature / module / optimizer)、LangChain/LangGraph、LlamaIndex,重点是"框架抽象了什么 vs. 你手写了什么" · 设计模式:workflow vs. agent 的分类学,五种可组合 workflow 模式,ReAct / plan-and-execute / reflection,"scaffold(脚手架)本身就是设计决策" · 记忆架构:短期/长期记忆、记忆作为工具动作、文件系统作为外化记忆、跨 Agent 记忆(MemGPT、Mem0、Generative Agents) · 多 Agent 系统:编排模式、handoff 与状态传递,以及一个很有态度的对照阅读——既读 AutoGen,也读《Why Do Multi-Agent LLM Systems Fail?》和 Neubig 的《Don't Sleep on Single-agent Systems》 模块 3:优化(Week 5) · 从提示词到微调的全景:GEPA、MIPROv2、OPRO、TextGrad(提示优化);LoRA/QLoRA、蒸馏、RLHF/DPO(权重优化);test-time scaling(推理算力) · 核心问题是决策框架:什么时候优化 prompt、什么时候优化 weights、什么时候堆推理算力 模块 4:数据与评估(Week 6–8)——最有分量的部分 · 数据:trace、demonstration、feedback 三类数据;训练数据 vs. 评估数据;数据飞轮;合成数据;从 Agent 轨迹构建数据集(SWE-smith) · 评估基础:为什么 eval 难;4 元组框架(request / environment / stopping criteria / scorer);好 benchmark 的性质;tinyBenchmarks · 评估基础设施:三类 grader、LLM-as-judge 的 prompt 设计与已知偏差、pairwise vs. pointwise、非确定性指标 pass@k vs. pass^k、harness 设计 模块 5:安全与前沿(Week 8–11) · 安全:工具访问的隐私风险、prompt injection(含间接注入)、红队、沙箱与权限模型、输出护栏、human-in-the-loop · Coding Agent:SWE-agent、Claude Code、OpenHands 的端到端架构对比 · 主动式 Agent:从 reactive 到 proactive,General User Models(GUM)、Next Action Prediction,以及"Agent 何时应主动、何时应等待"的 mixed-initiative 问题 · 开放问题:多模态/web/计算机使用 Agent、科学 Agent、长时运行架构、生产可观测性(tracing、monitoring、成本管理) # 作业设计:一手建、一手评 HW1:从零构建 Agent 系统(10%) 给定论文库,构建能检索并推理回答科学问题的 Agent。 · Part A:只用 litellm 手写 RAG + 工具调用 + ReAct 式 Agent 循环 · Part B:用 DSPy 重建关键组件,并反思框架抽象了什么 HW2:评估一个 Agent(10%) 给定一个预构建 Agent,设计完整评估套件:代码型 grader、至少一个 LLM-as-judge、用 4 元组框架构建 benchmark 任务、错误分析。
显示更多
0
23
148
30
转发到社区
From Digital Twins to Data Engine: Cutting the Real-Data Burden with Sim-Powered Robot Learning Teaching a robot a new task could take hundreds of teleoperated demonstrations. For foundation models, adapting to an entirely new robot can cost orders of magnitude more — dedicated hardware, trained operators, months of engineering. On @boosterobotics' dual-arm robot, we studied this at two levels: ✱ Specialist: Can task-aligned simulation mixed with a small set of real demonstrations reduce the real-data burden? ✅ Yes. With only 10 real demos, the policy made no contact at all in physical rollouts (0/20). Adding 50 simulated trajectories brought contact to 17/20. ✱ Foundation: Can data accumulated across tasks build a reusable starting point (a Booster-specific model prior)? ✅ Yes. After full-parameter continued pretraining, a model adapted with just 30 demonstrations per task beat the original given twice as many: 14/16 vs 10/16 in simulated evaluation. Before any task-specific adaptation, in zero-shot simulation, it was already roughly 3× closer to the target (17.27 cm → 5.78 cm). This work runs on Axis Suite, our Physical AI solution across different robot embodiments. Distributed contributors generate task-aligned sim data on Axis Hub at scale, reducing real-data needs for specialist adaptation while powering cross-embodiment generalist training. Read the full blog:
显示更多
0
110
374
46
转发到社区
SpaceXAI just released a new update for @Grok Build. Update to v1.0.11 Features: • Headless sessions are now browsable in the resume picker without mixing into default history. • Default permission mode for new interactive sessions is now configurable. • Turn duration now appears in session history after resume. • Turn footers (Worked for, cancelled, failed) now appear after /resume. • mkdir and touch no longer prompt in auto mode or safe-command lists. • Subagent messages are now allowed automatically in permission Auto mode. • Headless sessions can now auto-allow permission prompts via a startup hint. Bug Fixes: • Permission prompts for common command chains and subcommands are now more reliable. • Blocked prompts no longer appear in conversation history or scrollback after restart. • Background monitors no longer have a 10-hour default timeout. • Mouse input at the right margin no longer types characters into the prompt on certain terminals. • Pasting text ending in a newline no longer accidentally submits the prompt on some terminals. • Auto mode now shows a permission card when the classifier blocks an action on interactive sessions. • Image previews no longer leave ghost artifacts when using the Kitty protocol on Warp. • Expanded Execute tool output no longer snaps closed during live progress. • /voice now falls back to parec/arecord on older PipeWire installs. Performance: • Background command waits now finish as soon as the process exits. Download Grok Build: Update to the latest Alpha release: grok update --alpha Update to the latest Stable release: grok update
显示更多
0
27
182
41
转发到社区
My third (and last) post in my series on obfuscation (iO): local mixing Local mixing is a very different philosophy from the other two, no prior background in lattices required, and very different tradeoffs (eg. in terms of computational overhead, it's viable today; the entire question is verifying the security of what's ultimately a novel family of cryptography). I highly encourage following their work, they plan to publish much more soon.
显示更多
0
119
327
38
转发到社区
My investigation on the GTA 6 Leaker is done and I have emailed all the evidence to TAKE2 and Rockstar Games legal team As much as I wanted and loved to actually reveal who we are dealing with here, full identity and all, I believe that would be doxxing. Since the way I got to this guy was through a mix of OSINT and some people familiar with the matter via TG, I don’t want anyone to get into trouble, so for now I’ll leave out personal details and how I obtained the information One of the stronger findings links to what appears to be this guy’s personal bank account, the bank is located in Europe. It’s obviously not out of the question that the bank account in question might be stolen, but given the circumstances and some other correlations I saw, I believe we have him Going forward, I’ll refer to him as Leeker With the data I have provided to T2, I believe it will make their job easier when cooperating with law enforcement and help pin down the people involved (it’s not just one). Given the nature of the data I found, revealing it publicly could give the guys time to cover their tracks or alter evidence But I don’t want to leave people hanging so I’ll share some stuff below and answer as much as I can in the comments without it being damaging Stuff I have learned that can be shared: -> The leeker is actually a threat actor who has previous malicious activity under a different alias It also seems that there is genuine hatred toward this guy from the people around him. Much of my current information comes from his peers snitching after my first post -> This leak was apparently due to a breach or an insider. Nature of how is still unknown to me, But I’ve received two different stories about this so I'll share them anyway: 1. One story I got is that leeker bought a backdoor access or some sort from someone 2. Another person told me this is somewhat related to some incident inside Rockstar a while ago Keep in mind it’s pretty common for these crypto bros to boast and lie about their “gains” to each other, so I honestly don’t know which to believe here -> Not related at all to the ‘Mafia 808’ group -> I haven’t heard anything about Rockstar India being hacked unlike some other people are claiming -> This build we are seeing here is apparently from a year ago (as confirmed in ResetEra forums), but the theft I believe happened in April 2026 and the leeker has been preparing since then -> From what I can tell, and as I explained in one of the earlier posts, I don’t think they possess a live build. Most likely pre-recorded footage, but this is just my assumption The voting thing they did initially was to make people want to get the coin so the market cap can grow I highly recommend staying away from the coin. This is a literal pump and dump, don't buy into their "Fighting for gamers" bullshit. I won’t be surprised if they at some point start threatening to leak the story unless people donate to them in the poll. Even if that happens, don’t do it That’s all I can responsibly share right now. Questions that don’t risk the investigation or the people involved are welcome in the replies
显示更多
0
1.5K
8K
442
转发到社区