注册并分享邀请链接,可获得视频播放与邀请奖励。

与「gauges」相关的搜索结果

gauges 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 gauges 的内容
I love these gauges, what are your favorite sites or places to get yours from? #gauges# #alternativegirls# #lingeries# #suicidegirls# #suicidegirlhopeful# #allnatural# #girlswithtattoos# #onlyfans# #latina# #tatted# #onlineshopping#
显示更多
here's a prompt to improve your agent harness based on what we've learned at cursor. enjoy # Improve this agent harness's token efficiency You're working on an LLM agent harness: the system prompt, tool definitions, request assembly, context caching, compaction, and retrieval, and how work is split across agents. Make the agent's runs cheaper without making it worse at its job. - Objective: lower price-weighted token cost per completed task. - Constraint: no measurable drop in task quality. Measure per task, not per request. Every turn resends the prefix (tools, instructions, setup, and the conversation so far), so a change that shrinks each request but adds turns can cost more. Weight tokens by billing type: output, uncached input, and cached input are priced very differently. Work in this order: map the harness and measure the baseline, rank the opportunities, make the changes that are safe to make directly, put the rest behind flags or in proposals, then report. Figures below come from one team's production coding agent and its multi-agent experiments. Use them to gauge magnitude, not as targets. One round of these changes (prompt trimming, tool offloading, cache layout, sparse line numbers, subagent tuning) cut that team's overall token cost about 7% with no loss in quality. The larger percentages apply only to the part of the request each change touched. ## Principles 1. Change what the harness sends, not how hard the model tries. Don't ask the model to conserve tokens. A harness that told its model to "take care to preserve tokens and not be wasteful" found it grew reluctant to take on ambitious tasks and sometimes quit, saying it wasn't supposed to waste tokens. 2. Capable models need definitions, not commands. Lists of "DO NOT", "You must", and "Important", and guards against older models' habits, can usually be replaced with plain descriptions of what each tool does. One team cut about two-thirds of its system prompt this way, and the shorter prompt worked across model families. Instruct only on what the model can't know (the product, the environment, the user's processes) and on quirks you've seen in transcripts. 3. Static context is for what most turns need. Everything else should be discoverable when needed. Less up-front context also means less confusing or contradictory information. 4. Expect removals to win. Guardrails written for weaker models, coordination steps that became bottlenecks, and prompting for behavior the model now does on its own all cost tokens. 5. Real usage decides. Evals are a fast proxy, but they skew toward hard problems and miss the real mix of requests. ## 1. Map the harness and measure the baseline Find: - Where requests are assembled, the system prompt, and tool schemas. If a framework or SDK builds requests, find its hooks for message order, cache control, and tool loading. - How tool results are formatted, and how history is kept, trimmed, or summarized. - How subagents or parallel agents are spawned, if any. - Which models and provider APIs are used. From the provider's docs, get the prompt caching behavior (automatic or explicit breakpoints, TTL, minimum cacheable length) and the prices for output, uncached input, and cached input. - Existing logging, token accounting, and evals. If the harness doesn't record per-request token usage by billing type and cache hits, add that first. Everything later depends on it. Then render a few real requests (from logs, or by running representative tasks) and count tokens per section with the model's tokenizer or the API's usage fields. Produce: - Cost share by source × billing type. Sources: system prompt, tool definitions, skill/rule/integration descriptions, user messages, file reads, search results, command and other tool output, history, summaries, subagents. - Static tokens per request, cache hit rate, and turns per task. - Per tool: the share of runs that call it at least once, and its error rate. Read the rendered requests, not just the templates. Duplication, leaked volatile values, and misordered blocks only show up there. Rank opportunities by share of spend × fraction removable ÷ quality risk. ## 2. System prompt and injected context Label every instruction: - Keep: product or environment knowledge the model can't infer, fixes for quirks seen in this model's transcripts, and rules a mode depends on. - Rewrite: commands and emphasis into plain descriptions. Reminders into constraints: "No TODOs, no partial implementations" works better than "remember to finish implementations." Vague quantities into ranges: "generate 20–100 tasks" gets far more ambitious behavior than "generate many tasks." - Delete: things capable models do by default, guards against behavior you haven't seen from this model, text that repeats tool descriptions, and lines that could contradict a user request. Models trained to rank system instructions above user messages will side with the system prompt. - Move: anything per-user or per-request (date, environment, repo state, lists of skills or subagents, user rules) into a user-role setup message after the cache boundary. Audit other injected context the same way. As models improved, the team behind these figures dropped directory trees, pre-retrieved snippets, compressed copies of attached files, lint errors injected after every edit, forced expansion of short file reads, and caps on tool calls per turn. They kept small, high-value facts: OS, repo status, and open or recently viewed files. Skip checklists for open-ended work. The model optimizes the listed items and deprioritizes everything else. ## 3. Tool definitions Tool schemas ride along on every request. Most tools beyond the core set were each needed in under 20% of conversations, and moving them out of static context cut tool-description tokens 60%. Doing the same for integration tools (such as MCP servers), with names in context and full schemas in one folder per server that the agent can search with grep or jq, cut total tokens 46.9% in sessions that used them. - Keep in static context: high-frequency tools (for a coding agent: read, search, edit, shell), tools the model tries to call even when they're absent, and tools a mode depends on. - Offload the rest: leave a name or one-line pointer and make the full schema discoverable on demand. Group related tools so they load together, and put status (such as "needs re-authentication") where the agent will see it. - Tighten what remains: describe behavior and arguments, and drop usage lectures. - Pick the split by testing a few configurations and tracking tokens, cost, latency, tool-call errors, and task success. ## 4. Cache layout Order each request so the reusable prefix is as long as possible: `tool definitions → system instructions → [breakpoint] → setup message (skills, subagents, rules, environment) → [breakpoint] → conversation` - Keep the prefix byte-identical across turns. Use deterministic tool order and serialization, put timestamps and IDs after the boundary, and don't rewrite earlier messages except when compacting. - Use explicit breakpoints if the provider supports them. Otherwise rely on automatic prefix caching with the stable part first. Respect TTL and minimum-length rules. - Switching models mid-conversation throws away the cache (caches are per model and provider) and hands the new model a history it didn't write. When a different model is needed, run it as a subagent with fresh context. Explicit breakpoints plus moving per-request setup after them cut cold cache misses 20%. ## 5. Tool results and other context added during a run - Large outputs (commands, integrations, logs): write them to a file and return the path, size, and a short tail. The agent can tail, grep, or read ranges for more. Truncating loses data, and inlining bloats every later request. Treat long-running terminal sessions the same way. - High-volume formats: look for overhead repeated on every line or item. Numbering every 10th line of a file read instead of every line cut cache-read tokens 1.6% without hurting citation accuracy. Each number costs 3–5 tokens, and agents read tens of thousands of lines per session. Also check repeated absolute paths, verbose JSON keys, ANSI codes, progress bars, and repeated headers. - Good retrieval saves exploration turns. Adding semantic search alongside grep raised codebase question-answering accuracy 12.5% on average and cut the iterations users needed. - Tool errors waste tokens and leave confusing debris in context. Classify expected errors (invalid arguments, unexpected environment, provider error, timeout, user abort), treat unknown errors as harness bugs, and track rates per tool and per model. One focused effort along these lines cut unexpected tool errors 10×. ## 6. Long runs: compaction, subagents, and model mix - Compaction: keep the summarization prompt short and the summary compact, carry forward plan state and remaining tasks, and save the full history to a file the agent can search for details the summary dropped. A model trained to self-summarize from a one-line prompt wrote ~1k-token summaries with half the compaction error of a multi-thousand-token prompt that produced 5k+ token summaries. Untrained models may need more guidance, so test how short you can go. A more expensive summarization model made a negligible difference. - Scratchpads and running notes: rewrite them instead of appending. For repeated work in one environment, a small agent-maintained notes file with a line budget, loaded at start, is a promising way to shorten later runs. - Subagents: fresh context keeps the parent lean, but isolation adds coordination cost (duplicate or stale work). If the model already delegates on its own, remove prompting that pushes it to. Have subagents return short handoffs: what was done, findings, concerns, and deviations. A subagent should use a different model only when the user or harness says so. - Model mix: in large multi-agent runs, workers used at least 69% of tokens, and over 90% in most runs. A frontier planner with cheap workers matched a frontier model doing everything at about one-eighth the cost. Planner choice still changes worker spend. One planner that cost less on its own saw its workers use several times more tokens, and the run cost more overall. Measure the whole tree. - Routing and reasoning effort: send simple turns to a cheaper model or lower effort, and upgrade only when a stronger model is clearly better. A router built this way matched or beat single frontier models on user satisfaction at 41–68% lower cost. - Reasoning continuity: if the API returns reasoning items (including encrypted ones), pass them back on later turns and alert when they go missing. Dropping them cost one reasoning model 30% on a coding benchmark, and it burned tokens reconstructing its plan. ## 7. Fit the harness to each model Adapt to what each model was trained on instead of forcing one shape on all of them. If you've tuned the harness for a similar model, start from that version. - Edit format: use the one the model was trained on (for example, patch-style or search-and-replace). An unfamiliar format costs extra reasoning tokens and causes more mistakes. - Shell or tools: shell-first models fall back to `cat` or inline scripts. Name tools after their shell equivalents (such as `rg`), and if needed add: "If a tool exists for an action, prefer to use the tool instead of shell commands (e.g. read_file over `cat`)." - Literalness: some model families follow instructions literally and others tolerate imprecision. Some spiral on emphasized wording. Strip caps and emphasis for literal models. - Triggers: some models ignore a tool until told when to use it. A literal trigger works: "After substantive edits, use the to check recently edited files for linter errors. If you've introduced any, fix them if you can easily figure out how." - Progress updates: if a model reports progress through reasoning summaries, keep them to 1–2 sentences that note new findings or a change of tactic, and remove instructions about messaging mid-turn. - Quirks worth a targeted line: hedging or refusing as context fills ("context anxiety"), declaring completion early, stopping to ask permission, and calling tools that don't exist. Tie each added instruction to the transcript behavior it fixes. Re-audit when models change, since guidance one version needed can be dead weight for the next. ## 8. Validate - Offline: run a fixed set of realistic tasks before and after, ideally drawn from real usage and phrased the way users actually write (short and ambiguous). Compare task success, tokens, cost per task, turns, and tool errors. Don't ship a change that lowers success. - Online, if you have users: A/B test each change or small bundle. The primary metric is cost per completed task. Guardrails are task success signals, tool-call errors, latency, turns per task, and cache hit rate. For a coding agent, a good success signal is how much agent-written code survives over time. In general, check whether the user's next message moves on or reports a problem. - Ship only when cost drops and no guardrail regresses beyond noise. Record null results. ## What to change directly and what to propose - Change directly, each in its own revertible commit: token and cache telemetry, deterministic serialization and tool order, moving volatile content out of the cached prefix, explicit cache breakpoints, writing large outputs to files instead of truncating, passing back reasoning items that are being dropped, and fixes for recurring tool errors. - Change behind a flag so it can be tested: system prompt edits, tool offloading, output format changes, compaction changes, and subagent prompting. - Propose only: changes to which models run, routing, reasoning-effort defaults, or how work is split across agents. ## Traps - Asking the model to use fewer tokens or do less. - Truncating tool output. - Dropping reasoning items to save input tokens. - Volatile content in the cached prefix, or tool order that changes between requests. - Offloading a tool the model needs on the first turn or tries to call when it's missing. - Emphasis-heavy prompts (MUST, NEVER, IMPORTANT, all caps), especially with literal models. - Forcing a terser output format than the model was trained on. Fewer output tokens can mean less thinking and worse results. - Optimizing raw token counts instead of cost, per request instead of per task, or evals instead of real usage. - Switching models mid-conversation to save money. - Adding coordination layers that become bottlenecks. ## Report back with 1. The harness map and baseline: cost by source × billing type, with the biggest sources called out. 2. A ranked list of changes: layer, what changes, estimated savings and how you estimated them, quality risk, how to validate, and how to roll back. 3. The changes you made, including a system prompt diff with a keep, rewrite, delete, or move reason for each line. 4. A test plan for the flagged changes. 5. Gaps: anything you couldn't find or measure.
显示更多
0
92
1.4K
65
转发到社区
The S&P 500 fell slightly and the Dow ended virtually unchanged as investors ​waited for key earnings reports to gauge the health of a market rally fed by enthusiasm for AI
显示更多
Amazon Prime Day to gauge US consumer strain as focus shifts to basics
ve33大概率将退出历史舞台,Aerodrome在7月要上线的新版本对ve33的几项改进都是极具针对性的。 Aerodrome指出的ve33存在的最大问题就是滞后性,池激励每周投票更新一次,且投票趋势都是基于历史表现,比如上周Pool A的表现好,那么投票多数在下周还会集中在A,这对于新Pool非常不友好, 对某个新Pool的爆发潜力也是基本感知不到的。 新的设计取消了vetoken,改成sAERO(可随时转移了),取消了epoch每周投票这个设定,而是可以随时调整投票。AERO 奖励也是以streaming的形式连续发放到Pool gauge,同时fee也是实时积累返给sAERO,看起来改动不大,但实际上Pool、LP、holder三者的关系发生了巨大变化。 原来的设计,投票和激励获得都是有滞后性的(1周),而手握票权的veAERO,想要获得更多的fee分成,那么就要判断未来表现好的Pool,在老机制里,这基本是没什么办法的,只能去选历史表现好的Pool,因为你即使判断下周有一个新Pool会爆发,你去押注它的性价比也极低,这1周的窗口期选错就会损失大部分收益,基本没有人会去做这种前瞻性的预测,而是不停的追数据好的Pool,一旦一个Pool成为“历史赢家”,它会持续吃排放,即使用户开始怀疑它未来会变差,也很难快速切换(1周的硬性周期),明显这对于新Pool来说不友好。 把epoch改掉之后,灵活性大大提升,预测正确后的放大效应更强,预测错误后的止损速度也更快,这会让整体资本配置效率更高。 进一步思考,这种实时、可动态调整的设计天然更适合 AI Agent 进行自动化决策和频繁调仓,这也是Aerodrome一直在做的适合AI接入的dex这样一个故事,原来的vetoken模型则是很难的。
显示更多
BOJ's new trend gauge shows inflation exceeding target
I explained the Chinese real estate & debt crisis in much detail in 2024 on my Substack. Nothing has changed since. China is in what we call "the largest balance-sheet recession the world has ever seen". And it will take years to get out of it and assuming the CCP's investment-led growth model does not dig the next hole in the meantime - a likely. The FT published added some colour to it two days ago: "Housing is important to every economy. But to China, it’s extra important. According to the PBoC, 96% of urban households own a home, and 41% own at least two. The average household owns 1.5 properties. And as such, property constitutes around 70% of China’s private wealth. The comparable figure for the US is around 30%. So when Chinese property prices fall, the authors make a pretty compelling case that this has all sorts of particularly bad economic spillovers. And fall they have. The negative wealth effect is substantial, and “effects are amplified by elevated household debt, much of which consists of mortgage obligations”. This — and the weaker income expectations that the falls generate — goes some way to suppressing consumption. Moreover, declining land-sale revenues constrain local government budgets, “limiting their capacity to finance developmental projects and maintain existing public infrastructure”. And this is even before any credit impacts from rising non-performing loans and mortgages on bank balance sheets are considered. Tl;dr: bad bad bad. Of course, China isn’t the first soon-to-be-global-economic-hegemon-East-Asian-power staring down demographic oblivion to have piled its savings into a property boom. Back in 1991, the world was fretting over the rise and rise of Japan. And the Japanese were buying Japanese residential real estate at outlandish prices. Japan’s house prices peaked back in 1991 and spent the next 30 years on a downward trajectory. We’re only a few years into the Chinese property bust, and its ultimate trajectory is both unknown and unknowable. But Rogoff and Yang have pulled together some cool data they kindly shared with Alphaville, allowing us to make this chart below. So far, it looks like prices in Chinese cities are falling at around the same pace as they did over the first five-to-10 years of Japan’s bust. Japan’s property crash is associated with a lost decade (or two) of economic growth. In the 10 years leading up to 1991, Japanese real annual GDP growth averaged 4.4%. In the subsequent 10 years it averaged only 0.9% per annum. The same numbers for China, with 2021 marking its property zenith, are 7.0% per year and 4.6% per year (so far). If the IMF’s forecasts turn out right, this latter number will fall to around 4.0% per annum. While the levels are different, the before-and-after drop looks comparable. Was it housing wot dun it? Rogoff and Yang reckon that a 40% decline in house prices translates into a total consumption loss of 2-4% of GDP. Not nothing, but not a single answer explaining life, the universe and wiggles in the decadal pace of real economic growth. To get here, they construct a historical dataset comprising subnational data across 47 prefectures, and input and output data at granular industry levels. They then use this to examine the macroeconomic implications of Japan’s real estate bust. And the authors argue that: a housing bust can generate substantial adverse effects on the economy via real channels. . . . overbuilding during the boom can trigger a demand-driven recession with limited reallocation and low output. Unlike financial channels, which amplify shocks through leverage, bank balance sheets, credit constraints, or fire sales, real channels operate directly through investment, consumption, labour markets, or productivity. In Japan’s case, the housing market collapse depressed activity through three key real channels: investment, consumption, and sentiment. This is all pretty intuitive. But using city-level and household-level Chinese data plus some whizzy maths, they put meat on the bone for these three channels. They find that Chinese cities that overbuilt housing the most are less keen on new building, suppressing investment. Sounds legit. Chinese household consumption is estimated to be more responsive to house price changes than it was in either Japan or the US given its outsized role in private wealth. And it looks to the authors like people have scrambled to rebuild precautionary savings they thought they had amassed in property. Understandable. Then, on the sentiment side, Rogoff and Yang use an LLM to gauge market perceptions of the housing market. And by incorporating city-specific perceptions, they double the estimated effect of house price changes on consumption. Huh. While China is not Japan, 1991 was not 2021, and a *lot* of other things are/were going on, it’s interesting to see that the overall magnitude and pace of property price falls — as well as the aggregate drop in the pace of headline GDP growth — has (so far) been spookily similar. And as for the big question — are we there yet? "If China’s adjustment unfolds in a similar way as Japan’s, it would mean China has not gone half way through the transition. By contrast, if China’s path is eventually comparable to the United States, it appears to have already covered roughly two-thirds of the adjustment before reaching the bottom." So more to come.
显示更多
0
17
336
85
转发到社区
Hackers breached automatic tank gauge systems used to monitor fuel levels in underground storage tanks at gas stations across the United States. These systems were connected to the internet but lacked basic security protections such as passwords. U.S. officials suspect Iranian hackers. The intruders could view and alter the readings on tank monitors but could not change actual fuel amounts or cause physical damage or operational disruptions. The main concern is that attackers could mask a real fuel leak by falsifying data, creating safety or environmental hazards. No such incidents have been reported.
显示更多
0
23
713
42
转发到社区
A key strength: Gemini Robotics-ER 1.6 combines spatial reasoning, world knowledge, and agentic vision to allow robots to read a variety of instruments. See how it reads an analogy gauge right down to sub tick accuracy ↓
显示更多
While visiting the Ganges River, @nikolajcw met Ankit Agarwal, founder of Phool, a company turning temple flower waste into new products. Here’s what inspired him to start the mission. Watch the full episode of An Optimist’s Guide to the Planet 
显示更多