注册并分享邀请链接,可获得视频播放与邀请奖励。

与「CompleteDE」相关的搜索结果

CompleteDE 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 CompleteDE 的内容
Crew-13, SpaceX, and @NASA completed a full rehearsal of launch day activities
0
244
6K
589
转发到社区
Kling 4.0 is coming this October. Kling 4.0 Flash is live now for Ultra Yearly subscribers. Kling 4.0 brings video creation to a new level of visual realism, creative control, and narrative completeness. 🎞️ Upgraded Audio & Visuals: Stable dynamic motion, high-quality stereo audio, more accurate lip sync, up to 4K resolution, and 10-bit HDR output. 🔮 Omni Reference: Richer reference options with up to 15 multi-modal references, more consistent results, and enhanced video editing. 🎥 Seamless Storytelling: Multi-keyframe control supporting up to 10 keyframes, native 30-second generation. 🌍 Diverse Possibilities: Video extension, support for multiple languages, accents, and dialects. The stage is set. You call the shots!
显示更多
0
166
1.7K
230
转发到社区
Bitget will begin resuming withdrawals in orderly phases following the security incident identified on September 24. The vulnerability involved in the incident has been identified and remediated. Bitget's security and technical teams have since been conducting additional validation and security checks across the withdrawal infrastructure, while independent cybersecurity experts Mandiant and SlowMist continue to support the investigation. The temporary withdrawal pause remains a security measure and is not related to the availability of user assets. User account balances remain unaffected, and Bitget's Protection Fund covers the financial impact of this platform-wide incident. Withdrawals will resume in an orderly manner once these security checks are completed. The current withdrawal resumption schedule is as follows: > Sep 28, 8:00 (UTC): BTC (Bitcoin Network) > Sep 29, 8:00 (UTC): ETH (Ethereum, BSC, Arbitrum, Base, Optimism) > Sep 30, 8:00 (UTC): USDT (Ethereum, BSC, Solana, Tron) > Oct 2, 8:00 (UTC): Other Tokens / Fiat / P2P Our objective is to restore withdrawal services across all supported assets and networks as quickly and safely as possible. The incident remains contained, and no further unauthorized transfers are possible. User funds are unaffected throughout this process. Trading and deposits continue to operate. Users do not need to take any action ahead of the rollout. Withdrawal availability will be reflected directly on the Bitget platform, and users are advised to follow Bitget's official channels for updates.
显示更多
0
253
679
113
转发到社区
here's a prompt to improve your agent harness based on what we've learned at cursor. enjoy # Improve this agent harness's token efficiency You're working on an LLM agent harness: the system prompt, tool definitions, request assembly, context caching, compaction, and retrieval, and how work is split across agents. Make the agent's runs cheaper without making it worse at its job. - Objective: lower price-weighted token cost per completed task. - Constraint: no measurable drop in task quality. Measure per task, not per request. Every turn resends the prefix (tools, instructions, setup, and the conversation so far), so a change that shrinks each request but adds turns can cost more. Weight tokens by billing type: output, uncached input, and cached input are priced very differently. Work in this order: map the harness and measure the baseline, rank the opportunities, make the changes that are safe to make directly, put the rest behind flags or in proposals, then report. Figures below come from one team's production coding agent and its multi-agent experiments. Use them to gauge magnitude, not as targets. One round of these changes (prompt trimming, tool offloading, cache layout, sparse line numbers, subagent tuning) cut that team's overall token cost about 7% with no loss in quality. The larger percentages apply only to the part of the request each change touched. ## Principles 1. Change what the harness sends, not how hard the model tries. Don't ask the model to conserve tokens. A harness that told its model to "take care to preserve tokens and not be wasteful" found it grew reluctant to take on ambitious tasks and sometimes quit, saying it wasn't supposed to waste tokens. 2. Capable models need definitions, not commands. Lists of "DO NOT", "You must", and "Important", and guards against older models' habits, can usually be replaced with plain descriptions of what each tool does. One team cut about two-thirds of its system prompt this way, and the shorter prompt worked across model families. Instruct only on what the model can't know (the product, the environment, the user's processes) and on quirks you've seen in transcripts. 3. Static context is for what most turns need. Everything else should be discoverable when needed. Less up-front context also means less confusing or contradictory information. 4. Expect removals to win. Guardrails written for weaker models, coordination steps that became bottlenecks, and prompting for behavior the model now does on its own all cost tokens. 5. Real usage decides. Evals are a fast proxy, but they skew toward hard problems and miss the real mix of requests. ## 1. Map the harness and measure the baseline Find: - Where requests are assembled, the system prompt, and tool schemas. If a framework or SDK builds requests, find its hooks for message order, cache control, and tool loading. - How tool results are formatted, and how history is kept, trimmed, or summarized. - How subagents or parallel agents are spawned, if any. - Which models and provider APIs are used. From the provider's docs, get the prompt caching behavior (automatic or explicit breakpoints, TTL, minimum cacheable length) and the prices for output, uncached input, and cached input. - Existing logging, token accounting, and evals. If the harness doesn't record per-request token usage by billing type and cache hits, add that first. Everything later depends on it. Then render a few real requests (from logs, or by running representative tasks) and count tokens per section with the model's tokenizer or the API's usage fields. Produce: - Cost share by source × billing type. Sources: system prompt, tool definitions, skill/rule/integration descriptions, user messages, file reads, search results, command and other tool output, history, summaries, subagents. - Static tokens per request, cache hit rate, and turns per task. - Per tool: the share of runs that call it at least once, and its error rate. Read the rendered requests, not just the templates. Duplication, leaked volatile values, and misordered blocks only show up there. Rank opportunities by share of spend × fraction removable ÷ quality risk. ## 2. System prompt and injected context Label every instruction: - Keep: product or environment knowledge the model can't infer, fixes for quirks seen in this model's transcripts, and rules a mode depends on. - Rewrite: commands and emphasis into plain descriptions. Reminders into constraints: "No TODOs, no partial implementations" works better than "remember to finish implementations." Vague quantities into ranges: "generate 20–100 tasks" gets far more ambitious behavior than "generate many tasks." - Delete: things capable models do by default, guards against behavior you haven't seen from this model, text that repeats tool descriptions, and lines that could contradict a user request. Models trained to rank system instructions above user messages will side with the system prompt. - Move: anything per-user or per-request (date, environment, repo state, lists of skills or subagents, user rules) into a user-role setup message after the cache boundary. Audit other injected context the same way. As models improved, the team behind these figures dropped directory trees, pre-retrieved snippets, compressed copies of attached files, lint errors injected after every edit, forced expansion of short file reads, and caps on tool calls per turn. They kept small, high-value facts: OS, repo status, and open or recently viewed files. Skip checklists for open-ended work. The model optimizes the listed items and deprioritizes everything else. ## 3. Tool definitions Tool schemas ride along on every request. Most tools beyond the core set were each needed in under 20% of conversations, and moving them out of static context cut tool-description tokens 60%. Doing the same for integration tools (such as MCP servers), with names in context and full schemas in one folder per server that the agent can search with grep or jq, cut total tokens 46.9% in sessions that used them. - Keep in static context: high-frequency tools (for a coding agent: read, search, edit, shell), tools the model tries to call even when they're absent, and tools a mode depends on. - Offload the rest: leave a name or one-line pointer and make the full schema discoverable on demand. Group related tools so they load together, and put status (such as "needs re-authentication") where the agent will see it. - Tighten what remains: describe behavior and arguments, and drop usage lectures. - Pick the split by testing a few configurations and tracking tokens, cost, latency, tool-call errors, and task success. ## 4. Cache layout Order each request so the reusable prefix is as long as possible: `tool definitions → system instructions → [breakpoint] → setup message (skills, subagents, rules, environment) → [breakpoint] → conversation` - Keep the prefix byte-identical across turns. Use deterministic tool order and serialization, put timestamps and IDs after the boundary, and don't rewrite earlier messages except when compacting. - Use explicit breakpoints if the provider supports them. Otherwise rely on automatic prefix caching with the stable part first. Respect TTL and minimum-length rules. - Switching models mid-conversation throws away the cache (caches are per model and provider) and hands the new model a history it didn't write. When a different model is needed, run it as a subagent with fresh context. Explicit breakpoints plus moving per-request setup after them cut cold cache misses 20%. ## 5. Tool results and other context added during a run - Large outputs (commands, integrations, logs): write them to a file and return the path, size, and a short tail. The agent can tail, grep, or read ranges for more. Truncating loses data, and inlining bloats every later request. Treat long-running terminal sessions the same way. - High-volume formats: look for overhead repeated on every line or item. Numbering every 10th line of a file read instead of every line cut cache-read tokens 1.6% without hurting citation accuracy. Each number costs 3–5 tokens, and agents read tens of thousands of lines per session. Also check repeated absolute paths, verbose JSON keys, ANSI codes, progress bars, and repeated headers. - Good retrieval saves exploration turns. Adding semantic search alongside grep raised codebase question-answering accuracy 12.5% on average and cut the iterations users needed. - Tool errors waste tokens and leave confusing debris in context. Classify expected errors (invalid arguments, unexpected environment, provider error, timeout, user abort), treat unknown errors as harness bugs, and track rates per tool and per model. One focused effort along these lines cut unexpected tool errors 10×. ## 6. Long runs: compaction, subagents, and model mix - Compaction: keep the summarization prompt short and the summary compact, carry forward plan state and remaining tasks, and save the full history to a file the agent can search for details the summary dropped. A model trained to self-summarize from a one-line prompt wrote ~1k-token summaries with half the compaction error of a multi-thousand-token prompt that produced 5k+ token summaries. Untrained models may need more guidance, so test how short you can go. A more expensive summarization model made a negligible difference. - Scratchpads and running notes: rewrite them instead of appending. For repeated work in one environment, a small agent-maintained notes file with a line budget, loaded at start, is a promising way to shorten later runs. - Subagents: fresh context keeps the parent lean, but isolation adds coordination cost (duplicate or stale work). If the model already delegates on its own, remove prompting that pushes it to. Have subagents return short handoffs: what was done, findings, concerns, and deviations. A subagent should use a different model only when the user or harness says so. - Model mix: in large multi-agent runs, workers used at least 69% of tokens, and over 90% in most runs. A frontier planner with cheap workers matched a frontier model doing everything at about one-eighth the cost. Planner choice still changes worker spend. One planner that cost less on its own saw its workers use several times more tokens, and the run cost more overall. Measure the whole tree. - Routing and reasoning effort: send simple turns to a cheaper model or lower effort, and upgrade only when a stronger model is clearly better. A router built this way matched or beat single frontier models on user satisfaction at 41–68% lower cost. - Reasoning continuity: if the API returns reasoning items (including encrypted ones), pass them back on later turns and alert when they go missing. Dropping them cost one reasoning model 30% on a coding benchmark, and it burned tokens reconstructing its plan. ## 7. Fit the harness to each model Adapt to what each model was trained on instead of forcing one shape on all of them. If you've tuned the harness for a similar model, start from that version. - Edit format: use the one the model was trained on (for example, patch-style or search-and-replace). An unfamiliar format costs extra reasoning tokens and causes more mistakes. - Shell or tools: shell-first models fall back to `cat` or inline scripts. Name tools after their shell equivalents (such as `rg`), and if needed add: "If a tool exists for an action, prefer to use the tool instead of shell commands (e.g. read_file over `cat`)." - Literalness: some model families follow instructions literally and others tolerate imprecision. Some spiral on emphasized wording. Strip caps and emphasis for literal models. - Triggers: some models ignore a tool until told when to use it. A literal trigger works: "After substantive edits, use the to check recently edited files for linter errors. If you've introduced any, fix them if you can easily figure out how." - Progress updates: if a model reports progress through reasoning summaries, keep them to 1–2 sentences that note new findings or a change of tactic, and remove instructions about messaging mid-turn. - Quirks worth a targeted line: hedging or refusing as context fills ("context anxiety"), declaring completion early, stopping to ask permission, and calling tools that don't exist. Tie each added instruction to the transcript behavior it fixes. Re-audit when models change, since guidance one version needed can be dead weight for the next. ## 8. Validate - Offline: run a fixed set of realistic tasks before and after, ideally drawn from real usage and phrased the way users actually write (short and ambiguous). Compare task success, tokens, cost per task, turns, and tool errors. Don't ship a change that lowers success. - Online, if you have users: A/B test each change or small bundle. The primary metric is cost per completed task. Guardrails are task success signals, tool-call errors, latency, turns per task, and cache hit rate. For a coding agent, a good success signal is how much agent-written code survives over time. In general, check whether the user's next message moves on or reports a problem. - Ship only when cost drops and no guardrail regresses beyond noise. Record null results. ## What to change directly and what to propose - Change directly, each in its own revertible commit: token and cache telemetry, deterministic serialization and tool order, moving volatile content out of the cached prefix, explicit cache breakpoints, writing large outputs to files instead of truncating, passing back reasoning items that are being dropped, and fixes for recurring tool errors. - Change behind a flag so it can be tested: system prompt edits, tool offloading, output format changes, compaction changes, and subagent prompting. - Propose only: changes to which models run, routing, reasoning-effort defaults, or how work is split across agents. ## Traps - Asking the model to use fewer tokens or do less. - Truncating tool output. - Dropping reasoning items to save input tokens. - Volatile content in the cached prefix, or tool order that changes between requests. - Offloading a tool the model needs on the first turn or tries to call when it's missing. - Emphasis-heavy prompts (MUST, NEVER, IMPORTANT, all caps), especially with literal models. - Forcing a terser output format than the model was trained on. Fewer output tokens can mean less thinking and worse results. - Optimizing raw token counts instead of cost, per request instead of per task, or evals instead of real usage. - Switching models mid-conversation to save money. - Adding coordination layers that become bottlenecks. ## Report back with 1. The harness map and baseline: cost by source × billing type, with the biggest sources called out. 2. A ranked list of changes: layer, what changes, estimated savings and how you estimated them, quality risk, how to validate, and how to roll back. 3. The changes you made, including a system prompt diff with a keep, rewrite, delete, or move reason for each line. 4. A test plan for the flagged changes. 5. Gaps: anything you couldn't find or measure.
显示更多
0
92
1.4K
65
转发到社区
China’s first 2-tonne-payload tethered balloon has completed a successful trial run, carrying transmission tower materials to a 4,000m ridge.
0
53
900
184
转发到社区
LATEST: ⚡ StarkWare completed a quantum-resistant Bitcoin transaction on mainnet using Avihu Levy’s QSB method, showing it can work under existing rules but requiring up to $200 in off-chain computing.
显示更多
0
29
56
8
转发到社区
Link achieved 100% green building certification coverage by gross floor area across its regional portfolio in 2025/2026. We also completed BEAM Plus certification assessments for 46 properties and obtained HKGOC Wastewi$e certification for 74 properties in Hong Kong, while all properties in Chinese Mainland achieved the WELL Health-Safety Rating.
显示更多
LATEST: ⚡ South Korea's POSCO International completed a pilot tokenizing trade receivables on an Avalanche-based network, its second onchain trade finance test in a month after an earlier pilot on Injective.
显示更多
0
36
49
10
转发到社区
A TON OF THINGS HAPPENED IN THE STOCK MARKET TODAY. Here's a full recap: 1. The U.S. reportedly offered Iran a deal to halt the siege and lift sanctions in exchange for reopening the Strait of Hormuz and ending proxy attacks, according to Al Arabiya. Axios also reports that Rubio told several foreign counterparts the U.S. does not plan new strikes on Iran for now, with pressure shifting toward the naval blockade and new sanctions campaign instead. Crude Oil fell 4% and the 10-year treasury bond fell from 4.72% to 4.62%. 2. Global physical gold-backed ETFs $GLD attracted $6.4B of inflows last week, their largest weekly intake since January and the 3rd-largest weekly inflow on record. North America led with $4.4B, followed by Europe at $1.7B and Asia at $300M. This marked the 7th straight week of inflows, with global gold ETFs pulling in $16.4B over that stretch. Total AUM in global gold ETFs rose by $33B last week to $615B, the highest level since the second week of May. 3. Intuit $INTU reported Q4’26 revenue of $4.4B, beating estimates of $4.27B and up 14% YoY. Adjusted EPS came in at $4.03 versus $3.58 expected. Global Business Solutions revenue rose 14% YoY to $3.4B, the Online Ecosystem grew 17% YoY to $2.6B, Consumer revenue increased 14% YoY to $930M, and Credit Karma revenue rose 16% YoY to $743M. For FY27, Intuit guided revenue to $23.3B–$23.5B versus $23.72B expected, while adjusted EPS guidance of $22.88–$23.12 came in well below the $27.32 estimate. The company also raised its dividend 15% YoY to $1.38/share, bought back $5.5B of stock, and has $7.9B remaining on its authorization. Management said its strategy is to win as an AI-driven expert platform while staying disciplined on investments and scaling its big bets. 4. President Trump said the U.S. Navy has removed and/or detonated all mines from international waters in the Strait of Hormuz. He said Iran has been notified that any ship or boat placing new mines will be “immediately and systematically destroyed.” Trump added that Space Force is monitoring every square inch of the Strait, along with Pickaxe Mountain and the three previously destroyed nuclear sites, and said a “Zero Tolerance” policy on mine placement is now in full effect. 5. Canada is responding to U.S. tariffs with new tariffs of its own. The country is raising steel tariffs to 50% from 25%, while roughly 700 products will face new tariff rates of 15%, 25%, and 50%. The measures are set to take effect on September 8, marking another escalation in the U.S.–Canada trade dispute. 6. Anthropic is expected to tell IPO investors its total addressable market exceeds $30T, topping SpaceX’s $28.5T estimate, according to WSJ. The figure represents the potential value of work Anthropic believes AI models could eventually perform, not a direct revenue forecast. Anthropic generated $11.6B in Q2 revenue and could seek to raise as much as $100B at roughly a $2T valuation. IPO documents are expected within weeks, potentially setting up a September or early October listing. 7. OpenAI’s data-center head Chris Malone left the company last week, according to WSJ. Malone joined in March 2025 shortly after Stargate was announced and played a key role overseeing OpenAI’s massive data-center buildout. He previously led data-center strategy at Meta and earlier worked on data-center technology at Google. The departure comes just weeks after OpenAI also replaced its chief revenue officer, adding another senior leadership change as the company races to scale infrastructure, revenue, and compute capacity. 8. ClickHouse has surpassed $350M in annual recurring revenue, up 40% since May, as AI agents drive demand for database and observability infrastructure. OpenAI’s usage has reportedly grown roughly 10x over the past year to more than 30 petabytes of data per day, or around 30T events daily. OpenAI has also shifted parts of its log-management workload from Datadog to ClickHouse over the past year. ClickHouse was valued at $15B in January and says gross margins currently range from 50%–70%. Earlier this year, the company acquired Langfuse to expand deeper into monitoring AI applications and agents. Nebius $NBIS owned a 28% stake in ClickHouse as of May 2025, though that stake has likely been diluted by subsequent fundraising. 9. JPMorgan reiterated its Overweight rating on SpaceX $SPCX with a $240 price target, saying the company’s AI ambitions are coming into sharper focus and that it is increasingly positive on Grok. The firm highlighted SpaceX’s completed acquisition of Cursor on 8/14 as an important step in building enterprise AI capabilities. Cursor brings roughly $4B of ARR as of June 2026, with about 75% coming from businesses, which JPMorgan says should help streamline go-to-market and provide valuable model-training data. The firm also said Cursor data is already showing up in Grok’s supplemental training, with tangible improvements in recent model performance. 10. OpenAI says its new Broadcom-built Jalapeno AI chip outperformed Nvidia $NVDA GB300 in both throughput per watt and response latency during internal testing, according to Bloomberg. The chip is built specifically for inference, not training, and runs at roughly 700 watts. OpenAI plans to begin deploying Jalapeno for its models later this year, saying the performance gap widened on larger workloads, including Moonshot’s Kimi model, and that the chip has also performed well on unreleased OpenAI models. The key caveat is that Jalapeno was tested against GB300, not Nvidia’s newer Vera Rubin generation. OpenAI says a second-generation chip is already nearing tape-out, while work on a third generation has begun. 11. The top 10 most active options today by contracts traded were $NVDA with 1.8M contracts, $TSLA with 1.8M contracts, $AAPL with 636K contracts, $SPCX with 548K contracts, $INTC with 540K contracts, $AMZN with 498K contracts, $MU with 483K contracts, $AMD with 403K contracts, $PLTR with 361K contracts, and $SOFI with 359K contracts. 12. Raymond James raised its Nvidia $NVDA price target to $352 from $330 and reiterated a Strong Buy rating. The firm says Nvidia’s CPU opportunity is becoming more important, especially for agentic AI workloads, even though CPUs are only about 3% of sales today. Raymond James expects CPU revenue to reach roughly 5% of total revenue by CY28 and believes Nvidia could potentially become the world leader in CPU revenue within several years. The firm also argued the stock remains inexpensive, trading at less than 15x CY27 GAAP earnings, below the S&P 500 at 18.6x, despite sales and net income growth still expected to exceed 20% in CY28. Its new $352 target is based on a 22x multiple on CY28 estimates, which Raymond James views as conservative given Nvidia’s leadership, CUDA moat, GPU performance, free cash flow, and history of trading at much higher multiples. WALL STREET IS THE GREATEST SHOW ON EARTH.
显示更多
0
75
2.5K
153
转发到社区
Grok Summary of JPMorgan’s new note today on @SpaceX and @Grok • JPMorgan maintained its Overweight rating on SpaceX, saying it is “increasingly positive” about Grok. • The completed Cursor acquisition is an important step in expanding SpaceX’s enterprise AI capabilities. • JPMorgan estimates that Cursor generates roughly $4 billion in annual recurring revenue, with about 75% coming from business customers. Cursor also gives SpaceX an enterprise distribution channel and valuable coding data. • JPMorgan has already seen “tangible improvements” in recent Grok models after Cursor data was incorporated into their supplemental training. • The firm says Grok 4.6 combines frontier intelligence with meaningfully lower costs than its peers, supporting stronger adoption. • Grok Bot expands SpaceX into enterprise AI agents that can perform tasks across workplace applications, targeting the fast-growing AI productivity market. • JPMorgan expects monthly model releases throughout the rest of 2026, culminating with Grok 5 in December. The firm expects Grok 5 to deliver a major improvement in performance and reach. • JPMorgan believes Cursor’s coding expertise, combined with roughly 25 years of SpaceX’s proprietary engineering knowledge, could give Grok an advantage in software development and complex engineering applications. • The firm expects improving Grok monetization, particularly among enterprise customers, to become an increasingly important driver of SpaceX’s AI revenue. • JPMorgan expects Grok, Cursor and other enterprise AI products to become core long-term growth drivers alongside launch services and Starlink. SpaceX is steadily turning AI into a powerful long-term growth engine. Bullish.
显示更多
0
37
250
38
转发到社区