注册并分享邀请链接,可获得视频播放与邀请奖励。

与「SD」相关的搜索结果

SD 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 SD 的内容
鲸鱼 qianbaidu.eth 在过去半个小时里把 2575 万 U 转进 Binance 的同时,从 Binance 提取出了 7,500 枚 ETH ($2053 万)。 -------------------------------------------------------------------------- #Bitget# 来了就是VIP!Crypto、美股、CFD,全球先机一站布局
显示更多
原来这就是赛博活佛的含金量!!免费用啊家人们,这得省多少token!! DeepSeek-V4.1-Flash 最新模型,免费用两周,每天 200 次调用不限速!! 我一开始还以为看错了,反复确认了三遍 作为一个每天都在跑 Agent 任务的人,我最头疼的事不是模型不够强,是管理成本太高 DeepSeek 一个账号、Kimi 一个账号、GLM 一个账号、Qwen 一个账号,API Key 管了五六把,每个月充值要开四五个后台,对完账才知道这个月到底在 AI 上烧了多少钱 更离谱的是,大部分平台充进去的额度根本用不完,但你不充又没法调用,等于在给每个平台交月租 后来发现了 HiLinkup @HiLinkup_AI ——一个 API Key,直接接 DeepSeek、Qwen、Kimi、GLM、Doubao、MiniMax,文本图像视频音频全覆盖 我试了一下,换个 base URL 就行,代码一行不用改,OpenAI SDK 直接兼容 响应速度很快,关键是价格——大概是官方的 1 折左右!!,没有 TPU 上限 说白了就是:以前我每个月在五六个平台上分别充钱,现在一个账户统一计费,按量走,用多少算多少 关键是它现在每天送 1500 万 token 免费额度,日常跑 Agent 任务完全够用 唯一的门槛就是首充 5 块钱激活一下,充完还送 5 块,而且充进去的钱不消耗,每天只走免费额度 对于我们这种每天都在用国产模型跑各种任务的人来说,这个事的意义不只是省钱——是终于不用当五个平台的出纳了 一个 Key 管所有模型,按任务选型,不被任何一家绑死 说实话这种聚合平台的窗口期不会太长,各家模型迟早会收紧 API 授权,能用的时候赶紧先用起来!! 链接:
显示更多
0
30
86
16
转发到社区
casual pics cause i’ve been whoring it up all week
终于有人把这事做了!不用碰 Python,零依赖 JS/TS 直接拿 A股/港股/美股/基金/期货/期权全市场行情 stock-sdk,1.9K stars,内置 92 个 MCP 工具直接喂给 AI。 以前前端想做行情看板,行情工具基本全是 Python 生态,要么自己起后端,要么忍受财经接口 GBK/并发/批量各种坑。 stock-sdk 不一样,v2 架构大改,命名空间 API + 统一符号模型,零依赖,浏览器和 Node 18+ 双端直接跑。 核心能力: 1️⃣ 一套 API 覆盖 A 股、港股、美股、公募基金、期货、期权,实时行情、K 线、分钟线都能直接调 2️⃣ 资金流、龙虎榜、融资融券、大宗交易、涨跌停池、盘口异动、板块资金这些 A 股常用数据也做进去了 3️⃣ 自带 MACD、RSI、KDJ、BOLL、ATR、SAR 等技术指标,还能直接识别指标信号 4️⃣ 内置 Screener + 本地回测,不只是拿数据,筛股和策略验证也能直接做 5️⃣ 筹码分布也已经支持,A 股 / 港股 / 美股都能本地计算获利比例、平均成本和筹码区间 6️⃣ 最关键的是 MCP 已经原生做好了,Claude、Codex、Cursor 可以直接调用这些金融工具,不用你再包一层接口 我觉得它现在最有意思的一点,是已经从“股票数据包”慢慢变成了 Agent 的金融基础设施。 上面 AI 负责研究,下面 stock-sdk 负责行情、指标、选股和回测。 做股票 Agent、AI 投研,或者想给自己的量化系统补一套统一数据层,这个很值得留着。
显示更多
Who drafted right this season? 🙋‍♂️

 @LACvsBUF – Sunday 1pm ET on FOX/FOX One Stream on @NFLPlus
宅飲みの優勝がこれw w
0
40
1.1K
46
转发到社区
Google 开源了 "Agent 工作负载的 Kubernetes"「AX」 AX 是为 Agent 设计的声明式编排运行时,你用 YAML 声明一个 Agent 任务,AX 负责在集群中沙箱化、配置环境、管控网络并大规模运行它。 开源地址: 它解决什么问题 项目的立论很清晰:Agent 是一种既有的编排体系都不匹配的新型工作负载。 · 它不像微服务(无状态、常驻),Agent 会持续积累状态(对话记忆、工作区文件、工具会话); · 它不像批处理作业(跑完即弃),Agent 大部分时间在等待,等模型响应、等工具返回、等人类审批,期间沙箱空转烧钱; · 它运行的是不可信代码,需要严格隔离;它还调用外部模型 API 和 MCP 工具服务器,需要网络与凭据管控。 架构:四个二进制 + Redis 四个二进制分工:ax(开发者 CLI)、ax-server(无状态 gRPC API)、ax-controller(调和循环)、ax-task-runner(每个任务容器内的 PID 1)。 核心原语:Task / Workspace / Model (+ Gateway) Task 是最小执行单元:带 CPU/内存限制的隔离沙箱。AX 刻意把它做得细粒度、可自由组合:一个任务可以是全部工作,也可以是任务树的根节点。生命周期用 status.phase + Conditions 表达,支持挂起(检查点保存状态)与恢复。 Workspace 是最有产品想象力的一层。它把“环境准备”声明化:列出需要的 Git 仓库、MCP 服务器、技能包,runner 在命令启动前物化好。更激进的是 generative workspace:你可以只写一句自然语言目标("搭一个 Python 3 开发环境"),首次启动时 runner 会派一个引导 Agent(Antigravity,需 GEMINI_API_KEY,默认限时 10 分钟)去实际安装工具链并验证依赖。声明一次,任意任务复用。 Model 把“用哪个模型、什么参数、密钥在哪”抽成命名资源,密钥引用 K8s Secret。轮换密钥、锁版本、调温度只需一次 ax apply。 沙箱与运行时细节 每个任务容器以 ax-task-runner 为 PID 1:启动元数据服务(端口 80,HTTP/1.1+h2c,暴露 /healthz、/readyz 和任务/工作区自省端点——Agent 不需要 SDK 就能读到自己的配置)、按绑定顺序初始化工作区、然后 fork 出 spec.command 并持续监管。命令退出后 runner 仍驻留,所以 ax ssh 和元数据服务在命令结束后依然可用。停机采用 SIGTERM + 10 秒宽限 + 强杀的两级策略。 安全模型有几处值得注意的门控:guest 服务(任意进程执行与文件读写,ax ssh 的底层)默认关闭,只在 debug: true 时开启——ax ssh 连不上未开启的任务是刻意的安全设计,不是故障;Gateway 用显式 allowlist 管控出站流量,并可向入站请求注入凭据,避免把 API key 直接塞进 Agent 环境。 路线图透露的方向 五大方向:Actor 架构深化、空闲检测自动挂起、有状态任务 fork、任务级 SPIFFE 身份做零信任 mTLS、runner 层自动采集 OpenTelemetry 遥测与结构化 Agent 轨迹。
显示更多
here's a prompt to improve your agent harness based on what we've learned at cursor. enjoy # Improve this agent harness's token efficiency You're working on an LLM agent harness: the system prompt, tool definitions, request assembly, context caching, compaction, and retrieval, and how work is split across agents. Make the agent's runs cheaper without making it worse at its job. - Objective: lower price-weighted token cost per completed task. - Constraint: no measurable drop in task quality. Measure per task, not per request. Every turn resends the prefix (tools, instructions, setup, and the conversation so far), so a change that shrinks each request but adds turns can cost more. Weight tokens by billing type: output, uncached input, and cached input are priced very differently. Work in this order: map the harness and measure the baseline, rank the opportunities, make the changes that are safe to make directly, put the rest behind flags or in proposals, then report. Figures below come from one team's production coding agent and its multi-agent experiments. Use them to gauge magnitude, not as targets. One round of these changes (prompt trimming, tool offloading, cache layout, sparse line numbers, subagent tuning) cut that team's overall token cost about 7% with no loss in quality. The larger percentages apply only to the part of the request each change touched. ## Principles 1. Change what the harness sends, not how hard the model tries. Don't ask the model to conserve tokens. A harness that told its model to "take care to preserve tokens and not be wasteful" found it grew reluctant to take on ambitious tasks and sometimes quit, saying it wasn't supposed to waste tokens. 2. Capable models need definitions, not commands. Lists of "DO NOT", "You must", and "Important", and guards against older models' habits, can usually be replaced with plain descriptions of what each tool does. One team cut about two-thirds of its system prompt this way, and the shorter prompt worked across model families. Instruct only on what the model can't know (the product, the environment, the user's processes) and on quirks you've seen in transcripts. 3. Static context is for what most turns need. Everything else should be discoverable when needed. Less up-front context also means less confusing or contradictory information. 4. Expect removals to win. Guardrails written for weaker models, coordination steps that became bottlenecks, and prompting for behavior the model now does on its own all cost tokens. 5. Real usage decides. Evals are a fast proxy, but they skew toward hard problems and miss the real mix of requests. ## 1. Map the harness and measure the baseline Find: - Where requests are assembled, the system prompt, and tool schemas. If a framework or SDK builds requests, find its hooks for message order, cache control, and tool loading. - How tool results are formatted, and how history is kept, trimmed, or summarized. - How subagents or parallel agents are spawned, if any. - Which models and provider APIs are used. From the provider's docs, get the prompt caching behavior (automatic or explicit breakpoints, TTL, minimum cacheable length) and the prices for output, uncached input, and cached input. - Existing logging, token accounting, and evals. If the harness doesn't record per-request token usage by billing type and cache hits, add that first. Everything later depends on it. Then render a few real requests (from logs, or by running representative tasks) and count tokens per section with the model's tokenizer or the API's usage fields. Produce: - Cost share by source × billing type. Sources: system prompt, tool definitions, skill/rule/integration descriptions, user messages, file reads, search results, command and other tool output, history, summaries, subagents. - Static tokens per request, cache hit rate, and turns per task. - Per tool: the share of runs that call it at least once, and its error rate. Read the rendered requests, not just the templates. Duplication, leaked volatile values, and misordered blocks only show up there. Rank opportunities by share of spend × fraction removable ÷ quality risk. ## 2. System prompt and injected context Label every instruction: - Keep: product or environment knowledge the model can't infer, fixes for quirks seen in this model's transcripts, and rules a mode depends on. - Rewrite: commands and emphasis into plain descriptions. Reminders into constraints: "No TODOs, no partial implementations" works better than "remember to finish implementations." Vague quantities into ranges: "generate 20–100 tasks" gets far more ambitious behavior than "generate many tasks." - Delete: things capable models do by default, guards against behavior you haven't seen from this model, text that repeats tool descriptions, and lines that could contradict a user request. Models trained to rank system instructions above user messages will side with the system prompt. - Move: anything per-user or per-request (date, environment, repo state, lists of skills or subagents, user rules) into a user-role setup message after the cache boundary. Audit other injected context the same way. As models improved, the team behind these figures dropped directory trees, pre-retrieved snippets, compressed copies of attached files, lint errors injected after every edit, forced expansion of short file reads, and caps on tool calls per turn. They kept small, high-value facts: OS, repo status, and open or recently viewed files. Skip checklists for open-ended work. The model optimizes the listed items and deprioritizes everything else. ## 3. Tool definitions Tool schemas ride along on every request. Most tools beyond the core set were each needed in under 20% of conversations, and moving them out of static context cut tool-description tokens 60%. Doing the same for integration tools (such as MCP servers), with names in context and full schemas in one folder per server that the agent can search with grep or jq, cut total tokens 46.9% in sessions that used them. - Keep in static context: high-frequency tools (for a coding agent: read, search, edit, shell), tools the model tries to call even when they're absent, and tools a mode depends on. - Offload the rest: leave a name or one-line pointer and make the full schema discoverable on demand. Group related tools so they load together, and put status (such as "needs re-authentication") where the agent will see it. - Tighten what remains: describe behavior and arguments, and drop usage lectures. - Pick the split by testing a few configurations and tracking tokens, cost, latency, tool-call errors, and task success. ## 4. Cache layout Order each request so the reusable prefix is as long as possible: `tool definitions → system instructions → [breakpoint] → setup message (skills, subagents, rules, environment) → [breakpoint] → conversation` - Keep the prefix byte-identical across turns. Use deterministic tool order and serialization, put timestamps and IDs after the boundary, and don't rewrite earlier messages except when compacting. - Use explicit breakpoints if the provider supports them. Otherwise rely on automatic prefix caching with the stable part first. Respect TTL and minimum-length rules. - Switching models mid-conversation throws away the cache (caches are per model and provider) and hands the new model a history it didn't write. When a different model is needed, run it as a subagent with fresh context. Explicit breakpoints plus moving per-request setup after them cut cold cache misses 20%. ## 5. Tool results and other context added during a run - Large outputs (commands, integrations, logs): write them to a file and return the path, size, and a short tail. The agent can tail, grep, or read ranges for more. Truncating loses data, and inlining bloats every later request. Treat long-running terminal sessions the same way. - High-volume formats: look for overhead repeated on every line or item. Numbering every 10th line of a file read instead of every line cut cache-read tokens 1.6% without hurting citation accuracy. Each number costs 3–5 tokens, and agents read tens of thousands of lines per session. Also check repeated absolute paths, verbose JSON keys, ANSI codes, progress bars, and repeated headers. - Good retrieval saves exploration turns. Adding semantic search alongside grep raised codebase question-answering accuracy 12.5% on average and cut the iterations users needed. - Tool errors waste tokens and leave confusing debris in context. Classify expected errors (invalid arguments, unexpected environment, provider error, timeout, user abort), treat unknown errors as harness bugs, and track rates per tool and per model. One focused effort along these lines cut unexpected tool errors 10×. ## 6. Long runs: compaction, subagents, and model mix - Compaction: keep the summarization prompt short and the summary compact, carry forward plan state and remaining tasks, and save the full history to a file the agent can search for details the summary dropped. A model trained to self-summarize from a one-line prompt wrote ~1k-token summaries with half the compaction error of a multi-thousand-token prompt that produced 5k+ token summaries. Untrained models may need more guidance, so test how short you can go. A more expensive summarization model made a negligible difference. - Scratchpads and running notes: rewrite them instead of appending. For repeated work in one environment, a small agent-maintained notes file with a line budget, loaded at start, is a promising way to shorten later runs. - Subagents: fresh context keeps the parent lean, but isolation adds coordination cost (duplicate or stale work). If the model already delegates on its own, remove prompting that pushes it to. Have subagents return short handoffs: what was done, findings, concerns, and deviations. A subagent should use a different model only when the user or harness says so. - Model mix: in large multi-agent runs, workers used at least 69% of tokens, and over 90% in most runs. A frontier planner with cheap workers matched a frontier model doing everything at about one-eighth the cost. Planner choice still changes worker spend. One planner that cost less on its own saw its workers use several times more tokens, and the run cost more overall. Measure the whole tree. - Routing and reasoning effort: send simple turns to a cheaper model or lower effort, and upgrade only when a stronger model is clearly better. A router built this way matched or beat single frontier models on user satisfaction at 41–68% lower cost. - Reasoning continuity: if the API returns reasoning items (including encrypted ones), pass them back on later turns and alert when they go missing. Dropping them cost one reasoning model 30% on a coding benchmark, and it burned tokens reconstructing its plan. ## 7. Fit the harness to each model Adapt to what each model was trained on instead of forcing one shape on all of them. If you've tuned the harness for a similar model, start from that version. - Edit format: use the one the model was trained on (for example, patch-style or search-and-replace). An unfamiliar format costs extra reasoning tokens and causes more mistakes. - Shell or tools: shell-first models fall back to `cat` or inline scripts. Name tools after their shell equivalents (such as `rg`), and if needed add: "If a tool exists for an action, prefer to use the tool instead of shell commands (e.g. read_file over `cat`)." - Literalness: some model families follow instructions literally and others tolerate imprecision. Some spiral on emphasized wording. Strip caps and emphasis for literal models. - Triggers: some models ignore a tool until told when to use it. A literal trigger works: "After substantive edits, use the to check recently edited files for linter errors. If you've introduced any, fix them if you can easily figure out how." - Progress updates: if a model reports progress through reasoning summaries, keep them to 1–2 sentences that note new findings or a change of tactic, and remove instructions about messaging mid-turn. - Quirks worth a targeted line: hedging or refusing as context fills ("context anxiety"), declaring completion early, stopping to ask permission, and calling tools that don't exist. Tie each added instruction to the transcript behavior it fixes. Re-audit when models change, since guidance one version needed can be dead weight for the next. ## 8. Validate - Offline: run a fixed set of realistic tasks before and after, ideally drawn from real usage and phrased the way users actually write (short and ambiguous). Compare task success, tokens, cost per task, turns, and tool errors. Don't ship a change that lowers success. - Online, if you have users: A/B test each change or small bundle. The primary metric is cost per completed task. Guardrails are task success signals, tool-call errors, latency, turns per task, and cache hit rate. For a coding agent, a good success signal is how much agent-written code survives over time. In general, check whether the user's next message moves on or reports a problem. - Ship only when cost drops and no guardrail regresses beyond noise. Record null results. ## What to change directly and what to propose - Change directly, each in its own revertible commit: token and cache telemetry, deterministic serialization and tool order, moving volatile content out of the cached prefix, explicit cache breakpoints, writing large outputs to files instead of truncating, passing back reasoning items that are being dropped, and fixes for recurring tool errors. - Change behind a flag so it can be tested: system prompt edits, tool offloading, output format changes, compaction changes, and subagent prompting. - Propose only: changes to which models run, routing, reasoning-effort defaults, or how work is split across agents. ## Traps - Asking the model to use fewer tokens or do less. - Truncating tool output. - Dropping reasoning items to save input tokens. - Volatile content in the cached prefix, or tool order that changes between requests. - Offloading a tool the model needs on the first turn or tries to call when it's missing. - Emphasis-heavy prompts (MUST, NEVER, IMPORTANT, all caps), especially with literal models. - Forcing a terser output format than the model was trained on. Fewer output tokens can mean less thinking and worse results. - Optimizing raw token counts instead of cost, per request instead of per task, or evals instead of real usage. - Switching models mid-conversation to save money. - Adding coordination layers that become bottlenecks. ## Report back with 1. The harness map and baseline: cost by source × billing type, with the biggest sources called out. 2. A ranked list of changes: layer, what changes, estimated savings and how you estimated them, quality risk, how to validate, and how to roll back. 3. The changes you made, including a system prompt diff with a keep, rewrite, delete, or move reason for each line. 4. A test plan for the flagged changes. 5. Gaps: anything you couldn't find or measure.
显示更多
0
92
1.4K
65
转发到社区
unexpected bestie meetups---💀💀 part 1/2
0
501
26.5K
1.1K
转发到社区
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached. TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt. I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted. Is formal verification the future of coding (or at least, bug finding)?
显示更多
0
444
5.3K
292
转发到社区