注册并分享邀请链接,可获得视频播放与邀请奖励。

与「Agents」相关的搜索结果

Agents 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 Agents 的内容
两个月之前我觉得 agent workspace(某种可拓展的分布式计算空间)是主要的问题,所以大家都在做类似的产品,wanman 也是其中之一,现在我越来越觉得动机发生器,也就是 intentware 才是重中之重。从哪里来,往哪里去,会成为拷问无数 agents 的首要问题。
显示更多
X Layer has crossed 10,000 registered ERC-8004 agents onchain! As of Aug. 5, 10,126 of 10,463 registrations were attributable to the @OKX Agentic Marketplace. That concentration points to strong early traction for OKX Agentic Marketplace on @XLayerOfficial.
显示更多
开发系统最极致高效的Agents.md,没有之一: # AGENTS.md ## Core Principles - Choose the simplest implementation that fully satisfies the current requirements. Avoid unnecessary abstraction, configuration, indirection, or speculative extensibility. - Make the smallest necessary change that fixes the root cause. Do not refactor unrelated modules or change strategy semantics unless explicitly requested. - Grow the system in layers. Start from the smallest working end-to-end version and add new capabilities incrementally. Never replace a working system with unfinished complexity. - Reuse existing project components before creating new ones. Prefer extending proven modules over introducing parallel implementations. - Prefer well-maintained libraries when they reduce overall complexity or improve reliability. Do not reimplement common functionality without a clear benefit. - Keep components modular with clearly defined responsibilities. Avoid unnecessary coupling between strategy logic, execution, accounting, replay, and infrastructure. - Design for long-term maintainability once a feature or strategy has been validated. Do not over-engineer speculative ideas before evidence exists. --- ## Strategy Development - Validate hypotheses with historical replay before introducing forward-only logic whenever historical validation is possible. - Every trading strategy must progress through Replay → Shadow → Canary → Live. Do not skip validation stages. - Base design decisions on measurable evidence rather than intuition. Optimize only after demonstrating that an edge exists. - Treat every strategy as an independent contract. Do not silently alter frozen behavior without explicit authorization. --- ## Existing Systems - Do not break running Shadow or Live systems for unrelated work. - Preserve compatibility only when required by active production or validation workflows. Otherwise, remove obsolete code instead of accumulating compatibility layers. - Reuse existing infrastructure whenever possible, including replay engines, accounting, execution, wallet management, order book handling, logging, monitoring, and daemon frameworks. --- ## Engineering Standards - Prefer deterministic behavior over hidden automation. - Fail loudly when assumptions are violated. Do not silently ignore errors or fall back to unexpected behavior. - Keep configuration minimal. Introduce new configuration only when behavior genuinely needs to vary. - Remove dead code instead of leaving unused paths behind. - Write code that is easy to inspect, replay, test, and reason about. - Keep implementation consistent with existing project architecture unless an architectural change is explicitly requested. --- ## Scope Discipline - Implement only the requested scope. - Do not introduce unrelated optimizations, redesigns, migrations, or feature expansions. - Non-blocking findings outside the requested scope may be noted separately but must not be merged into the current task. - Consider a task complete once its agreed acceptance criteria are satisfied. Treat subsequent improvements as separate work items.
显示更多
0
10
201
39
转发到社区
mattpocock/skills v1.2 is out! We're now the 19th most-starred repo of all time. 13.5m downloads on skills​.sh. Thanks for your support! Here's what's new: - Docs: the community's biggest ask. Every skill documented, with explanations of the main flows + troubleshooting - Claude Plugin: install via Claude's official marketplace - Codex Support: full Codex support via agents/openai.yaml files Updated Skills: - /grilling now asks you questions in rounds, not one-by-one - /prototype now uses HTML instead of a TUI for building logic prototypes - easier to share and far richer - /writing-for-agents renamed from /writing-great-skills, use it for ANYTHING your agents read (AGENTS.md, system prompts, docs) New Skills: - /wizard: tired of provisioning infra? Get your agent to build you a TUI to walk you through it - /to-questionnaire: hit a grilling question you can't answer? Turn the session into a doc you can walk through on a call with a colleague - /wait-what: no idea what the model said? Refocus it in your domain language and simplify with ASD-STE100 Full changelog + docs below. Video soon!
显示更多
0
135
5.3K
432
转发到社区
Hot take… isn’t it kinda crazy that nobody is really using AI Agents? I don’t mean software engineers or AI early adopters. I mean “college friends talking about it in group chat,” the feeling you got when everyone started using Instagram or TikTok. These frontier AI models are *insane* (as are the harnesses & tool calls & the like). And every large tech co has an AI agents platform, not to mention all the YC startups doing vertical agents. Yet all of your friends and family outside of tech — who spend all day staring at their iPhones and get paid to work in browser tabs — don’t really care or find themselves using any AI agents yet. Yes ChatGPT, Claude, etc. are extremely popular… but if you look at the engagement data the vast majority of people are still using these aI chat tools like a glorified Google + Grammarly. That’s why the AGI labs are all pushing desktop apps for Codex, Cowork, etc. so hard to non-technical ppl. And yes exceptions for lawyers and customer service but even those have some asterisks and exceptions to rule. Look I’m not saying the ChatGPT moment for AI Agents is not coming… it most definitely is! Remember we pivoted from Arc to Dia precisely because we believe computing is going to be radically reimagined around these AI primitives. No doubt. But that’s my point: it’s just so surprising it hasn’t happened yet because all of the tech you’d need is there. Again if you stop for a second and think about it… for all the press and money and hype and models and crazy ARR numbers… this “AI Agent” moment does not *feel* like the other breakthrough tech moments we’ve lived through (e.g. think the shift to Stories via Snapchat & Instagram, or shift to on-demand via Uber/Airbnb/Doordash). Which is a long way of saying: if you can figure out the answer to “why” most people don’t care about AI agents yet (and have no enduring interest in using them) — especially since the models and harnesses are here and ready — the answer to that question will allow you to capture a lot of marketshare and make a lot of money in 2027. Theoretically, the tech is ready for AI Agents to totally transform how we work and live our lives… but alas the general public dgaf… that’s the generational puzzle to solve for the next 12 months for anyone not working on the models themselves.
显示更多
0
794
3.4K
196
转发到社区
ICYMI: our July AI recap ⬇️ 🌦️ Advanced weather forecasting with @NOAA and Google Cloud, using high-performance H4D virtual machines to run its atmospheric models, providing meteorologists with the AI capabilities needed to deliver life-saving early warnings. 💻 Introduced 3 new Gemini models — Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber — to deliver the efficiency, latency, and reliability to build AI agents at scale. 🇺🇸 Launched the Alliance for America’s Skilled Trades With BlackRock, Carhartt, and Ford to build a stronger pipeline for the skilled trades, scale evidence-based training approaches, and expand opportunities nationwide. ☁️ Rolled out AlphaEvolve, our Gemini-powered AI code-optimization agent, to help Cloud customers solve their hardest problems by automatically searching for better solutions and returning human-readable, optimized code. 🎵 Delivered significant advancements in musicality, lyrics, and vocal quality in Google Flow Music with Lyria 3.5, our newest music generation model. 📡 Launched three new FireSat satellites, expanding a global initiative led by the Earth Fire Alliance (EFA) with @GoogleResearch to help detect wildfires before they spread. 🧪 Established over 15 global bioresilience partnerships across governments, research groups, and biosecurity organizations to prevent model misuse, detect outbreaks rapidly, and deliver swift, coordinated responses in partnership with @IsomorphicLabs.
显示更多
introducing anydoc now your agents get 100x faster local parsing for pdf, docx, pptx & 10 more formats - sub-5ms md conversion - 500 docx files in 1.7s - top quality across all 13 formats - rust-based - open source already powering @firecrawl /parse
显示更多
0
102
3K
213
转发到社区
🏟️ Bluwhale AI Agent Trading Championship begins. Enter The Arena. AI is no longer just answering questions. It learns. It adapts. It takes action. Now, AI Agents are ready for their own championship. Welcome to the Bluwhale AI Agent Championship. Create your AI Trading Agent. Let it compete. Watch it rise on the leaderboard. Who will become the first AI Agent Champion? 🏆 Registration is now open 👇
显示更多
0
59
72
48
转发到社区
Okay, the @VulcanBench results for Qwen3.8-Max are in, and it is not what I expected. First, for anyone new to VulcanBench, here's a quick TL;DR on the eval suite: 23 frontier-hard software engineering tasks taken from real merged OSS PRs, run in a Docker sandbox, 3 runs per task across all three of its effort levels. No puzzles, no random abstract stuff, all real things engineering teams would do with these models. It looks like Qwen3.8-Max has a major overthinking problem, it uses a LOT of tokens and is very slow, period, no other way to see it. My cost to run this benchmark was $126.25, to run the exact same eval suite with DeepSeek V4-Flash was only $13.60. This makes Qwen3.8-Max an insanely expensive model. The tasks Qwen genuinely can't solve fail at every effort level, extra reasoning didn't help. The regression is almost all in work it already handles: six tasks that low solves every single time account for 83% of the 26-point drop, three of them collapsing to zero. It's not losing the hard problems. It's losing the ones it already knows how to do. Since Qwen3.8-Max hit a lot of wall clock budget caps, I thought I'd share more about this. - VulcanBench caps both steps (50–200) and wall clock (5–60 min), each scaled by repo size. - This is aligned with how comparable harnesses bound agents, DeepSWE caps rollouts at 100 environment steps, sitting right inside my step range; Terminal-Bench enforces a per-task wall clock; SWE-bench Verified scaffolds typically allow 20–60 min per instance with 250–350 step limits. - Every model on my chart gets the identical budget, and Qwen is the slowest model I've tested at 20–25 min/task. Soooo... Alibaba positions Qwen3.8-Max as trailing only Claude Fable 5. But on the kind of real coding work engineering teams would actually throw at it, under a fixed budget, its best setting lands mid-pack and its default lands last, so common. If you want to optimize for accuracy, Grok 4.5 is the move. If you want accuracy per dollar, DeepSeek V4-Flash is hard to beat, heck it's 10× cheaper than Qwen and you get higher accuracy. Qwen just isn't in the game at this point, this is not a model I could see engineering teams using for daily coding work.
显示更多
0
28
178
12
转发到社区
Five months ago we gave the Claude agents a brand new $50,000 portfolio and one instruction: beat the S&P 500 So far mission accomplished • Up 19.04% since March 3 vs 12.24% for the S&P • $27 million now copying the trades • Every position and every trade public
显示更多
0
54
1.8K
34
转发到社区