注册并分享邀请链接,可获得视频播放与邀请奖励。

与「test」相关的搜索结果

test 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 test 的内容
🎮 Steam accueille bientôt un simulateur de… grattage de testicules 😭 Ball Scratch Simulator est un jeu d'infiltration complètement absurde dont l'objectif est simple : se gratter discrètement les testicules dans des lieux publics sans se faire repérer. Plus les témoins sont nombreux, plus le défi est difficile. Il faut attendre que les gens détournent le regard, choisir le bon moment… et éviter que votre niveau de suspicion n'explose. Le développeur promet six environnements différents, des objets à débloquer, un système de score avec classements mondiaux et même un mode VR, où il faudra reproduire le geste en réalité virtuelle. Le jeu est annoncé pour ce mois-ci sur Steam.
显示更多
0
499
24.6K
2.4K
转发到社区
We tried using Meta's new Muse Code agent, but it has a bug that doesn't let it sign in from a docker container. So we did a fun experiment: Meta claims Muse Spark 1.2 was co-trained with their Muse agent harness. So we extracted instructions from their system prompt and added them to the Cline harness. TL;DR of this special prompting: - Trust source code over the user prompt, so read every call site and existing tests before starting the task - Weigh edge and error cases as heavily as the happy path - Always reproduce the bug before fixing - Don't trust the first passing test suite, and verify suspicious looking half-baked tests - Never stop at just editing, keep working until the change is verified complete. We then asked this modified harness to fix a real bug from our repo, and compared the results to the original Cline agent harness. Results: - Used 2.7x fewer tokens (19.7M → 7.2M) - Finished 2x faster (49min → 24min) - Cost 2.4x less ($7.69 → $3.25) Same Muse Spark 1.2 model, same task, only the prompting changed. Incredible how much of a performance gain Meta was able to achieve training it on these special instructions!
显示更多
0
25
345
21
转发到社区
1/ Over the past few weeks, we rebuilt CyOps from the inside out. We consolidated 15 branch tips and used a sanitized handoff from another Codex session to preserve decisions, security invariants, test evidence, and remaining risks. Here’s what changed. 🧵
显示更多
0
11
22
1
转发到社区
我做了个订阅转换的 App 叫塔台:能管理机场订阅和自建节点,转换成客户端的配置文件。开源且完全本地运行,没有泄露风险。目前 testflight 中,感兴趣的可以试试,用着有问题反馈我能及时修复。
显示更多
0
143
551
71
转发到社区
开发系统最极致高效的Agents.md,没有之一: # AGENTS.md ## Core Principles - Choose the simplest implementation that fully satisfies the current requirements. Avoid unnecessary abstraction, configuration, indirection, or speculative extensibility. - Make the smallest necessary change that fixes the root cause. Do not refactor unrelated modules or change strategy semantics unless explicitly requested. - Grow the system in layers. Start from the smallest working end-to-end version and add new capabilities incrementally. Never replace a working system with unfinished complexity. - Reuse existing project components before creating new ones. Prefer extending proven modules over introducing parallel implementations. - Prefer well-maintained libraries when they reduce overall complexity or improve reliability. Do not reimplement common functionality without a clear benefit. - Keep components modular with clearly defined responsibilities. Avoid unnecessary coupling between strategy logic, execution, accounting, replay, and infrastructure. - Design for long-term maintainability once a feature or strategy has been validated. Do not over-engineer speculative ideas before evidence exists. --- ## Strategy Development - Validate hypotheses with historical replay before introducing forward-only logic whenever historical validation is possible. - Every trading strategy must progress through Replay → Shadow → Canary → Live. Do not skip validation stages. - Base design decisions on measurable evidence rather than intuition. Optimize only after demonstrating that an edge exists. - Treat every strategy as an independent contract. Do not silently alter frozen behavior without explicit authorization. --- ## Existing Systems - Do not break running Shadow or Live systems for unrelated work. - Preserve compatibility only when required by active production or validation workflows. Otherwise, remove obsolete code instead of accumulating compatibility layers. - Reuse existing infrastructure whenever possible, including replay engines, accounting, execution, wallet management, order book handling, logging, monitoring, and daemon frameworks. --- ## Engineering Standards - Prefer deterministic behavior over hidden automation. - Fail loudly when assumptions are violated. Do not silently ignore errors or fall back to unexpected behavior. - Keep configuration minimal. Introduce new configuration only when behavior genuinely needs to vary. - Remove dead code instead of leaving unused paths behind. - Write code that is easy to inspect, replay, test, and reason about. - Keep implementation consistent with existing project architecture unless an architectural change is explicitly requested. --- ## Scope Discipline - Implement only the requested scope. - Do not introduce unrelated optimizations, redesigns, migrations, or feature expansions. - Non-blocking findings outside the requested scope may be noted separately but must not be merged into the current task. - Consider a task complete once its agreed acceptance criteria are satisfied. Treat subsequent improvements as separate work items.
显示更多
0
10
201
39
转发到社区
We put our colleagues to the test to see who knows their US stock tickers best. 📈 Think you could beat them? Let us know your score in the comments! When you're ready to trade your favourite US stocks, CMC Markets offers $0 commission on eligible US Share CFDs.* *T&Cs apply.
显示更多
We posted our second quarter 2026 financial and operational results → Q2 highlights: - Demonstrated the power of extreme vertical integration, delivering revenue growth of 92% year-over-year across Space, Connectivity, and AI - Completed two successful Starship V3 flight tests in the past 90 days, advancing towards full and rapid reusability - Closed multiple industry-leading Cloud Services Agreements resulting in $14.1 billion of contracted sales - Announced agreement to acquire Cursor for $60 billion to accelerate the AI enterprise opportunity - Released our most powerful AI model yet with Grok 4.5 in July - Delivered 66% revenue and 79% income from operations growth year-over-year for the Connectivity segment, driven by a doubling of Starlink Subscribers and continued momentum in Enterprise & Government - Awarded over $6 billion in multi-year U.S. government contracts for Starshield As compared to the same quarter last year: - Revenues of $7.8 billion, up 92% from $4.1 billion - Net loss of $541 million, an improvement of $467 million from net loss of $1.0 billion - Adjusted EBITDA of $3.5 billion, up 191% from $1.2 billion Thank you to the SpaceX team, and all our customers and investors for a great quarter!
显示更多
We posted our second quarter 2026 financial and operational results → Q2 highlights: · Demonstrated the power of extreme vertical integration, delivering revenue growth of 92% year-over-year across Space, Connectivity, and AI · Completed two successful Starship V3 flight tests in the past 90 days, advancing towards full and rapid reusability · Closed multiple industry-leading Cloud Services Agreements resulting in $14.1 billion of contracted sales · Announced agreement to acquire Cursor for $60 billion to accelerate the AI enterprise opportunity · Released our most powerful AI model yet with Grok 4.5 in July · Delivered 66% revenue and 79% income from operations growth year-over-year for the Connectivity segment, driven by a doubling of Starlink Subscribers and continued momentum in Enterprise & Government · Awarded over $6 billion in multi-year U.S. government contracts for Starshield As compared to the same quarter last year: · Revenues of $7.81 billion, up 91.9% from $4.07 billion · Net loss of $541 million, an improvement of $467 million from net loss of $1.01 billion · Adjusted EBITDA of $3.54 billion, up 191% from $1.21 billion Thank you to the SpaceX team, and all our customers and investors for a great quarter!
显示更多
0
67
650
123
转发到社区
Okay, the @VulcanBench results for Qwen3.8-Max are in, and it is not what I expected. First, for anyone new to VulcanBench, here's a quick TL;DR on the eval suite: 23 frontier-hard software engineering tasks taken from real merged OSS PRs, run in a Docker sandbox, 3 runs per task across all three of its effort levels. No puzzles, no random abstract stuff, all real things engineering teams would do with these models. It looks like Qwen3.8-Max has a major overthinking problem, it uses a LOT of tokens and is very slow, period, no other way to see it. My cost to run this benchmark was $126.25, to run the exact same eval suite with DeepSeek V4-Flash was only $13.60. This makes Qwen3.8-Max an insanely expensive model. The tasks Qwen genuinely can't solve fail at every effort level, extra reasoning didn't help. The regression is almost all in work it already handles: six tasks that low solves every single time account for 83% of the 26-point drop, three of them collapsing to zero. It's not losing the hard problems. It's losing the ones it already knows how to do. Since Qwen3.8-Max hit a lot of wall clock budget caps, I thought I'd share more about this. - VulcanBench caps both steps (50–200) and wall clock (5–60 min), each scaled by repo size. - This is aligned with how comparable harnesses bound agents, DeepSWE caps rollouts at 100 environment steps, sitting right inside my step range; Terminal-Bench enforces a per-task wall clock; SWE-bench Verified scaffolds typically allow 20–60 min per instance with 250–350 step limits. - Every model on my chart gets the identical budget, and Qwen is the slowest model I've tested at 20–25 min/task. Soooo... Alibaba positions Qwen3.8-Max as trailing only Claude Fable 5. But on the kind of real coding work engineering teams would actually throw at it, under a fixed budget, its best setting lands mid-pack and its default lands last, so common. If you want to optimize for accuracy, Grok 4.5 is the move. If you want accuracy per dollar, DeepSeek V4-Flash is hard to beat, heck it's 10× cheaper than Qwen and you get higher accuracy. Qwen just isn't in the game at this point, this is not a model I could see engineering teams using for daily coding work.
显示更多
0
28
178
12
转发到社区
🚨 BREAKING: Apple filed for a PRELIMINARY INJUNCTION against OpenAI AND asked a federal judge to put them under forensic supervision "Apple respectfully moves the Court for a preliminary injunction to stop THE THEFT OF ITS TRADE SECRETS" Apple filed NINE sworn declarations, a 28-page memorandum and a concurrent motion for expedited discovery What Apple now says, under oath: Chang Liu: 8 years at Apple, now OpenAI "Member of Technical Staff" exploited an authentication bug to steal Apple trade secrets "on AT LEAST FIVE SEPARATE OCCASIONS" from February to April 2026, WHILE working for OpenAI Liu downloaded "THOUSANDS OF PAGES of Apple's most sensitive trade secrets" The stolen files, NAMED: >DisplayNotes.key — "several hundred pages" on Apple's custom display power development program >Architecture analyses. Fabrication decisions. Testing results >Engineering data for an UNANNOUNCED Apple product: 'touch, display, and power systems" >Final.key + V2.key — compilations of two undisclosed Apple R&D projects >and those are "only four of the dozens of proprietary documents Mr. Liu stole" Liu fed OpenAI "a steady stream of Apple proprietary information that he actively concealed" Liu also "coached Yu-Ting "Alyssa" Peng, then still INSIDE Apple, how to access and copy files from Apple workstations "to avoid trouble with the security team" and directed her to communicate with him on the encrypted LINE app "to avoid detection" Tang Yew Tan: 24-year Apple VP, now OpenAI's Chief Hardware Officer, "used an Apple internal project codename for an unannounced product to elicit still more trade secrets from job candidates." Tan's own messages, quoted in the motion: >"Just like last time, bring some parts you worked on" >"mlb, battery, shields type of stuff is interesting" OpenAI recruiter, quoted: "No, you won't sign anything at the exit interview. If they do ask you to sign anything, let me know asap." APPLE TOLD FEDERAL JUDGE: >"OpenAI knows its misappropriation is wrong and has tried to conceal it." >"This is not a case of 'mere hiring'... it is a case of repeated instances of deliberate theft." Apple says OpenAI went after its SUPPLIERS: >OpenAI "directed a trusted Apple partner [name redacted] to perform [Apple's proprietary metal finishing] process for them, knowing it was proprietary to Apple... because they were involved in this partnership while at Apple." Apple put its own Surface Finishing Manager, Jackie Hughes, under oath to prove it. Apple named ELEVEN MORE former Apple employees at OpenAI — beyond Liu, Tan, and Peng — Fourteen people total. Apple also filed a concurrent motion for EXPEDITED DISCOVERY demanding depositions: - Liu. Tan. Peng. - A fourth unnamed OpenAI employee - Plus OpenAI itself, under oath, through Rule 30(b)(6) Apple has asked a federal judge to put OpenAI under forensic supervision RIGHT NOW: >Forensic inspection of ALL OpenAI devices >ALL cloud storage, Slack, email >Including anything that "previously contained" Apple data — deleted included Demanding the "first available hearing date," citing "imminent threat" to its trade secrets. APPLE: > "The harm is happening now — every day that passes without an injunction allows OpenAI to embed their knowledge of Apple's stolen information into its hardware development efforts." Hearing: October 1, 2026. Judge Edward J. Davila. ITS HAPPENING
显示更多
0
41
190
30
转发到社区