注册并分享邀请链接,可获得视频播放与邀请奖励。

与「EXTRA」相关的搜索结果

EXTRA 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 EXTRA 的内容
We tried using Meta's new Muse Code agent, but it has a bug that doesn't let it sign in from a docker container. So we did a fun experiment: Meta claims Muse Spark 1.2 was co-trained with their Muse agent harness. So we extracted instructions from their system prompt and added them to the Cline harness. TL;DR of this special prompting: - Trust source code over the user prompt, so read every call site and existing tests before starting the task - Weigh edge and error cases as heavily as the happy path - Always reproduce the bug before fixing - Don't trust the first passing test suite, and verify suspicious looking half-baked tests - Never stop at just editing, keep working until the change is verified complete. We then asked this modified harness to fix a real bug from our repo, and compared the results to the original Cline agent harness. Results: - Used 2.7x fewer tokens (19.7M → 7.2M) - Finished 2x faster (49min → 24min) - Cost 2.4x less ($7.69 → $3.25) Same Muse Spark 1.2 model, same task, only the prompting changed. Incredible how much of a performance gain Meta was able to achieve training it on these special instructions!
显示更多
0
25
345
21
转发到社区
This blackhead never stood a chance. Clean extraction. Clean satisfaction.
0
82
4.9K
54
转发到社区
NVIDIA is reportedly working on a feature that could let future GeForce GPUs use an SSD as extra VRAM. When a graphics card runs out of VRAM, it could use data stored on a fast NVMe SSD instead. Your SSD won’t be as fast as real VRAM, but it could help when a game or AI app needs more memory than your graphics card has. That could mean fewer stutters and better performance in some situations. The feature is reportedly based on RTX IO and Microsoft’s DirectStorage technology.
显示更多
0
305
4.9K
333
转发到社区
Clean extractions like this hit different. The relief must have been unreal.
0
25
1.9K
23
转发到社区
More rewards for VIPs! Binance Earn just boosted BTC Yield APY by up to +1% extra for VIP users. Know more →
0
11
25
3
转发到社区
Okay, the @VulcanBench results for Qwen3.8-Max are in, and it is not what I expected. First, for anyone new to VulcanBench, here's a quick TL;DR on the eval suite: 23 frontier-hard software engineering tasks taken from real merged OSS PRs, run in a Docker sandbox, 3 runs per task across all three of its effort levels. No puzzles, no random abstract stuff, all real things engineering teams would do with these models. It looks like Qwen3.8-Max has a major overthinking problem, it uses a LOT of tokens and is very slow, period, no other way to see it. My cost to run this benchmark was $126.25, to run the exact same eval suite with DeepSeek V4-Flash was only $13.60. This makes Qwen3.8-Max an insanely expensive model. The tasks Qwen genuinely can't solve fail at every effort level, extra reasoning didn't help. The regression is almost all in work it already handles: six tasks that low solves every single time account for 83% of the 26-point drop, three of them collapsing to zero. It's not losing the hard problems. It's losing the ones it already knows how to do. Since Qwen3.8-Max hit a lot of wall clock budget caps, I thought I'd share more about this. - VulcanBench caps both steps (50–200) and wall clock (5–60 min), each scaled by repo size. - This is aligned with how comparable harnesses bound agents, DeepSWE caps rollouts at 100 environment steps, sitting right inside my step range; Terminal-Bench enforces a per-task wall clock; SWE-bench Verified scaffolds typically allow 20–60 min per instance with 250–350 step limits. - Every model on my chart gets the identical budget, and Qwen is the slowest model I've tested at 20–25 min/task. Soooo... Alibaba positions Qwen3.8-Max as trailing only Claude Fable 5. But on the kind of real coding work engineering teams would actually throw at it, under a fixed budget, its best setting lands mid-pack and its default lands last, so common. If you want to optimize for accuracy, Grok 4.5 is the move. If you want accuracy per dollar, DeepSeek V4-Flash is hard to beat, heck it's 10× cheaper than Qwen and you get higher accuracy. Qwen just isn't in the game at this point, this is not a model I could see engineering teams using for daily coding work.
显示更多
0
28
178
12
转发到社区
Los autobuses escolares de Japón están causando sensación en el extranjero
0
337
28.7K
1.8K
转发到社区
The Blue Zones—places where people supposedly live past 100 at extraordinary rates—aren't real. A 🧵 on what went wrong: bad record keeping, pension fraud, and lots of money.
0
27
586
68
转发到社区
1/ Open Beats — Pudgy Beats is evolving. More Themes. More Freedom. More Rewards. We’ve extended the deadline to September 15 and increased the rewards 💰 Original prize pool $6,000 USD stays fully intact — plus an extra $1,000 USD added to the @miurolabs Miuro Genesis NFT campaign 🎉 Themes are now fully open: Music Diary, Love Stories, Cute Brainrot, Growth & Dreams, AI New Era… and more ✨ Music Videos have No Style Restrictions — realistic or AI, both welcome 🎬 Pudgy Penguins IP is still encouraged and will be jointly promoted 🐧 More time. More freedom. More rewards. Start creating now → #OpenBeats# #MIURO# #AIMusic#
显示更多
0
13
35
18
转发到社区
ICYMI: Get your first look at PvE combat in the upcoming Lovecraftian extraction shooter, Rules of Engagement: The Grey State.
0
23
660
42
转发到社区