注册并分享邀请链接,可获得视频播放与邀请奖励。

与「Eva」相关的搜索结果

Eva 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 Eva 的内容
Okay, the @VulcanBench results for Qwen3.8-Max are in, and it is not what I expected. First, for anyone new to VulcanBench, here's a quick TL;DR on the eval suite: 23 frontier-hard software engineering tasks taken from real merged OSS PRs, run in a Docker sandbox, 3 runs per task across all three of its effort levels. No puzzles, no random abstract stuff, all real things engineering teams would do with these models. It looks like Qwen3.8-Max has a major overthinking problem, it uses a LOT of tokens and is very slow, period, no other way to see it. My cost to run this benchmark was $126.25, to run the exact same eval suite with DeepSeek V4-Flash was only $13.60. This makes Qwen3.8-Max an insanely expensive model. The tasks Qwen genuinely can't solve fail at every effort level, extra reasoning didn't help. The regression is almost all in work it already handles: six tasks that low solves every single time account for 83% of the 26-point drop, three of them collapsing to zero. It's not losing the hard problems. It's losing the ones it already knows how to do. Since Qwen3.8-Max hit a lot of wall clock budget caps, I thought I'd share more about this. - VulcanBench caps both steps (50–200) and wall clock (5–60 min), each scaled by repo size. - This is aligned with how comparable harnesses bound agents, DeepSWE caps rollouts at 100 environment steps, sitting right inside my step range; Terminal-Bench enforces a per-task wall clock; SWE-bench Verified scaffolds typically allow 20–60 min per instance with 250–350 step limits. - Every model on my chart gets the identical budget, and Qwen is the slowest model I've tested at 20–25 min/task. Soooo... Alibaba positions Qwen3.8-Max as trailing only Claude Fable 5. But on the kind of real coding work engineering teams would actually throw at it, under a fixed budget, its best setting lands mid-pack and its default lands last, so common. If you want to optimize for accuracy, Grok 4.5 is the move. If you want accuracy per dollar, DeepSeek V4-Flash is hard to beat, heck it's 10× cheaper than Qwen and you get higher accuracy. Qwen just isn't in the game at this point, this is not a model I could see engineering teams using for daily coding work.
显示更多
0
28
178
12
转发到社区
U.S. Suspect Clings to Moving Car to Evade Police, Falls Off After Nearly 400 Meters
💥 TOP 100 DUNKS OF 2025-26 💥 100. Nikola Jokić 99. Jalen Green 98. Chet Holmgren 97. Jeff Green 96. Jalen Williams 95. Peyton Watson 94. Giannis Antetokounmpo 93. Jaxson Hayes 92. Evan Mobley 91. Precious Achiuwa Our Dunk Week countdown begins!
显示更多
0
42
91
14
转发到社区
Introducing Ori Eval: the easiest way to write your first eval. There's no definitive best model, only the best model for each task. Ori Eval leverages OpenRouter's APIs for each task in your codebase, and then evaluates the results. curl -fsSL
显示更多
0
24
557
35
转发到社区
Players who just missed the NFL Top 100... 110: Tyler Warren 109: Tee Higgins 108: Brian Branch 107: Devon Witherspoon 106: Jameson Williams 105: Dexter Lawrence 104: Chris Lindstrom 103: CJ Stroud 102: Lane Johnson 101: Mike Evans
显示更多
0
10
21
0
转发到社区
.@Buccaneers wide receiver Emeka Egbuka reflects on life after Mike Evans out in Tampa Bay 🏴‍☠️ @Sara_Walsh | @Geraldini93
The fast-moving Old Trails Fire in Spokane, WA, has surpassed 3,500 acres and swept directly into northwest Spokane residential neighborhoods, igniting multiple homes and threatening thousands of structures. Mandatory Level 3 "Go Now" evacuation orders remain in effect as heavy traffic gridlock delays fleeing residents across major evacuation corridors.
显示更多
0
23
461
168
转发到社区
Brock Purdy is loving throwing to Mike Evans 🤩 @49ers | @brockpurdy13 | @MikeEvans13_ | @OmarDRuiz
Here is my AI investing guide. Sitting here August 2026, my current best thoughts are as follows: 1. LPS (Land Power Shell) is still the most obvious and fastest path to cash on cash returns. Lots of value can be assembled and traded quickly at this layer. And as data centers get more pushback, energized land can explode in value. Very bullish here. I’ve stepped into this layer very aggressively. My partner @anitavlallian and I have acquired almost 6GW coming online in a ramp from today thru 2029 of grid power and behind the meter. 2. Silicon - I helped get @GroqInc off the ground in 2015 and we licensed it to @nvidia for $20B Dec2025. I won’t invest or incubate anything in this layer now. The perf demands of the chips are too high, manufacturing precision is too complex and supply chain influence to get adjacent components like memory isn’t possible for a startup anymore. Lots of capital will be wasted here chasing Groq and Cerebras’ success. Note that both startups made sense a decade ago when these constraints were much more modest. 3. Clouds - Clouds are very very lucrative but very hard to build and very expensive and technically complicated to maintain. And as alignment becomes a more important issue, I expect the clouds will be asked to build robust KYC and attest to it. This makes the risk:reward ratio skewed. I don’t want to be responsible when the USG says a cloud allowed a bad actor to do something bad because of poor KYC. 4. Models are complicated. The big open question is how much of the revenue being generated by them today is because of tokenmaxxing and poor model behavior. If it’s a lot, then the annualized revenues will diminish meaningfully even as token consumption inflects upwards. This is the big economic question at this layer. 5. Harnesses are where the action is and why I started @8090solutions two years ago. In a nutshell, the harness helps enterprises owns their proprietary context (what Alex Karp calls their ‘alpha’). This is an enterprise’s data, workflows, evals, and business rules. A harness that gives this to an enterprise is what creates very low model-agnostic switching costs, which further reinforces my views of #4# above. 6. Applications will be another long term winner along with harnesses. This is where the differentiation between “off the shelf” and “custom time and materials” melts away. Every company, with the right harness, can now imbue their alpha into the software that runs their company. I expect this to mean that “off the shelf” is largely replaced with custom software creating a huge opportunity to write these solutions for companies. Build once and sell repeatedly is a laggard GTM motion for a SaaS world that isn’t needed here. Think custom by design, alpha embedded, proprietary by nature. Fin. Good luck to all the players!
显示更多
0
351
6K
625
转发到社区
A high-stakes rescue moment brought to life with BACH 1.0 from spotting the danger to calling for evacuation before the avalanche hits. With Lip Sync and expressive character performance, every warning, every reaction, and every moment of urgency feels more real. Try BACH Free.
显示更多