Okay, the
@VulcanBench results for Qwen3.8-Max are in, and it is not what I expected.
First, for anyone new to VulcanBench, here's a quick TL;DR on the eval suite:
23 frontier-hard software engineering tasks taken from real merged OSS PRs, run in a Docker sandbox, 3 runs per task across all three of its effort levels.
No puzzles, no random abstract stuff, all real things engineering teams would do with these models.
It looks like Qwen3.8-Max has a major overthinking problem, it uses a LOT of tokens and is very slow, period, no other way to see it.
My cost to run this benchmark was $126.25, to run the exact same eval suite with DeepSeek V4-Flash was only $13.60. This makes Qwen3.8-Max an insanely expensive model.
The tasks Qwen genuinely can't solve fail at every effort level, extra reasoning didn't help. The regression is almost all in work it already handles: six tasks that low solves every single time account for 83% of the 26-point drop, three of them collapsing to zero. It's not losing the hard problems. It's losing the ones it already knows how to do.
Since Qwen3.8-Max hit a lot of wall clock budget caps, I thought I'd share more about this.
- VulcanBench caps both steps (50–200) and wall clock (5–60 min), each scaled by repo size.
- This is aligned with how comparable harnesses bound agents, DeepSWE caps rollouts at 100 environment steps, sitting right inside my step range; Terminal-Bench enforces a per-task wall clock; SWE-bench Verified scaffolds typically allow 20–60 min per instance with 250–350 step limits.
- Every model on my chart gets the identical budget, and Qwen is the slowest model I've tested at 20–25 min/task.
Soooo... Alibaba positions Qwen3.8-Max as trailing only Claude Fable 5. But on the kind of real coding work engineering teams would actually throw at it, under a fixed budget, its best setting lands mid-pack and its default lands last, so common.
If you want to optimize for accuracy, Grok 4.5 is the move. If you want accuracy per dollar, DeepSeek V4-Flash is hard to beat, heck it's 10× cheaper than Qwen and you get higher accuracy.
Qwen just isn't in the game at this point, this is not a model I could see engineering teams using for daily coding work.
显示更多
this is my personal singularity moment
this post may sound like a paid ad. I only wish. I'm concerned, more so than happy. the world is changing, and, among the scenarios where AI goes terribly wrong, inequality is the most realistic, yet, the one Anthropic seems to be the least concerned about. I'm glad OpenAI is taking the opposite stance: *personal AGI for everyone*. I think this is a commendable position in the times we live. but who am I in the queue of the bread?
anyway, Fable is here, so I'll just report my first-hour experience
first of all, all my pet prompts are solved.
→ λ-calculus puzzles
→ bug questions
→ one-shot apps
all are trivial to it.
I don't have anything harder other than my
ongoing work
so, in the last several days, I've been toying with HVM5, a new interaction net evaluator with a faster loop.
after writing the first version, I left 32 GPT-5 agents working for ~20 hours each. this resulted in up to 2x speedups, but the file size increased by 2-fold and quality decreased significantly.
I then simplified the whole thing into an even simpler core, and left Opus 4.8 and GPT 5.5 optimizing it for 8 hours. Opus got a legit 6% - 34% speedup in most benches. GPT got better results, but, sadly, an unusable file.
I then asked Fable to optimize it.
2 hours later, it landed a 1770% speedup in one case, 100%+ in other 4, and 22% in average. yes, in 2 hours it outperformed me, opus 4.8 and a swarm of gpt 5.5 agents, by one order of magnitude.
that could not possibly be legit. "it must be hardcoding the benchmarks" (GPT trauma). so I read its explanation and what it did was, indeed, the most high impact optimization one could try first. seems like HVM5 was wasting a lot of time garbage-collecting unused branches of pattern-match nodes. I had optimized that for static mats, but not for dynamic mats. skill issue. Fable figured how to do it for these, resulting in a massive speedup in some benches
but wait, is that *correct*? I'm not sure yet, it is credible, but this is the kind of thing that is very easy to get wrong on interaction nets. the problem is, when I was ready to start auditing Fable's solution so I could tell whether it was buggy or legit, it interrupted me to tell me it had found a massive bug on the code *I* had written.
... wait, what?
so... for garbage collection purposes, I stored a bit on lambda term pointers that meant "the variable bound by this lambda has been freed, so, its lambda must free whatever argument it is applied to". that's fine. yet, on duplicator nodes, I also used the same bit to mean "one of the duplicated variables was freed, so, treat this dup as a passthrough no-op". so, if a lambda entered a duplicator, it would mistake the lambda's collection bit for its own, resulting in corrupted interaction!
that's a mouthful, why I'm writing this?
just so you can appreciate the sheer absurdity of what just happened. I didn't ask it to find bugs. I asked it for an optimization. and even if I did ask it to find bugs, this bug is so astonishingly subtle and specific, identifying it takes mastering the domain to an extent that it beyond even me. I'd easily need hours or days to fix it, *if* I ever came across it. chances are it would just go unnoticed. and Fable found it and fixed it like it was nothing, while it was busy adding a 17x speedup to a file that neither I, nor Opus 4.8, nor a fleet of GPT 5.5 managed to barely make 2x faster.
oh and there is also another tab where it is also ripping through Bend's codebase and finishing everything I had to do
I don't know what to say anymore
this isn't about Anthropic or OpenAI, this is about our collective future as a species. the world is changing, and we need to be aware of it, and discuss how to handle this change.
receipt below . . .
显示更多
SANTIAGO 🚨
What a night & it felt like yesterday
The energy and passion you all showed was insanely memorable.
I am so thankful because you all allowed me to be myself.
It’s the last show i had on LEG 2 MAGICMAN2 world tour
I had a wonderful time traveling around sharing my stories to yall face to face
Im very thankful and blessed to be able to do so
Thank you for listening & i hope again:
find yourself and live the life you want to live.
Another thing i want to add upon that is:
If you really think about it, People are all reactions.
what you do actually determines everything.
When you are moving too fast, they tell you to slow down.
When you are moving too slow, they say you lazy.
When you are dreaming, they wake you up.
When you believe in something, they question and judge.
And the list goes on.
And THIS is the pattern.
My question to you all who i love and care about is:
why bother to care ?
Be you, and you are the main character of your movie “life”.
Live the magic, live the dream and make it happen.
That’s what i always want to share to you all.
I love you and i care, I hope you hear me.
You got it.
It’s too late if you don’t start.
.
And fyi
Im working on my next album.
Up until this point, it’s chaos.🤯
Putting all the puzzles into the right places,
New environment, new people, new chemistry,
It’s going to be ridiculous.🏃🏻💨🚀
NO ONE READY FOR THIS.
.
#
MAGICMAN2WORLDTOUR#
#
Santiago#
#
JACKSONWANG#
#
王嘉爾#
#
MovistarArena#
显示更多
0
0
1.1K
16.7K
6.4K
转发到社区
Google has a new system called Cloud Fraud Defense, which is the next version of reCAPTCHA, and has started rolling out to users
When the system detects risky web activity, it no longer shows the old picture puzzles where you pick out buses or traffic lights. Instead, it displays a QR code that you scan with your Android phone, but to pass the test your phone must have Google Play Services installed and running.
This change has been active since October 2025 based on support pages and old web records, and it blocks users of privacy-focused Android phones such as GrapheneOS, CalyxOS, and /e/OS because these phones remove Google services on purpose to provide stronger privacy and security.
The result is that millions of websites now treat these privacy phones as risky, so users must either add Google Play Services or stay locked out.
This is similar to Google’s 2023 Web Environment Integrity idea that wanted websites to check if devices were trustworthy through Google software.
That plan received heavy criticism from developers and privacy groups and was dropped, but the new QR code method does something very similar in a simpler way.
Website owners who use this system are now blocking people who chose to remove Google from their phones for better privacy.
显示更多
Temple Maker 64 is a new game coming soon to Steam for PC with a retro look and feel like old Nintendo 64 Zelda games such as Ocarina of Time.
It is a third-person dungeon creator that lets players design and build their own custom dungeons.
A simple 3D editor to create dungeons from the ground up
-Build rooms on different floors
-Add traps and puzzles
-Place monsters and bosses
-Include classic items like swords, bows, bombs, and grappling hooks
Once your dungeon is ready you can publish it online so the community can try it, you can also explore and play dungeons made by other players.
The game is made by one person and left his full-time software engineering job about a year and a half ago to work on this full time.
Temple Maker 64 does not have a release date yet.
显示更多