注册并分享邀请链接,可获得视频播放与邀请奖励。

与「sllow」相关的搜索结果

sllow 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 sllow 的内容
インスタに載せてないアザーカット😉 #叫 ##shakebe# ## #sllow# #かぷちゅーる#
0
4
451
21
转发到社区
I love slow mornings 💖🪷🌹
0
35
3.1K
56
转发到社区
i love a good slow burn reveal
0
176
13.4K
548
转发到社区
This Mole CLI release is almost all reliability work, on purpose. Refusals now explain themselves and print the fix. Dry run is fully dry. Stricter cleanup boundaries, hard deadlines on slow scans, Parallels VMs in disk analysis. Safer everywhere.
显示更多
Okay, the @VulcanBench results for Qwen3.8-Max are in, and it is not what I expected. First, for anyone new to VulcanBench, here's a quick TL;DR on the eval suite: 23 frontier-hard software engineering tasks taken from real merged OSS PRs, run in a Docker sandbox, 3 runs per task across all three of its effort levels. No puzzles, no random abstract stuff, all real things engineering teams would do with these models. It looks like Qwen3.8-Max has a major overthinking problem, it uses a LOT of tokens and is very slow, period, no other way to see it. My cost to run this benchmark was $126.25, to run the exact same eval suite with DeepSeek V4-Flash was only $13.60. This makes Qwen3.8-Max an insanely expensive model. The tasks Qwen genuinely can't solve fail at every effort level, extra reasoning didn't help. The regression is almost all in work it already handles: six tasks that low solves every single time account for 83% of the 26-point drop, three of them collapsing to zero. It's not losing the hard problems. It's losing the ones it already knows how to do. Since Qwen3.8-Max hit a lot of wall clock budget caps, I thought I'd share more about this. - VulcanBench caps both steps (50–200) and wall clock (5–60 min), each scaled by repo size. - This is aligned with how comparable harnesses bound agents, DeepSWE caps rollouts at 100 environment steps, sitting right inside my step range; Terminal-Bench enforces a per-task wall clock; SWE-bench Verified scaffolds typically allow 20–60 min per instance with 250–350 step limits. - Every model on my chart gets the identical budget, and Qwen is the slowest model I've tested at 20–25 min/task. Soooo... Alibaba positions Qwen3.8-Max as trailing only Claude Fable 5. But on the kind of real coding work engineering teams would actually throw at it, under a fixed budget, its best setting lands mid-pack and its default lands last, so common. If you want to optimize for accuracy, Grok 4.5 is the move. If you want accuracy per dollar, DeepSeek V4-Flash is hard to beat, heck it's 10× cheaper than Qwen and you get higher accuracy. Qwen just isn't in the game at this point, this is not a model I could see engineering teams using for daily coding work.
显示更多
0
28
178
12
转发到社区
The small bridges and flowing water in Zhouzhuang hold the quietest Jiangnan. Do you enjoy this slow-paced life?
Rain, mist and a city wrapped in clouds. This is another side of Chongqing worth slowing down for. Would you visit on a rainy day? 📸 wubiye #Chongqing# #ChongqingChina# #RainyDay# #TravelChina# #Cityscape#
显示更多
We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it. Opus went off for ~2 hours and wrote 5500 lines of code that (procedurally) rendered the story. It's kind of janky but fun. But it's a bit mindboggling that the LLM has to place and orchestrate various polygon assets in (x,y,z) coordinates and write code that animates it all, and that it even does anything at all. I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free". There might be a lot more. But I'm excited about creating hyper custom worlds that you can imagine dropping players into, e.g. here to participate in the LoTR story as a spectator NPC, or one of the characters, or etc. Something like an ephemeral GTA of X on demand. Last thought is that the domain of worlds/games exposes a weakness in LLMs: they can't easily audit their work because they aren't able to efficiently and natively perceive videos or play games within them. Here, Opus 5 had to very slowly and painstakingly take screenshots at different points, and it messed up a few times and created a bunch of jank. An example of raw capability (multimodal, gameplay) that I think is still quite lacking.
显示更多
0
1.2K
23.7K
1.8K
转发到社区
C.J. Gardner-Johnson injured The Bills new safety went down, later clutched at his leg near his Achilles. Couldn’t put any pressure on his right leg as he was helped off the field slowly. Appeared non-contact Bad news for one of the Bills’ best defenders at training camp so far
显示更多
0
65
503
40
转发到社区
【📡#サミマ# Info.】 🎪ゲストレイヤー紹介🎪 尊みを感じて桜井 (@angelia_lapin ) 8月1日15:00~20:00  ブース出展 8月2日 15:00~20:00  ブース出展 記念すべき第一回 #サミマフォトセッション# に登場!! 8月1日18:30〜19:00 8月2日16:30〜17:00 8月2日19:00〜19:30 上記にて出演予定📸 📅2026年8月1日(土)・2日(日) 📍SLOW ART CENTER NAGOYA 🎫7月22日(水) 21:00〜LivePocketにてチケット販売開始(先着) #WCS# #コスサミ2026#
显示更多
0
0
357
28
转发到社区