注册并分享邀请链接,可获得视频播放与邀请奖励。

与「Dr柴田亮」相关的搜索结果

Dr柴田亮 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 Dr柴田亮 的内容
Thanks The Billboarders🎉🎉🎉‼️‼️‼️また逢いましょうッ😉💫💫💫🎉🎉🎉‼️‼️‼️(全快笑顔😜😜😜) #Pf加藤景子# #Ba河本奏輔# #Gt加藤素朗# #Dr柴田亮# #SaxDAVIDNEGRETE# #感謝#
显示更多
0
53
938
192
转发到社区
HIROYUKI TAKAMI LV TOUR2023 〜Heavenly Soar〜 Vo.貴水博之🎉Gt.清水武仁🎉Dr.柴田尚🎉Key.守尾崇🎉2023メンバーでお届けしマス😉一緒に楽しんで盛り上がろうッ🎉イーーーエイッ😉🎉🎉🎉‼️‼️‼️ 新曲試聴は下の画像をタップ☝️ ツアースケジュールはこちら↓ #貴水博之#
显示更多
0
45
894
235
转发到社区
恭喜 @real_dr_pump @SuperBILI cash-cat:native 回到他自己家里了😂。 只能说狗 @sss_crypto fud,天理难容!
We tried using Meta's new Muse Code agent, but it has a bug that doesn't let it sign in from a docker container. So we did a fun experiment: Meta claims Muse Spark 1.2 was co-trained with their Muse agent harness. So we extracted instructions from their system prompt and added them to the Cline harness. TL;DR of this special prompting: - Trust source code over the user prompt, so read every call site and existing tests before starting the task - Weigh edge and error cases as heavily as the happy path - Always reproduce the bug before fixing - Don't trust the first passing test suite, and verify suspicious looking half-baked tests - Never stop at just editing, keep working until the change is verified complete. We then asked this modified harness to fix a real bug from our repo, and compared the results to the original Cline agent harness. Results: - Used 2.7x fewer tokens (19.7M → 7.2M) - Finished 2x faster (49min → 24min) - Cost 2.4x less ($7.69 → $3.25) Same Muse Spark 1.2 model, same task, only the prompting changed. Incredible how much of a performance gain Meta was able to achieve training it on these special instructions!
显示更多
0
25
345
21
转发到社区
Okay, the @VulcanBench results for Qwen3.8-Max are in, and it is not what I expected. First, for anyone new to VulcanBench, here's a quick TL;DR on the eval suite: 23 frontier-hard software engineering tasks taken from real merged OSS PRs, run in a Docker sandbox, 3 runs per task across all three of its effort levels. No puzzles, no random abstract stuff, all real things engineering teams would do with these models. It looks like Qwen3.8-Max has a major overthinking problem, it uses a LOT of tokens and is very slow, period, no other way to see it. My cost to run this benchmark was $126.25, to run the exact same eval suite with DeepSeek V4-Flash was only $13.60. This makes Qwen3.8-Max an insanely expensive model. The tasks Qwen genuinely can't solve fail at every effort level, extra reasoning didn't help. The regression is almost all in work it already handles: six tasks that low solves every single time account for 83% of the 26-point drop, three of them collapsing to zero. It's not losing the hard problems. It's losing the ones it already knows how to do. Since Qwen3.8-Max hit a lot of wall clock budget caps, I thought I'd share more about this. - VulcanBench caps both steps (50–200) and wall clock (5–60 min), each scaled by repo size. - This is aligned with how comparable harnesses bound agents, DeepSWE caps rollouts at 100 environment steps, sitting right inside my step range; Terminal-Bench enforces a per-task wall clock; SWE-bench Verified scaffolds typically allow 20–60 min per instance with 250–350 step limits. - Every model on my chart gets the identical budget, and Qwen is the slowest model I've tested at 20–25 min/task. Soooo... Alibaba positions Qwen3.8-Max as trailing only Claude Fable 5. But on the kind of real coding work engineering teams would actually throw at it, under a fixed budget, its best setting lands mid-pack and its default lands last, so common. If you want to optimize for accuracy, Grok 4.5 is the move. If you want accuracy per dollar, DeepSeek V4-Flash is hard to beat, heck it's 10× cheaper than Qwen and you get higher accuracy. Qwen just isn't in the game at this point, this is not a model I could see engineering teams using for daily coding work.
显示更多
0
28
178
12
转发到社区
🚨NEWS: A migrant surgeon in the NHS has been fired after he connected the wrong organs during a surgery Dr Abdel Rahman connected a patients small bowel back to their stomach creating a feedback loop meaning the patient did not defecate for 3 weeks before being rushed in for an emergency procedure to fix it We should train our own surgeons instead of importing people from the third world
显示更多
0
595
22.6K
4K
转发到社区
Some of our earliest stories were captured one frame at a time Explore the role of the Bell & Howell 70-DR in capturing defining moments in our history
0
54
475
47
转发到社区
“[Fauci’s] focus was on promoting himself, on getting on TV, and not doing public health." HHS Secretary @RobertKennedyJr blasts Dr. Fauci for his "giddy narcissism", “self-promotion” and pushing the public’s safety to the wayside. “The worst part is that giddy narcissism that you see throughout the diary is juxtaposed against really bad public health decisions.” @IngrahamAngle
显示更多
0
160
623
162
转发到社区
WATCH: Rep. @timburchett calls out Dr. Fauci for appearing on Rachel Maddow's MSNBC show and dubbing it his "favorite show" while serving under Biden: KILMEADE: “If Rachel Maddow was your favorite show, what’s your political ideology?” BURCHETT: “Far left as you can get, brother. You wouldn’t get elected to dogcatcher in Tennessee.” “It’s just pathetic. It shows his mindset. I bet he couldn’t even spell Fox.” @WillCainShow @kilmeade
显示更多
0
157
1.1K
191
转发到社区