注册并分享邀请链接,可获得视频播放与邀请奖励。

Kun Chen 的个人资料封面
Kun Chen 的头像

Kun Chen (@kunchenguid)

@kunchenguid
0 正在关注    0 粉丝
opus 5 is a VERY interesting release for a few reasons 1. it showed that the general benchmarks we use today are almost completely useless now opus 5 is nowhere near fable in practical use, not even close. anyone who’s used it meaningfully can tell this very quickly after a few tasks. yet opus beats fable on many benchmarks i now trust domain specific benchmarks built with private datasets a lot more than the popular ones. perhaps the future is everyone running their own evals because the public ones are really not telling us much 2. it seems with the 5 series, anthropic is trying a new way of training models previously, the same generation of sonnet and opus were often released at the same time or sonnet comes out before opus, which indicates sonnet and opus were trained by separate pipelines in parallel with the 5 series, it was very clear that they trained mythos first, and then distilled it into sonnet and opus. it seems this approach has a big influence on the models seeing sonnet 5 being a flop and opus 5 getting pretty mixed reviews already, i’m not sure this is working out 3. “how pleasant is it to work with the model” used to be a strength in claude, but now it’s not. honestly, grok is my favorite right now on the “pleasant” dimension. kimi is not bad either it feels like both anthropic and openai are giving RLHF less care, in favor of scalable RL that’s machine verifiable this almost looks like AI is directing humans to build a world that’s more friendly for machines rather than humans, and most humans don’t even realize they are being manipulated to help with that almost every new generation of frontier models now talk more jargons, need more steering to do what you want, and are just less fun to work with if this continues, AI will start to speak their own language that looks like English but average humans can’t understand. they will choose to do things that their human user never asked for. are we already failing at alignment?
显示更多
0
255
3.2K
199
转发到社区
pro tip - when you use OpenAI's gpt models in Codex, it uses a server-side encrypted compaction that seems to work better than anything else out there, which allows Codex to just keep hammering on long running tasks like there's infinite context window that's great, but if you run gpt in other harnesses like Pi, most of them don't inherit that by default, resulting in worse performance in long running tasks but - because of how extensible Pi is, i just found this cool extension from @alexisgallagher that enables the same server-side compaction in Pi - benchmark seems to support the argument that OpenAI server side compaction is indeed superior - so if you are using gpt models in Pi, install that extension to improve long running task performance. firstmate benefits a lot from this if you are using other harnesses, be aware of this difference and see if you can find a similar solution
显示更多
0
19
577
39
转发到社区
grok 4.5 made me give grok build a serious run today here's my honest first impression (non affiliated neutral view point): 1. grok build is a very good harness firstmate stretches harness capabilities to their limits, and i've been testing it with claude code, codex, opencode, pi and grok build so far, grok build and claude code are the only two harnesses that can automatically wake up when the background polling process finishes codex hard fails on this kind of background polling task, with no escape hatch. opencode and pi can both do it with custom plugins, but not out of the box grok build also feels really clean, smooth, and transparent. you can see what background tasks are running, what hooks got triggered at what step, context window, token usage etc all out of the box and somehow the UI does not feel cluttered at all 2. grok 4.5 is a very good model i've been using opus 4.8 as my primary firstmate, and today i did a full switch to grok 4.5. so far i don't think anything degraded, while token throughput is a lot faster, although time-to-first-byte for each response seems long - i wonder if prompt caching is done properly or no 3. the quota that comes with X premium is quite generous, and there's no session level limit so overall i'm quite pleased by this and plan to switch a lot of my tasks to grok. the landscape just got a lot more interesting...
显示更多
0
49
363
29
转发到社区
many people asked me to make a video about my complete agentic engineering workflow excited to share it's finally here!!! it took me about 20 hours in total to record this 45 minutes of walkthrough - it covers everything i do to ship production quality code at an average 40+ PRs/day velocity hope this can be a useful reference to everyone exploring good ways to use AI. and would appreciate a reshare with anyone you think might benefit from this! enjoy!
显示更多
0
40
688
69
转发到社区