注册并分享邀请链接,可获得视频播放与邀请奖励。

九原客 的个人资料封面
九原客 的头像

九原客 (@9hills)

@9hills
0 正在关注    0 粉丝
一种新的Agent时代的开源方式,只接受prompt request不接受pull request。 😄
我不太理解A和O对自家风控那么自信么,一点也不会有假阳性? A社其实表现还好点,它就是封+退款,没啥小动作。 O 就比较鸡贼,根据ip质量风控降智我是见过的,现在还有降额度?误封的都不知道。 你不想挣钱就封号+退款呗,搞这么拧巴。
显示更多
We've investigated a few messages about codex usage limits being different. That's not something we change without engaging the community and being transparent. What we did see is that when talking to affected users many were using sub2api. Converting a subscription into api traffic to then re-serve or share across many users is not something we support and this type of usage gets flagged by our fraud-prevention systems. You are completely fine if you use your subscription through Sign in With ChatGPT, either through the official clients or through one of the many OSS clients (Pi, OpenCode, ...) that support signing in with your account and using your included usage.
显示更多
Qwen3.8-27B 本地部署指南 1. Q4 基本没有质量损失,甚至可以用Q3 2. 用Unsloth的GGUF。 3. 思考开low就行了。 4. 开 dflash2
Main result. At xhigh every quant landed between 88.0 and 90.0% pass@1 on the full suite: AWQ INT4: 90.0% NVFP4: 89.3% GGUF Q4_K_M: 89.3% FP8: 88.7% NInfer: 88.0% Yes, the 4 bit quants scored above the FP8 baseline. McNemar says its a statistical tie, first and last place differ by three tasks out of 150. I started this run to show you how quantization eats quality. There is nothing to show. The gap between quants is smaller than the gap between reasoning presets.
显示更多
用了十年的翻墙服务已经扛不住,变得很慢而且关闭了订阅链接。 紧急花49.99刀买了三个月的搬瓦工CN2GIA,给Agent 一条指令就配好了。 但是有点贵啊,有没有便宜的推荐
显示更多
这个情况我愿称之为 Harness 屎山。 包括最近有人提到Claude经常跨session对话等问题,当累积的 Harness 足够多(Prompt、Rule、Tools 等等),非预期的影响就更可能会出现。 Harness 也要 KISS。
显示更多
越来越看不懂 claude code 了,为啥 bypass permission 优先用 shell 脚本来改文件? 看过源码就知道,他那个 Read/Write/Edit 工具有并发修改检测,还有内容追踪和回滚的功能。用 shell 命令来改文件,完全无法追踪和恢复。 咋想的?
显示更多
有几篇论文我也看过。 目前多 Agent 有两种形态,主子 Agent 和 Agent Team。后者我认为是一个很复杂的协作问题,就和当年分布式系统一样。 主子 Agent 本质没有协作,只是任务分发。它在总上下文超过了单 Agent 的有效上下文窗口时,是成功的。
显示更多
做了个知识库skill,相当于给代码库建llm wiki。简单benchmark了下,KB on 效果全面落后 KB off,最困惑的是成本也变高了。 之前测试过很多传的很广的skill、code graph ,可能并不能带来正收益。 我现在pi的插件和skill 只剩寥寥几个。
显示更多
orca还有一个可能是小众的需求,支持linux。
Coding Agent IDE 我只推荐3个: 1. 如果需要兼容各类Agent,那么选择 Orca,你能想到的功能它都全,支持切换Chat模式和终端模式。内置编排技能可做 Agent Team。 2. 最简单易用就是 Codex App,只是锁定了 Harness。 3. 如果喜欢终端,那么选择 Herdr,功能相对简单但也足够。
显示更多
对DSH的一个评测,仅供参考。
同一个 DeepSeek V4 Flash,我们进行了 9 个不同 harness 的头对头测试。 Maka 排名第二,通过率 77.5%,官方 DSH 的 Minimal Mode 排名第四,通过率 73.0%。 最高的是 Codex,最糟糕的是 Claude Code,通过率低且成本最高。
显示更多
GLM-5.3 发布!这一轮国内模型的爆发是从 GLM-5.2 开始,我们也都知道后训练还能大幅提升,所以 5.3 值得期待。 当然你也可以看到 Kimi K3 靠庞大的体积,依然在 Coding 领域领先。
显示更多
还要离线跑起来 DeepSWE,感觉头大
好难啊,要在客户这里跑起来 Harbor + DeepSWE。 没有顺畅的国际互联网访问,跑不起来啊。 不想干 ToB 了。
换个思路,这是个做Agent Harness的研究人员用的,很容易就定制一种的新的harness idea发paper。 有用没用另说。
...这是哪一代agent才能设计出来的构思😹 deepseek pro已经够拉了,不要在agent上再拉一遍啊 如果做插件就得专注先把core做好
现在告诉我说Pro新版还有盗版和正版? 没心情测试,过几天再说吧….
上次我说要两倍起。 结果是2-12倍!现在coding agent的大头就是cached input,涨的最多。 感觉综合成本应该会是3-4倍上涨。
真的无语了,DeepSeek V4 Pro 高峰期照着 12 倍去涨价…
貌似没人关心 Qwen3.8-2.4T-A95B 开源,其实 Qwen3.8-Max还行,就是太贵了。 DeepSeek-V4-Pro 感觉有点问题,不如用 Flash。让后训练再跑一会~ 目前国产模型我的选择: 超大杯:Kimi K3 大杯:GLM5.2 中杯:DeepSeek-V4-Flash 国外模型选择: 超大杯:Opus 5 大杯:GPT-5.6-Sol 中杯:Grok 4.5(4.6还在测试)
显示更多
利用 Claude 推理块加密跨模型可复用的漏洞,让 Haiku 恢复 Opus 的加密思维链。 有点意思,不过并不能证明完美恢复,类似于提示词越狱。
We can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.
显示更多
the bitter lession 还在发力。 我的观点是 harness infra(沙盒、工具等等) 肯定是持续做厚,但是 harness context(比如什么 loop、graph、agent team、外置 memory 等等)价值是下降的,甚至有些东西的价值一开始就不成立。
显示更多
We ran DeepSeek V4 Flash through 4 more agent harnesses (Hermes Agent, Pi Agent, Prime Agent, Deep Agents) on 30 challenging agentic tasks. Pi Agent was the cheapest harness and passed the most tasks 🧵🧵
显示更多