注册并分享邀请链接,可获得视频播放与邀请奖励。

与「Max_Ver」相关的搜索结果

Max_Ver 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 Max_Ver 的内容
GHOST 👻 CAR Ride onboard with Max Verstappen as he vies for pole with Kimi Antonelli ⚔️ #F1# #BelgianGP#
GHOST 👻 CAR Ride onboard with Max Verstappen as he vies for pole with Kimi Antonelli ⚔️ #F1# #BelgianGP#
0
69
1.5K
115
转发到社区
Ethan Pimstone sets a new max vertical jump world record with an incredible 1.33m (52.5”) leap. 🚀 The hang time on this jump is absolutely unreal. 😳 (Via: ethanpimstone1/IG)
显示更多
0
28
924
87
转发到社区
[📽] 대성(DAESUNG) ‘한도초과(HANDO-CHOGUA)’ (Max Ver.) VIDEO 🔗 #대성# #DAESUNG# #DLITE# #한도초과# #HANDO_CHOGUA# #Max_Ver#
0
5
1.5K
475
转发到社区
NVIDIA 发布 Skill2Env:用“集体技能”强化智能体 NVIDIA 研究者们把社区公开的 Agent Skills 编译成可执行 RL 训练环境的数据流水线:3.4k 个 Skills 变成 8k 个带程序化测试和行为量规的终端任务;仅 300 步 RL 训练就让 Qwen3.8-27B 在 Terminal-Bench 2.1 上提升 4.7 个百分点,且模型行为显著向源 Skills 的方法论对齐。 开源项目: 核心洞察:公开 Agent Skills 是一个被忽视的监督来源 Agent Skills 是“教智能体做某件事”的文件夹:一个 SKILL.md 加上可选的脚本、参考资料和资产。论文指出,把公开 Skill 语料当作数据来读,它同时提供三样东西: · 任务分布的采样:人们真正想让智能体处理的任务分布(有人愿意花时间写下工作流,说明这活儿值得自动化); · 真实世界的锚点:指向真实的仓库、数据集、工具和工件; · 结果测试表达不了的质量标准:领域专长、默认参数、常见坑、“好结果长什么样”。 # 数据流水线:四阶段编译,验证靠构造 1. Plan(分解):容器化的 Codex 规划器读取完整 Skill 包、联网调研相关公共资产,把 Skill 拆解成若干可验证的 workflow,每个附带元计划(场景、初始世界、预埋缺陷、难点来源、解法草案、验证策略)、资产建议和“任务轴池”(任务原型 × 验证器模式 × 人物画像)。 2. Diversify(多样化):宿主从轴池采样一组组合,加上复杂度、指令语气、请求者专业水平。关键设计是轴池以 workflow 为条件:研究型 workflow 配“证据可追溯”验证和研究者画像,而不是从全轴乘积空间乱抽,这让多样化保持 sensible。 3. Create(构造):全新创建者 Codex agent 在 Docker 内工作,尽可能用真实素材(钉在特定 commit 的开源仓库、真实版本化文档、官方 API 规范);需要联网服务的场景改造成本地替身(stub 服务器、录制回放 fixture、PATH 上的假 CLI、种子数据库),求解时绝不依赖网络。创建顺序被严格固定:先建世界 → 写指令 → 写测试 → 写量规 → 最后才写参考解,测试先于解法冻结,保证解法必须迁就评分契约而非反过来。 4. Verify(验证):宿主端无模型参与的接收门:静态检查(布局、符号链接、Dockerfile 安全、基础镜像按内容摘要钉死)+ 两个容器内试跑:Oracle(参考解)必须全指标满分,NOP(什么都不做的 agent)必须全指标零分。任一失败即拒绝。 值得注意的一个反直觉选择:不做 teacher 模型预验证(不像部分工作用强模型试解、解不出就丢弃任务)。理由有二:这会把任务难度上限压到验证器能力,且成本翻倍;而 group-based RL 的在线动态过滤(rollout 无优势的 prompt 自动不产生梯度)天然淘汰过难/过易任务。 # 数据画像:广、贵、且忠实于源 规模与成本:7,971 个任务,用 GPT-5.6 Sol(xhigh 推理档)生成,API 花费超 9 万美元。(脚注:出于法律原因,公开发布的数据集改用 Kimi-K3-max 在同一流水线下生成。) 领域分布:13 个领域中,软件工程仅占 22.5%,AI/ML 10.5%,商业/金融/法律/HR 10.5%,营销 9.3%……论文对比了 TMax-15K、Terminal-Bench、DeepSWE 等,Skill2Env 是唯一全覆盖 13 域、且非技术知识工作占大头的语料。 忠实度探针(很聪明的设计):用任务指令+量规作查询、对 3.4k 个 SKILL.md 做 TF-IDF 检索,73.2% 的任务 top-1 命中真实源 Skill,94.6% 进 top-10(随机 0.03%)。单用量规也有 68.5% top-1,证明量规携带的是 Skill 专属方法论而非泛泛建议。 SFT 数据:用 GLM-5.3 对每个任务 rollout 两次,得到 15,968 条轨迹,平均奖励 0.74,中位轨迹 19 次模型调用 + 23 次工具调用。 S2EBench:考虑到公开基准饱和,从 SkillHub 另外生成、逐条人工审核(指令无歧义、忠实于源 Skill、测试公允)后的 79 任务私有 held-out 基准。 # RL 实验:基础设施 + 极简配方 基础设施(论文明确说“现代 agentic RL 首先是基础设施挑战”):Molt(PyTorch 原生全异步训练,Ray + vLLM + FSDP2)+ Polar(agent rollout 层:rootless Apptainer 沙箱、代理回传 token ID 和采样时 log-prob、prefix merging 把 harness 的多次补全缝合成训练轨迹)。 配方(刻意走“简单路线”):GRPO 组归一优势 + DPPO 的 binary-KL 信任域掩码(δ=0.05,超出阈值的 token 直接丢弃,无需参考模型,还能防训练-推理失配);G=8 rollouts/组,批 64,lr 1e-6 恒定,无 KL 惩罚、无熵奖励、无 SFT 热启动,每任务 65k 上下文。 量规校准奖励:开量规时,额外由 GPT-6 Astra 做 LLM-as-Judge(带“宪法”:惩罚无脑循环、reward hacking、答非所问;hacking 实证 = -5 分),总奖励 r = r_V + λs/5(λ=0.2),即 judge 最多把程序化奖励拉动 ±0.2。量规是校准可执行结果奖励,而非取代它,这是与“Rubrics as Rewards”一系的定位差异。 # 四项发现(论文最有信息量的部分) 发现 1:小规模 RL 即有跨域迁移。 仅 300 步、只用 2,400 任务子集训一个 epoch:S2EBench pass@1 +4.3(均分 +18.5),Terminal-Bench 2.1 +4.7(49.4→54.1)。训练集与 TB 无重叠(13-gram Jaccard < 0.8),且训练集从未针对 TB 调过,论文将其解读为规划、工具使用、收尾能力的通用提升而非任务族记忆。这让 27B 本地模型显著缩小了与云端前沿模型的差距。 发现 2:量规校准 RL 在基准上落后于纯结果 RL,一个诚实的负结果。 量规版在 TB 2.1 只有 50.1(纯结果版 54.1);训练中量规版的程序化奖励长期停在 0.5–0.6,judge 分项从头到尾无上升趋势,两个奖励在训练分布上互相拉扯。论文不把它当作对量规奖励的终审判决(两者优化不同目标,而基准只考结果那一半),并给出两个疑因:λ=0.2 的加性形式让失败任务仍能拿正奖励、judge 看不到文件系统等设定均未调优;以及更本质的,Skill 写下的方法论可能本来就不是最大化基准通过率的分布。 发现 3:行为确实向 Skill 对齐,量规的价值所在。 200 个任务的成对偏好测试(judge 拿源 SKILL.md 当标准,比较匿名化的 base 与 RL 轨迹):纯结果 RL 已被偏好 54.5% vs 33.5%;量规版被偏好 73.0% vs 24.0%。这说明量规奖励买到的东西在结果基准上看不见,但对“怎么做事”影响实质,对网页开发、报告综合、开放研究这类难验证任务尤其重要。 发现 4:GLM-5.3 蒸馏 SFT 反而伤害 Qwen。 在 GLM-5.3 轨迹上做 SFT:27B 上 TB 2.1 掉到 45.8;4B 上直接崩塌(TB 18.7→3.4,出现思维/工具调用死循环)。归因:教师的 interleaved-thinking + 工具调用风格与学生自身 post-training 不兼容,模仿覆盖了学生依赖的行为模式却带不来教师的能力。与 TMax 报告的“SFT 混合数据劣化已后训练的 Qwen”互相印证。因此论文所有 RL 结果都从未修改的原始 checkpoint 出发。
显示更多
Grok 4.6 just tied for #1# on the Artificial Analysis Agentic Index • Grok 4.6 (high) — 59 • Claude Opus 5 (max) — 59 Outperforming Claude Fable 5 and GPT-5.6 Sol We’re entering the agentic era, and this is exactly the kind of benchmark that matters so much: tool use, planning, autonomy and complex problem solving Grok 4.6 is now sitting at the very top And that matters even more as Grok powers Grok Build and Grok Bot, where the model has to go beyond answering questions and actually take actions, use tools and complete real work Grok’s agentic capabilities are getting seriously powerful
显示更多
0
81
455
47
转发到社区
Agentic work is where @grok 4.6 lands hardest, taking the top spot on the Artificial Analysis Agentic Index at 59, tied with Claude Opus 5 Max. The index measures tool use, planning, autonomy and complex problem solving rather than single answers Grok 4.6 completes tasks in ~53 turns and ~0.5bn input tokens on average, against ~103 turns and ~2.0bn for Claude Opus 5 Max Cost of $0.84 per task, putting it on the intelligence versus cost per task Pareto frontier Enterprises buying agents pay per completed task, not per benchmark point. Turn efficiency is what determines whether a long-running workflow is affordable at volume. Two labs now sit at the top of this index with very different cost structures. Buyers get real choice on price for the first time in agentic deployment.
显示更多
10 PLATAFORMAS QUE REGALAN CRÉDITOS GRATIS PARA APIs DE IA AHORA MISMO Sin suscripciones. Sin listas de espera. Te dan gratis hasta 300$ solo por crear una cuenta. Modelos como Gemini, GPT, Claude y muchos más sin pagar ni un solo token. Guárdate esta lista para tu próximo proyecto. 1. Google Cloud (300 $, 90 días) Es la mayor bonificación de toda la lista. Google regala 300 $ en créditos durante 90 días para utilizar gran parte de sus servicios, incluidos los modelos Gemini desde Vertex AI y Agent Platform. Importante: • Los 300 $ no sirven para la API de Gemini en Google AI Studio. • No pueden utilizarse para modelos de terceros ofrecidos como APIs gestionadas. • No puedes usar GPUs mientras la cuenta esté en el periodo de prueba. 2. Oracle Cloud (300 $, 30 días + Always Free) También ofrece hasta 300 $ en créditos, aunque solo durante 30 días. Lo mejor es su nivel Always Free, que sigue funcionando incluso cuando termina el crédito inicial, sin costes. Es una de las mejores opciones si prefieres alojar tus propios modelos en lugar de consumir APIs de terceros. 3. Microsoft Azure (200$, 30 días) Azure ofrece 200 $ en créditos para gastar durante los primeros 30 días. Es una de las formas más sencillas de acceder a los modelos GPT mediante Azure OpenAI. Además incluye: • Más de 20 servicios gratuitos durante 12 meses. • Más de 65 servicios Always Free. Eso sí, el crédito caduca a los 30 días y no puede recuperarse. 4. AWS (hasta 200$, 6 meses) La prueba gratuita más larga de la lista. Recibes 100 $ al registrarte y puedes conseguir hasta 100 $ adicionales explorando distintos servicios de AWS. Si quieres usar Claude desde Amazon Bedrock, esta es probablemente la mejor opción. 5. Google AI Studio (gratis para siempre) Es independiente de los 300$ de Google Cloud. Incluye un plan gratuito permanente con tokens para: • Gemini 3.6 Flash • Gemini 3.5 Flash • Flash-Lite • Modelos de embeddings Además ofrece 5.000 consultas mensuales con Search Grounding. La contrapartida es que, en el plan gratuito, Google puede utilizar tus datos para mejorar sus modelos. 6. Cloudflare Workers AI (10.000 Neurons al día) Cloudflare ofrece 10.000 Neurons gratuitos cada día, que se reinician automáticamente a las 00:00 UTC. No es una prueba temporal: el límite se renueva diariamente. Da acceso a unos 80 modelos distintos. Ten en cuenta que Kimi K2.6, K2.7-Code y GLM-5.2 requieren un método de pago. 7. Groq (plan gratuito) Permite utilizar varios modelos sin necesidad de tarjeta. Incluye: • 30 peticiones por minuto. • 14.400 peticiones al día con Llama 3.1 8B. El principal límite está en el rendimiento por tokens: • Entre 1.200 y 15.000 tokens por minuto, según el modelo. • Llama 3.3 70B está limitado a 1.000 peticiones diarias. Ideal para agentes sencillos, aunque menos recomendable para tareas con mucho contexto. 8. OpenRouter Una única API para acceder a cientos de modelos diferentes. El plan gratuito permite 50 peticiones diarias en los modelos compatibles. Si compras 10 $ en créditos una sola vez, el límite aumenta permanentemente a 1.000 peticiones al día. Es una de las mejores inversiones si utilizas varias APIs. 9. Cerebras (5 $ en créditos) Solo por crear una cuenta recibes 5 $ en créditos para probar todos sus modelos. Puede parecer poco, pero gracias a la velocidad de inferencia permite hacer bastantes pruebas. Actualmente, Cerebras Code Pro y Code Max no están disponibles por alta demanda. 10. Mistral (plan gratuito) Mistral ofrece acceso gratuito a sus modelos mediante: • Chat. • Búsqueda. • Programación con agentes desde la terminal. Tiene límites de uso, pero no requiere tarjeta ni ningún pago.
显示更多
0
20
105
23
转发到社区
Okay, the @VulcanBench results for Qwen3.8-Max are in, and it is not what I expected. First, for anyone new to VulcanBench, here's a quick TL;DR on the eval suite: 23 frontier-hard software engineering tasks taken from real merged OSS PRs, run in a Docker sandbox, 3 runs per task across all three of its effort levels. No puzzles, no random abstract stuff, all real things engineering teams would do with these models. It looks like Qwen3.8-Max has a major overthinking problem, it uses a LOT of tokens and is very slow, period, no other way to see it. My cost to run this benchmark was $126.25, to run the exact same eval suite with DeepSeek V4-Flash was only $13.60. This makes Qwen3.8-Max an insanely expensive model. The tasks Qwen genuinely can't solve fail at every effort level, extra reasoning didn't help. The regression is almost all in work it already handles: six tasks that low solves every single time account for 83% of the 26-point drop, three of them collapsing to zero. It's not losing the hard problems. It's losing the ones it already knows how to do. Since Qwen3.8-Max hit a lot of wall clock budget caps, I thought I'd share more about this. - VulcanBench caps both steps (50–200) and wall clock (5–60 min), each scaled by repo size. - This is aligned with how comparable harnesses bound agents, DeepSWE caps rollouts at 100 environment steps, sitting right inside my step range; Terminal-Bench enforces a per-task wall clock; SWE-bench Verified scaffolds typically allow 20–60 min per instance with 250–350 step limits. - Every model on my chart gets the identical budget, and Qwen is the slowest model I've tested at 20–25 min/task. Soooo... Alibaba positions Qwen3.8-Max as trailing only Claude Fable 5. But on the kind of real coding work engineering teams would actually throw at it, under a fixed budget, its best setting lands mid-pack and its default lands last, so common. If you want to optimize for accuracy, Grok 4.5 is the move. If you want accuracy per dollar, DeepSeek V4-Flash is hard to beat, heck it's 10× cheaper than Qwen and you get higher accuracy. Qwen just isn't in the game at this point, this is not a model I could see engineering teams using for daily coding work.
显示更多
0
28
178
12
转发到社区