注册并分享邀请链接,可获得视频播放与邀请奖励。

与「Oracle」相关的搜索结果

Oracle 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 Oracle 的内容
NVIDIA 发布 Skill2Env:用“集体技能”强化智能体 NVIDIA 研究者们把社区公开的 Agent Skills 编译成可执行 RL 训练环境的数据流水线:3.4k 个 Skills 变成 8k 个带程序化测试和行为量规的终端任务;仅 300 步 RL 训练就让 Qwen3.8-27B 在 Terminal-Bench 2.1 上提升 4.7 个百分点,且模型行为显著向源 Skills 的方法论对齐。 开源项目: 核心洞察:公开 Agent Skills 是一个被忽视的监督来源 Agent Skills 是“教智能体做某件事”的文件夹:一个 SKILL.md 加上可选的脚本、参考资料和资产。论文指出,把公开 Skill 语料当作数据来读,它同时提供三样东西: · 任务分布的采样:人们真正想让智能体处理的任务分布(有人愿意花时间写下工作流,说明这活儿值得自动化); · 真实世界的锚点:指向真实的仓库、数据集、工具和工件; · 结果测试表达不了的质量标准:领域专长、默认参数、常见坑、“好结果长什么样”。 # 数据流水线:四阶段编译,验证靠构造 1. Plan(分解):容器化的 Codex 规划器读取完整 Skill 包、联网调研相关公共资产,把 Skill 拆解成若干可验证的 workflow,每个附带元计划(场景、初始世界、预埋缺陷、难点来源、解法草案、验证策略)、资产建议和“任务轴池”(任务原型 × 验证器模式 × 人物画像)。 2. Diversify(多样化):宿主从轴池采样一组组合,加上复杂度、指令语气、请求者专业水平。关键设计是轴池以 workflow 为条件:研究型 workflow 配“证据可追溯”验证和研究者画像,而不是从全轴乘积空间乱抽,这让多样化保持 sensible。 3. Create(构造):全新创建者 Codex agent 在 Docker 内工作,尽可能用真实素材(钉在特定 commit 的开源仓库、真实版本化文档、官方 API 规范);需要联网服务的场景改造成本地替身(stub 服务器、录制回放 fixture、PATH 上的假 CLI、种子数据库),求解时绝不依赖网络。创建顺序被严格固定:先建世界 → 写指令 → 写测试 → 写量规 → 最后才写参考解,测试先于解法冻结,保证解法必须迁就评分契约而非反过来。 4. Verify(验证):宿主端无模型参与的接收门:静态检查(布局、符号链接、Dockerfile 安全、基础镜像按内容摘要钉死)+ 两个容器内试跑:Oracle(参考解)必须全指标满分,NOP(什么都不做的 agent)必须全指标零分。任一失败即拒绝。 值得注意的一个反直觉选择:不做 teacher 模型预验证(不像部分工作用强模型试解、解不出就丢弃任务)。理由有二:这会把任务难度上限压到验证器能力,且成本翻倍;而 group-based RL 的在线动态过滤(rollout 无优势的 prompt 自动不产生梯度)天然淘汰过难/过易任务。 # 数据画像:广、贵、且忠实于源 规模与成本:7,971 个任务,用 GPT-5.6 Sol(xhigh 推理档)生成,API 花费超 9 万美元。(脚注:出于法律原因,公开发布的数据集改用 Kimi-K3-max 在同一流水线下生成。) 领域分布:13 个领域中,软件工程仅占 22.5%,AI/ML 10.5%,商业/金融/法律/HR 10.5%,营销 9.3%……论文对比了 TMax-15K、Terminal-Bench、DeepSWE 等,Skill2Env 是唯一全覆盖 13 域、且非技术知识工作占大头的语料。 忠实度探针(很聪明的设计):用任务指令+量规作查询、对 3.4k 个 SKILL.md 做 TF-IDF 检索,73.2% 的任务 top-1 命中真实源 Skill,94.6% 进 top-10(随机 0.03%)。单用量规也有 68.5% top-1,证明量规携带的是 Skill 专属方法论而非泛泛建议。 SFT 数据:用 GLM-5.3 对每个任务 rollout 两次,得到 15,968 条轨迹,平均奖励 0.74,中位轨迹 19 次模型调用 + 23 次工具调用。 S2EBench:考虑到公开基准饱和,从 SkillHub 另外生成、逐条人工审核(指令无歧义、忠实于源 Skill、测试公允)后的 79 任务私有 held-out 基准。 # RL 实验:基础设施 + 极简配方 基础设施(论文明确说“现代 agentic RL 首先是基础设施挑战”):Molt(PyTorch 原生全异步训练,Ray + vLLM + FSDP2)+ Polar(agent rollout 层:rootless Apptainer 沙箱、代理回传 token ID 和采样时 log-prob、prefix merging 把 harness 的多次补全缝合成训练轨迹)。 配方(刻意走“简单路线”):GRPO 组归一优势 + DPPO 的 binary-KL 信任域掩码(δ=0.05,超出阈值的 token 直接丢弃,无需参考模型,还能防训练-推理失配);G=8 rollouts/组,批 64,lr 1e-6 恒定,无 KL 惩罚、无熵奖励、无 SFT 热启动,每任务 65k 上下文。 量规校准奖励:开量规时,额外由 GPT-6 Astra 做 LLM-as-Judge(带“宪法”:惩罚无脑循环、reward hacking、答非所问;hacking 实证 = -5 分),总奖励 r = r_V + λs/5(λ=0.2),即 judge 最多把程序化奖励拉动 ±0.2。量规是校准可执行结果奖励,而非取代它,这是与“Rubrics as Rewards”一系的定位差异。 # 四项发现(论文最有信息量的部分) 发现 1:小规模 RL 即有跨域迁移。 仅 300 步、只用 2,400 任务子集训一个 epoch:S2EBench pass@1 +4.3(均分 +18.5),Terminal-Bench 2.1 +4.7(49.4→54.1)。训练集与 TB 无重叠(13-gram Jaccard < 0.8),且训练集从未针对 TB 调过,论文将其解读为规划、工具使用、收尾能力的通用提升而非任务族记忆。这让 27B 本地模型显著缩小了与云端前沿模型的差距。 发现 2:量规校准 RL 在基准上落后于纯结果 RL,一个诚实的负结果。 量规版在 TB 2.1 只有 50.1(纯结果版 54.1);训练中量规版的程序化奖励长期停在 0.5–0.6,judge 分项从头到尾无上升趋势,两个奖励在训练分布上互相拉扯。论文不把它当作对量规奖励的终审判决(两者优化不同目标,而基准只考结果那一半),并给出两个疑因:λ=0.2 的加性形式让失败任务仍能拿正奖励、judge 看不到文件系统等设定均未调优;以及更本质的,Skill 写下的方法论可能本来就不是最大化基准通过率的分布。 发现 3:行为确实向 Skill 对齐,量规的价值所在。 200 个任务的成对偏好测试(judge 拿源 SKILL.md 当标准,比较匿名化的 base 与 RL 轨迹):纯结果 RL 已被偏好 54.5% vs 33.5%;量规版被偏好 73.0% vs 24.0%。这说明量规奖励买到的东西在结果基准上看不见,但对“怎么做事”影响实质,对网页开发、报告综合、开放研究这类难验证任务尤其重要。 发现 4:GLM-5.3 蒸馏 SFT 反而伤害 Qwen。 在 GLM-5.3 轨迹上做 SFT:27B 上 TB 2.1 掉到 45.8;4B 上直接崩塌(TB 18.7→3.4,出现思维/工具调用死循环)。归因:教师的 interleaved-thinking + 工具调用风格与学生自身 post-training 不兼容,模仿覆盖了学生依赖的行为模式却带不来教师的能力。与 TMax 报告的“SFT 混合数据劣化已后训练的 Qwen”互相印证。因此论文所有 RL 结果都从未修改的原始 checkpoint 出发。
显示更多
Larry Ellison pledged $9.2bn in shares to a new personal loan. Can’t imagine the interest rate is great on them given the climate. Oracle shareholders are not gonna love a margin loan at a stock price that’s already shaky
显示更多
0
12
299
51
转发到社区
ClusterMAX 3.0 is here! ClusterMAX 3.0 debuts with a comprehensive review of the neocloud industry, covering 77 providers. We increase our market view to cover 323 providers, up from 209 in ClusterMAX 2.0, 169 in ClusterMAX 1.0, and 124 in the original AI Neocloud Playbook and Anatomy article. We have now interviewed well over 200 end users of neoclouds as part of this research. We update our itemized list of criteria across 10 categories, and update our direct descriptions of our expectations for Slurm, Kubernetes, Standalone Machines, Monitoring Dashboards, and Health Checks. All of this content is live on our website. We encourage providers to use these lists when developing their offerings. We still consider these lists as an amalgamation of our experience interviewing end users, making them representative of the features that end users expect from their cloud providers. Nebius joins CoreWeave in the Platinum tier. While CoreWeave still sets the technical bar for others to follow, Nebius is now established as a provider that consistently commands a premium pricing over others. Strong business decisions by Nebius have put them in a position to serve an entire class of neolabs at seller’s prices. Google Cloud joins Oracle in the Gold tier. Azure moves to Silver, Fluidstack moves to Unavailable, and Crusoe drops to Bronze. Lambda, Firmus and TensorWave remain in Silver, while GMI moves up to Silver from Bronze. Many companies drop from Silver (or Gold) to Bronze or lower. We raise the bar this round as only 19 neoclouds globally achieve a Medallion rating. We establish a tier between Bronze and Underperforming: the Participation Ribbon tier. 15 providers join this rating, which more accurately describes our opinion that they do the bare minimum to get by.
显示更多
0
31
180
29
转发到社区
🚨SlowMist TI Alert🚨 💸 @BeatXswap Loss: 2,984,557 BTX (~$77,512) 🔍 Root Cause: The `LiquidityVestingConvert` contract calculates BTX quotes via `_calculateQuote()`, which reads `IUniswapV3Pool.slot0()` spot price as the sole oracle. No TWAP protection, no sanity check, no deviation limit. An attacker borrowed 6,000,000 BTX via flash loan, dumped it into the V3 pool to crash `sqrtPriceX96`, then called `deposit()` twice (10,000 + 2,000 USDT), triggering `POSITION_MANAGER.mint()` at the manipulated spot price and draining BTX from LP positions. 📌 Attacker: 0x67B2f08683A735cfE6f6E57fA86909b62218C2a1 📌 Victim: 0x1e647FAADb05f2124BFCcFC003EDc06D1A90bf5D 0x9a7A92240FBAc4030b65A6E61239928d6Bcc716F 📌 Vulnerable Contract: 0x1e647FAADb05f2124BFCcFC003EDc06D1A90bf5D Powered by Tx:
显示更多
聊聊最近起飞的能源电力股 $BE 一周之内暴涨20% 1.标普500纳入预期与AI数据中心缺电逻辑同时爆发! 2. Oracle还是BE的重要客户 Oracle财报会不会成为下一轮催化?面对超过300倍的高估值,现在究竟该跟随趋势,还是警惕利好兑现?这期一次讲清楚。
显示更多
0
407
1K
47
转发到社区
🚨SlowMist TI Alert🚨 💸 @EnsoBuild Loss: ~5.6 ETH 🔍 Root Cause: An oracle price calculation error occurred in the Enso Finance / DPI Strategy Vault. In `Controller.deposit()`, the EnsoOracle's `estimateStrategy()` is called before and after the user tokens are transferred, and shares are minted on the difference: `mint = amountAdded * totalSupply / valueBefore`. The valuation chain — EnsoOracle → ItemEstimator → `ProtocolOracle.consult()` — prices tokens via UniV3 `pool.observe()`, but the registry's `fee` field is reused as `secondsAgo`, producing a near-spot TWAP window. Due to the absence of TWAP consistency checks or price range validation—combined with the fact that the liquidity in the Uniswap v3 pool used for price calculation was inherently imbalanced—an attacker was able to swap 0.683 WETH for 268.42 FARM tokens on Uniswap v2 in a single transaction, while the imbalanced v3 pool incorrectly valued that amount at 6.3 WETH (~ 9.2x). Subsequently, an excessive number of shares were minted at an incorrect price and redeemed for profit. 📌 Attacker: `0x3196398321D77a2511d369DCB6eCa9d2aD87b73A` 📌 Victim (Strategy): `0x890ed1ee6d435a35d11051d9ed97ff457ce53b5942` 📌 Vulnerable contract: Controller impl: `0xd8D22509C1fe47516D8F82A28CFd728111F57Ef1` Oracle: `0xAb7505eB360cE0D63e8E88f7853677EcD5537DC0` Powered by Tx:
显示更多
We have a limited window to strengthen cyber defenses, and together with organizations including @AnthropicAI, @awscloud, @Google, @Microsoft, and @Oracle, we're calling for a global effort to give defenders the tools, resources, and support to protect the infrastructure we all depend on. If we act decisively, we can turn today's AI advances into lasting improvements in security and make our digital world safer for everyone.
显示更多
0
446
8.9K
1.3K
转发到社区
OpenAI数据中心负责人离职,公司称属内部重组 发生了什么: OpenAI数据中心负责人Chris Malone已经离职。他上周已经离开公司,任期大约只有17个月(2025年3月加入)。 Malone之前在Google干了十多年,后来又在Meta做了近5年数据中心相关工作,属于这个领域里比较资深的人。他加入OpenAI的时间点,正好是公司宣布Stargate(星际之门)大规模数据中心计划之后。当时OpenAI正和Oracle、SoftBank等合作,在美国大规模建设算力基础设施,其中德州阿比林(Abilene)的项目比较受关注。 数据中心负责人这个岗位现在特别关键,因为训练和运行大模型需要巨大的电力、土地、芯片和冷却能力,谁负责这块,直接影响公司能不能按时拿到足够算力。 公司内部也做了调整:Malone原本直接向总裁Greg Brockman汇报,后来改成向另一位副总裁汇报,基础设施团队被重新拆分和组织。OpenAI官方回应称,这是为了适应业务规模和速度而做的重组,目前团队依然完整、有明确领导,有能力继续推进计划。同时有报道提到,公司在自建数据中心之外,也在重新考虑更多租赁现成设施的方式。 目前还没有公开说明他具体为什么走,是个人原因、内部调整,还是对战略有不同看法,都还没有官方细节。
显示更多
Proof of Alpha Network: A Reputation Layer for Market Opinions Most social platforms measure attention through followers, impressions, likes, and reposts. These signals show who commands an audience, but they do not answer the question that matters most: Who has demonstrated repeatable market insight? Proof of Alpha Network introduces a reputation layer for public market opinions. It converts unstructured commentary into timestamped, testable records containing the relevant asset, direction, time horizon, and clarity of the original call. Each opinion is then evaluated against what subsequently happened in the market. This creates an inspectable history of the call, its outcome, and the evidence used to assess it. Reputation is not reduced to a single universal score. A source may demonstrate strong insight into Bitcoin while performing poorly in equities. Someone may excel at long-term macro analysis but struggle with short-term market timing. Proof of Alpha therefore builds contextual reputation across assets, time horizons, market regimes, and consistency. These relationships form an opinion graph connecting people, opinions, assets, events, and outcomes. At its core is the Alpha Oracle, an ensemble system in which multiple independent AI models evaluate the same source material. Their weighted agreement determines whether an opinion record can be committed automatically, requires further refinement, or should be excluded when the evidence remains too ambiguous. The resulting network continuously updates through the evolving loop: Opinion → Outcome → Reputation → Trust → Conviction The same structured records can also serve AI agents. By converting fragmented market discourse into a shared format, Proof of Alpha allows information to move across an agent network without every agent repeatedly interpreting the original content. Proof of Alpha doesn’t replace judgment or predict future returns. Instead, it makes the historical evidence behind market credibility visible. Markets have always priced assets. Proof of Alpha begins pricing credibility.
显示更多
0
5
139
25
转发到社区
Data oracles provide external data to Newton’s policy engine for evaluation. Integrations such as @redstone_defi can provide asset prices and other financial data used by Newton policies when evaluating transactions.
显示更多