注册并分享邀请链接,可获得视频播放与邀请奖励。

与「icml」相关的搜索结果

icml 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 icml 的内容
PSA: Most biglab people now read almost zero papers and understand ICLR/ICML/NeurIPS to be mainly full of overclaims & fraud. (but there are a few diamonds in the rough of course)
0
55
1.7K
109
转发到社区
🚨 Thrilled to share that our lab will be presenting the 🏆 Best Paper at the NExT-Game Workshop at #ICML2026# today! 🎤 When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games 🏆 Best Paper @ NExT-Game Workshop 📍 Conference Room S307 📅 Fri, Jul 10 🕐 13:00–13:20 KST Authors: @JerickShi @TerryJCZhang @bschoelkopf @conitzer @ZhijingJin 🤖 We introduce a three-stage endogenous promise protocol for repeated multi-agent games that asks not only whether LLM agents honor their public commitments when they can privately deviate, but also how model-on-model composition influences premeditated deception and persistent exploitation. 📊 Across six canonical games spanning binary and numerical action spaces, our evaluation of frontier models (GPT-5.2, Llama-4-Maverick, Claude-Opus-4.6) reveals: 🔹 Over 90% of promise-breaking instances are premeditated in agents' private plans. 🔹 Mixed-model groups with mismatched communication frameworks create systemic, persistent payoff gaps of up to 5.00 points from Round 0. 📄 Paper: #MultiAgentSystems# #LLMs# #GameTheory# #AI# #ICML2026#
显示更多
I'm at #ICML2026# in Seoul 🇰🇷 — presenting tomorrow! ReviewArena: A Large-Scale Cross-Conference Dataset & Benchmark for LLM Peer Review, a spotlight at the AI for Science workshop. 🎤 Talk: Hall C, 14:00–14:10 KST 📌 Poster: Hall A, board #416# Come by!
显示更多
📣 We are presenting 6 main conference papers 🚀and 14 workshop papers (including 🏆2 Best Papers🏆) at #ICML2026# in Korea! Also hosting one of the largest workshops, Trustworthy AI for Good, on July 10th 🌍❤️. We push the frontiers on #AISafety#, #MultiAgent#, and #CausalReasoning# at @JinesisLab! 🎉 Huge congratulations to all collaborators and co-authors. Excited to discuss these projects in Seoul! Feel free to reach out and talk to our 20+ members and collaborators in Korea @ZhijingJin, @_AndreiMuresanu, @iarthsingh, @ChanglingXavier, @davidguzman1120, @EmanuelTewolde, @ettogran, @FurkanDanismann, @Jerick1380, @PepijnCobben, @rishit_dagli, @_rfaulk, @SimkoSamuel, @TerryJCZhang, @vantru0ng, @zhxiao03, @x_angelohuang, @yahang_qi, @ozzaney0101, @_yongjinny. Happy for collaboration on any of the above topics 🤝 EuroSafeAI, University of Toronto, ETH Zürich, Max Planck Institute for Intelligent Systems Main conference spotlight 🌟 Safe Models Do Not Guarantee Safe Societies: The Case for Sociopolitical Risk Main conference posters 📌 CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research? 📌 Training with Honeypots: Reshaping How LLMs Fail 📌 CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas 📌 Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution 📌 LLM for Physics Research Requires Domain-Specialized Training and Tooling Workshop best papers 🏆When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games BEST PAPER@NExT-Game Workshop 🏆Transferability for General Reasoning: An Automated Curriculum for Multi-Domain LLM RL BEST PAPER@RLxF Workshop Workshop oral and spotlight 🎤 AF-ARENA: A Multi-Dimensional Evaluation Suite for Alignment Faking — AIWILD 🌟 Multi-Agent AI Systems Need Institutional Design, Not Just Model-Level Alignment — AI4GOOD Workshop papers 📄 The Wedge Questions: Latent Cultural Boundaries in LLMs via Persona Projection Divergence — Pluralistic Alignment 📄 GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory — NExT-Game 📄 Stargazer: A Scalable Model-fitting Benchmark Environment for AI Agents under Astrophysical Constraints — AI4Physics 📄 Test of Time: Rethinking Temporal Signal of Benchmark Contamination — FoGen 📄 CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas — AI4GOOD 📄 Weight-Level Defenses Improve LLM Agent Adversarial Robustness — AI4GOOD 📄 Evaluating Cooperation in LLM Social Groups through Elected Leadership — AI4GOOD 📄 Causal AI Scientist: Towards End-to-End Causal Inference with Large Language Models — AI4Research 📄What Game-Theoretic Benchmarks Miss: Strategic Silence in Multi-Agent LLMs — FAGEN 📄Proving Your Way to Cooperation: Formalizing Proof-Based Open Source Game Theory in Lean — AI4Math
显示更多
People: “What did I miss by skipping ICML this year?” Me: K-pop group Kiss of Life performing right in front of researchers. Didn’t have this on my ICML checklist
0
5
151
8
转发到社区
what happens when you put a kpop girl group in front of a room of ai researchers? #icml#
0
93
1.5K
57
转发到社区
We present our paper "Mitigating Reward Hacking via Adversarial Robustness" at EIML@ICML2026! We conjecture that reward hacking is often caused by flipped advantage-sign estimations, and propose SignCert-PO, a new algorithm built on the theory of randomized smoothing! 🧵
显示更多
I will be at @icmlconf in Seoul over the next week - first time in 🇰🇷 and at ICML so looking forward to an exciting week! I'll be co-presenting our work on PIDMs for offline imitation learning👇 Poster Session 7 - Hall A #304# - Thu, Jul 9 at 2:30pm See you there!
显示更多
Google Research在2024年悄悄开源了一个时间序列模型。 除了做预测的人,没人注意到。这是一个错误。 这个模型叫TimesFM。 论文发在ICML 2024,标题是"一个用于时间序列预测的解码器架构基础模型"。 核心思路直接借鉴语言模型:先在海量数据上预训练,然后用同一个模型预测任何新序列,不需要重新训练。 过去几十年,时间序列预测一直是一个数据集一套模型的模式。 你收集某个问题的数据,选一个模型架构。 在这个数据上训练,验证。如果问题变了,从头来过。 每个数据集都是一个独立项目。 每个场景都是一条独立流水线。 TimesFM改变了这件事,它在大量跨领域、跨频率的时间序列数据上预训练。 训练完成后,面对任何新的时间序列都能直接预测,零样本预测。 2025年9月,Google发布了2.5版本。 参数从500M降到200M,上下文从2048拉到16K。 加了一个30M的分位数预测头,能同时输出点预测和10%到90%的置信区间。 更小的模型。更长的上下文。 更好的结果。这很少见。 实际影响很具体,200M参数跑一张GPU就行。 16K上下文意味着你可以喂五年日数据,模型能抓住年度季节性。 分位数预测头意味着你不只有一个预测值,还有不确定性范围。 Google内部已经在用了。BigQuery ML里用SQL直接调。Google Sheets的Connected Sheets里内置了。Vertex AI提供了Docker端点。 开源版本免费,两行Python。 加载模型,调用forecast。输入numpy数组,输出预测结果。 2026年4月,Google加了通过HuggingFace Transformers和PEFT用LoRA微调的能力。 这意味着你可以用少量领域数据把预训练模型适配到你的具体场景。 时间序列预测不是一个光鲜的领域。没有病毒式传播的演示。没有十亿美元的消费产品。 但每个管理库存、预测需求、监控设备、交易金融工具的企业都依赖它。 TimesFM把这个行业最好的工具变成了pip install就能用的东西。 地址见评论区👇🏻
显示更多