📣 We are presenting 6 main conference papers 🚀and 14 workshop papers (including 🏆2 Best Papers🏆) at #
ICML2026# in Korea! Also hosting one of the largest workshops, Trustworthy AI for Good, on July 10th 🌍❤️.
We push the frontiers on #
AISafety#, #
MultiAgent#, and #
CausalReasoning# at
@JinesisLab! 🎉
Huge congratulations to all collaborators and co-authors. Excited to discuss these projects in Seoul! Feel free to reach out and talk to our 20+ members and collaborators in Korea
@ZhijingJin,
@_AndreiMuresanu,
@iarthsingh,
@ChanglingXavier,
@davidguzman1120,
@EmanuelTewolde,
@ettogran,
@FurkanDanismann,
@Jerick1380,
@PepijnCobben,
@rishit_dagli,
@_rfaulk,
@SimkoSamuel,
@TerryJCZhang,
@vantru0ng,
@zhxiao03,
@x_angelohuang,
@yahang_qi,
@ozzaney0101,
@_yongjinny.
Happy for collaboration on any of the above topics 🤝
EuroSafeAI, University of Toronto, ETH Zürich, Max Planck Institute for Intelligent Systems
Main conference spotlight
🌟 Safe Models Do Not Guarantee Safe Societies: The Case for Sociopolitical Risk
Main conference posters
📌 CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research?
📌 Training with Honeypots: Reshaping How LLMs Fail
📌 CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas
📌 Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution
📌 LLM for Physics Research Requires Domain-Specialized Training and Tooling
Workshop best papers
🏆When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games
BEST PAPER
@NExT-Game Workshop
🏆Transferability for General Reasoning: An Automated Curriculum for Multi-Domain LLM RL
BEST PAPER
@RLxF Workshop
Workshop oral and spotlight
🎤 AF-ARENA: A Multi-Dimensional Evaluation Suite for Alignment Faking — AIWILD
🌟 Multi-Agent AI Systems Need Institutional Design, Not Just Model-Level Alignment — AI4GOOD
Workshop papers
📄 The Wedge Questions: Latent Cultural Boundaries in LLMs via Persona Projection Divergence — Pluralistic Alignment
📄 GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory — NExT-Game
📄 Stargazer: A Scalable Model-fitting Benchmark Environment for AI Agents under Astrophysical Constraints — AI4Physics
📄 Test of Time: Rethinking Temporal Signal of Benchmark Contamination — FoGen
📄 CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas — AI4GOOD
📄 Weight-Level Defenses Improve LLM Agent Adversarial Robustness — AI4GOOD
📄 Evaluating Cooperation in LLM Social Groups through Elected Leadership — AI4GOOD
📄 Causal AI Scientist: Towards End-to-End Causal Inference with Large Language Models — AI4Research
📄What Game-Theoretic Benchmarks Miss: Strategic Silence in Multi-Agent LLMs — FAGEN
📄Proving Your Way to Cooperation: Formalizing Proof-Based Open Source Game Theory in Lean — AI4Math