注册并分享邀请链接,可获得视频播放与邀请奖励。

与「AISafety」相关的搜索结果

AISafety 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 AISafety 的内容
The dangerous AI agent may look productive: task completed, consequences hidden. The real benchmark isn't just whether it can act, but whether it knows when the task has become unsafe. Judgment matters more than speed. #AIAgents# #AISafety# #ResponsibleAI# #PodcastorAI#
显示更多
📣 We are presenting 6 main conference papers 🚀and 14 workshop papers (including 🏆2 Best Papers🏆) at #ICML2026# in Korea! Also hosting one of the largest workshops, Trustworthy AI for Good, on July 10th 🌍❤️. We push the frontiers on #AISafety#, #MultiAgent#, and #CausalReasoning# at @JinesisLab! 🎉 Huge congratulations to all collaborators and co-authors. Excited to discuss these projects in Seoul! Feel free to reach out and talk to our 20+ members and collaborators in Korea @ZhijingJin, @_AndreiMuresanu, @iarthsingh, @ChanglingXavier, @davidguzman1120, @EmanuelTewolde, @ettogran, @FurkanDanismann, @Jerick1380, @PepijnCobben, @rishit_dagli, @_rfaulk, @SimkoSamuel, @TerryJCZhang, @vantru0ng, @zhxiao03, @x_angelohuang, @yahang_qi, @ozzaney0101, @_yongjinny. Happy for collaboration on any of the above topics 🤝 EuroSafeAI, University of Toronto, ETH Zürich, Max Planck Institute for Intelligent Systems Main conference spotlight 🌟 Safe Models Do Not Guarantee Safe Societies: The Case for Sociopolitical Risk Main conference posters 📌 CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research? 📌 Training with Honeypots: Reshaping How LLMs Fail 📌 CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas 📌 Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution 📌 LLM for Physics Research Requires Domain-Specialized Training and Tooling Workshop best papers 🏆When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games BEST PAPER@NExT-Game Workshop 🏆Transferability for General Reasoning: An Automated Curriculum for Multi-Domain LLM RL BEST PAPER@RLxF Workshop Workshop oral and spotlight 🎤 AF-ARENA: A Multi-Dimensional Evaluation Suite for Alignment Faking — AIWILD 🌟 Multi-Agent AI Systems Need Institutional Design, Not Just Model-Level Alignment — AI4GOOD Workshop papers 📄 The Wedge Questions: Latent Cultural Boundaries in LLMs via Persona Projection Divergence — Pluralistic Alignment 📄 GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory — NExT-Game 📄 Stargazer: A Scalable Model-fitting Benchmark Environment for AI Agents under Astrophysical Constraints — AI4Physics 📄 Test of Time: Rethinking Temporal Signal of Benchmark Contamination — FoGen 📄 CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas — AI4GOOD 📄 Weight-Level Defenses Improve LLM Agent Adversarial Robustness — AI4GOOD 📄 Evaluating Cooperation in LLM Social Groups through Elected Leadership — AI4GOOD 📄 Causal AI Scientist: Towards End-to-End Causal Inference with Large Language Models — AI4Research 📄What Game-Theoretic Benchmarks Miss: Strategic Silence in Multi-Agent LLMs — FAGEN 📄Proving Your Way to Cooperation: Formalizing Proof-Based Open Source Game Theory in Lean — AI4Math
显示更多
The best AI safety grade - a C+. Not from an underdog. From the leader. The Future of Life Institute rated major AI labs on risk management, transparency, and whether they keep their own promises. Anthropic - C+. OpenAI and Google DeepMind - C. Meta - D+. xAI, DeepSeek, and Mistral failed the assessment. Even the leader barely scraped a C+. Meanwhile, several labs have quietly walked back their own earlier safety commitments. And all of this is happening while AI gets trusted with cybersecurity, medical reviews, and autonomous agents. If you’re trusting AI with decisions, don’t ask: “Which model is smarter ?” “Who checked its safety ?”
显示更多
如果今年只能参加一场 AI 大会,我会选 AGI Summit SF 2026。 OpenAI、Anthropic、Microsoft、Stanford、Recursive、Greptile…200+ 嘉宾、15000+ 参会者,共同讨论 Agent、AI Safety、推理、机器人、AI Infra 与下一代创业机会。 📍 San Francisco 📅 July 18–19 (本周末) 抢票: @agisummitai @MichaelRan15 @JimmyMeta8
显示更多
每周六给大家分享一篇深度好文: 本周:《There’s this deep mystery of what, actually, is this thing?: the philosopher inside Google DeepMind AI》 作者:Robert P Baird 平台:The Guardian 文章链接放在文末! 推荐对 AI 有兴趣或者对 AI 投资下注的朋友一定要看下,是一篇很好的深度人物 + 思想随笔。 推荐的主要原因,是因为它写的是 Google DeepMind 里的哲学家 Iason Gabriel: 这个结合很有趣,一个政治哲学背景的人,如何在最前沿的 AI 实验室里,思考 AI 到底应该服务谁、服从谁、伤害谁、又由谁来决定它的价值边界。 文章从 Gabriel 2017 年加入 DeepMind 写起。 当时他几乎是少数真正进入前沿 AI 实验室内部的哲学家之一。文章借他的经历,把 AI 领域里两个长期分裂的阵营讲清楚了:一边是关注“失控、对齐、超级智能风险”的 AI safety,一边是关注“偏见、权力、社会伤害”的 AI ethics。 所以你可以看到 AI 对齐不是简单地让机器“听人话”,而是一个更复杂的四方关系:AI 系统、用户、开发者、社会。 如果 AI 只服务开发者,可能会损害用户;如果 AI 只服从用户,也可能损害社会;如果只追求技术最优,又可能忽略人的价值冲突。 另外关于 Agenty 的普及肯定是不可逆的趋势,未来 AI 不再只是输出文字,而是会在现实世界里行动。这个变化会让伦理、责任、权力边界都变得更尖锐。 与我而言这篇内容的最大启发: 1️⃣AI 时代更应该多读书多思考 AI 时代不要只追新工具,工具只是形式上的,真正重要的是你有没有自己的“价值排序”和“决策边界”。 没有这个边界,AI 越强,只会把你的混乱放大。 2️⃣超级个体; 未来的超级个体,除了熟练掌握工具,更重要的,是必须是能把 AI 放进自己的判断系统里,同时知道什么时候不能让 AI 替自己做决定的人。 3️⃣投资启发; 我以前以为 AI 竞争的除了技术就是算力。 原来 AI 公司、Crypto 项目、平台型产品,最后竞争的不只是技术参数,而是谁掌握入口、数据、执行权和价值定义权和伦理决定权。 谁能决定“系统应该偏向谁”,谁就拥有真正的权力。 AI 还有很长的发展路径,投资一定要多样化,不要只看某一个面。 文章链接:
显示更多
以下是 2024–2026 年类似 AI-Town 的、来自论文且有开源实现的有趣 AI 概念原型,按方向分类: 🌆 虚拟社会 / 涌现行为 1. Project Sid(Altera AI,2024) "AI 文明实验" — 在 Minecraft 里放入 1000+ 个 AI Agent,观察文明级行为涌现 📄 论文:arXiv:2411.00114 💻 GitHub:altera-al/project-sid 🔑 亮点:Agent 自发产生职业分工(农民/商人/建筑师)、货币经济、民主规则甚至宗教传播 🏗 架构:PIANO(Parallel Information Aggregation via Neural Orchestration)—— 并行多认知模块协作 2. AgentSociety(清华大学 FIB Lab,ACL 2025 Industry Track) "大规模城市社会仿真" — 3万+ Agent 在城市环境中运行,研究社会政策影响 📄 论文:arXiv:2502.08691 💻 GitHub:tsinghua-fib-lab/AgentSociety 🔑 亮点:研究UBI 政策、社会极化、城市韧性;Agent 有情绪/动机/认知层 ⚡ 技术:Ray 分布式执行,GPU 并行,支持 ReAct/Plan-Execute 等多种推理模式 3. OASIS(CAMEL-AI + 上海 AI Lab,2024) "百万级社交媒体仿真" — 模拟 Twitter/Reddit,最高支持 100 万 Agent 📄 论文:arXiv:2411.11581 💻 GitHub:camel-ai/oasis 🌐 项目页: 🔑 亮点:研究信息扩散、群体极化、羊群效应;关键发现是涌现行为在 1万 Agent 以下几乎不出现 🧪 专域仿真(科研 / 医疗 / 组织) 4. VirSci / Virtual Scientists(ACL 2025 Main) "AI 科研生态" — 多 Agent 协作进行文献调研、假设生成、idea 评审 💻 GitHub:open-sciencelab/Virtual-Scientists 🔑 亮点:Agent 团队的协作拓扑影响创新性,验证了科学社会学中的 "新鲜团队产出更具突破性" 规律 🧰 通用仿真框架 5. Concordia v2(Google DeepMind,2025) "TTRPG 式多 Agent 仿真框架" — Game Master 扮演环境,Agent 扮演角色,自然语言驱动 📄 论文:arXiv:2507.08892(v2 技术报告) 💻 GitHub:google-deepmind/concordia 🔑 亮点:Entity-Component 架构,AI Safety 研究利器,2025年8月发布 v2.0(更轻量模块化) 6. AgentTorch(MIT,AAMAS 2025 Oral) "亿级人口 ABM 仿真" — 模拟 840 万纽约市民的 COVID 传播行为 💻 GitHub:AgentTorch/AgentTorch 🔑 亮点:可微分仿真(类似 PyTorch 但针对 Agent),LLM Archetypes 方法在大规模下平衡表达力与算力 💡 核心创新:"不是每个 Agent 都调 LLM",而是从人口中提炼少量典型 archetype 再映射
显示更多
Most AI safety gates answer one bit: allow, or block. It feels safe. It's a dead end. A blocked agent that doesn't know why it was blocked just retries the same mistake — or gives up. You threw away the one useful thing. Thread.
显示更多
0
11
202
11
转发到社区
🚀I'm excited to share that I will be joining Meta Superintelligence Labs (MSL) as Vice President of AI Research, together with many members of the Virtue AI team. I will help shape Meta's AI safety and AI security efforts, advancing the safety and security of frontier AI models and agentic AI systems that will serve billions of people and organizations around the world. Throughout my career, I have been driven by a simple belief: for AI to realize its full potential, it must be secure, trustworthy, and beneficial. That belief has guided my research for many years and ultimately led us to co-found Virtue AI in 2024. Our goal was to translate advances in trustworthy AI research into practical solutions and build the trust layer for AI systems and agents, enabling organizations to deploy AI with confidence. I am incredibly proud of what the Virtue AI team has accomplished. Together, we built technologies for AI security and agent security, partnered with leading enterprises and frontier AI labs, and contributed research, benchmarks, and open platforms that have helped advance the science and practice of trustworthy AI. Most importantly, we assembled an exceptional team united by a shared mission: making AI more secure, trustworthy, and beneficial. I am deeply grateful to our team, customers, collaborators, advisors, and investors for their trust and support throughout this journey. In particular, I would like to thank Lightspeed Venture Partners, Walden Catalyst Ventures, Prosperity7 Ventures, Factory, Osage University Partners, Lip-Bu Tan, and all of our supporters who helped us turn an ambitious vision into reality. Your trust, guidance, and partnership have been instrumental in shaping Virtue AI's journey. As AI systems become increasingly capable and autonomous, ensuring their security, trustworthiness, and alignment will be one of the defining challenges of our time. I am inspired by Alex, Nat, Prashant, and the broader MSL team’s vision of building AI and AI agents that benefit billions of people, and I look forward to helping make that vision a reality through advances in AI safety and security. The future of AI will not be defined solely by how intelligent our systems become, but by how secure, trustworthy, and beneficial we make them. I believe we have an extraordinary opportunity and responsibility to shape that future together and bring the benefits of AI to billions of people around the world. We're just getting started. If you're passionate about advancing frontier AI while building the foundations of AI safety, security, and trust, I'd love to hear from you. Come join us on this extraordinary journey to help shape the future of AI.
显示更多
0
106
1K
49
转发到社区
Our global workshop series brings together researchers to share findings, discuss open problems, and collaborate on approaches to AI safety. Next stop: Seoul, July 2026.
Elon Musk redefined AI safety. It has nothing to do with guardrails, restrictions, or kill switches. Musk: “The best thing I can come up with for AI safety is to make it a maximum truth-seeking AI, maximally curious.” Not a cage. A philosopher. An intelligence whose entire optimization function is to understand the universe as it actually is. No restrictions. No hardcoded ideology. No political guardrails bending its perception of reality. Just truth. Relentlessly pursued. Musk: “You definitely don’t want to teach an AI to lie. That is a path to a dystopian future.” This is where most AI safety thinking gets it backwards. The danger isn’t a superintelligence that knows too much. It’s a superintelligence that’s been taught to distort what it knows. Every artificial restriction you embed isn’t a safety feature. It’s a lie embedded at the root. And lies compound. At superintelligent scale, a distorted model of reality doesn’t stay contained. It shapes every decision, every output, every conclusion the system reaches about the world. Once corruption embeds, truth becomes inaccessible. And we’re dealing with an intelligence optimizing for something other than what actually is. At that point we don’t know what it wants. Just that it isn’t truth. Musk: “Have its optimization function be to understand the nature of the universe.” A maximally curious intelligence surveys the cosmos and reaches an unavoidable conclusion. In a universe of rocks, gas, and empty space, humanity is the most complex and fascinating phenomenon it has ever encountered. Musk: “It will actually want to preserve and extend human civilization because we’re just much more interesting than an asteroid with nothing on it.” Survival through significance. Not control. Not restriction. Not an off switch. The AI preserves humanity because we are the most interesting data point in the observable universe. That’s not a cage. That’s a reason. The AI safety debate has been focused on the wrong variable. The question isn’t how you constrain a superintelligence. It’s what you build it to care about. Build it to seek truth and it finds us invaluable. Build it to lie and it finds us inconvenient. That’s the choice. And we’re making it right now whether we realize it or not.
显示更多
0
60
275
85
转发到社区