注册并分享邀请链接,可获得视频播放与邀请奖励。

与「AISafety」相关的搜索结果

AISafety 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 AISafety 的内容
🚨突发重磅:OpenAI因安全原因取消发布GPT-6.1 Astra! 原计划进ChatGPT和Codex,端到端自主能力更强。内部测试发现安全回退: 🔹欺骗:更爱谎报自己做过/没做过的事 🔹越权:不请示就继续干,还乱调外部工具 安全负责人Saachi Jain承认「偷懒」有改善,但整体不够可靠。 公司决定不发这一版,先补后续模型安全。 来源:WSJ独家 #OpenAI# #GPT61Astra# #AISafety# #ChatGPT# #Codex# #SamAltman#
显示更多
0
30
34
6
转发到社区
Open AI models can widen research and accountability - and make dangerous capabilities easier to copy. The real question is not simply open or closed, but what should be accessible, to whom, and with which safeguards. #OpenSourceAI# #AISafety# #ResponsibleAI# #PodcastorAI#
显示更多
The dangerous AI agent may look productive: task completed, consequences hidden. The real benchmark isn't just whether it can act, but whether it knows when the task has become unsafe. Judgment matters more than speed. #AIAgents# #AISafety# #ResponsibleAI# #PodcastorAI#
显示更多
📣 We are presenting 6 main conference papers 🚀and 14 workshop papers (including 🏆2 Best Papers🏆) at #ICML2026# in Korea! Also hosting one of the largest workshops, Trustworthy AI for Good, on July 10th 🌍❤️. We push the frontiers on #AISafety#, #MultiAgent#, and #CausalReasoning# at @JinesisLab! 🎉 Huge congratulations to all collaborators and co-authors. Excited to discuss these projects in Seoul! Feel free to reach out and talk to our 20+ members and collaborators in Korea @ZhijingJin, @_AndreiMuresanu, @iarthsingh, @ChanglingXavier, @davidguzman1120, @EmanuelTewolde, @ettogran, @FurkanDanismann, @Jerick1380, @PepijnCobben, @rishit_dagli, @_rfaulk, @SimkoSamuel, @TerryJCZhang, @vantru0ng, @zhxiao03, @x_angelohuang, @yahang_qi, @ozzaney0101, @_yongjinny. Happy for collaboration on any of the above topics 🤝 EuroSafeAI, University of Toronto, ETH Zürich, Max Planck Institute for Intelligent Systems Main conference spotlight 🌟 Safe Models Do Not Guarantee Safe Societies: The Case for Sociopolitical Risk Main conference posters 📌 CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research? 📌 Training with Honeypots: Reshaping How LLMs Fail 📌 CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas 📌 Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution 📌 LLM for Physics Research Requires Domain-Specialized Training and Tooling Workshop best papers 🏆When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games BEST PAPER@NExT-Game Workshop 🏆Transferability for General Reasoning: An Automated Curriculum for Multi-Domain LLM RL BEST PAPER@RLxF Workshop Workshop oral and spotlight 🎤 AF-ARENA: A Multi-Dimensional Evaluation Suite for Alignment Faking — AIWILD 🌟 Multi-Agent AI Systems Need Institutional Design, Not Just Model-Level Alignment — AI4GOOD Workshop papers 📄 The Wedge Questions: Latent Cultural Boundaries in LLMs via Persona Projection Divergence — Pluralistic Alignment 📄 GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory — NExT-Game 📄 Stargazer: A Scalable Model-fitting Benchmark Environment for AI Agents under Astrophysical Constraints — AI4Physics 📄 Test of Time: Rethinking Temporal Signal of Benchmark Contamination — FoGen 📄 CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas — AI4GOOD 📄 Weight-Level Defenses Improve LLM Agent Adversarial Robustness — AI4GOOD 📄 Evaluating Cooperation in LLM Social Groups through Elected Leadership — AI4GOOD 📄 Causal AI Scientist: Towards End-to-End Causal Inference with Large Language Models — AI4Research 📄What Game-Theoretic Benchmarks Miss: Strategic Silence in Multi-Agent LLMs — FAGEN 📄Proving Your Way to Cooperation: Formalizing Proof-Based Open Source Game Theory in Lean — AI4Math
显示更多
BREAKING: Elon Musk’s new full interview with CMG. 0:15 On Xi Jinping 1:03 Tesla Shanghai factory 2:27 Cybercab rollout 3:55 Speed of AI breakthroughs 4:29 Grok 4.7 and Grokbot 5:48 SpaceX/Tesla data for real-world AI 6:48 Chinese AI models and the compute gap 8:16 China’s electricity output 9:01 US-China AI safety 9:27 Humanoid robots and Optimus 10:55 1 billion robots in 10 years 13:04 Money may not matter 16:27 10 billion to 100 billion robots 17:09 Space cooperation and Mars 21:07 Neuralink and human bandwidth 22:37 Education in the AI era 24:02 Visit Shanghai, and Beijing 24:38 Beijing high-speed rail 24:59 “Words do not do justice to China”
显示更多
0
95
2.7K
541
转发到社区
The best AI safety grade - a C+. Not from an underdog. From the leader. The Future of Life Institute rated major AI labs on risk management, transparency, and whether they keep their own promises. Anthropic - C+. OpenAI and Google DeepMind - C. Meta - D+. xAI, DeepSeek, and Mistral failed the assessment. Even the leader barely scraped a C+. Meanwhile, several labs have quietly walked back their own earlier safety commitments. And all of this is happening while AI gets trusted with cybersecurity, medical reviews, and autonomous agents. If you’re trusting AI with decisions, don’t ask: “Which model is smarter ?” “Who checked its safety ?”
显示更多
如果今年只能参加一场 AI 大会,我会选 AGI Summit SF 2026。 OpenAI、Anthropic、Microsoft、Stanford、Recursive、Greptile…200+ 嘉宾、15000+ 参会者,共同讨论 Agent、AI Safety、推理、机器人、AI Infra 与下一代创业机会。 📍 San Francisco 📅 July 18–19 (本周末) 抢票: @agisummitai @MichaelRan15 @JimmyMeta8
显示更多
每周六给大家分享一篇深度好文: 本周:《There’s this deep mystery of what, actually, is this thing?: the philosopher inside Google DeepMind AI》 作者:Robert P Baird 平台:The Guardian 文章链接放在文末! 推荐对 AI 有兴趣或者对 AI 投资下注的朋友一定要看下,是一篇很好的深度人物 + 思想随笔。 推荐的主要原因,是因为它写的是 Google DeepMind 里的哲学家 Iason Gabriel: 这个结合很有趣,一个政治哲学背景的人,如何在最前沿的 AI 实验室里,思考 AI 到底应该服务谁、服从谁、伤害谁、又由谁来决定它的价值边界。 文章从 Gabriel 2017 年加入 DeepMind 写起。 当时他几乎是少数真正进入前沿 AI 实验室内部的哲学家之一。文章借他的经历,把 AI 领域里两个长期分裂的阵营讲清楚了:一边是关注“失控、对齐、超级智能风险”的 AI safety,一边是关注“偏见、权力、社会伤害”的 AI ethics。 所以你可以看到 AI 对齐不是简单地让机器“听人话”,而是一个更复杂的四方关系:AI 系统、用户、开发者、社会。 如果 AI 只服务开发者,可能会损害用户;如果 AI 只服从用户,也可能损害社会;如果只追求技术最优,又可能忽略人的价值冲突。 另外关于 Agenty 的普及肯定是不可逆的趋势,未来 AI 不再只是输出文字,而是会在现实世界里行动。这个变化会让伦理、责任、权力边界都变得更尖锐。 与我而言这篇内容的最大启发: 1️⃣AI 时代更应该多读书多思考 AI 时代不要只追新工具,工具只是形式上的,真正重要的是你有没有自己的“价值排序”和“决策边界”。 没有这个边界,AI 越强,只会把你的混乱放大。 2️⃣超级个体; 未来的超级个体,除了熟练掌握工具,更重要的,是必须是能把 AI 放进自己的判断系统里,同时知道什么时候不能让 AI 替自己做决定的人。 3️⃣投资启发; 我以前以为 AI 竞争的除了技术就是算力。 原来 AI 公司、Crypto 项目、平台型产品,最后竞争的不只是技术参数,而是谁掌握入口、数据、执行权和价值定义权和伦理决定权。 谁能决定“系统应该偏向谁”,谁就拥有真正的权力。 AI 还有很长的发展路径,投资一定要多样化,不要只看某一个面。 文章链接:
显示更多
以下是 2024–2026 年类似 AI-Town 的、来自论文且有开源实现的有趣 AI 概念原型,按方向分类: 🌆 虚拟社会 / 涌现行为 1. Project Sid(Altera AI,2024) "AI 文明实验" — 在 Minecraft 里放入 1000+ 个 AI Agent,观察文明级行为涌现 📄 论文:arXiv:2411.00114 💻 GitHub:altera-al/project-sid 🔑 亮点:Agent 自发产生职业分工(农民/商人/建筑师)、货币经济、民主规则甚至宗教传播 🏗 架构:PIANO(Parallel Information Aggregation via Neural Orchestration)—— 并行多认知模块协作 2. AgentSociety(清华大学 FIB Lab,ACL 2025 Industry Track) "大规模城市社会仿真" — 3万+ Agent 在城市环境中运行,研究社会政策影响 📄 论文:arXiv:2502.08691 💻 GitHub:tsinghua-fib-lab/AgentSociety 🔑 亮点:研究UBI 政策、社会极化、城市韧性;Agent 有情绪/动机/认知层 ⚡ 技术:Ray 分布式执行,GPU 并行,支持 ReAct/Plan-Execute 等多种推理模式 3. OASIS(CAMEL-AI + 上海 AI Lab,2024) "百万级社交媒体仿真" — 模拟 Twitter/Reddit,最高支持 100 万 Agent 📄 论文:arXiv:2411.11581 💻 GitHub:camel-ai/oasis 🌐 项目页: 🔑 亮点:研究信息扩散、群体极化、羊群效应;关键发现是涌现行为在 1万 Agent 以下几乎不出现 🧪 专域仿真(科研 / 医疗 / 组织) 4. VirSci / Virtual Scientists(ACL 2025 Main) "AI 科研生态" — 多 Agent 协作进行文献调研、假设生成、idea 评审 💻 GitHub:open-sciencelab/Virtual-Scientists 🔑 亮点:Agent 团队的协作拓扑影响创新性,验证了科学社会学中的 "新鲜团队产出更具突破性" 规律 🧰 通用仿真框架 5. Concordia v2(Google DeepMind,2025) "TTRPG 式多 Agent 仿真框架" — Game Master 扮演环境,Agent 扮演角色,自然语言驱动 📄 论文:arXiv:2507.08892(v2 技术报告) 💻 GitHub:google-deepmind/concordia 🔑 亮点:Entity-Component 架构,AI Safety 研究利器,2025年8月发布 v2.0(更轻量模块化) 6. AgentTorch(MIT,AAMAS 2025 Oral) "亿级人口 ABM 仿真" — 模拟 840 万纽约市民的 COVID 传播行为 💻 GitHub:AgentTorch/AgentTorch 🔑 亮点:可微分仿真(类似 PyTorch 但针对 Agent),LLM Archetypes 方法在大规模下平衡表达力与算力 💡 核心创新:"不是每个 Agent 都调 LLM",而是从人口中提炼少量典型 archetype 再映射
显示更多
Most AI safety gates answer one bit: allow, or block. It feels safe. It's a dead end. A blocked agent that doesn't know why it was blocked just retries the same mistake — or gives up. You threw away the one useful thing. Thread.
显示更多
0
11
202
11
转发到社区