注册并分享邀请链接,可获得视频播放与邀请奖励。

与「WorldModel」相关的搜索结果

WorldModel 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 WorldModel 的内容
HappyOyster Adventure Mode just turned this world into the ultimate race track🏎️💨 #HappyOyster# #HappyOysterAI# #WorldModel# #AlibabaATH# #Adventure#
做大世界模型研究的人很需要这个工具:stable-worldmodel 它做的就是这一块。这个仓库提供可复现的 world model 研究与评测平台,文档、测试、PyPI 包和论文入口都在一处,适合想把实验、对比和结果复核放到同一套流程里的人。 GitHub 现在约 1.5k stars,今天新增 318 stars。 仓库地址:
显示更多
0
26
23
1
转发到社区
Introducing Khora, our first multiplayer world model. The demo is LIVE. Join an 8-player real-time deathmatch in one shared world now! Built by Ophilus in collaboration with @RhOS_AI , Khora is designed to scale toward Interactive Worlds with unlimited players, with more generated maps and richer gameplay interactions to come. Special thanks to @alibaba_cloud and FITCLOUD for infra and MaaS support, and to @DeemosTech for technical input.
显示更多
0
28
132
50
转发到社区
Meet the first Sound World Model from Sonilo. Generated video can look incredible, but without the right sound, it still feels distant. Sonilo takes a video and creates the music and sound effects it needs, timed to the scene, motion, mood, and environment. And in head-to-head benchmarks, Sonilo outperforms leading models across both music and sound effects. With Sound Effects 1.0, we’re bringing music and sound effects together into one sound layer for videos. Music carries emotion. Sound effects create presence. Together, they make worlds feel real, immersive, and alive! Sound Effects 1.0 is live. #SoundWorldModel# #Sonilo# #AIAudio#
显示更多
0
66
168
63
转发到社区
Introducing Cosmos 3 Edge: our open frontier world model built to run on-device. Cosmos 3 Edge helps robots learn and act, autonomous vehicles understand road scenes and predict intent, and vision AI agents reason across live video for smart infrastructure. With 4B parameters and a 2B Nemotron-based reasoner, you can run it on DGX Spark, NVIDIA Jetson, and more.
显示更多
0
38
1.3K
135
转发到社区
最近 Anthropic 招募了一篇顶尖人才!我们可以从他们招了谁、进了哪个组,反推出 Anthropic 未来 12–18 个月的下注重心。人事就是路线图,我让 Indigo-mind Agent 从我的知识库收录中按信号强弱排了一下👀 一、核心引擎:用 AI 造 AI + 下一代预训练/效率 信号最强!Karpathy 进预训练组、专门“用 Claude 加速 Claude 的预训练研究”;Jelani Nelson(流式算法/降维、大规模高效算法)也进同一条预训练线。两个顶级人都压在预训练+效率上,说明: • Anthropic 判断预训练远没撞墙(和 Yann Dubois 笔记一致),下一波增益在效率 + 把研究循环自动化(RSI)。 • Nelson 的专长(大规模高效算法、降维)= 压榨每一分算力;Karpathy = 让 AI 接管研究 loop。合起来就是"递归自我改进的复利飞轮"从口号变建制。 → 这是它的主引擎,其它都是围着它转。 二、算力/能源基础设施扩张 招 Ross Nordeen(xAI 数据中心整体规划、选址、能源策略、算力扩容)说明 Anthropic 在自建大规模算力 + 能源布局,不再只靠合作方。又一家前沿实验室进入“抢电、抢算力、烧 Capex”的军备赛。 三、AI for Science(尤其生物)- 皇冠上的垂直 Jumper(AlphaFold 诺奖)+ Neklyudov(生成建模 for 蛋白折叠/分子动力学)+ 自建 wet lab + Allen/HHMI 合作 + 收购 Coefficient Bio。这是一整套认真的 AI-for-bio 下注,而且自建湿实验室 = 把"数字→物理"生物闭环补上。和 Demis/DeepMind 抢同一颗明珠,前沿实验室在往生命科学的物理层走。 四、Agentic 检索/记忆/上下文 Bryan McCann(搜索、检索、LM 集成)直接对口把模型连到外部上下文——这正是我的记忆/持续学习那条簇的产品侧(Engram/Karl Mehta/Satya 的 exhaust)。加上 Bailis 的系统/数据库底子,指向 Agentic 检索 + 上下文工程的产品化。 五、企业/产品(可能含 fintech / agentic commerce) Tom Blomfield(Monzo/GoCardless,支付基础设施 + 消费级产品)+ Peter Bailis(Workday CTO,企业软件)。这批是产品与商业化肌肉——尤其 Blomfield 的支付背景,值得留意 Anthropic 是否往 Agentic 支付/商务方向走。 六、软实力长线:对齐 + AI 经济学/治理 • Lederman(哲学家 → 对齐 + 模型"人格 character" + AI 福祉); • Chad Jones(Anthropic Institute,Jack Clark,研究 AI 对经济/社会/法治的系统性影响)。 这不是赚钱线,是“负责任守门人”定位 + 政策影响力——正好和 Demis 那篇"前沿 AI 治理框架"是一套打法:用安全/治理/经济学研究占据话语权。 这份人事表画出的 Anthropic 是——主引擎压在"预训练效率 + RSI"(不追消费/机器人/World model),认真做科学垂直(生物方向),自建算力能源,外加治理/经济学的软实力。明显没重仓的:机器人、World model、消费社交——它在收窄、做深,而不是铺开。
显示更多
最近 Anthropic 招募了一篇顶尖人才!我们可以从他们招了谁、进了哪个组,反推出 Anthropic 未来 12–18 个月的下注重心。人事就是路线图,我让 Indigo-mind Agent 从我的知识库收录中按信号强弱排了一下👀 一、核心引擎:用 AI 造 AI + 下一代预训练/效率 信号最强!Karpathy 进预训练组、专门“用 Claude 加速 Claude 的预训练研究”;Jelani Nelson(流式算法/降维、大规模高效算法)也进同一条预训练线。两个顶级人都压在预训练+效率上,说明: • Anthropic 判断预训练远没撞墙(和 Yann Dubois 笔记一致),下一波增益在效率 + 把研究循环自动化(RSI)。 • Nelson 的专长(大规模高效算法、降维)= 压榨每一分算力;Karpathy = 让 AI 接管研究 loop。合起来就是"递归自我改进的复利飞轮"从口号变建制。 → 这是它的主引擎,其它都是围着它转。 二、算力/能源基础设施扩张 招 Ross Nordeen(xAI 数据中心整体规划、选址、能源策略、算力扩容)说明 Anthropic 在自建大规模算力 + 能源布局,不再只靠合作方。又一家前沿实验室进入“抢电、抢算力、烧 Capex”的军备赛。 三、AI for Science(尤其生物)- 皇冠上的垂直 Jumper(AlphaFold 诺奖)+ Neklyudov(生成建模 for 蛋白折叠/分子动力学)+ 自建 wet lab + Allen/HHMI 合作 + 收购 Coefficient Bio。这是一整套认真的 AI-for-bio 下注,而且自建湿实验室 = 把"数字→物理"生物闭环补上。和 Demis/DeepMind 抢同一颗明珠。对你判断"AI 卖铲人/垂直纵深"是方向信号:前沿实验室在往生命科学的物理层走。 四、Agentic 检索/记忆/上下文 Bryan McCann(搜索、检索、LM 集成)直接对口把模型连到外部上下文——这正是我的记忆/持续学习那条簇的产品侧(Engram/Karl Mehta/Satya 的 exhaust)。加上 Bailis 的系统/数据库底子,指向 Agentic 检索 + 上下文工程的产品化。 五、企业/产品(可能含 fintech / agentic commerce) Tom Blomfield(Monzo/GoCardless,支付基础设施 + 消费级产品)+ Peter Bailis(Workday CTO,企业软件)。这批是产品与商业化肌肉——尤其 Blomfield 的支付背景,值得留意 Anthropic 是否往 Agentic 支付/商务方向走。 六、软实力长线:对齐 + AI 经济学/治理 • Lederman(哲学家 → 对齐 + 模型"人格 character" + AI 福祉); • Chad Jones(Anthropic Institute,Jack Clark,研究 AI 对经济/社会/法治的系统性影响)。 这不是赚钱线,是“负责任守门人”定位 + 政策影响力——正好和 Demis 那篇"前沿 AI 治理框架"是一套打法:用安全/治理/经济学研究占据话语权。 这份人事表画出的 Anthropic 是——主引擎压在"预训练效率 + RSI"(不追消费/机器人/World model),认真做科学垂直(生物方向),自建算力能源,外加治理/经济学的软实力。明显没重仓的:机器人、World model、消费社交——它在收窄、做深,而不是铺开。
显示更多
AI Is Moving Beyond “Generating Videos” — Toward “Generating Worlds” Over the past two years, AI video models have advanced at an astonishing pace. From Runway and Pika to Sora and Veo, AI-generated videos have become increasingly realistic and more consistent with the physical laws of the real world. Many people believe the next objective is simply to generate videos that are longer, sharper, and more lifelike. But if we take a step back, we can see that the real transformation is not happening in video itself. It is happening in world models. What Is a World Model? In 1943, psychologist Kenneth Craik proposed an idea that would influence artificial intelligence research for decades. He argued that the human brain does not merely react to the outside world. Instead, it maintains an internal model of how the world works. Because we have this internal model, we can predict the outcome of an action before we actually take it. Before crossing a road, we estimate whether a car will pass by. Before catching a ball, we predict its trajectory. These abilities come from continuously simulating the world in our minds, rather than relying entirely on trial and error. This idea later became known by a more formal term: World Model. A world model does not describe a single image or a fixed video clip. It is an internal representation capable of continuously simulating the rules and dynamics of the real world. Why Is AI Research Turning Toward World Models? Because predicting “what comes next” is becoming increasingly central to how AI systems work. Language models predict the next token. Image models predict the next step in the denoising process. Video models predict the next frame. A world model, however, attempts to predict something broader: What should the world look like in the next moment? In 2018, David Ha and Jürgen Schmidhuber proposed in their paper World Models that an intelligent agent could first learn a model of the world, and then use that internal model to plan its actions. The Dreamer series later demonstrated that many complex tasks could be learned by training agents inside an “imagined world.” At the same time, the development of video models such as Sora and Veo led researchers to another realization: A model capable of continuously generating video has already learned, at least implicitly, many of the rules governing the real world. As a result, these two research directions have gradually begun to converge. But Video Is Not Yet a World This is where the distinction is often misunderstood. For a world model to support meaningful real-time interaction, it must solve several critical problems. Most video models today are essentially answering one question: What should the next frame look like? A true world model needs to answer much more: What happens if I take one step forward? If I walk behind a building and then return, will the building still be there? If I suddenly change the camera angle, will the entire space remain consistent? If I enter a command such as: “Summon a dragon.” Will the world respond immediately? In other words, a world model must do more than generate content. It must understand space. It must understand time. It must understand causality. And it must understand interaction. Moving from watching to participating is where the real difficulty of world models begins. World Models Are Entering the Interactive Era One of the latest attempts in this direction is Alaya World, recently open-sourced by Alaya World, or @alayastd. Instead of generating a fixed video clip, it generates a world that users can explore in real time. Users can begin with text, an image, or a video, enter the generated scene, move freely through it, and introduce new prompts at any moment during generation. The world responds immediately. According to the publicly released information, Alaya World provides: Real-time streaming generation at 720p and 24 FPS Stable continuous exploration for more than one minute The ability to switch prompts and trigger skills or events during generation Model weights and inference code released under the Apache 2.0 License Training code and datasets planned for future release What makes these capabilities important is not simply the technical specifications. It is that the generated “world” can now support continuous interaction. The official demo shows that users can genuinely control, transform, and explore the generated environment. AI Is Evolving From a Tool Into an Environment Over the past few years, most discussions around AI have focused on content generation. Generating text. Generating images. Generating videos. But world models raise a fundamentally different question: Can AI generate an environment that people can inhabit, explore, and continuously evolve? If the answer is yes, the impact will extend far beyond video generation. Game development, robotics training, embodied intelligence, digital twins, virtual production, and many other fields could be transformed by the development of world models. World models are still at a very early stage. Yet from Craik’s proposal of an internal mental model more than eighty years ago to the emergence of today’s interactive world-generation systems, a clear evolutionary path is beginning to take shape. Perhaps what AI is ultimately learning has never been limited to images, videos, or language. Perhaps it is learning the world itself. References GitHub: Technical Report:
显示更多
This month, Alaya Lab is releasing a series of research projects and open-source initiatives. Today, we're excited to introduce Alaya World - an open-source interactive video world model. Project Page: Github: ✨ Highlights: - 720p 24 FPS streaming generation - Navigation and prompt-driven interactions (e.g., spell casting and summoning) - Stable long-horizon generation (>1 minute) - State-of-the-art performance - Inference code is available today. Training code and datasets are coming soon.
显示更多
0
11
23
3
转发到社区
AI video is becoming playable. Robbyant just released LingBot-World 2.0, also called LingBot-World-Infinity: an open-source interactive world model that can generate a world in real time as you move, act and trigger events. You step in, and it keeps going: 👇
显示更多
0
20
31
12
转发到社区