Can robot foundation models scale?
Xiaomi-Robotics-1 explores this question with over 100,000 hours of real-world manipulation data.
Pre-trained on large-scale real-world trajectories and post-trained with cross-embodiment robot data, Xiaomi-Robotics-1 demonstrates consistent scaling across data and model size, strong generalization in unseen environments, and efficient adaptation to new tasks.
🔗
#
Robotics# #
EmbodiedAI# #
FoundationModels# #
RobotLearning# #
XiaomiAI# #
XiaomiRobotics#
显示更多
Video foundation models (e.g., Seedance 1.0 → 1.5 → 2.0; V-JEPA 1.0 → 2.0 → 2.1) keep getting stronger, while public descriptions of how their training data is built keep getting shorter.
⚒️ We built VidaForge, an open-source, five-stage video data pipeline that turns raw video collections into training-ready datasets.
In an academic lab, studying video data recipes begins with a lot of tedious engineering: handling broken videos and transcode failures, keeping large jobs resumable, tracking every clip, and packaging the result for training.
Across ingestion, segmentation, selection, annotation, and training dataset packaging, data moves through:
raw videos
→ standardized videos
→ clips
→ curated clips
→ annotated clips
→ training datasets
At every step, we can inspect what happened to each video or clip. This lets us see which samples changed under a data recipe and trace a training dataset back through the pipeline.
VidaForge currently connects processed data to two video foundation model pretraining paths: Wan video generation through NeMo-AutoModel, and self-supervised video representation learning through the official V-JEPA2 repository.
We ran VidaForge end to end on 200K videos, producing over 700K clips. From these clips, we built Selected-200K, Mixed-200K, and Rejected-200K datasets for Wan2.1-1.3B and V-JEPA2.1-1B pretraining.
In these early runs, the three datasets produced different training behavior: data selection appeared in Wan2.1-1.3B eval curves and V-JEPA2.1-1B training stability.
Pipeline, experiments, open data, and project resources in the thread below ↓
显示更多
Apple runs its own AI models, its own silicon, its own map platform. And it still needs a guy in a red polo walking Malaysian sidewalks with a $50,000 LIDAR backpack strapped to his shoulders.
Five states in Malaysia. Pedestrian-only zones where the Apple Maps cars can't drive. Areas around markets, courtyards, mosques where the roads stop but the sidewalks keep going. Somebody has to physically walk it.
What's actually on the backpack. A 360 spherical camera on the top pole. A LIDAR unit for depth. GPS at the base. Combined cost of the rig runs around $50,000 in sensors. Contractors get hired through outsourced recruiting sites, not the Apple careers page. Apple doesn't want its logo on this part.
This is the layer at the bottom of the AI stack that nobody writes threads about. Foundation models don't wake up knowing Kuala Lumpur's alleys.
Somebody feeds them the alleys. That somebody is a guy in a red polo stacking bag as an Apple contractor while everyone else fights over the newest Claude release.
There's a whole layer of AI-adjacent jobs that pay quietly while the discourse stays loud somewhere else. Pedestrian mapping. RLHF review. Data annotation. Voice recording for TTS training. And one more that most of AI-Twitter would refuse to admit even exists.
显示更多
Wan-Streamer v0.1
End-to-end Real-time Interactive Foundation Models
Anthropic 更新了 API 文档,提供了 Foundation Models 的 SDK。
感觉中文圈因为没法用上 Siri AI ,所以对他讨论不多。
事实上, Apple 在 WWDC 2026 给出的 AI 方案,协调了 - iPhone 硬件销售需求(根基)
- 模型厂商需求(上游)
- 第三方 App 开发者
- 消费者体验
的利益,让大家都能在 Apple 生态中获取自己所需。
Apple 的给的答卷,很正确,很务实,也很 John Ternus。
显示更多
GPT-3 level robot foundation models already exist
It’s just not obvious because GPT-3 was only partially successful most of the time. That’s ok for some knowledge work but not for the work we need robots to do
GPT-5 level intelligence for robots we could see within a year because research compounds
We won’t necessarily have something that looks exactly like the ChatGPT moment for robotics, but the FigureAI livestream was close to that for a lot of people
显示更多
Nicolai Tangen, CEO of Norges Bank Investment Management pressed IBM CEO Arvind Krishna directly on whether AI is a bubble (Save this).
And Krishna responded with what has become known inside financial circles as the $8 trillion math problem.
A single gigawatt of AI data center capacity filled with accelerators, liquid cooling, and power infrastructure costs roughly $60 to $80 billion to build and populate.
The industry has committed to more than 100 gigawatts of buildout globally.
That is $6 to $8 trillion in capital expenditure and because AI grade hardware depreciates on a five-year cycle, that entire sum must be effectively replaced and refreshed every five years.
To service the interest on $8 trillion in capital at a conservative 10% borrowing rate, the AI ecosystem would need to generate approximately $800 billion in annual profit, a number that currently exceeds the combined net income of every large technology company in the world.
Goldman Sachs estimates $7.6 trillion in aggregate AI CapEx between 2026 and 2031 alone, and Reuters Breakingviews has flagged that even if the capital is available, physical bottlenecks power permits, land, cooling infrastructure, and electrical grid connections mean that half of the planned data center projects are being cancelled or delayed before they ever go live.
Krishna also raised a second, structurally distinct concern that markets have largely ignored.
He argued that the largest foundation models, GPT, Gemini, Claude, Llama are converging toward commodity status.
When a product is a commodity, switching costs collapse.
When switching costs collapse, pricing power evaporates and margins compress regardless of how much capital was spent building the capability.
Morningstar's equity research team conducted a review of 132 technology companies in 2026 and found that AI had caused moat rating downgrades across roughly 40 major stocks concentrated in enterprise software, IT services, and SaaS with Adobe, Salesforce, Workday, and ADP among the companies whose competitive moats have materially weakened.
The implication is that the companies spending the most on AI model development may be building an asset that is simultaneously the most expensive to produce and the most difficult to monetize with durable margins.
This bear case is serious but it is also incomplete and that is what makes Krishna's framing so important to understand precisely.
When pressed further, Krishna explicitly said he does not believe there is an AI bubble in the technology itself only in a subset of the infrastructure capital that is being deployed against speculative assumptions rather than proven demand.
He draws the same analogy, the fiber optic overbuild of the late 1990s. Dozens of companies went bankrupt laying cable that nobody was using.
And yet that exact "wasted" infrastructure became the physical backbone of every cloud company, every streaming service, every mobile network, and every modern AI training cluster that followed.
The builders lost, the infrastructure won.
And the companies that were built on top of it, Amazon, Google, Netflix, Salesforce compounded for two decades.
The question, as Krishna framed it, is not whether AI is real.
It is which capital deployment earns a return versus which gets stranded and crucially, whether you own the stranded assets or the companies built on top of them.
On winners, Krishna was direct that distribution is the moat on the consumer side, and enterprise is wide open.
The data supports this, Meta with 3.3 billion daily active users across Facebook, Instagram, and WhatsApp is building AI into a distribution network that no startup can replicate at any cost.
Meanwhile, the productivity evidence arriving in real time is beginning to challenge the bear case's revenue projections.
Jensen Huang just showed on stage at Computex that GitHub commits, the universal measure of global software output nearly tripled in the first months of 2026, effectively converting $3 trillion in developer salaries into $9 trillion in productive output.
That is measurable, real time economic value already flowing through the system and it feeds directly back into token demand in a compounding loop that Krishna's static CapEx math does not fully capture.
显示更多
CNBC: Meta starts major cuts with 8,000 layoffs as AI shakes the tech giant.
Along with the layoffs, around 7,000 employees will be shifted into new AI-focused roles.
Meta is not only trimming costs, it is changing its internal shape around AI infrastructure, foundation models, and AI monetization, which means the company wants more people building the systems that train models, the models themselves, and the products that turn those models into revenue.
---
cnbc .com/2026/05/20/meta-layoffs-zuckerberg-says-success-isnt-a-given-in-memo.html
显示更多
Interesting position paper on agentic AI as a foreseeable pathway to AGI.
(bookmark it)
There has been strong debate on whether a larger single model get us there or a multi-agent system.
The authors argue that agentic AI systems, not bigger foundation models on their own, are the most foreseeable route to AGI.
Formalizes what "agentic" actually contributes beyond the base model: memory, reasoning, tool use, self-improvement, alignment.
Each is a separable axis with its own bottlenecks (long-horizon coherence, credit assignment, safety auditing).
They argues that none of those bottlenecks get solved by another order of magnitude on pretraining compute.
Paper:
Learn to build effective AI agents in our academy:
显示更多
Step into a 1,000m² AI universe.
Registration for Qwen Conference 2026 is officially open! Join us at Singapore’s Sands Expo on May 26 to experience the New Era in person. From foundation models to hands-on AI coding—this is your journey from code to impact.
显示更多