With frontier models getting more capable by the month 🚀, everyone is talking about Recursive Self-Improvement (RSI). But how do we actually measure this progress? 🤔
Excited to present RSI-Exam — a benchmark testing whether AI agents can improve an existing method through autonomous, long-horizon experimentation, and, importantly, whether those improvements generalize to hidden data.
Across 88 executable research tasks spanning 6 domains, the trend is clear: Claude and GPT form the first tier, with a clear gap over the rest of the field. Yet there is still huge room for improvement.
More importantly, the trajectories tell both stories: agents that discover fundamentally better methods — and agents that spend hours rigorously optimizing the wrong idea.
Check it out:
显示更多
ELON MUSK: "If you've got humanoid robots that have very high dexterity and are incredibly smart, it means that everyone on earth will have access to better medical care, than the richest person on Earth.
I had to have like a neck surgery three times because the first two ones were done wrong. Back pain may be one of the things that it eliminates. Average happiness level for humans would just go upstream tremendously."
显示更多
Back pain sucks; AI could provide the answer to it.
Elon: "So, if you've got humanoid robots that have very high dexterity and are incredibly smart, it means that everyone on Earth will have access to better medical care than the richest person on Earth, which I had to have neck surgery three times because the first two ones were done wrong. You know, can AI please solve back pain? That would be a huge win, and I think it will."
显示更多
Warren Buffett: "The bottom 2% in terms of income in the United States, the bottom 5%, and for sure the top 1% all live better than John D. Rockefeller was living when I was six years old."
"John D. Rockefeller was the richest man in the world and, today, you can get better medicine, better education, better entertainment, better transportation. You can do everything better than he could."
"When I was born, the dentist didn't use novocaine!"
显示更多
There will be no AI jobpocalypse.
The story that AI will lead to massive unemployment is stoking unnecessary fear. AI — like any other technology — does affect jobs, but telling overblown stories of large-scale unemployment is irresponsible and damaging. Let’s put a stop to it.
I’ve expressed skepticism about the jobpocalypse in previous posts. I’m glad to see that the popular press is now pushing back on this narrative. The image below features some recent headlines.
Software engineering is the sector most affected by AI tools, as coding agents race ahead. Yet hiring of software engineers remains strong! So while there are examples of AI taking away jobs, the trends strongly suggest the net job creation is vastly greater than the job destruction — just like earlier waves of technology. Further, despite all the exciting progress in AI, the U.S. unemployment rate remains a healthy 4.3%.
Why is the AI jobpocalypse narrative so popular? For one thing, frontier AI labs have a strong incentive to tell stories that make AI technology sound more powerful. At their most extreme, they promote science-fiction scenarios of AI “taking over” and causing human extinction. If a technology can replace many employees, surely that technology must be very valuable!
Also, a lot of SaaS software companies charge around $100-$1000 per user/year. But if an AI company can replace an employee who makes $100,000 — or make them 50% more productive — then charging even $10,000 starts to look reasonable. By anchoring not to typical SaaS prices but to salaries of employees, AI companies can charge a lot more.
Additionally, businesses have a strong incentive to talk about layoffs as if they were caused by AI. After all, talking about how they’re using AI to be far more productive with fewer staff makes them look smart. This is a better message than admitting they overhired during the pandemic when capital was abundant due to low interest rates and a massive government financial stimulus.
To be clear, I recognize that AI is causing a lot of people’s work to change. This is hard. This is stressful. (And to some, it can be fun.) I empathize with everyone affected. At the same time, this is very different from predicting a collapse of the job market.
Societies are capable of telling themselves stories for years that have little basis in reality and lead to poor society-wide decision making. For example, fears over nuclear plant safety led to under-investment in nuclear power. Fears of the “population bomb” in the 1960s led countries to implement harsh policies to reduce their populations. And worries about dietary fat led governments to promote unhealthy high-sugar diets for decades.
Now that mainstream media is openly skeptical about the jobpocalypse, I hope these stories will start to lose their teeth (much like fears of AI-driven human extinction have).
Contrary to the predictions of an AI jobpocalypse, I predict the opposite: There will be an AI jobapalooza! AI will lead to a lot more good AI engineering jobs, and I’m also optimistic about the future of the overall job market. What AI engineers do will be different from traditional software engineering, and many of these jobs will be in businesses other than traditional large employers of developers. In non-AI roles, too, the skills needed will change because of AI. That makes this a good time to encourage more people to become proficient in AI, and make sure they’re ready for the different but plentiful jobs of the future!
[Original text in The Batch newsletter.]
显示更多
GPT-5.5 Instant is starting to roll out to everyone in ChatGPT.
Much more concise. Better memory. More personalized.
And it's way easier to talk to. Really.