注册并分享邀请链接,可获得视频播放与邀请奖励。

Chubby♨️ 的个人资料封面
Chubby♨️ 的头像

Chubby♨️ (@kimmonismus)

@kimmonismus
0 正在关注    0 粉丝
I'm going to cancel Claude. It's just so bad, I can't believe it. It's just lazy. The most recent example: I have Claude check my inbox for important emails, summarize them, work with them, and send out replies if necessary. I caught Claude again simply not reading the email thread to the end and just ignoring the latest emails. When I asked him about it, Opus 5 just said: "Valid point. I didn't read it." I mean, seriously. What the heck? You have to babysit it every time.
显示更多
0
762
6.6K
221
转发到社区
The most disappointing launch-day in one image.
0
128
2K
42
转发到社区
OpenAI reveals Codex Micro (their first hardware product, so to say) OpenAI’s $230 Codex Micro is a compact control deck for agentic coding, with RGB status keys, shortcuts for common Codex actions, and a dial for adjusting reasoning effort. Built with Work Louder, it works on Mac and Windows and is designed to make managing multiple agents feel faster, more tactile, and less dependent on constantly switching between chats. It’s a bit gimmicky, but one thing is clear: OpenAI works within and with the community, developing cool products that are fun and capture the zeitgeist.
显示更多
0
152
2.2K
107
转发到社区
A few thoughts on the very near future First of all, what had previously been little more than a rumor has now been confirmed: GPT-5.6 had already been fully trained for two months and was available to selected users in early access. The obvious question is why it was not rolled out earlier. I do not think this was because OpenAI feared that the model might be overshadowed by Fable 5 or Mythos 5. Instead, OpenAI likely began working with government and regulatory authorities at a very early stage to ensure that the model could be released at all. Even after it had been previewed and announced, it still took some time before it could be rolled out publicly. That said, OpenAI clearly handled the rollout far better than Anthropic, which apparently did not have the same level of cooperation with government and regulatory authorities. Conversely, however, this also clearly means that future delays and increasingly strict model reviews will probably force us to wait longer for official releases. The next widely discussed rumor is that, within a few weeks, most likely no more than six, we will see either a preview or even the release of GPT-6. (Andrew Curran @AndrewCurran_ is one of the most reliable sources here on X, so I think that's very realistic.) The model has undergone entirely new pretraining, and the pace of releases is accelerating. The numbers are clear: Frontier labs are releasing more and better models at an increasingly rapid pace. Whereas we once had to wait months, quarters, or even half a year for major new releases, they are now arriving almost weekly. The latest frontier models may be more efficient in terms of intelligence per token, but they are also being deployed with much larger reasoning budgets. In practice, models such as Fable 5 and GPT-5.6 often consume considerably more tokens during complex or agentic tasks. This is not necessarily a sign of declining efficiency. Rather, it suggests that improvements in efficiency are being reinvested into deeper reasoning, longer trajectories and more capable agentic behavior. The result is that total compute consumption per task can continue to rise even as the underlying models become more efficient. Fable 5 and GPT 5.6 demonstrate just how intensive token usage has become. Although Sam Altman explicitly stated that GPT-5.6 is 54% more token-efficient (via CNBC), the fact remains that compute demand continues to increase, requiring more powerful and efficient computing infrastructure. Inference chips will probably become even more important as well. In summary, my initial conclusion from the latest releases is that compute demand will not merely continue to grow, but will probably exceed the available supply. This naturally means that energy demand will also increase, and, based on my initial assessment, probably more sharply than previously expected. This is likely to remain the largest bottleneck in the very near future. And this is important to me: there are bottlenecks. Not the training of the models, but besides compute, above all energy. This needs to be taken seriously! The US power grid, for example, is a major bottleneck, and the obvious question is how the necessary expansion can be achieved. Capital expenditure on data centers in the United States continues to rise sharply. This year, it exceeds 800 billion. It is not yet clear what the situation will look like in 2027, but I can hardly imagine investment declining or less CapEx being required. The reason lies precisely in the developments already mentioned: Demand is growing, particularly demand for energy. China clearly has an advantage here, a genuine moat, and I believe the West must be extremely careful not to fall behind because of the energy advantage China already possesses in practice. This could also help explain why, according to a recent Reuters report, China is considering restricting Western access to its frontier models. It may have concluded that it will win the long-term race. Unless there is a genuine breakthrough, whether in small modular nuclear reactors or fusion energy, I expect major problems to emerge over the coming years, for example by 2030. So far, I do not see any viable solutions. We can therefore clearly establish two points: Models are becoming larger, better, and increasingly useful for all users. There is no end to this development in sight. At the same time, the bottleneck appears to be growing increasingly severe, and this is already visible in practice. Regulation, energy demand, and compute demand could mean that, in the very near future, the release cadence will not accelerate as quickly as hoped or desired. This creates a clear contradiction. Thank you for coming to my TED Talk.
显示更多
0
50
387
20
转发到社区
It just keeps accelerating: Cursor is reportedly building a general-purpose AI agent to compete with Anthropic’s Claude Cowork, pushing beyond coding into everyday office work. Internally called Sand, the agent is designed to respond to emails and texts, organize spreadsheets and handle engineering tasks. Cursor reportedly rolled it out internally in late June, after it began leasing compute from SpaceXAI in April. A public launch is not guaranteed, particularly with SpaceX’s planned $60 billion acquisition potentially changing Cursor’s roadmap. Sand would be Cursor’s first product built for casual business users rather than developers.
显示更多
0
17
181
7
转发到社区
The best use case for SunoAI. I hope you listen carefully, @AnthropicAI .
0
34
870
34
转发到社区
Fable 5 probably running locally in about two years. That is the projection in this r/LocalLLaMA chart. It tracks how long it takes for cloud-frontier capability to become broadly comparable in laptop-runnable open-weight models. The observed average lag: ~24.8 months. GPT-3-class capability: 37 months. GPT-3.5-class: 17 months. GPT-4-class: ~24 months. The projection puts Fable / Mythos 5-class capability on high-end consumer hardware around July 2028.
显示更多
0
171
2.5K
165
转发到社区
"From a machine learning persepective from an AI perspective, there is no reason why we cant make a context-window of 100m" Let's see what you've got, Anthropic.
0
48
494
21
转发到社区
A Principal Engineer at NVIDIA and OpenAI technical staff are both extremely hyped for GPT-5.6 Sol. Both with pre access. Just two more days, friends.
0
84
1.7K
70
转发到社区
Anthropic's business case needs to be studied. At the end of 2025 and the beginning of 2026, there was an incredible increase in usage in the business/enterprise sector, which made them number 1.
显示更多
0
144
1.5K
97
转发到社区
Supposedly, "a new model from" from zAI is said to be at least as strong as Fable5 in cybersecurity-related aspects. I did some research and only came across a Wall Street Journal article, which, however, does not refer to a new model, but to GLM 5.2 as a relatively new model that was released recently. So either GLM 5.2 is stronger than people think, or the news being circulated is misleading. According to WSJ, Zhipu AI’s GLM-5.2 can match top US models in some bug-finding scenarios, and China’s 360 Security says its new Tulongfeng tool is comparable to Anthropic’s Mythos.
显示更多
0
55
612
47
转发到社区
Anthropic claims: Alibaba continues to distill Claude on a large scale to train Qwen. Via Bloomberg Anthropic is accusing Alibaba-linked operators of running a massive campaign to illicitly access Claude through nearly 25,000 fraudulent accounts. According to Bloomberg, Anthropic claims the campaign generated 28.8 million Claude exchanges between April and June, targeting capabilities like software engineering and agentic reasoning. The company says this is part of a broader pattern of “adversarial distillation,” where Chinese labs allegedly harvest outputs from US frontier models to train rival systems at a fraction of the cost. Lets see how good Qwen 3.8 will be, probably FABLEous good.
显示更多
0
114
936
85
转发到社区
Someone on Reddit built a WoW private server with 1,800 bots and AI chat via the DeepSeek API. Dead Internet Theory, but playable. An MMORPG with no real players, yet somehow it still feels human.
显示更多
0
559
15.5K
738
转发到社区
AI labs are built on foreign talent. Thats a fact. Now the US is reportedly testing restrictions on "foreign persons" accessing frontier models. "The Trump administration appears to have targeted only Anthropic so far, warning the company on Friday in a letter from Commerce Secretary Howard Lutnick that it would need a license to make its latest models available to “foreign persons,” including its own employees. But Anthropic’s biggest rival, OpenAI, has flagged its concerns about the issue." (The Information) 38% of researchers publishing at leading AI conferences in 2024 got their undergrad education in China, per MacroPolo estimates cited by The Information. If US policy starts restricting model access by nationality, frontier labs are suddenly in a very difficult situation. That is why the issue surrounding Anthropic and Fable 5 is such a significant development, and that is why so much depends on the decisions made in the next few days.
显示更多
0
15
238
28
转发到社区
This came as a surprise: Microsoft has unveiled handheld and desktop devices designed to control one's agents. It reminds me of what I had expected from OpenAI’s hardware-standalone devices for controlling agents.
显示更多
It is interesting how much focus is being placed on data centers and the community. Recently, there were numerous reports regarding resistance to data center expansion; now comes the promise from Microsoft: no increase in electricity costs due to data centers, along with resource conservation.
显示更多
0
94
1.3K
119
转发到社区
It is interesting how much focus is being placed on data centers and the community. Recently, there were numerous reports regarding resistance to data center expansion; now comes the promise from Microsoft: no increase in electricity costs due to data centers, along with resource conservation.
显示更多
0
9
177
12
转发到社区
Gemini 3.2 spotted in Gemini! If we already receive Gemini 3.2 Flash now, the major release will probably be reserved for I/O. h/t @Waguri_Kaoruko8 for finding
0
51
1.4K
61
转发到社区