Deepseek-V4-Pro-0813 doesn't seem as insanely impressive as I expected.
This seems to be a pattern with Chinese AI labs:
smaller models like Qwen 27b, Deepseek-V4-Flash, and GLM-5.2 perform ridiculously well for their size,
but maybe due to a lack of training compute, their larger models feel a bit unaligned.
As they secure more compute, they will definitely get better over time.
显示更多
Chinese AI models have one massive advantage:
Their training data is mostly in Chinese, and a single Chinese character packs significantly more meaning than an English letter.
This means they can compress a lot more data per token.
It gives them up to a 4x token compression efficiency—an inherent advantage that English-based US models simply cannot replicate.
This is exactly why GLM, a mere 750B model, can compete with 2T-level frontier models.
显示更多
Flexing a Rolls-Royce and a Patek Philippe is such an old trend.
In 2026, do this instead:
Today, I bought a total of 2TB of VRAM.
Today I bought a Patek Philippe 42mm aquanaut stainless to buy more bitcoin.
Follow me for more investment advice.