Deepseek-V4-Pro-0813 doesn't seem as insanely impressive as I expected.
This seems to be a pattern with Chinese AI labs:
smaller models like Qwen 27b, Deepseek-V4-Flash, and GLM-5.2 perform ridiculously well for their size,
but maybe due to a lack of training compute, their larger models feel a bit unaligned.
As they secure more compute, they will definitely get better over time.
显示更多