deepseek v4 is terrible at hallucinations...
>deepseek v4 pro scores 94%.
>deepseek v4 flash is even worse at 96%.
for comparison:
>glm-5.2: 28%
>mimo v2.5 pro: 25%
>minimax m3: 16%
lower is better.
this means deepseek often answers confidently even when it should admit it doesn’t know.
use glm-5.2 for execution...
but never trust the output without tests, verification or a second model reviewing the work!!!
deepseek v4 is coming in 2 weeks...
5 things we know:
1. v4 is expected to launch officially in mid-july
2. there will be pro and flash versions
3. deepseek is introducing peak and off-peak api pricing
4. prices will double during 7 peak hours per day
5. deepseek promises feature optimizations and performance improvements
5 things that are still speculation:
1. native vision and multimodal input
2. a new checkpoint rather than the current preview model
3. engram memory being included
4. a major intelligence jump toward glm-5.2 or kimi k2.7
5. the price increase being caused by new hardware and backend infrastructure
the most likely outcome?
a more stable, faster and production-ready version of v4 preview...
possibly with new capabilities unlocked, but not necessarily a dramatically smarter model.
deepseek v4 is coming in 2 weeks...
5 things we know:
1. v4 is expected to launch officially in mid-july
2. there will be pro and flash versions
3. deepseek is introducing peak and off-peak api pricing
4. prices will double during 7 peak hours per day
5. deepseek promises feature optimizations and performance improvements
5 things that are still speculation:
1. native vision and multimodal input
2. a new checkpoint rather than the current preview model
3. engram memory being included
4. a major intelligence jump toward glm-5.2 or kimi k2.7
5. the price increase being caused by new hardware and backend infrastructure
the most likely outcome?
a more stable, faster and production-ready version of v4 preview...
possibly with new capabilities unlocked, but not necessarily a dramatically smarter model.