DeepSeek-V4-Flash-0731 is Ollama's fastest growing model ever in token usage. We are scaling capacity in US & Europe.
On Ollama, this model runs with high performance (100tps+) and zero data retention. Your data stays yours.
ollama run deepseek-v4-flash:0731-cloud
Kimi K3 is now available on Ollama’s cloud.
To use it with Claude Code, run:
ollama launch claude --model kimi-k3:cloud
Currently Kimi K3 requires a Pro or Max subscription, and consumes extra usage credits. We’re quickly working on adding capacity to expand access.
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights:
Tech report:
Tech blog:
Ollama is proud to sign @satyanadella's letter.
Our mission from day one has been to make open models accessible to every developer to unlock the next frontier in America and across the globe.
Open-weight models are essential to a healthy AI ecosystem. Together with others across our industry, we are outlining a path for open-weight models to strengthen American competitiveness and expand economic opportunity, while protecting national security.