
DeepSeek V4.1 Flash
MIT-licensed open weights, native vision, a 1M context, and $0.15 per million input tokens off-peak.
Quick verdict
- Best for
- Developers who want frontier-class open weights
- Pricing
- Open weights free; API $0.15 / $0.60 per million tokens off-peak (double at peak), cache hits $0.003
- Not ideal if
- Self-hosting a ~750B model is out of reach for most teams
What is DeepSeek V4.1 Flash?
DeepSeek V4.1 Flash became available on September 10, 2026 under the API name deepseek-flash, with MIT-licensed open weights for self-hosting. It uses a redesigned architecture of roughly 748 billion total parameters (a 552B backbone plus 196B Engram parameters), adds native image input through a DeepSeek-ViT encoder, supports a 1M-token context and outputs up to 384K tokens, and DeepSeek says it surpasses V4 Pro on performance, cost, and speed. Note that DeepSeek raised API prices across its range on September 25.
Key features
- MIT-licensed open weights
- Native vision input up to about 1344×1344
- 1M-token context, 384K-token output
- Off-peak and peak API pricing
- Available via DeepSeek API and OpenRouter
Pros
- The most permissive licence in the 500B+ class at release
- Very low API prices even after the September increase
- Vision without a separate model
Cons
- Self-hosting a ~750B model is out of reach for most teams
- Peak-hour pricing is double
- Chinese-hosted API considerations for data
Best for
Read more
Related comparisons
Alternatives to DeepSeek V4.1 Flash
DeepSeek
Chinese open-source AI rivalling GPT-5 at a fraction of the cost.
GPT-6.1 Sol
OpenAI's efficiency model: near-Astra results at one-fifth of the token price.
Kimi
Moonshot AI's open-weight chat model, known for huge context windows and strong agentic benchmarks.


