Quick answer
DeepSeek V4.1 Flash became available on September 10, 2026 under the API name deepseek-flash, with MIT-licensed open weights. The architecture is roughly 748 billion total parameters: a 552B backbone plus 196B "Engram" parameters. It adds native image input through a DeepSeek-ViT encoder at resolutions up to about 1344×1344, supports a 1M-token context with outputs up to 384K tokens, and DeepSeek says it beats its own V4 Pro on performance, cost, and speed. API pricing is $0.15 per million input tokens and $0.60 per million output off-peak, double at peak, with cache hits at $0.003. Weights are published for self-hosting; the model is also on OpenRouter. DeepSeek raised prices across its API on September 25, after the launch.
The MIT licence is the headline. Most large open models ship with custom licences that restrict regions, user counts, or revenue. V4.1 Flash has none of that in the 500B-plus class, which is why it matters even to teams that will never download a 750-billion-parameter checkpoint.
What is new
- MIT licence: no revenue clauses, no regional restrictions, use it in products freely
- Native vision: images alongside text in one model, no separate vision pipeline
- 1M-token context and 384K-token output, long enough for whole codebases and long agent runs
- Redesigned architecture with a 552B backbone and 196B Engram parameters, DeepSeek's term for a memory-style component
- Off-peak pricing: half price outside China's peak hours, which falls in many Western working days
What it costs now
At $0.15 input / $0.60 output per million tokens off-peak, V4.1 Flash is more than ten times cheaper than the $2/$10 tier that Sonnet 5.5, GPT-6.1 Sol, and Gemini 4 Argon now share. The September 25 increase of 2.3 to 4.5 times across DeepSeek's API applies to the range, so check the current rate card rather than launch coverage. Self-hosting is free in licence terms and expensive in hardware: a model this size needs multiple high-end GPUs even with quantisation, so most teams will use the API or a host such as OpenRouter.
Who it suits
- High-volume workloads where cost per token dominates
- Products that need a permissive licence for an open model
- Vision plus text at a price no Western vendor matches
- Not a fit where data residency or Chinese-hosted services are excluded by policy, unless self-hosted or via a Western host
Bottom line
V4.1 Flash is the most permissive frontier-class open model available and among the cheapest to call. Use the API or a host for volume work, and keep the licence in mind the next time a vendor tells you open weights come with conditions.



