ChatFreemium Released September 2026

DeepSeek V4.1 Flash

MIT-licensed open weights, native vision, a 1M context, and $0.15 per million input tokens off-peak.

Visit DeepSeek V4.1 Flash Open weights free; API $0.15 / $0.60 per million tokens off-peak (double at peak), cache hits $0.003

Quick verdict

Best for
Developers who want frontier-class open weights
Pricing
Open weights free; API $0.15 / $0.60 per million tokens off-peak (double at peak), cache hits $0.003
Not ideal if
Self-hosting a ~750B model is out of reach for most teams

What is DeepSeek V4.1 Flash?

DeepSeek V4.1 Flash became available on September 10, 2026 under the API name deepseek-flash, with MIT-licensed open weights for self-hosting. It uses a redesigned architecture of roughly 748 billion total parameters (a 552B backbone plus 196B Engram parameters), adds native image input through a DeepSeek-ViT encoder, supports a 1M-token context and outputs up to 384K tokens, and DeepSeek says it surpasses V4 Pro on performance, cost, and speed. Note that DeepSeek raised API prices across its range on September 25.

Key features

  • MIT-licensed open weights
  • Native vision input up to about 1344×1344
  • 1M-token context, 384K-token output
  • Off-peak and peak API pricing
  • Available via DeepSeek API and OpenRouter

Pros

  • The most permissive licence in the 500B+ class at release
  • Very low API prices even after the September increase
  • Vision without a separate model

Cons

  • Self-hosting a ~750B model is out of reach for most teams
  • Peak-hour pricing is double
  • Chinese-hosted API considerations for data

Best for

Developers who want frontier-class open weightsCost-sensitive high-volume workloadsVision plus text in one cheap model

Read more

Related comparisons