DeepSeek, the Hangzhou-based lab whose V3 release shipped at a fraction of frontier training costs, announced DeepSeek-V4.1-Flash on X on September 10, calling it the smallest entry in a freshly designed architecture family. The model adds native visual understanding, meaning it handles images without a separate vision module bolted on.

DeepSeek describes the underlying design as asymmetric, a structure the company says delivers more intelligence at lower cost by pairing a smaller KV cache with the rest of the architecture. The company frames that smaller cache as the direct source of inference savings for users running the model at volume, though it did not publish benchmark figures, latency numbers, or pricing to support the comparison.

V4.1-Flash is live now on the DeepSeek API, and the company retired both V4-Flash and V4-Flash-Vision-Exp in the same move, consolidating its lower-cost tier into a single multimodal release.

For teams evaluating cheap inference tiers, the retirement of two prior models signals DeepSeek is standardizing its Flash line rather than running parallel vision and text variants, worth tracking once independent throughput or cost benchmarks surface.

DeepSeek announced the release in a six-post thread on X on September 10, 2026.