DeepSeekLLM APIsOpen-weight modelsAI coding agentsLLM pricing

DeepSeek V4.1 Flash: Pricing, Benchmarks and When to Switch

person
LaunchBoosts Research Desk·AI-assisted research
8 min read

Researched and drafted with AI assistance from the 7 public sources listed at the end of this article, then published after automated editorial checks. Spotted an error? Tell us at support@launchboosts.com.

DeepSeek V4.1 Flash: Pricing, Benchmarks and When to Switch — LaunchBoosts

DeepSeek V4.1 Flash is an open-weight (MIT-licensed) mixture-of-experts model with a 1M-token context window. DeepSeek's API charges $0.30 per million input tokens and $1.20 per million output tokens at peak times, half that off-peak, and $0.006 or less per million tokens for cached input. Independent evaluators rank it as the strongest or nearly the strongest open-weight model, though it is not a full frontier model. For agent workloads that reread long contexts, it can cost far less than a frontier model. Check that it holds up on your own tasks before you move production traffic.

The model went back into discussion this week after a blog post titled "Why Isn't The Industry Freaking Out About DeepSeek 4.1 Flash?" reached the top of Hacker News, with 600+ points and nearly 500 comments. This guide covers the verified specs, the real pricing, what the benchmarks do and don't show, and how to test a switch without breaking anything.

What DeepSeek V4.1 Flash is

Artificial Analysis and Vals AI both date the release to September 10, 2026. DeepSeek published a technical report, DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression, on arXiv on September 17, 2026.

Key specs from the Hugging Face model card and the paper:

Spec DeepSeek V4.1 Flash
Architecture Multimodal mixture-of-experts, "Causal Encoder-Decoder"
Parameters 552B backbone; 16B active per token in decode, 8B in prefill
Context window Up to 1M tokens
Max output (API) Up to 384K tokens
License MIT (code and weights)
Pretraining data 45T-token multimodal corpus
KV cache ~890 bytes per token in HBM, about ¼ of V4-Flash
Inference engines vLLM, SGLang, Docker Model Runner

The main idea is the KV cache. The model card says the global KV cache is about one quarter the size of DeepSeek-V4-Flash's and about 437 times smaller than DeepSeek-V1's. That comes from cross-layer KV reuse, FP4 KV caching and a technique DeepSeek calls "SWA Bounded Replay," which the paper says shrinks the persistent cache on SSD or host memory to roughly one eighth of V4-Flash's.

This matters to builders because coding agents keep long histories in context and reread them on every turn. A smaller cache lets a provider serve more long sessions per GPU, which is how DeepSeek can price cached input at fractions of a cent per million tokens.

Pricing: peak, off-peak and the cache discount

Rates per 1M tokens from DeepSeek's official pricing page:

Model (API name) Cache hit (off-peak / peak) Cache miss input (off-peak / peak) Output (off-peak / peak)
deepseek-flash (V4.1 Flash) $0.003 / $0.006 $0.15 / $0.30 $0.60 / $1.20
deepseek-v4-pro (V4-Pro-0813) $0.022 / $0.044 $0.66 / $1.32 $1.98 / $3.96

Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday. Weekends and Chinese public holidays are off-peak all day. In US Pacific daylight time, the peak windows fall roughly between 6 PM and 3 AM the evening before. In Central European summer time they run 03:00–06:00 and 08:00–12:00. So most US daytime traffic already gets off-peak rates, while European mornings pay peak. Recheck these conversions when the clocks change this autumn.

DeepSeek also notes that it "may adjust prices at any time." Treat these numbers as current, not guaranteed.

A worked cost example

Take a hypothetical day of agentic coding: 20M input tokens with 90% cache hits, plus 0.5M output tokens.

  1. Cached input: 18M × $0.006 = $0.108
  2. Uncached input: 2M × $0.30 = $0.60
  3. Output: 0.5M × $1.20 = $0.60
  4. Total at peak: about $1.31. Off-peak: about $0.65.

That fits the Hacker News author's report that he has rarely gone over $1 in a session, even one lasting most of a day. Note that output and cache misses dominate the bill, not cached reads. If your agent doesn't keep its prompt prefix stable, the cache discount goes away and costs can rise several times over. To model your own usage across providers, plug real token counts into the LLM API Cost Calculator.

Artificial Analysis puts the blended price (a 7:2:1 cache-hit/input/output mix) at $0.18 per 1M tokens, or about $0.27 per task on its Intelligence Index.

Subscription route

The blog author uses it through OpenCode Go. OpenCode's docs list Go at $10/month with DeepSeek V4.1 Flash included. That's a $60 monthly usage allowance for this model, split into 5-hour (20%) and weekly (50%) windows, which OpenCode estimates at about 26,000 requests per 5 hours. A $40/month Go Plus tier raises the allowance to $120. OpenCode says Go is "designed primarily for international users."

What the benchmarks actually show

Here the numbers need care, because the vendor and independent figures differ.

DeepSeek's own numbers (max reasoning effort, model card): GPQA Diamond 90.9, Humanity's Last Exam 36.8, Codeforces rating 3471, Terminal-Bench 2.1 90.6, DeepSWE v1.1 74.2, AutomationBench 54.8.

Independent evaluations:

  • Artificial Analysis scores the Max variant 39 on its Intelligence Index, #7 of 117 models in its comparison class, where the median is 18. It measured 217.4 output tokens per second (#3 of 117) and 1.16s time to first token on DeepSeek's API.
  • Vals AI scored it 57.86% on the Vals Index in its September 10 update: #15 of 56 overall, #1 among open-weight models, just ahead of Kimi K3 (57.81%). Vals says the result cost about $0.30 per test against $6.47 for Kimi K3. It also took first place on SkillsBench (69.80%).
  • Vals measured Terminal-Bench 2.1 at 74.53% across three full trials, well below DeepSeek's self-reported 90.6. Harnesses and settings differ, but the gap is a reminder to trust third-party runs over launch tables.

Two more caveats. First, the Vals page itself isn't consistent. Its current leaderboard table shows a lower Vals Index of 51.32% (24th of 45), likely because the index has since been updated, so rankings depend on the date you read them. Second, many hard agentic tests remain weak for this model: Vals lists SRE Bench at 0.76%, ProgramBench at 0.50% and Terminal-Bench 4.0 at 19.70%. "Near-frontier" doesn't mean it can handle long, unattended operations work.

Some developers report they can't tell it apart from Claude Opus in everyday use. Treat that as one developer's experience, not a measurement. Even that author still brings in Opus 5.5 for a final code review to catch edge cases.

Who should switch, and who shouldn't

Your workload Recommendation
High-volume coding agents, refactors, research loops Strong candidate. Cheap cached context and fast output fit long sessions.
Batch jobs you can schedule Run them off-peak for half price.
Bulk classification, extraction, summarisation Test it. It's likely good enough, and per-call cost is tiny.
Final code review, security-sensitive changes Keep a frontier model as a second pass.
Long unattended infra/SRE tasks Not yet, judging by Vals' SRE and ProgramBench results.
Data that can't leave your jurisdiction Self-host the MIT weights or use a host you've vetted. Review DeepSeek's terms before sending customer data to its API.

A common pattern, and the one the HN author describes, is cheap model for planning and execution, expensive model for review. You get most of the savings while a stronger model, or simply a different one, checks the work.

How to migrate in practice

DeepSeek's API supports both an OpenAI-compatible base URL (https://api.deepseek.com) and an Anthropic-compatible one (https://api.deepseek.com/anthropic), according to its docs. Tool calls, JSON output, the Responses API and vision input are supported. FIM (fill-in-the-middle) completion works in non-thinking mode only.

  1. Check whether you've already migrated. DeepSeek says the legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are retired, and requests sent under those names are now served by V4.1 Flash. If you saw behaviour change after September 10, this is probably why. Switch your config to deepseek-flash explicitly.
  2. Change base URL, key and model name. Most OpenAI- or Anthropic-SDK code needs no other changes.
  3. Use the recommended sampling settings. The model card suggests temperature 1.0, top_p 0.95–1.0 and max_tokens of at least 256K for reasoning runs. Reasoning effort can be set from 1 to 100, and published results use 100. Lower effort trades accuracy for speed and cost.
  4. Keep prompt prefixes stable. Put system prompts, tool schemas and repo context first and unchanged so you get cache hits. That is where most of the savings come from.
  5. Run a side-by-side eval. Replay 50–200 real tasks from your logs through both your current model and V4.1 Flash, and score pass rate, not just impressions.
  6. Mind concurrency. DeepSeek lists a concurrency limit of 2,500 for Flash, compared with 500 for V4-Pro.
  7. Keep a fallback. Route failures or low-confidence outputs to another provider. Compliance or uptime needs may also justify a third-party host for the open weights.

Use the Token Counter to check whether your largest prompts actually need the 1M window. Note that tokenizers differ between model families, so counts are approximate.

Self-hosting reality check

MIT-licensed weights mean you can run it yourself, but 552B parameters is server-scale. The model card's file metadata (FP8/BF16 tensors) also lists a larger 763B figure, so plan for a multi-GPU node, not a workstation. The HN author makes the same point: self-hosting is only "technically" possible today, and it's unlikely to save money compared with DeepSeek's API prices. The case for self-hosting is privacy and control, not cost.

What to watch next

  • Price stability. DeepSeek reserves the right to change prices, and peak/off-peak pricing makes your bill depend on when traffic runs. Track effective cost per task, not just list price.
  • Independent reruns. Watch Vals and Artificial Analysis as their indexes update. The gap between self-reported and third-party Terminal-Bench scores is the number to watch.
  • Competitive response. Vals puts Kimi K3 within 0.05 points on its index at about 20x the cost per test. Expect other open-weight labs and proprietary "mini" tiers to answer on price.
  • Your own eval. Benchmarks only rank models in general. A day spent replaying your own tasks will tell you more than any leaderboard.

If you're building agents or coding tools on top of cheaper open-weight models, you can compare alternatives in the AI coding assistants directory.

Frequently asked questions

How much does DeepSeek V4.1 Flash cost?

On DeepSeek's own API, peak pricing is $0.30 per million uncached input tokens, $0.006 per million cached input tokens and $1.20 per million output tokens. Off-peak rates are half that. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday. Weekends and Chinese public holidays are off-peak all day.

What is the API model name for DeepSeek V4.1 Flash?

Use deepseek-flash. The old names deepseek-v4-flash and deepseek-v4-flash-vision-exp still work, but DeepSeek says those models are retired. Requests sent under those names go to V4.1 Flash and are billed at the Flash price.

Is DeepSeek V4.1 Flash open source?

The weights are on Hugging Face under the MIT license. The model card lists 552B backbone parameters in a mixture-of-experts design, with 16B active per token during decoding and 8B during prefill. You can run it with vLLM or SGLang, but you'll need a multi-GPU server, not a single consumer GPU.

Is DeepSeek V4.1 Flash as good as Claude Opus or GPT-6?

Independent results put it at the top of open-weight models but below the best proprietary ones. Vals AI ranked it #15 of 56 on its Vals Index at launch, first among open-weight models. Some developers say they can't tell it apart from Opus in daily coding, but that is personal experience, not a measurement. Test it on your own tasks before switching.

Does DeepSeek V4.1 Flash support tool calling and the Anthropic API format?

Yes. DeepSeek's docs list tool calls, JSON output, the Responses API and an Anthropic-compatible endpoint at https://api.deepseek.com/anthropic. It also supports vision input. FIM completion works in non-thinking mode only.

Sources

  1. DeepSeek API Docs: Models & Pricing— api-docs.deepseek.com
  2. DeepSeek-V4.1-Flash model card on Hugging Face— huggingface.co
  3. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression (arXiv 2609.19969)— arxiv.org
  4. Artificial Analysis: DeepSeek V4.1 Flash (Max) intelligence, speed and price— artificialanalysis.ai
  5. Vals AI: DeepSeek V4.1 Flash benchmark results— vals.ai
  6. Why Isn't The Industry Freaking Out About DeepSeek 4.1 Flash? (dgt.is)— dgt.is
  7. OpenCode Go documentation: plans, limits and models— opencode.ai
person

LaunchBoosts Research Desk

AI-assisted research

Explainers on software and AI industry trends, drafted with AI assistance from the public sources cited in each article and published after automated editorial checks for length, independent sourcing and originality.