llmrelay
// other

DeepSeek V4.1 Flash

Current DeepSeek Flash. Official GA 10 September 2026. Flat $0.19/$0.54 here, no peak surcharge.

DeepSeek V4.1 Flash is the current Flash model. DeepSeek published it as GA on 10 September 2026 (api-docs.deepseek.com/updates). Their first-party API name is deepseek-flash. On llmrelay the id is deepseek-v4.1-flash — that is the string you put in the request. We serve it through a third-party hosted deployment, not DeepSeek’s first-party API.

DeepSeek’s own pricing page lists this version as 1M context and 384k max output, with thinking and non-thinking modes. Off-peak cache-miss list is $0.15 / $0.60; peak is $0.30 / $1.20 (peak hours 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday). We bill a flat $0.19 / $0.54, 24/7, with no peak surcharge. That sits a little above their off-peak list and well below peak. We do not pass cache-hit discounts through.

The 09-08 community id deepseek-v4.1-flash-expires-on-0910 was a 48-hour preview. It is gone from the upstream catalogue. V4 Flash (deepseek-v4-flash) is retired at DeepSeek; they temporarily route that name to V4.1 Flash. On this site the old slug stays up so existing links keep working. New work should call deepseek-v4.1-flash.

Official list
llmrelay
Input / M tokens
$0.15
$0.19
Output / M tokens
$0.60
$0.54
Context window
1,000,000 tokens
Max output
384,000 tokens

What it costs you per month

Real-world budget scenarios. Numbers are simple sums — official list price vs llmrelay's 50% off tier.

Usage scenario
Official
llmrelay
Light coding (1M in / 200K out per month)
$0.27
$0.30save $-0.03
Heavy Cursor / Cline user (50M in / 5M out per month)
$10.50
$12.20save $-1.70
Production RAG (500M in / 20M out per month)
$87.00
$105.80save $-18.80

What it's good at

Best for

Pick DeepSeek V4.1 Flash when

  • +You want the current DeepSeek Flash, not the retired V4 Flash snapshot. Same 1M / 384k window, current weights.
  • +High-volume generation on a budget. At $0.19/M input it is cheaper than Claude Sonnet 5 ($1.00/M) and a fraction of Opus 5 or GPT-6 Astra.
  • +Long output in one shot. 384k max output is 3× the 128k ceiling on Claude and GPT here.
  • +You want one rate 24/7. DeepSeek first-party splits peak and off-peak; we do not.

Choose something else when

  • !You need DeepSeek’s first-party API, their cache-hit $0.003 rate, or native vision billed as DeepSeek. This endpoint is a third-party text deployment. We have not sold vision on it.
  • !The work is a hard reasoning job where being wrong is expensive. Use Claude Opus 5 ($2.50/$12.50) or GPT-6 Astra ($5/$25).
  • !You already validated a pipeline against deepseek-v4-flash and do not want to re-check the new id. That slug still works here; it maps to the same upstream model.
  • !You expected the expired preview id deepseek-v4.1-flash-expires-on-0910. That name is dead. Use deepseek-v4.1-flash.

Questions people ask about DeepSeek V4.1 Flash

How can I use DeepSeek V4.1 Flash?

Officially: DeepSeek’s API as deepseek-flash (api.deepseek.com). On llmrelay, set base URL to https://api.llmrelay.dev/v1 and model to deepseek-v4.1-flash. Same prepaid key as Opus 5 and GPT-6 Astra. Cursor, Cline, or any OpenAI-compatible SDK works if it lets you set a base URL.

What is the model id?

On llmrelay: deepseek-v4.1-flash. DeepSeek’s own docs tell first-party callers to use deepseek-flash. The 09-08 preview id with expires-on-0910 is retired. deepseek-v4-flash still works here as a compatibility alias to the same model.

How much does DeepSeek V4.1 Flash cost?

We charge a flat $0.19 per million input tokens and $0.54 per million output, 24/7. DeepSeek’s published cache-miss list for this version is $0.15/$0.60 off-peak and $0.30/$1.20 at peak. We do not currently pass cache-hit discounts through. This is a third-party hosted deployment, not DeepSeek’s first-party API.

Is V4.1 Flash the same as the 09-08 expires-on-0910 preview?

No. That id was a 48-hour beta. DeepSeek GA’d V4.1 Flash on 10 September 2026. Call deepseek-v4.1-flash here. The expired name is not in the catalogue.

Should I use V4.1 Flash or Claude Sonnet 5?

V4.1 Flash if you want $0.19/M input and 384k output. Sonnet 5 ($1.00/$5.00) if you want Anthropic tuning, Claude Code, or a longer production track record on this key.

What is the context window on DeepSeek V4.1 Flash?

DeepSeek’s pricing table lists 1,000,000 tokens input and up to 384,000 output, with thinking and non-thinking modes. Same window class as the retired V4 Flash.

Try DeepSeek V4.1 Flash at half the price

Free to create an account, no subscription, $10 minimum top-up — enough to run this model against a real task and judge quality yourself.

Get API key →

Compare DeepSeek V4.1 Flash against the alternatives

The comparisons this model appears in, the models nearest it on price, and the full rate card.