DeepSeek V4.1 Flash
Current DeepSeek Flash. Official GA 10 September 2026. Flat $0.19/$0.54 here, no peak surcharge.
DeepSeek V4.1 Flash is the current Flash model. DeepSeek published it as GA on 10 September 2026 (api-docs.deepseek.com/updates). Their first-party API name is deepseek-flash. On llmrelay the id is deepseek-v4.1-flash — that is the string you put in the request. We serve it through a third-party hosted deployment, not DeepSeek’s first-party API.
DeepSeek’s own pricing page lists this version as 1M context and 384k max output, with thinking and non-thinking modes. Off-peak cache-miss list is $0.15 / $0.60; peak is $0.30 / $1.20 (peak hours 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday). We bill a flat $0.19 / $0.54, 24/7, with no peak surcharge. That sits a little above their off-peak list and well below peak. We do not pass cache-hit discounts through.
The 09-08 community id deepseek-v4.1-flash-expires-on-0910 was a 48-hour preview. It is gone from the upstream catalogue. V4 Flash (deepseek-v4-flash) is retired at DeepSeek; they temporarily route that name to V4.1 Flash. On this site the old slug stays up so existing links keep working. New work should call deepseek-v4.1-flash.
What it costs you per month
Real-world budget scenarios. Numbers are simple sums — official list price vs llmrelay's 50% off tier.
What it's good at
- +Current Flash generation
- +384k output
- +1M context
- +Flat 24/7 rate
Best for
- — High-volume generation
- — Agent workflows
- — Long-output tasks
Pick DeepSeek V4.1 Flash when
- +You want the current DeepSeek Flash, not the retired V4 Flash snapshot. Same 1M / 384k window, current weights.
- +High-volume generation on a budget. At $0.19/M input it is cheaper than Claude Sonnet 5 ($1.00/M) and a fraction of Opus 5 or GPT-6 Astra.
- +Long output in one shot. 384k max output is 3× the 128k ceiling on Claude and GPT here.
- +You want one rate 24/7. DeepSeek first-party splits peak and off-peak; we do not.
Choose something else when
- !You need DeepSeek’s first-party API, their cache-hit $0.003 rate, or native vision billed as DeepSeek. This endpoint is a third-party text deployment. We have not sold vision on it.
- !The work is a hard reasoning job where being wrong is expensive. Use Claude Opus 5 ($2.50/$12.50) or GPT-6 Astra ($5/$25).
- !You already validated a pipeline against deepseek-v4-flash and do not want to re-check the new id. That slug still works here; it maps to the same upstream model.
- !You expected the expired preview id deepseek-v4.1-flash-expires-on-0910. That name is dead. Use deepseek-v4.1-flash.
Questions people ask about DeepSeek V4.1 Flash
How can I use DeepSeek V4.1 Flash?
Officially: DeepSeek’s API as deepseek-flash (api.deepseek.com). On llmrelay, set base URL to https://api.llmrelay.dev/v1 and model to deepseek-v4.1-flash. Same prepaid key as Opus 5 and GPT-6 Astra. Cursor, Cline, or any OpenAI-compatible SDK works if it lets you set a base URL.
What is the model id?
On llmrelay: deepseek-v4.1-flash. DeepSeek’s own docs tell first-party callers to use deepseek-flash. The 09-08 preview id with expires-on-0910 is retired. deepseek-v4-flash still works here as a compatibility alias to the same model.
How much does DeepSeek V4.1 Flash cost?
We charge a flat $0.19 per million input tokens and $0.54 per million output, 24/7. DeepSeek’s published cache-miss list for this version is $0.15/$0.60 off-peak and $0.30/$1.20 at peak. We do not currently pass cache-hit discounts through. This is a third-party hosted deployment, not DeepSeek’s first-party API.
Is V4.1 Flash the same as the 09-08 expires-on-0910 preview?
No. That id was a 48-hour beta. DeepSeek GA’d V4.1 Flash on 10 September 2026. Call deepseek-v4.1-flash here. The expired name is not in the catalogue.
Should I use V4.1 Flash or Claude Sonnet 5?
V4.1 Flash if you want $0.19/M input and 384k output. Sonnet 5 ($1.00/$5.00) if you want Anthropic tuning, Claude Code, or a longer production track record on this key.
What is the context window on DeepSeek V4.1 Flash?
DeepSeek’s pricing table lists 1,000,000 tokens input and up to 384,000 output, with thinking and non-thinking modes. Same window class as the retired V4 Flash.
Try DeepSeek V4.1 Flash at half the price
Free to create an account, no subscription, $10 minimum top-up — enough to run this model against a real task and judge quality yourself.
Get API key →Compare DeepSeek V4.1 Flash against the alternatives
The comparisons this model appears in, the models nearest it on price, and the full rate card.
- V4 Flash vs V4.1 Flash — which id to callV4 Flash is the old id.
- V4.1 Flash vs Sonnet 5 — cheap bulk vs Anthropic defaultV4.1 Flash is the current DeepSeek Flash, GA 10 September 2026.
- DeepSeek V4 Flash pricing and specs$0.19/M in, $0.54/M out — half list. Previous Flash id. Still callable here. DeepSeek retired it; new work should use V4.1 Flash.
- Claude Sonnet 5 pricing and specs$1.00/M in, $5.00/M out — half list. The next-gen Sonnet. Priced to replace GPT for daily work.
- Claude Haiku 4.5 pricing and specs$0.50/M in, $2.50/M out — half list. The fastest and cheapest Claude. Perfect for classification and routing.
- GPT-5.6 Luna pricing and specs$0.50/M in, $3.00/M out — half list. The small GPT-5.6. Cheap enough for high-volume routing and classification.