DeepSeek V4 Flash
Fast, cheap, and capable. Flat pricing, no peak/off-peak surcharge, available 24/7.
DeepSeek V4 Flash is a fast, cost-effective reasoning model with a 1M-token context window and the ability to produce outputs as long as 384,000 tokens. It is built for high-volume work and agent workflows where speed and cost matter more than flagship-tier reasoning depth. We serve it through a third-party hosted deployment, not DeepSeek’s first-party API.
Our pricing is flat: $0.19/M input and $0.54/M output, billed the same 24/7 with no peak/off-peak surcharge. For reference, DeepSeek’s own published list price for this model runs $0.22/$0.66 off-peak and $0.44/$1.32 at peak — so our flat rate lands below even the off-peak reference and you never schedule workloads around time zones.
The 384k output ceiling is 3× what most models offer (128k), making V4 Flash the right tool for generating entire documents, large code refactors, or multi-file agent outputs in one request. At $0.19/M input it is cheaper than Claude Sonnet 5 ($1.00/M) and significantly cheaper than any flagship model.
What it costs you per month
Real-world budget scenarios. Numbers are simple sums — official list price vs llmrelay's 50% off tier.
What it's good at
- +Extended 384k output
- +Very fast
- +1M context window
Best for
- — High-volume generation
- — Agent workflows
- — Long-output tasks
Pick DeepSeek V4 Flash when
- +You need extremely long output. The 384k output ceiling lets you generate entire documents or large codebases in one shot, without splitting across multiple requests.
- +High-volume generation on a budget. At $0.19/M input V4 Flash is cheaper than Claude Sonnet 5 and a fraction of the cost of flagship models, while still carrying a 1M context window.
- +Agent workflows with extended reasoning. V4 Flash supports thinking mode with adjustable effort levels, making it suitable for multi-step agent loops.
- +You want predictable flat pricing. We bill one rate 24/7 with no peak/off-peak tiers, so you never pay a peak premium and never need to schedule around UTC time zones.
Choose something else when
- !You need the deepest reasoning. V4 Flash is optimised for speed and cost, not depth. For hard problems where being wrong is expensive, use Claude Opus 5 or GPT-5.6 Sol.
- !You are locked into Claude or OpenAI tooling. V4 Flash integrates with OpenAI-compatible APIs, but if your workflow is tightly coupled to Claude Code or OpenAI-native agent frameworks, switching for cost savings alone may not be worth the friction.
- !Output volume is low. If you are not generating long outputs, the 384k ceiling is wasted capability. Use a cheaper small model like Claude Haiku or GPT-5.6 Luna for classification and routing.
- !The model is too new to trust for production. V4 Flash exited preview in July 2026. If you need battle-tested reliability, stick with models that have been running at scale for longer.
Questions people ask about DeepSeek V4 Flash
How much does DeepSeek V4 Flash cost?
We charge a flat $0.19 per million input tokens and $0.54 per million output tokens, billed the same 24/7. For reference, DeepSeek’s own published list price for this model is $0.22/$0.66 off-peak and $0.44/$1.32 at peak — our flat rate sits below even the off-peak reference, with no time-of-day surcharge. Note: we serve V4 Flash via a third-party hosted deployment, not DeepSeek’s first-party API.
What is the 384k output ceiling useful for?
Generating entire documents, large code refactors, or multi-file agent outputs in one request. Most models cap at 128k output, which forces you to split large generations into multiple calls. V4 Flash eliminates that constraint.
Should I use DeepSeek V4 Flash or Claude Sonnet 5?
V4 Flash if you need extended output (384k vs 128k) or want lower input cost ($0.19/M vs $1.00/M). Sonnet 5 if you prefer Anthropic’s tuning, need wider tool ecosystem support, or want a model with a longer production track record.
What is the context window on DeepSeek V4 Flash?
1,000,000 tokens input, with up to 384,000 output tokens. That is the same input window as Claude Opus and GPT-5.6 models, but with 3× the output ceiling.
Try DeepSeek V4 Flash at half the price
Free to create an account, no subscription, $10 minimum top-up — enough to run this model against a real task and judge quality yourself.
Get API key →Compare DeepSeek V4 Flash against the alternatives
The comparisons this model appears in, the models nearest it on price, and the full rate card.
- Claude Haiku 4.5 pricing and specs$0.50/M in, $2.50/M out — half list. The fastest and cheapest Claude. Perfect for classification and routing.
- GPT-5.6 Luna pricing and specs$0.50/M in, $3.00/M out — half list. The small GPT-5.6. Cheap enough for high-volume routing and classification.
- Full price listEvery model we serve, at 50% of official list. No subscription, no volume gate.
- Cost calculatorPut your own monthly token volume in and see the bill against vendor list price.
- Tool setup guidesBase-URL and key steps for Cursor, Cline, Claude Code, Continue and others.