llmrelay
// other

Kimi K3

Moonshot flagship for long-context coding and knowledge work. 1M context, half the published list.

Kimi K3 is Moonshot's current flagship. Moonshot publishes a 1,048,576-token context window and positions it for long-horizon coding and end-to-end knowledge work. We serve it through a third-party hosted deployment, not Moonshot's first-party API.

Our price is $1.50 per million input tokens and $7.50 per million output tokens. Moonshot's own published list for cache-miss traffic is ¥20 / ¥100 per M (about $3 / $15). That makes our rate half the published list, billed the same 24/7 with no cache-hit / cache-miss split on our side.

Live-checked 2026-08-22 through api.llmrelay.dev: it solved the bat-and-ball problem ($0.05), wrote a working palindrome helper, and emitted a correct get_weather tool call for Tokyo. We have not verified vision, video, or web_search on this endpoint, so we do not sell those as features.

Official list
llmrelay
Input / M tokens
$3.00
$1.50
Output / M tokens
$15.00
$7.50
Context window
1,048,576 tokens
Max output
1,048,576 tokens

What it costs you per month

Real-world budget scenarios. Numbers are simple sums — official list price vs llmrelay's 50% off tier.

Usage scenario
Official
llmrelay
Light coding (1M in / 200K out per month)
$6.00
$3.00save $3.00
Heavy Cursor / Cline user (50M in / 5M out per month)
$225.00
$112.50save $112.50
Production RAG (500M in / 20M out per month)
$1800.00
$900.00save $900.00

What it's good at

Best for

Pick Kimi K3 when

  • +You need a 1M context window outside the Claude / GPT families. Moonshot publishes 1,048,576 tokens, the same class of window as Claude Opus and GPT-5.6.
  • +Long-document or whole-repo work on a budget. At $1.50/M input it is cheaper than Claude Opus 5 ($2.50/M) and GPT-5.6 Sol ($2.50/M), and half Moonshot's published list.
  • +Tool-using agents. Our live probe returned a well-formed function call (city=Tokyo) rather than a hallucinated weather report.
  • +You want a flat rate. We bill one price 24/7. Moonshot's own list splits cache-hit (¥2/M) and cache-miss (¥20/M); we do not.

Choose something else when

  • !You need Moonshot's first-party API, vision, or web_search. This endpoint is a third-party hosted text + tool-call deployment. We have not verified image, video, or search on it.
  • !You need the deepest reasoning. For hard problems where being wrong is expensive, use Claude Opus 5 or GPT-5.6 Sol.
  • !You are already in the Claude or OpenAI toolchain. Switching families for a cheaper 1M window is not worth it if Cursor / Claude Code / Codex is doing the work.
  • !You want the cheapest tokens on this site. DeepSeek V4 Flash is $0.19/$0.54. Kimi K3 is a flagship-priced model, not a bulk model.

Questions people ask about Kimi K3

How much does Kimi K3 cost?

We charge $1.50 per million input tokens and $7.50 per million output tokens. Moonshot's published list for cache-miss traffic is ¥20 / ¥100 per M (about $3 / $15). Note: we serve K3 via a third-party hosted deployment, not Moonshot's first-party API.

What is the model id?

Kimi-K3. That is the id returned by GET /v1/models and the string you pass as model. The URL slug is kimi-k3.

Should I use Kimi K3 or Claude Opus 5?

Opus 5 if you want Anthropic tuning, Claude Code, or the model we treat as flagship. Kimi K3 if you want a 1M-context flagship outside that family at $1.50/M input instead of $2.50/M.

What is the context window on Kimi K3?

Moonshot publishes 1,048,576 tokens. A separate max-output figure is not on their pricing page. We have verified text chat and tool calls on this endpoint; we have not verified vision or web_search.

Try Kimi K3 at half the price

Free to create an account, no subscription, $10 minimum top-up — enough to run this model against a real task and judge quality yourself.

Get API key →

Compare Kimi K3 against the alternatives

The comparisons this model appears in, the models nearest it on price, and the full rate card.