llmrelay
// gpt

GPT-5.6 Luna

The small GPT-5.6. Cheap enough for high-volume routing and classification.

GPT-5.6 Luna is the small, fast model in the GPT-5.6 line. It is built for work you run at massive scale - classification, routing, guardrails, batch tagging - where speed and cost matter more than reasoning depth.

Luna carries the same 1M+ context window as Sol and Terra, which is unusual for a small model. That gives it long-context classification capability without the flagship price, making it the right tool for tasks like document triage or large-corpus tagging.

We sell it at half OpenAI’s list price - $0.50 per million input tokens against $1.00. At that rate the model cost stops being the constraint for high-volume production work.

Official list
llmrelay
Input / M tokens
$1.00
$0.50
Output / M tokens
$6.00
$3.00
Context window
1,050,000 tokens
Max output
128,000 tokens

What it costs you per month

Real-world budget scenarios. Numbers are simple sums — official list price vs llmrelay's 50% off tier.

Usage scenario
Official
llmrelay
Light coding (1M in / 200K out per month)
$2.20
$1.10save $1.10
Heavy Cursor / Cline user (50M in / 5M out per month)
$80.00
$40.00save $40.00
Production RAG (500M in / 20M out per month)
$620.00
$310.00save $310.00

What it's good at

Best for

Pick GPT-5.6 Luna when

  • +Classification or routing at massive scale. Intent detection, sentiment tagging, category assignment - Luna is cheap enough to run on every request without the bill becoming the bottleneck.
  • +Guardrails and content moderation. Check user input for policy violations, prompt injection, or PII leakage before handing it to a more expensive model downstream.
  • +Batch tagging where volume matters more than nuance. Labelling millions of documents or messages is a Luna job, not a Sol job.
  • +You need long-context classification on a budget. The 1M window at $0.50/M input lets you put entire documents in front of the model for triage or routing decisions.

Choose something else when

  • !Complex reasoning or multi-step work. Luna is optimised for speed and cost, not depth. If the task needs synthesis or planning, use Terra or Sol.
  • !Paying even a small model for work that is just labels. Fable is not a writing SKU; stay on Luna for classification.
  • !Agent loops where one bad decision poisons the chain. Luna will make more mistakes than Terra or Sol, and on agentic work that cost compounds. Spend more per call to avoid restarts.
  • !Work where you need to review nuance rather than label categories. If the answer is not cleanly one of N buckets, Luna is the wrong tool.

Questions people ask about GPT-5.6 Luna

How much does GPT-5.6 Luna cost?

OpenAI’s list price is $1.00 per million input tokens and $6.00 per million output tokens. On llmrelay it is $0.50 and $3.00 - exactly half list. Billing is prepaid per token from one credit pool, with no subscription.

Should I use GPT-5.6 Luna or Claude Haiku 4.5?

Both are small, fast models priced for volume. Luna costs $0.50/M input against Haiku’s $0.50, so they are equivalent on price. Luna has a 1M context window; Haiku has 200k. If your classification task needs long context, Luna is the pick. If 200k is enough, Haiku is slightly faster.

Is Luna good enough for real work?

For the right work, yes. It is demonstrably good at classification, sentiment analysis, short-form Q&A, and routing. It is not good at reasoning through ambiguity or multi-step synthesis. The question is not whether Luna is capable in some absolute sense - it is whether your specific task is a labelling job or a thinking job.

What is the context window on GPT-5.6 Luna?

1,050,000 tokens input, with up to 128,000 output tokens. That is the same window as GPT-5.6 Sol and Terra, which is unusually large for a small model. It makes Luna the right tool for long-document classification or triage.

Try GPT-5.6 Luna at half the price

Free to create an account, no subscription, $10 minimum top-up — enough to run this model against a real task and judge quality yourself.

Get API key →

Compare GPT-5.6 Luna against the alternatives

The comparisons this model appears in, the models nearest it on price, and the full rate card.