GPT-6 Luna pricing: the volume tier at a tenth of Haiku's rates

Источник: Unbiased

GPT-6 Luna pricing: the volume tier at a tenth of Haiku's rates

Source: Unbiased

GPT-6 Luna costs $0.10 input / $0.50 output per million tokens, one tenth of Claude Haiku 4.5's $1/$5. The volume tier is where unit economics live. List-price math beside our measured Claude bills.

•Updated: October 6, 2026

Skip the read: get a measured recommendation in a few quick questions. Run the Stack Finder

$0.10 / $0.50

GPT-6 Luna per million tokens in / out, OpenAI Standard rates for short-context requests

$1 / $5

Claude Haiku 4.5, Anthropic's volume tier: ten times Luna's rates on both sides

$0.0016

our measured Haiku 4.5 bill for a real invoice-extraction task on July 28, 2026. Volume-tier work is this cheap done right

GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output as of October 1, 2026, OpenAI's volume tier. Claude Haiku 4.5, the closest Anthropic rival, lists at $1 and $5, ten times Luna on both sides. Luna's cached input bills at $0.01 per million, a tenth of its input rate.

How much does GPT-6 Luna cost?

$0.10 per million input tokens, $0.50 per million output, and $0.01 per million for cached input (OpenAI pricing). These are Standard rates for short-context requests, verified October 1, 2026; the API model id is gpt-6-luna. At these prices the cache discount is a fraction of a fraction of a cent per request, which is the point: this is the tier where you stop thinking per request and start thinking per million requests.

Luna vs Haiku 4.5: the bake-off that matters

On the rate card this is not close. Take a short extraction request with 2,000 input tokens and 300 output tokens: on Luna that is $0.0002 of input plus $0.00015 of output, about $0.00035. The same token counts on Haiku 4.5 cost $0.0035. Per million such requests: about $350 versus $3,500. That is arithmetic on list prices, not a measurement; we have not metered Luna. The caveat runs both directions: at this tier, quality cliffs cost more than rates do, and each family fails differently. Run both on your worst hundred cases before believing either rate card, and remember a volume-tier failure that escalates to a frontier retry can cost more than the task you tried to save on.

The same two tasks on every model we could meter

We sent an identical dashboard-generation prompt and an identical invoice-extraction prompt to every model we could call directly, on July 28, 2026, and recorded the actual bills. No estimates in the measured rows. GPT-6 Luna and Claude Sonnet 5.5 came out after this run and are not in it.

Measured 2026-07-28, one shot each, default settings, list prices. *Fable extraction is same-tokens list math from our earlier run; every other figure is a metered bill. Claude Sonnet 5.5 was released after this run. Prompts published verbatim in the Fable 5 breakdown. All outputs completed the task; quality-per-dollar comparisons need the published benchmark scores, not this table alone.

The volume tier's real job

Luna and Haiku exist because most production traffic is routine: classify this, extract that, summarize this thread. The craft is in the split: send the routine 80% to the volume tier, escalate the hard 20%, and never pay frontier rates for boilerplate. Making that split per request, with receipts, is the entire product we sell: which is exactly why this page tells you the volume tiers are genuinely good. They are what good routing routes to.

GPT-6 Luna pricing questions, answered

Luna, on list price: $0.10 input and $0.50 output per million tokens against Haiku's $1 and $5. On the same token counts that is a tenth of the cost. At a million short tasks a month, list math puts Luna near $350 and Haiku near $3,500, if quality is equal on your traffic, which is the thing to verify.

It is worth testing for boilerplate, snippets, and simple transformations. We have not measured Luna on code; our July runs show even Haiku 4.5 producing a working one-shot dashboard, which is the bar to check it against. For multi-file agentic work, failure-and-retry costs erase volume-tier savings fast; that traffic belongs a tier or two up.

Cached input bills at $0.01 per million tokens, a tenth of Luna's $0.10 input rate. A 100,000-token context read from cache costs $0.001 instead of $0.01. Cheap to read does not mean good at reasoning over long inputs; quality on long-context tasks is the thing to test, not the price.

When failure rates climb: every failed cheap run that escalates to a frontier retry bills you twice, and review time is a cost too. Measure cost per successful task, not cost per request; the gap between those two numbers is where volume tiers quietly lose.

Not sure which model fits?

The Stack Finder asks a few quick questions about your workload and gives you a straight recommendation. No account required.

Try the Stack Finder

What this article says

Something is unclear? Ask about the article — I will explain in plain words.

Do not want to dig deeper? We will sort it out for you.