GPT-6.1 Sol vs GPT-6 Astra: One Point of Index, Five Times the Rate Card

Источник: OrcaRouter

GPT-6.1 Sol vs GPT-6 Astra: One Point of Index, Five Times the Rate Card

Source: OrcaRouter

GPT-6.1 Sol comes within 0.84 index points of GPT-6 Astra while costing 4.5x less per evaluated task. Where the $10/$50 flagship still earns its premium.

•Updated: September 29, 2026

GPT-6.1 Sol and GPT-6 Astra are one family's mid-tier and flagship, released four weeks apart — Sol on September 29, 2026, Astra on September 3, 2026 — and until this week the comparison between them rested entirely on Ope​nAI's own numbers. It no longer does. Artificial Analysis has evaluated both at maximum reasoning effort on the same index version, and the result is a 0.84-point gap in Astra's favour sitting next to a cost-per-evaluation-run gap of 4.5 times in GPT-6.1 Sol's. On the rate cards the ratio is starker still: Astra is 5 times GPT-6.1 Sol on input and output and 10 times on cached input. The question this page answers is where that premium buys something you can point at, and where it is simply the price of a name.

What the two models are, and what the board measured

Both are text-and-image in, text out, with a 1,050,000-token context window and 128,000 maximum output tokens, and both expose a five-rung effort ladder — low, medium, high, xhigh and max — having dropped the none setting that GPT-6 Sol still accepts. Neither supports fine-tuning. The differences that matter are the ones the price sheet states: Astra is the flagship tier, GPT-6.1 Sol is the refresh of the tier beneath it, and OpenAI's own positioning line for the newer model is "near-Astra performance for complex work at a lower cost."

Artificial Analysis scores that claim on its Intelligence Index — ten evaluations in version v4.3.2, the same version for both models — and records the release date it holds for each. It evaluates both in the maximum-effort configuration that the model page names in its heading. What follows is that comparison, plus the cost arithmetic the index does not do for you.

• GPT-6.1 Sol — released 2026-09-29, Intelligence Index 51.8 at max effort, $2.00 input / $0.10 cached / $2.50 cache write / $10.00 output per million tokens. • GPT-6 Astra — released 2026-09-03, Intelligence Index 52.7 at max effort, $10.00 input / $1.00 cached / $12.50 cache write / $50.00 output per million tokens. • Neither — accepts the none reasoning effort; both reprice the whole request at 2× input and cache rates and 1.5× output once input crosses 272,000 tokens. Astra's long-context tier is $20.00 in / $2.00 cached / $75.00 out; GPT-6.1 Sol's is $4.00 / $0.20 / $15.00. • Both — support web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search through the Responses API.

0.84 index points, 4.5 times the evaluation bill

Two figures sit side by side on the board and together they are the whole argument. GPT-6 Astra at maximum effort scores 52.7. GPT-6.1 Sol at maximum effort scores 51.8. Beside each score the board publishes what it cost to run the index behind it: $3.26 per task for Astra, $0.72 for GPT-6.1 Sol. Read as a single sentence — 0.84 points for 4.5 times the money — it is the most decision-relevant number either company has published about these two models.

Both models also publish a score for each point on the effort ladder, which lets the comparison be run at equal spend rather than at equal setting:

• max effort — GPT-6.1 Sol 51.8 at $0.72 per run vs GPT-6 Astra 52.7 at $3.26. • xhigh — 51.0 at $0.39 vs 52.4 at $2.31. • high — 50.2 at $0.32 vs 50.9 at $1.73. • medium — 47.8 at $0.21 vs 49.6 at $1.54. • low — 42.1 at $0.13 vs 45.8 at $0.82.

Line the rungs up across models and a fact falls out that neither vendor advertises: Astra's own high-to-max spread is 1.8 index points, more than twice the entire gap between Astra at max and GPT-6.1 Sol at max. In other words, choosing Astra over GPT-6.1 Sol buys you less score than choosing Astra's max effort over Astra's high effort — and the second of those choices costs a fraction of the first. Anyone whose workload is genuinely sensitive to those 1.8 points should be checking the effort setting before they check the model column.

Where Astra's premium still buys something

The composite is close; the individual evaluations are not. Scored row by row at maximum effort, GPT-6 Astra holds the advantage on six of ten, and the ones it holds are the ones that involve acting rather than answering.

• AutomationBench-AA, multi-step workflows across tools — 0.685 for Astra against 0.649 for GPT-6.1 Sol. Astra by 3.6 points. • Terminal-Bench 4.0, terminal work in real environments — 0.591 against 0.561. Astra by 3.0 points. • SciCode — 0.565 against 0.542. Astra by 2.3 points. • Humanity's Last Exam — 0.547 against 0.529. Astra by 1.8 points. • AA-Omniscience — 43.4 against 41.5, and on the same set Astra's hallucination rate is 0.513 against GPT-6.1 Sol's 0.543. Astra is both more accurate and more willing to say it does not know. • AA-Briefcase v1.1, professional deliverables — Elo 1568.9 against 1564.2. Astra by 4.7.

The pattern is consistent enough to plan around: when the task requires the model to keep a long chain of actions straight — drive a terminal, run a multi-tool workflow, produce a document-shaped deliverable — Astra's margin is real and repeatable across harnesses. When the task is a knowledge question answered in one turn, the margin largely disappears.

Where the cheaper model wins, including the cache line

GPT-6.1 Sol is not a uniformly slightly-worse Astra. On the evaluations that reward reading and reasoning over long material it leads, and on the rate card it leads by a margin the index never sees.

• GDPval-AA, professional work graded against expert output — Elo 1575.1 for GPT-6.1 Sol against 1541.9 for Astra. A 33-point lead for the model that costs a fifth as much, and the single largest independent win in the row. • AA-LCR v1.1, long-context reasoning — 0.830 against 0.807. GPT-6.1 Sol leads. • GDP.pdf — a dead tie at 0.310, as is CritPt at 0.317. • Cached input — $0.10 per million tokens against Astra's $1.00. This is the row that decides agent economics and it is nowhere in the index. • Output speed — 66.8 tokens per second against 56.95 on the same evaluation run, with 67.2 million output tokens spent against Astra's 60.0 million. Neither model is fast; GPT-6.1 Sol is the faster of the two.

An agent month, priced on both rate cards

Take a workload that resends a stable prefix: 40 million input tokens a month with 85% of them served from cache, plus 4 million output tokens. On GPT-6.1 Sol's rate card that is $3.40 for the cached reads, $12.00 for the uncached input and $40.00 for the output — $55.40 for the month. On GPT-6 Astra's, the same traffic is $34.00 of cached reads, $60.00 of uncached input and $200.00 of output — $294.00. The workload is identical; the bill is 5.3 times larger.

Two thresholds bend that arithmetic, and both are worth pricing before the migration decision rather than after. The first is 272,000 input tokens in a single request: cross it and the whole request reprices at twice the input and cache rates and 1.5× the output rate, which moves GPT-6.1 Sol to $4.00 / $0.20 / $15.00 and Astra to $20.00 / $2.00 / $75.00. The second is the effort ladder itself — dropping GPT-6.1 Sol from max to xhigh cuts its index-run cost from $0.72 to $0.39 for 0.8 of a point, and dropping Astra from max to high cuts $3.26 to $1.73 for 1.8. In a cache-heavy agent loop, the effort setting is a bigger cost lever than the model column, which is the opposite of how these two models are usually compared.

Which to run, and what you can actually call

On the evidence, the split is cleaner than the marketing suggests. Route the work that has to survive a long chain of actions — terminal sessions, tool-heavy automation, document deliverables where a hallucinated paragraph is expensive — to GPT-6 Astra, and consider running it at high rather than max unless you have measured that the extra 1.8 points land on your tasks. Route everything that is knowledge-heavy, cache-heavy or read-mostly to GPT-6.1 Sol: it is better at professional-work grading, better at long-context reasoning, faster, and five to ten times cheaper on the lines that dominate a real bill.

What you can call today is the honest part. OrcaRouter serves GPT-6 Astra at OpenAI's own list price — the $10.00 input, $50.00 output, $1.00 cached-read and $12.50 cache-write rates, with the 272K repricing rule passed through exactly as the vendor publishes it — alongside GPT-6 Sol and GPT-6 Luna. It does not serve GPT-6.1 Sol: the public catalogue returns "model not found" for openai/gpt-6.1-sol, and we are not going to describe a model we cannot route. Because provider list price is passed through with no markup added, the day that model lands it costs OpenAI's rate on our side without a repricing step; today the usable half of this matchup is the flagship, behind one API key, with the routing DSL available if you want to compose a cheap model and an expensive one into a single call rather than choosing between them per request.

The one thing to watch is the row this comparison is missing. Neither vendor number nor independent index tells you what your own traffic will do on these two rungs, and the only cheap way to find out is to run the same task set through both with failover holding the floor. That is a one-key experiment when both halves are callable, and a half-experiment until they are.