General-purpose routing
- Luna Medium — volume
Routine work and high-volume subagents. - Luna High — cost-aware default
My preferred general-purpose production tier. - Sol Medium — escalation
Harder analysis, architecture, planning, and difficult tasks.
Bottom line
General workloads run Luna Medium → Luna High → Sol Medium. Coding workloads run Luna Medium → Luna High → Luna Max → Sol High. Sol xhigh is the high-stakes quality and review tier; Sol Max is a manual maximum-capability option for exceptional tasks. Codex credit rates reinforce the Luna-first economics.
Higher means a stronger Artificial Analysis Intelligence Index score; farther left means a lower average benchmark task cost. Add the selected Claude points for cross-provider context. The dashed line marks configurations that are not dominated on those two metrics.
A configuration is “dominated” when another displayed option costs no more and scores at least as high. Enabling Claude recalculates the displayed frontier across the selected points. Frontier status is not automatically a recommendation: latency, token use, behavior, and workload fit remain outside this chart.
The general-purpose and coding benchmarks support different escalation paths.
Sol xhigh for high-stakes quality and review: architecture, security-sensitive changes, consequential review, conflicting approaches, and exceptionally difficult work. Sol Max for exceptional maximum-capability work where marginal cost matters less than capability.
Use the chart to eliminate clearly inefficient tiers, then apply workload evidence to the remaining choices.
These configurations are beaten on both displayed metrics or lack a score needed for comparison.
Luna Medium owns the low-cost end. Luna High is the most compelling upgrade in the entire set.
Luna xhigh, Luna Max, Sol High, Sol xhigh, and Sol Max remain on the two-metric frontier, but each asks you to pay more for smaller score gains.
The Artificial Analysis Coding Index and Coding Agent Index support Luna as the cost-aware coding family, with Luna Max as the premium-value reasoning step and Sol High as the standard difficult-task escalation.
The weighted coding component of the Artificial Analysis Intelligence Index, including Terminal-Bench v2.1 and SciCode. Higher is stronger. Do not compare these scores directly to the Coding Agent Index or the general Intelligence Index; they use different scales and task sets.
| Configuration | Coding Index | Operational role | |
|---|---|---|---|
| Sol xhigh | 78.3 | High-stakes quality and review | |
| Sol Max | 77.4 | Maximum-capability option | |
| Sol High | 77.2 | Difficult-task escalation | |
| Terra Max | 76.7 | Comparison data | |
| Sol Medium | 76.3 | General escalation | |
| Terra xhigh | 70.6 | Comparison data | |
| Sol Low | 69.7 | Comparison data | |
| Terra High | 67.1 | Comparison data |
The composite pass@1 result across DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA. Higher is stronger. This is a separate index from the Coding Index above; their scores are not directly comparable.
| Configuration | Coding Agent Index | Operational role | |
|---|---|---|---|
| Sol Max | 61 | Maximum-capability option | |
| Sol xhigh | 59 | High-stakes quality and review | |
| Sol High | 58 | Difficult-task escalation | |
| Terra Max | 57 | Comparison data | |
| Sol Medium | 55 | General escalation | |
| Terra xhigh | 53 | Comparison data | |
| Terra High | 51 | Comparison data | |
| Terra Medium | 44 | Comparison data | |
| Terra Low | 34 | Comparison data |
Luna Medium → Luna High → Luna Max → Sol High
Sol xhigh for high-stakes quality and review. Sol Max for exceptional maximum-capability work.
Luna Medium covers tests, documentation, small fixes, tool-heavy work, and scoped implementation. Luna High is the normal production coding default. Luna Max is the premium-value reasoning step before changing model families. Sol High is the standard difficult-task escalation.
Optional cross-provider context. Claude stays behind the chart toggle and this comparison table; it does not change the GPT-5.6 routing ladder.
| Workload | Claude model | Configuration | Cost / task | Index | Closest GPT-5.6 comparison |
|---|---|---|---|---|---|
| General | Claude Sonnet 5 | Max reasoning | $1.53 | 53 | Sol Medium scores 54 at $0.31 |
| General | Claude Opus 4.8 | Max reasoning | $1.80 | 56 | Sol High scores 56 at $0.45 |
| General | Claude Fable 5 | Max, Opus 4.8 fallback | $2.75 | 60 | Sol Max scores 59 at $1.04 |
| Coding | Claude Code, Opus 4.8 | Medium | $3.26 | 67 | Luna High scores 68 at $0.96 |
| Coding | Claude Code, Opus 4.8 | Max | $7.70 | 73 | Luna Max scores 75 at $1.57 |
| Coding | Claude Code, Fable 5 | Max with fallback | $11.75 | 77 | Terra Max scores 77 at $2.76 |
General Claude task costs are transcribed from the supplied comparison data. Coding values reflect the current Artificial Analysis Claude Code versus Codex comparison. These are selected configurations, not complete provider-family sweeps, and they do not alter the GPT-5.6 routing recommendations.
OpenAI meters GPT-5.6 Codex usage by input, cached input, and output tokens. These rates explain how model choice consumes included usage and how purchased credits are charged after a plan limit is reached.
| Model | Input | Cached input | Output | Family rate |
|---|---|---|---|---|
| GPT-5.6 Luna | 25 credits | 2.5 credits | 150 credits | 1× baseline |
| GPT-5.6 Terra | 62.5 credits | 6.25 credits | 375 credits | 2.5× Luna |
| GPT-5.6 Sol | 125 credits | 12.5 credits | 750 credits | 5× Luna |
Rates are credits per 1 million tokens, not credits per message. Approximate task usage is: uncached input × input rate + cached input × cached rate + output × output rate, with each token count divided by 1 million.
For every GPT-5.6 family, one output token costs six times one uncached input token and 60 times one cached input token. Hidden reasoning tokens are billed as output tokens. Reasoning effort does not change the family's per-token rate, but higher effort can generate more billed output even when the visible answer stays short.
Cached input receives a 90% discount. Stable instructions and reusable context can therefore become relatively cheap, while new tool output, changing context, generated text, and retries continue to drive usage. Cache behavior should be measured rather than assumed.
If their input, cache, and output proportions are similar, Terra must finish with roughly 40% as many tokens as Luna, including retries, to tie its credit cost. Sol must use roughly 20% as many as Luna, or 50% as many as Terra. Otherwise the premium must be justified by reliability, speed, or the cost of failure.
reasoning.mode: "pro" on the selected Luna, Terra, or Sol model, independently of reasoning effort. Priority processing is a paid service tier for lower and more consistent API latency, not a faster model variant.Use benchmark charts as directional evidence, then validate routing against real repositories and production tasks. Keep the five measures separate: their scores and units are not interchangeable.
GPT-5.6 configurations and the selected Claude comparison points, sorted by average Artificial Analysis benchmark task cost. “Index points / 1¢” is a simple raw benchmark-score-to-cost ratio, not completed production tasks per dollar. Pareto frontier does not mean recommended. It only means that no displayed configuration is both cheaper and equal or stronger on this benchmark.
| Provider | Model | Reasoning | Cost / task | Index | Assessment | Index points / 1¢ |
|---|
Benchmark and Claude comparison data is sourced from Artificial Analysis. Codex credit rates and reasoning-token billing are sourced from OpenAI. Review the live sources because benchmark results, pricing, and product terms can change.
The Artificial Analysis Intelligence Index (v4.1) incorporates nine evaluations: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. See Artificial Analysis methodology for details. Coding Agent Index figures come from the live Claude Code versus Codex comparison and remain separate from the general Intelligence Index data.