Bottom line

Luna High is my cost-aware default.

General workloads run Luna Medium → Luna High → Sol Medium. Coding workloads run Luna Medium → Luna High → Luna Max → Sol High. Sol xhigh is the high-stakes quality and review tier; Sol Max is a manual maximum-capability option for exceptional tasks. Codex credit rates reinforce the Luna-first economics.

General-purpose value chart

Higher means a stronger Artificial Analysis Intelligence Index score; farther left means a lower average benchmark task cost. Add the selected Claude points for cross-provider context. The dashed line marks configurations that are not dominated on those two metrics.

Luna Terra Sol Recommended stack Strictly dominated
GPT-5.6 cost per task versus Artificial Analysis Intelligence Index Scatter plot comparing Luna, Terra, and Sol configurations. Recommended configurations are Luna Medium, Luna High, and Sol Medium.

A configuration is “dominated” when another displayed option costs no more and scores at least as high. Enabling Claude recalculates the displayed frontier across the selected points. Frontier status is not automatically a recommendation: latency, token use, behavior, and workload fit remain outside this chart.

Choose the stack by workload

The general-purpose and coding benchmarks support different escalation paths.

General workloads

General-purpose routing

  1. Luna Medium — volume
    Routine work and high-volume subagents.
  2. Luna High — cost-aware default
    My preferred general-purpose production tier.
  3. Sol Medium — escalation
    Harder analysis, architecture, planning, and difficult tasks.
Coding workloads

Coding routing

  1. Luna Medium — volume
    Tests, documentation, small fixes, tool-heavy work, and scoped implementation.
  2. Luna High — cost-aware default
    Normal production coding and implementation.
  3. Luna Max — premium-value intermediate
    Harder debugging, refactoring, and implementation that warrants more reasoning before changing model families.
  4. Sol High — difficult-task escalation
    Complex implementation, unfamiliar repositories, larger refactors, and work that failed under Luna.

Sol xhigh for high-stakes quality and review: architecture, security-sensitive changes, consequential review, conflicting approaches, and exceptionally difficult work. Sol Max for exceptional maximum-capability work where marginal cost matters less than capability.

General-purpose takeaways

Use the chart to eliminate clearly inefficient tiers, then apply workload evidence to the remaining choices.

Eliminate dominated tiers

These configurations are beaten on both displayed metrics or lack a score needed for comparison.

  • Luna None: same 5¢ cost as Luna Medium, but 11 index points lower.
  • Terra None: no reported index, so it cannot be evaluated.
  • Terra Low and Sol None: Luna High is cheaper and materially stronger.
  • Sol Low and Terra High: Luna xhigh is cheaper at the same score.
  • Terra xhigh and Max: beaten by cheaper, higher-scoring Sol tiers.

The economic center

Luna Medium owns the low-cost end. Luna High is the most compelling upgrade in the entire set.

  • Medium → High: +8 index for +4¢.
  • High → xhigh: +3 index for +5¢.
  • After that, every additional point gets progressively more expensive.

Raw value versus role

Luna xhigh, Luna Max, Sol High, Sol xhigh, and Sol Max remain on the two-metric frontier, but each asks you to pay more for smaller score gains.

  • Luna xhigh/Max: credible intermediate escalation steps.
  • Sol Medium: not the third-best point by raw index per dollar; it is the operational family-level escalation.
  • Sol High/Max: maximum capability only when marginal gains justify the steep cost.

Coding changes the model mix

The Artificial Analysis Coding Index and Coding Agent Index support Luna as the cost-aware coding family, with Luna Max as the premium-value reasoning step and Sol High as the standard difficult-task escalation.

Artificial Analysis Coding Index

The weighted coding component of the Artificial Analysis Intelligence Index, including Terminal-Bench v2.1 and SciCode. Higher is stronger. Do not compare these scores directly to the Coding Agent Index or the general Intelligence Index; they use different scales and task sets.

Configuration Coding Index Operational role  
Sol xhigh78.3High-stakes quality and review
Sol Max77.4Maximum-capability option
Sol High77.2Difficult-task escalation
Terra Max76.7Comparison data
Sol Medium76.3General escalation
Terra xhigh70.6Comparison data
Sol Low69.7Comparison data
Terra High67.1Comparison data

Artificial Analysis Coding Agent Index

The composite pass@1 result across DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA. Higher is stronger. This is a separate index from the Coding Index above; their scores are not directly comparable.

Configuration Coding Agent Index Operational role  
Sol Max61Maximum-capability option
Sol xhigh59High-stakes quality and review
Terra Max57Comparison data
Sol Medium55General escalation
Terra xhigh53Comparison data
Terra High51Comparison data
Terra Medium44Comparison data
Terra Low34Comparison data

Coding route line

Luna Medium → Luna High → Luna Max → Sol High

Sol xhigh for high-stakes quality and review. Sol Max for exceptional maximum-capability work.

Luna Medium covers tests, documentation, small fixes, tool-heavy work, and scoped implementation. Luna High is the normal production coding default. Luna Max is the premium-value reasoning step before changing model families. Sol High is the standard difficult-task escalation.

Coding interpretation

  • Luna remains the cost-aware family for routine and production coding.
  • Luna Max is the premium-value reasoning step.
  • Sol High is the standard difficult-task escalation.
  • Sol xhigh is the high-stakes quality and review tier.
  • Sol xhigh leads the displayed Coding Index, while Sol Max leads the Coding Agent Index. No single reasoning level wins every coding benchmark.
Where Terra fits in coding
  • Benchmark comparison, not standard routing: Terra Max, Terra xhigh, and Terra High appear in both coding tables as measured configurations, but they are not the recommended next step after Luna Max.
  • Select Terra on evidence: teams may still select Terra when their own workload testing demonstrates a behavioral, output-length, latency, or reliability advantage.
  • Terra Medium token-efficiency option: a workload-specific exception, not a standard routing step. On the general benchmark, Terra Medium matched Luna High's Intelligence Index with roughly half the total output tokens, but at higher per-task cost. Consider it only where lower token usage, latency, or throughput outweighs benchmark cost.

Claude comparison data

Optional cross-provider context. Claude stays behind the chart toggle and this comparison table; it does not change the GPT-5.6 routing ladder.

Workload Claude model Configuration Cost / task Index Closest GPT-5.6 comparison
GeneralClaude Sonnet 5Max reasoning$1.5353Sol Medium scores 54 at $0.31
GeneralClaude Opus 4.8Max reasoning$1.8056Sol High scores 56 at $0.45
GeneralClaude Fable 5Max, Opus 4.8 fallback$2.7560Sol Max scores 59 at $1.04
CodingClaude Code, Opus 4.8Medium$3.2667Luna High scores 68 at $0.96
CodingClaude Code, Opus 4.8Max$7.7073Luna Max scores 75 at $1.57
CodingClaude Code, Fable 5Max with fallback$11.7577Terra Max scores 77 at $2.76

General Claude task costs are transcribed from the supplied comparison data. Coding values reflect the current Artificial Analysis Claude Code versus Codex comparison. These are selected configurations, not complete provider-family sweeps, and they do not alter the GPT-5.6 routing recommendations.

Codex subscription credit economics

OpenAI meters GPT-5.6 Codex usage by input, cached input, and output tokens. These rates explain how model choice consumes included usage and how purchased credits are charged after a plan limit is reached.

Model Input Cached input Output Family rate
GPT-5.6 Luna25 credits2.5 credits150 credits1× baseline
GPT-5.6 Terra62.5 credits6.25 credits375 credits2.5× Luna
GPT-5.6 Sol125 credits12.5 credits750 credits5× Luna

Rates are credits per 1 million tokens, not credits per message. Approximate task usage is: uncached input × input rate + cached input × cached rate + output × output rate, with each token count divided by 1 million.

Output is the expensive side

For every GPT-5.6 family, one output token costs six times one uncached input token and 60 times one cached input token. Hidden reasoning tokens are billed as output tokens. Reasoning effort does not change the family's per-token rate, but higher effort can generate more billed output even when the visible answer stays short.

Caching changes long sessions

Cached input receives a 90% discount. Stable instructions and reusable context can therefore become relatively cheap, while new tool output, changing context, generated text, and retries continue to drive usage. Cache behavior should be measured rather than assumed.

The break-even bar is high

If their input, cache, and output proportions are similar, Terra must finish with roughly 40% as many tokens as Luna, including retries, to tie its credit cost. Sol must use roughly 20% as many as Luna, or 50% as many as Terra. Otherwise the premium must be justified by reliability, speed, or the cost of failure.

Subscription and API billing caveats
  • Included usage comes first: monthly ChatGPT plans provide an allowance. Purchased credits are used after the applicable included limit is reached, so credits are not a direct cash charge for every in-plan message.
  • Messages are variable: OpenAI says GPT-5.6 usage averages 5–40 credits per Codex message, but model, context, reasoning, tools, retrieval, and caching can move a task outside a simple estimate.
  • API usage is separate: Codex authenticated with an API key is billed at API token prices rather than ChatGPT subscription credits.
  • Pro and Priority are controls, not models: Pro is enabled with reasoning.mode: "pro" on the selected Luna, Terra, or Sol model, independently of reasoning effort. Priority processing is a paid service tier for lower and more consistent API latency, not a faster model variant.
  • Rates can change: plan eligibility, usage limits, speed modes, and credit rates are product terms, not stable model properties. Check the live rate card before making budget commitments.

Benchmark limits and methodology

Use benchmark charts as directional evidence, then validate routing against real repositories and production tasks. Keep the five measures separate: their scores and units are not interchangeable.

  • Artificial Analysis Intelligence Index: general-purpose weighted index. The chart and source table on this page use this index.
  • Artificial Analysis Coding Index: the weighted coding component of the Intelligence Index, including Terminal-Bench v2.1 and SciCode.
  • Artificial Analysis Coding Agent Index: composite pass@1 result across DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA.
  • Benchmark task cost: average pay-per-token API cost per benchmark task, as reported by Artificial Analysis.
  • Codex subscription credit usage: OpenAI ChatGPT/Codex credits per 1M tokens, separate from API task cost.
  • Do not compare absolute scores across these indexes as though they share the same scale. The Intelligence Index, Coding Index, and Coding Agent Index use different task suites and weighting.
  • Benchmark task cost and Codex subscription credits are different billing views. Their family price ratios align, but their displayed units are not interchangeable.
  • The Claude comparison includes selected configurations rather than every Anthropic effort level, so it supports point comparisons rather than a complete provider-wide frontier.
  • Benchmark score per dollar is not the same as successfully completed production tasks per dollar.
  • Output length, latency, retries, context growth, task failure, and downstream human review can change real-world economics.
  • Operational role specialization can justify a model that is not the next point on a raw benchmark frontier.
  • Intelligent routing uses the cheapest model that reliably completes the task, then escalates when evidence warrants it.
  • Benchmark results, pricing, and product terms can change. Review the live sources before making budget commitments.

General-purpose source data

GPT-5.6 configurations and the selected Claude comparison points, sorted by average Artificial Analysis benchmark task cost. “Index points / 1¢” is a simple raw benchmark-score-to-cost ratio, not completed production tasks per dollar. Pareto frontier does not mean recommended. It only means that no displayed configuration is both cheaper and equal or stronger on this benchmark.

Provider Model Reasoning Cost / task Index Assessment Index points / 1¢

Sources

Benchmark and Claude comparison data is sourced from Artificial Analysis. Codex credit rates and reasoning-token billing are sourced from OpenAI. Review the live sources because benchmark results, pricing, and product terms can change.