Artificial Analysis Intelligence Index v4.2 · September 5, 2026 snapshot

Model reasoning family explorer

Compare GPT-6 Astra, GPT-5.6 Sol, Terra, and Luna, and gpt-oss-120b. In 3D, Intelligence Index is Y, cost per task is X, and output-token use is Z. The 2D view projects cost against intelligence and uses bubble area for token use. Each family has one color and a connected reasoning-effort trajectory where adjacent levels are available.

Selected variants
25
Astra 6 · Sol 5 · Terra 6 · Luna 6 · oss 2
Complete 3D / 2D points
18
Cost + Intelligence + token use available
Highest Intelligence Index
55
GPT-6 Astra max · $2.57/task · 49M tokens
Lowest task cost with full data
$0.02
gpt-oss-120b low · Index 14.9 · 2.8K tokens
Families

3D model family view. X = weighted cost per Intelligence Index task. Y = Artificial Analysis Intelligence Index. Z = output tokens from the Intelligence Index evaluation. Drag to rotate and scroll to zoom. Lines only connect adjacent reasoning-effort levels when both points have complete chart data.

Intelligence × cost × token use

Color identifies the model family; labels identify the reasoning level.

2D cost / intelligence projection. X = cost per Intelligence Index task. Y = Intelligence Index. Bubble area represents output-token use, and each family trajectory uses the same color as the 3D view.

Cost per task vs. intelligence

Lower and farther right is more expensive; higher is more intelligent. Bubble area indicates output-token use.

Astra's curve is unusually compact

Astra rises from 49 at $0.63 on low to 55 at $2.57 on max while output-token use increases from 5.4M to 49M. Its medium and high settings sit close to the frontier between intelligence, task cost, and token use.

Value is a screening ratio, not a verdict

Intelligence Index divided by hosted task cost puts gpt-oss-120b low at about 736 Index points/$ and high at 318, compared with about 700 for GPT-5.6 Luna xhigh. That is useful for screening, but not production value: include successful-task rate, retries, latency, token growth, provider limits, and self-hosting or operational costs.

Complete 25-variant dataset

All selected models are retained, including published N/A fields.

Artificial Analysis · Intelligence Index v4.2

ModelIntelligenceCost / taskOutput tokensOutput speedChart statusSource