Last updated September 21, 2026

Models

Intelligence, capability, cost, and open versus closed.

Summary

This page tracks current progress on major AI models and how the field got here since 2022: vendor-published scores, list prices, and whether weights are open. The three sections below cover those themes with charts and commentary.

  • Intelligence & Capability. Frontier launch scores keep climbing on coding and capability evals, even as those tests get harder.
  • Capability vs cost. Open-weight models, especially from China, keep undercutting closed US list prices at useful capability.
  • Open vs closed. China leads open-weight volume among tracked labs. Published open sizes now sit in the same range as the disclosed closed figures on this page.

Formatted citation and BibTeX are in the Copy menu at the top of this page. The compiled table and original charts may be reused with attribution under a Creative Commons Attribution 4.0 license.

Jump to Section Contents

Intelligence & Capability

Three evals commonly cited for distinct capabilities: coding (SWE-bench), agent terminal work (Terminal-Bench), and hard general knowledge (Humanity's Last Exam). Tabs switch the chart to that capability. Vertical ticks mark version changes. A dashed mark is the most recent score posted; moves on the chart are usually models improving or evals getting tougher.

Launch scores over time

Legend

Figure 1. Each line is that company's best launch score so far on the selected eval. A lab with one scored launch is a single point. Source: vendor announcements.

Download CSV

Milestones

Dated launches that made a new kind of work practical. The same number on the axis and in the list is the same launch. Charts restart each year, except 2022 and 2023, which share a sequence because the industry was still early and gaining traction.

When those shifts shipped

Figure 2. Milestones on a linear time axis. Company and model names are in the list below. Source: vendor launch announcements.

Download CSV
2026 4
  1. Days-long autonomous tasks become a shipping launch claim

    Anthropic, Claude Fable 5. Briefly suspended under USA export controls days after launch.

  2. Multi-agent execution becomes a launch feature

    OpenAI, GPT-5.6. First release split into three durable capability tiers, Sol (flagship), Terra, and Luna, each advancing on its own cadence; Sol Ultra multi-agent mode tops Terminal-Bench 2.1 at 91.9%.

  3. Productized multi-bot teammates with a shared computer

    xAI, Grok Bot. Grok Bot beta. Always-on AI teammates that share a cloud computer, sign into the user's apps, message each other in threads or group chats, and hand work back only when approval is needed.

  4. Computer use and coding become a launch claim on GPT-6 Astra

    OpenAI, GPT-6 Astra. OpenAI's GPT-6 generation, launched first to Daybreak partners then paid ChatGPT; vendor-claimed step change on computer use, coding, and saturated ARC-AGI-3.

2025 3
  1. Open-weight reasoning reaches frontier-level public attention

    DeepSeek, DeepSeek-R1. MIT-licensed o1-class reasoning model whose release triggered a global market shock and a wave of RL replication.

  2. Instant and extended reasoning merge in one model

    Anthropic, Claude 3.7 Sonnet. First hybrid reasoning model, blending instant replies and extended thinking in one model; shipped with Claude Code.

  3. Long-running agent work becomes a launch feature

    Anthropic, Claude Sonnet 4.5. Billed as the best coding model on release, with 30-plus-hour autonomous agent runs.

2024 3
  1. Million-token context becomes a shipping model feature

    Google, Gemini 1.5 Pro. Made million-token context real, an order of magnitude beyond anything shipping at the time.

  2. Native multimodal interaction becomes a shipping product

    OpenAI, GPT-4o. Natively multimodal omni model; the launch demoed real-time voice, but the API shipped with text and image in, text out.

  3. Test-time reasoning becomes a paid product tier

    OpenAI, o1. First mainstream reasoning model, introducing chain-of-thought test-time compute as a product.

2022 / 2023 2
  1. Consumer chat reaches the mainstream

    OpenAI, GPT-3.5 / ChatGPT. The ChatGPT launch that took LLMs mainstream and kicked off the current AI race.

  2. Tool use becomes a shipping API feature

    OpenAI, Function calling. Developers could describe functions to GPT-4 and GPT-3.5-turbo. The model returned a JSON call with arguments.

Capability vs cost

Capability alone does not drive AI adoption, because the work also has to be cheap enough to run at scale. Plotting a launch's score against its list price is what people call a Pareto frontier: the best tradeoffs between capability and cost you can get at once.

Visualizing the Pareto frontier

Each tab is one as-of date on the Terminal-Bench version in force then. The dashed line is that day’s Pareto frontier. The shaded area is everything off that line: a higher list price for the same score, or a lower score for the same price.

Score vs list input price

Other models sit in the shaded area: less capable at that price, or more expensive for that score.

Legend

Figure 3. As of Sep 3, 2026. Scores: Terminal-Bench 4.0 (tbench.ai), models released by that date. Prices: vendor list $/1M input. Source: tbench.ai · vendor announcements.

Download CSV
Cheapest and highest-scoring launches on the line, by as-of date.
As ofTerminal-Bench versionLow endHigh endNotes
Sep 3, 20264.0$1 · 17.3%GPT-5.6 Luna$10 · 58.2%GPT-6 AstraTB4.0 shipped Aug 28. Absolute % not comparable to Jul 31.
Jul 31, 20263.0$1 · 14.3%GPT-5.6 Luna$5 · 42.7%Claude Opus 5TB3.0 shipped Jul 30. Absolute % not comparable to Jun 30.
Jun 30, 20262.1$1 · 75.7%GPT-5.6 Luna$10 · 83.8%Claude Fable 5End Q2. TB2.1 current since May 6.

Open vs closed for tracked major models

Tracked labs are still shipping fast in 2026. Open-weight launches keep landing, and China accounts for most of the open count in 2025–2026.

United States
11 open 40 closed
tracked launches, 2025-2026
China
27 open 5 closed
tracked launches, 2025-2026
France
3 open 0 closed
tracked launches, 2025-2026

Tracked releases by launch year

Legend

lighter = open, darker = closed

Figure 4. Open and closed launches by year. Year label shows total tracked launches. Open includes later weight releases recorded as dated events. Source: vendor launch announcements.

Download CSV

Open launches in 2026

24 open-weight launches - same count as the 2026 open column above (includes open-license-restricted).

Open-weight scale

Most closed models, including Claude, Gemini, and ChatGPT, do not publish a parameter count. The points here are lab-published totals, plus xAI sizes Elon has stated. Open-weight counts have climbed from tens of billions into the trillions, and now sit in the same range as the Grok figures.

Parameter counts

Legend

  • Open
  • Closed
  • Stated / planned
  • Trend

Figure 5. Parameter count over time. Landmark models are labeled. Source: vendor announcements, model cards, and Elon Musk.

Download CSV

Launch table

114 launches, newest first. Frontier labs to start; All companies for the rest.

Launch snapshots

Launch-state snapshots of major AI models
Launch benchmarks
2026-09-22 Claude Opus 5.5 Anthropic closed 1M $4 / $20 Terminal-Bench 4.0 66.4%; Humanity's Last Exam 67.7% with tools
2026-09-22 GPT-6 Luna OpenAI closed 1M $0.10 / $0.50 n/a
2026-09-22 GPT-6 Sol OpenAI closed 1M $2 / $10 n/a
2026-09-21 Grok 4.7 xAI closed 500k $2 / $6 Terminal-Bench 4.0 38.0%
2026-09-21 MiMo-V2.6-Flash Xiaomi open 1M $0.14 / $0.28 n/a
2026-09-21 MiMo-V2.6-Pro Xiaomi open 1M $0.44 / $0.87 Artificial Analysis Intelligence Index 46 (vendor-claimed on Xiaomi page)
2026-09-10 DeepSeek-V4.1-Flash DeepSeek open n/a n/a n/a
2026-09-03 GPT-6 Astra OpenAI closed 1M $10 / $50 Terminal-Bench 4.0 57.9%; GPQA Diamond 96.0%; Humanity's Last Exam 57.2% (with tools)
2026-09-02 Gemini 3.8 Flash Google closed 1M n/a n/a
2026-09-01 Claude Fable 5.1 Anthropic closed 1M $10 / $50 Terminal-Bench 4.0 55.8%; GPQA Diamond 93.7%; Humanity's Last Exam 65.0% (with tools)
2026-08-27 Qwen3.8-Flash-Next Alibaba open n/a n/a n/a
2026-08-26 GLM-5.3-Flash Z.ai (Zhipu) open n/a n/a n/a
2026-08-14 GLM-5.3 Z.ai (Zhipu) open 200k n/a n/a
2026-08-13 Gemini 3.7 Flash Google closed 1M n/a n/a
2026-08-12 Grok 4.6 xAI closed 500k $2 / $6 CursorBench 3.2.0 69.9%; DeepSWE v1.1 65.9%
2026-08-11 Nemotron 3.5 Lightning NVIDIA open n/a n/a n/a
2026-08-02 Qwen3.8-Max Alibaba open n/a n/a n/a
2026-07-24 Claude Opus 5 Anthropic closed 1M $5 / $25 Humanity’s Last Exam 56.3%; Terminal-Bench 4.0 52.3%
2026-07-21 Gemini 3.6 Flash Google closed 1M $1.50 / $7.50 MLE-Bench 63.9%; DeepSWE 49%
2026-07-16 Grok 4.5 xAI closed 500k $2 / $6 SWE-bench Pro 64.7%; Terminal-Bench 2.1 83.3%
2026-07-16 Kimi K3 Moonshot AI open n/a n/a n/a
2026-07-15 Inkling Thinking Machines Lab open 1M n/a n/a
2026-07-09 GPT-5.6 OpenAI closed 1M $4 / $20 GPQA Diamond 94.6%; Terminal-Bench 2.1 88.8%; Terminal-Bench 4.0 37.3%; SWE-bench Pro 64.6%
2026-07-02 Hy3 Tencent open 256k n/a n/a
2026-06-30 Claude Sonnet 5 Anthropic closed 1M $2 / $10 Humanity’s Last Exam 34.6% (no tools); SWE-bench Verified 72.7%; SWE-bench Pro 63.2%
2026-06-16 GLM-5.2 Z.ai (Zhipu) open 200k n/a n/a
2026-06-09 Claude Fable 5 Anthropic closed 1M $10 / $50 SWE-bench Pro 80.3%; Terminal-Bench 2.1 88.0%; Terminal-Bench 4.0 42.0%; Humanity's Last Exam 63.8% (with tools)
2026-06-04 Nemotron 3 Ultra NVIDIA open 1M n/a n/a
2026-06-02 MiniMax-M3 MiniMax open 1M n/a n/a
2026-05-28 Claude Opus 4.8 Anthropic closed 1M $5 / $25 n/a
2026-05-22 Mistral Medium 3.5 Mistral AI open n/a n/a n/a
2026-05-19 Gemini 3.5 Flash Google closed 1M $1.50 / $9 n/a
2026-05-18 Qwen3.7-Max Alibaba closed n/a n/a n/a
2026-04-24 DeepSeek-V4 DeepSeek open 1M n/a n/a
2026-04-23 GPT-5.5 OpenAI closed 1M $5 / $30 Terminal-Bench 2.0 82.7%; SWE-bench Pro 58.6%
2026-04-20 Kimi K2.6 Moonshot AI open n/a n/a n/a
2026-04-16 Claude Opus 4.7 Anthropic closed 1M $5 / $25 CursorBench 70%; BigLaw Bench (Harvey) 90.9% (high effort)
2026-04-15 Qwen3.6 Alibaba open 131k n/a n/a
2026-04-08 Muse Spark Meta closed n/a n/a Humanity’s Last Exam 58% (Contemplating)
2026-04-07 GLM-5.1 Z.ai (Zhipu) open 200k n/a n/a
2026-04-02 Gemma 4 Google open 256k n/a MMLU-Pro 85.2% (31B); LiveCodeBench v6 80.0% (31B); AIME 2026 89.2% (31B, no tools)
2026-03-16 Mistral Small 4 Mistral AI open n/a n/a n/a
2026-03-11 Nemotron 3 Super NVIDIA open 1M n/a n/a
2026-03-05 GPT-5.4 OpenAI closed 1M $2.50 / $15 n/a
2026-02-17 Claude Sonnet 4.6 Anthropic closed 1M $3 / $15 n/a
2026-02-16 Qwen3.5 Alibaba open 131k n/a n/a
2026-02-14 Doubao Seed 2.0 ByteDance closed n/a n/a n/a
2026-02-11 GLM-5 Z.ai (Zhipu) open 200k n/a SWE-bench Verified 77.8%; GPQA Diamond 86.0%
2026-02-05 Claude Opus 4.6 Anthropic closed 1M $5 / $25 SWE-bench Verified 80.9%; Humanity’s Last Exam 53.0% (with tools)
2026-01-27 Kimi K2.5 Moonshot AI open n/a n/a n/a
2025-12-15 Nemotron 3 NVIDIA open 1M n/a n/a
2025-12-11 GPT-5.2 OpenAI closed 400k $1.75 / $14 n/a
2025-12-02 Amazon Nova 2 Amazon closed 1M n/a n/a
2025-12-02 Mistral Large 3 Mistral AI open 256k $0.50 / $1.50 n/a
2025-12-01 DeepSeek-V3.2 DeepSeek open 128k $0.28 / $0.42 GPQA Diamond 82.4%; SWE-bench Verified 73.1%
2025-11-24 Claude Opus 4.5 Anthropic closed 200k $5 / $25 SWE-bench Verified 80.9%
2025-11-18 Gemini 3 Pro Google closed 1M $2 / $12 GPQA Diamond 91.9%; Humanity’s Last Exam 37.5% (no tools)
2025-11-17 Grok 4.1 xAI closed 256k n/a n/a
2025-11-13 Ernie 5.0 Baidu closed n/a n/a n/a
2025-11-12 GPT-5.1 OpenAI closed 400k $1.25 / $10 SWE-bench Verified 76.3%
2025-10-22 MiniMax-M2 MiniMax open 205k $0.30 / $1.20 n/a
2025-10-15 Claude Haiku 4.5 Anthropic closed 200k $1 / $5 SWE-bench Verified 73.3%; Terminal-Bench 41.75% (32K thinking)
2025-09-30 GLM-4.6 Z.ai (Zhipu) open 200k n/a n/a
2025-09-29 Claude Sonnet 4.5 Anthropic closed 200k $3 / $15 SWE-bench Verified 77.2%
2025-09-23 Qwen3-Max Alibaba closed 262k n/a n/a
2025-08-21 DeepSeek-V3.1 DeepSeek open 128k n/a n/a
2025-08-18 Nemotron Nano 2 NVIDIA open 131k n/a n/a
2025-08-07 GPT-5 OpenAI closed 400k $1.25 / $10 SWE-bench Verified 74.9%; GPQA Diamond 88.4% (GPT-5 pro)
2025-08-05 Claude Opus 4.1 Anthropic closed 200k $15 / $75 SWE-bench Verified 74.5%
2025-08-05 gpt-oss-120b OpenAI open 128k n/a GPQA Diamond 80.1%
2025-07-28 GLM-4.5 Z.ai (Zhipu) open 128k $0.60 / $2.20 SWE-bench Verified 64.2%; GPQA 79.1%
2025-07-11 Kimi K2 Moonshot AI open 128k $0.60 / $2.50 SWE-bench Verified 65.8%
2025-07-09 Grok 4 xAI closed 256k $3 / $15 Humanity’s Last Exam 25.4% (no tools); GPQA 87.5%
2025-06-16 MiniMax-M1 MiniMax open 1M n/a n/a
2025-05-22 Claude 4 (Opus 4 / Sonnet 4) Anthropic closed 200k $15 / $75 SWE-bench Verified 72.5% (Opus 4); Terminal-bench 43.2% (Opus 4); GPQA Diamond 79.6% (Opus 4)
2025-04-28 Qwen3 Alibaba open 131k n/a AIME 2024 85.7 (235B-A22B); AIME 2025 81.5 (235B-A22B); LiveCodeBench 70.7 (235B-A22B)
2025-04-16 o3 OpenAI closed 200k $10 / $40 GPQA Diamond 83.3%; SWE-bench Verified 69.1%
2025-04-14 GPT-4.1 OpenAI closed 1M $2 / $8 SWE-bench Verified 54.6%; GPQA Diamond 66.3%
2025-04-07 Llama-3.1-Nemotron-Ultra-253B NVIDIA open 131k n/a GPQA (reasoning on) 76.0%
2025-04-05 Llama 4 (Scout/Maverick) Meta open 10M n/a GPQA Diamond 69.8% (Maverick)
2025-03-25 Gemini 2.5 Pro Google closed 1M $1.25 / $10 GPQA Diamond 84.0%; SWE-bench Verified 63.8%
2025-03-16 Ernie 4.5 Baidu open n/a n/a n/a
2025-03-12 Gemma 3 Google open 128k n/a GPQA Diamond 42.4% (27B)
2025-02-27 GPT-4.5 OpenAI closed 128k $75 / $150 GPQA Diamond 71.4%; SWE-bench Verified 38.0%
2025-02-24 Claude 3.7 Sonnet Anthropic closed 200k $3 / $15 SWE-bench Verified 62.3%; GPQA Diamond 78.2% (extended thinking)
2025-02-17 Grok 3 xAI closed 131k $3 / $15 GPQA Diamond 84.6% (Think)
2025-01-22 Doubao 1.5 Pro ByteDance closed n/a n/a n/a
2025-01-20 DeepSeek-R1 DeepSeek open 128k $0.55 / $2.19 GPQA Diamond 71.5%; SWE-bench Verified 49.2%
2024-12-26 DeepSeek-V3 DeepSeek open 128k $0.27 / $1.10 GPQA Diamond 59.1%
2024-12-11 Gemini 2.0 Flash Google closed 1M n/a GPQA Diamond 62.1%
2024-12-03 Amazon Nova Amazon closed 300k n/a n/a
2024-11-05 Hunyuan-Large Tencent open 256k n/a n/a
2024-09-19 Qwen2.5 Alibaba open 131k n/a MMLU 85+; HumanEval 85+; MATH 80+
2024-09-12 o1 OpenAI closed 128k $15 / $60 GPQA Diamond 78.0%
2024-08-13 Grok-2 xAI open n/a n/a n/a
2024-07-23 Llama 3.1 405B Meta open 128k n/a MMLU 88.6%; HumanEval 89.0%; GSM8K 96.8%
2024-06-20 Claude 3.5 Sonnet Anthropic closed 200k $3 / $15 GPQA Diamond 59.4%
2024-06-14 Nemotron-4 340B NVIDIA open 4k n/a Arena Hard 54.2; MT-Bench (GPT-4-Turbo judge) 8.22
2024-05-13 GPT-4o OpenAI closed 128k $5 / $15 GPQA 53.6%
2024-05-06 DeepSeek-V2 DeepSeek open 128k $0.14 / $0.28 MMLU 77.8 (Chat RL); HumanEval 81.1 (Chat RL); GSM8K 92.2 (Chat RL)
2024-04-18 Llama 3 (8B/70B) Meta open 8k n/a MMLU 82.0% (70B); HumanEval 81.7% (70B); GSM8K 93.0% (70B)
2024-03-04 Claude 3 (Haiku/Sonnet/Opus) Anthropic closed 200k $15 / $75 GPQA Diamond 50.4% (Opus)
2024-02-26 Mistral Large Mistral AI closed 32k $8 / $24 MMLU 81.2%; HumanEval 45.1%
2024-02-15 Gemini 1.5 Pro Google closed 1M n/a MMLU 81.9%; GSM8K 91.7%; HumanEval 71.9%
2023-12-06 Gemini 1.0 (Ultra/Pro/Nano) Google closed 32k n/a MMLU 90.0% (Ultra, CoT@32); GSM8K 94.4% (Ultra); HumanEval 74.4% (Ultra)
2023-11-03 Grok-1 xAI open 8k n/a MMLU 73.0%; GSM8K 62.9%; HumanEval 63.2%
2023-09-27 Mistral 7B Mistral AI open 8k n/a MMLU 60.1%; HumanEval 30.5%; GSM8K 52.2% (8-shot)
2023-07-18 Llama 2 Meta open 4k n/a MMLU 68.9% (70B); GSM8K 56.8% (70B); HumanEval 29.9% (70B)
2023-07-11 Claude 2 Anthropic closed 100k $11.02 / $32.68 HumanEval (Codex P@1) 71.2%; GSM8K 88.0%; Bar exam (MBE) 76.5%
2023-05-10 PaLM 2 Google closed 8k n/a n/a
2023-03-16 Ernie Bot Baidu closed n/a n/a n/a
2023-03-14 Claude 1 Anthropic closed 9k n/a n/a
2023-03-14 GPT-4 OpenAI closed 8k $30 / $60 MMLU 86.4%; HumanEval 67.0%; GSM8K 92.0%
2022-11-30 GPT-3.5 / ChatGPT OpenAI closed 4k n/a n/a

Table 1. Date, weights, context window, list price, and vendor launch scores. Default view is Frontier labs (Anthropic, OpenAI, Google, xAI, Meta, Amazon, DeepSeek, Mistral AI). 114 models tracked; scroll or switch filters for the full set. Showing 114 of 114 models Source: vendor announcements.

Download CSV

Closing

As examined earlier, major coding evals had to get harder twice this summer to keep up with model gains. Because those versions are different tests, you cannot read one cost slope through the as-of tabs.

Open-weight supply, primarily from China, is still the pricing pressure on the latest snapshot. How far that supply climbs toward closed frontier capability remains unsettled.

Sources

Model and eval providers below are from their owners. Hugging Face's Summer 2026 open-models note helped on open-weight coverage. The compiled table and original charts may be reused with attribution to ryansmedstad.com under a Creative Commons Attribution 4.0 license. Formatted citation and BibTeX are in the Copy menu at the top of this page.