AI Model Release Tracker

Last updated August 5, 2026

This tracks how major AI models looked the day they launched: context window, price, modalities, weights availability, and launch benchmark claims, plus the dated changes that followed. Most trackers overwrite that history with current specs. This one keeps the launch-state snapshot alongside it, hand-curated from vendor announcements.

89

Models tracked

17

Companies

6.3/yr

Avg releases, US big four (2026 pace)

$2.00

Median input price, active models (per 1M tokens)

16 days

US frontier release cadence, past year

Release timeline by company

One row per company, ordered by its first release. Each dot marks a launch date. The "2025" and "2026 pace" columns on the right show release counts for all releases, regardless of the granularity toggle below.

  • United States
  • China
  • France
  • filled = closed weights
  • outline = open weights today
AI model releases by company and launch date
CompanyModelRelease dateWeights
OpenAI GPT-3.5 / ChatGPT 2022-11-30 closed
OpenAI GPT-4 2023-03-14 closed
OpenAI GPT-4o 2024-05-13 closed
OpenAI o1 2024-09-12 closed
OpenAI GPT-4.5 2025-02-27 closed
OpenAI GPT-4.1 2025-04-14 closed
OpenAI o3 2025-04-16 closed
OpenAI gpt-oss-120b 2025-08-05 open
OpenAI GPT-5 2025-08-07 closed
OpenAI GPT-5.1 2025-11-12 closed
OpenAI GPT-5.2 2025-12-11 closed
OpenAI GPT-5.4 2026-03-05 closed
OpenAI GPT-5.5 2026-04-23 closed
OpenAI GPT-5.6 2026-07-09 closed
Anthropic Claude 1 2023-03-14 closed
Anthropic Claude 2 2023-07-11 closed
Anthropic Claude 3 (Haiku/Sonnet/Opus) 2024-03-04 closed
Anthropic Claude 3.5 Sonnet 2024-06-20 closed
Anthropic Claude 3.7 Sonnet 2025-02-24 closed
Anthropic Claude 4 (Opus 4 / Sonnet 4) 2025-05-22 closed
Anthropic Claude Opus 4.1 2025-08-05 closed
Anthropic Claude Sonnet 4.5 2025-09-29 closed
Anthropic Claude Haiku 4.5 2025-10-15 closed
Anthropic Claude Opus 4.5 2025-11-24 closed
Anthropic Claude Opus 4.6 2026-02-05 closed
Anthropic Claude Sonnet 4.6 2026-02-17 closed
Anthropic Claude Opus 4.7 2026-04-16 closed
Anthropic Claude Opus 4.8 2026-05-28 closed
Anthropic Claude Fable 5 2026-06-09 closed
Anthropic Claude Sonnet 5 2026-06-30 closed
Anthropic Claude Opus 5 2026-07-24 closed
Google PaLM 2 2023-05-10 closed
Google Gemini 1.0 (Ultra/Pro/Nano) 2023-12-06 closed
Google Gemini 1.5 Pro 2024-02-15 closed
Google Gemini 2.0 Flash 2024-12-11 closed
Google Gemma 3 2025-03-12 open
Google Gemini 2.5 Pro 2025-03-25 closed
Google Gemini 3 Pro 2025-11-18 closed
Google Gemma 4 2026-04-02 open
Google Gemini 3.5 Flash 2026-05-19 closed
Google Gemini 3.6 Flash 2026-07-21 closed
Meta Llama 2 2023-07-18 open
Meta Llama 3 (8B/70B) 2024-04-18 open
Meta Llama 3.1 405B 2024-07-23 open
Meta Llama 4 (Scout/Maverick) 2025-04-05 open
Meta Muse Spark 2026-04-08 closed
Mistral AI Mistral 7B 2023-09-27 open
Mistral AI Mistral Large 2024-02-26 closed
Mistral AI Mistral Large 3 2025-12-02 open
DeepSeek DeepSeek-V2 2024-05-06 open
DeepSeek DeepSeek-V3 2024-12-26 open
DeepSeek DeepSeek-R1 2025-01-20 open
DeepSeek DeepSeek-V3.1 2025-08-21 open
DeepSeek DeepSeek-V3.2 2025-12-01 open
DeepSeek DeepSeek-V4 2026-04-24 open
xAI Grok-1 2023-11-03 open
xAI Grok-2 2024-08-13 open
xAI Grok 3 2025-02-17 closed
xAI Grok 4 2025-07-09 closed
xAI Grok 4.1 2025-11-17 closed
xAI Grok 4.5 2026-07-16 closed
Alibaba Qwen2.5 2024-09-19 open
Alibaba Qwen3 2025-04-28 open
Alibaba Qwen3-Max 2025-09-23 closed
Alibaba Qwen3.5 2026-02-16 open
Alibaba Qwen3.6 2026-04-15 open
Alibaba Qwen3.7-Max 2026-05-18 closed
Alibaba Qwen3.8-Max 2026-07-19 open
Moonshot AI Kimi K2 2025-07-11 open
Z.ai (Zhipu) GLM-4.5 2025-07-28 open
Z.ai (Zhipu) GLM-4.6 2025-09-30 open
Z.ai (Zhipu) GLM-5 2026-02-11 open
Thinking Machines Lab Inkling 2026-07-15 open
NVIDIA Nemotron-4 340B 2024-06-14 open
NVIDIA Llama-3.1-Nemotron-Ultra-253B 2025-04-07 open
NVIDIA Nemotron Nano 2 2025-08-18 open
NVIDIA Nemotron 3 2025-12-15 open
MiniMax MiniMax-M1 2025-06-16 open
MiniMax MiniMax-M2 2025-10-22 open
MiniMax MiniMax-M3 2026-06-02 open
ByteDance Doubao 1.5 Pro 2025-01-22 closed
ByteDance Doubao Seed 2.0 2026-02-14 closed
Baidu Ernie Bot 2023-03-16 closed
Baidu Ernie 4.5 2025-03-16 open
Baidu Ernie 5.0 2025-11-13 closed
Tencent Hunyuan-Large 2024-11-05 open
Tencent Hy3 2026-07-02 open
Amazon Amazon Nova 2024-12-03 closed
Amazon Amazon Nova 2 2025-12-02 closed

Frontier benchmark scores by company

The running best launch score for each company with two or more scoring releases on the selected benchmark, as claimed in vendor launch materials, linear scale from 0 to 100. Solid segments run from one release to the next; dotted segments show the frontier being held, with no new release raising it, through the chart's right edge (or, for a superseded benchmark version, up to the version break). These are vendor launch claims, not a third-party re-run, so cross-vendor comparability is imperfect. SWE-bench and Terminal-Bench each cover two benchmark versions; the dotted vertical divider marks where vendor reporting switched from one to the other. Terminal-Bench and Humanity's Last Exam carry the 2026 frontier, since SWE-bench Verified fell out of vendor launch reporting in 2026.

  • Anthropic
  • DeepSeek
  • Google
  • Moonshot AI
  • OpenAI
  • Z.ai (Zhipu)
  • xAI

Single-score models shown as points without a trend line: SWE-bench Verified: Kimi K2, Gemini 2.5 Pro; SWE-bench Pro: Grok 4.5 .

Frontier-raising launch SWE-bench score by company, model, and release date
CompanyModelRelease dateVersion SWE-bench score at launch
Anthropic Claude 3.7 Sonnet 2025-02-24 SWE-bench Verified 62.3%
Anthropic Claude 4 (Opus 4 / Sonnet 4) 2025-05-22 SWE-bench Verified 72.5%
Anthropic Claude Opus 4.1 2025-08-05 SWE-bench Verified 74.5%
Anthropic Claude Sonnet 4.5 2025-09-29 SWE-bench Verified 77.2%
Anthropic Claude Opus 4.5 2025-11-24 SWE-bench Verified 80.9%
Z.ai (Zhipu) GLM-4.5 2025-07-28 SWE-bench Verified 64.2%
Z.ai (Zhipu) GLM-5 2026-02-11 SWE-bench Verified 77.8%
OpenAI GPT-4.5 2025-02-27 SWE-bench Verified 38%
OpenAI GPT-4.1 2025-04-14 SWE-bench Verified 54.6%
OpenAI o3 2025-04-16 SWE-bench Verified 69.1%
OpenAI GPT-5 2025-08-07 SWE-bench Verified 74.9%
OpenAI GPT-5.1 2025-11-12 SWE-bench Verified 76.3%
DeepSeek DeepSeek-R1 2025-01-20 SWE-bench Verified 49.2%
DeepSeek DeepSeek-V3.2 2025-12-01 SWE-bench Verified 73.1%
Moonshot AI Kimi K2 2025-07-11 SWE-bench Verified 65.8%
Google Gemini 2.5 Pro 2025-03-25 SWE-bench Verified 63.8%
Anthropic Claude Fable 5 2026-06-09 SWE-bench Pro 80.3%
OpenAI GPT-5.5 2026-04-23 SWE-bench Pro 58.6%
OpenAI GPT-5.6 2026-07-09 SWE-bench Pro 64.6%
xAI Grok 4.5 2026-07-16 SWE-bench Pro 64.7%

Capability and industry milestones (11)

A short list of releases and industry moments that changed what was available at launch, drawn from the dataset's notable lines.

  • Nov 2022 GPT-3.5 / ChatGPT launches, taking LLMs mainstream and kicking off the current AI race.
  • Mar 2023 GPT-4 ships and defines the state of the art for over a year; its announced image input stays gated until GPT-4V that fall.
  • Jul 2023 Llama 2 becomes the first commercially usable open-weight frontier-adjacent model, seeding the modern open ecosystem.
  • Feb 2024 Gemini 1.5 Pro makes million-token context real, an order of magnitude beyond anything shipping at the time.
  • Sep 2024 o1 introduces chain-of-thought test-time compute as a product, the first mainstream reasoning model.
  • Oct 2024 Claude 3.5 Sonnet ships computer use in beta alongside an upgraded model.
  • Jan 2025 DeepSeek-R1 releases as an MIT-licensed o1-class reasoning model, triggering a global market shock and a wave of RL replication.
  • Feb 2025 Claude 3.7 Sonnet ships as the first hybrid reasoning model, blending instant replies and extended thinking, alongside Claude Code.
  • Apr 2025 GPT-4.1 brings a 1M-token context window to OpenAI’s API-first lineup.
  • Sep 2025 Claude Sonnet 4.5 ships with 30-plus-hour autonomous agent runs, marking agentic coding going mainstream.
  • Jul 2026 Thirty-four companies and organizations, including OpenAI, Meta, Microsoft, Mistral, NVIDIA, and Hugging Face, sign a joint letter arguing open-weight models are essential to American AI leadership; Anthropic is notably absent.

The companies behind the models

Reported valuations for the private labs in this dataset, from their public funding announcements. End labels double as the legend.

Valuations as reported in each round's press coverage or company announcement at the time of the raise; see Sources and references below for links to each company.

Valuation events for private AI labs
CompanyDateValuationRoundDetail
Anthropic 2024-03-27 $18.4B Amazon investment (final tranche) Completed Amazon's $4B strategic investment announced Sep 2023
Anthropic 2025-03-03 $61.5B Series E Lightspeed-led round at $61.5B post-money
Anthropic 2025-09-02 $183B Series F ICONIQ-led, co-led by Fidelity and Lightspeed
Anthropic 2026-02-12 $380B Series G Led by GIC and Coatue; co-led by D. E. Shaw Ventures, Dragoneer, Founders Fund, ICONIQ, MGX
Anthropic 2026-05-28 $965B Series H Led by Altimeter, Dragoneer, Greenoaks, Sequoia; includes ~$15B previously committed hyperscaler capital; overtook OpenAI as most valuable AI startup
OpenAI 2023-04-28 $29B Tender offer Employee share sale led by Thrive, Sequoia, a16z after the ChatGPT breakout
OpenAI 2024-10-02 $157B Venture round Thrive-led round with Microsoft, Nvidia, SoftBank participating
OpenAI 2025-03-31 $300B SoftBank-led round Largest private tech raise to date; tied to Stargate infrastructure buildout
OpenAI 2025-10-02 $500B Secondary share sale Employee tender at $500B made OpenAI the most valuable private company; followed the for-profit restructuring
OpenAI 2026-03-31 $852B Strategic round $122B committed capital; Amazon (~$50B), Nvidia (~$30B), SoftBank (~$30B) anchoring
xAI 2024-05-26 $24B Series B Valor, Vy Capital, Sequoia, a16z among backers
xAI 2024-12-23 $50B Series C Included Nvidia and AMD as strategic investors
xAI 2025-03-28 $80B xAI-X merger All-stock acquisition of X (Twitter) valuing xAI at $80B and X at $33B
xAI 2025-09-30 $200B Venture round Reported ~$10B raise around the Grok 4 era at ~$200B
xAI 2026-01-06 $230B Series E Nvidia- and Cisco-backed; Valor, StepStone, Fidelity, QIA, MGX, Baron participating
Moonshot AI 2024-02-21 $2.5B Series B Alibaba-led round, among the largest for a Chinese AI startup at the time
Moonshot AI 2024-08-05 $3.3B Extension Tencent and Gaorong joined at $3.3B
Moonshot AI 2026-05-07 $20B Venture round Led by Meituan's Long-Z Investments with Tsinghua Capital, China Mobile, CPE Yuanfeng; China's top-funded LLM startup
Mistral AI 2023-12-11 $2B Series A a16z-led round (~EUR 385M) six months after founding
Mistral AI 2024-06-11 $6.2B Series B General Catalyst-led ~EUR 600M at ~EUR 5.8B valuation
Mistral AI 2025-09-09 $13.7B Series C ASML-led EUR 1.7B (ASML took ~11% for EUR 1.3B) at EUR 11.7B post-money; Nvidia, a16z, General Catalyst, Bpifrance participating
Z.ai (Zhipu) 2024-12-17 $2.8B Venture round ~CNY 3B round with state-affiliated and corporate backers
Z.ai (Zhipu) 2026-01-08 $7B Hong Kong IPO First LLM company to list publicly (HKEX 2513); raised ~HK$4.17B at ~US$7B; shares later surged to a ~$93B market cap by July 2026

Valuations are reported figures from private funding rounds as covered in the press; public parent companies (Google, Meta, Alibaba) are excluded rather than assigned artificial division valuations.

Single-round companies not plotted: DeepSeek, MiniMax, Thinking Machines Lab.

Founded Ownership Notable
Anthropic 2021 Private $965B (May 2026) Founded by ex-OpenAI researchers (Amodei siblings); major backing from both Amazon and Google; widely expected to IPO after the Series H
OpenAI 2015 Private $852B (Mar 2026) Confidentially filed a draft S-1 in June 2026; reportedly targeting a ~$1T IPO valuation
xAI 2023 Private $230B (Jan 2026) Elon Musk's lab; merged with X Corp in March 2025, so valuations after that date are for the combined entity
DeepSeek 2023 Private $55B (Jun 2026) Bankrolled solely by founder Liang Wenfeng's High-Flyer hedge fund until June 2026, no outside investors for its first three years; July 2026 reports of a follow-on at ~$74B ahead of a planned IPO
Moonshot AI 2023 Private $20B (May 2026) Kimi K2 made it the breakout open-weight lab of 2025; reports of a follow-on near $30B+ and a Hong Kong IPO were unconfirmed as of July 2026
Mistral AI 2023 Private $13.7B (Sep 2025) Backed by ASML and the French state's Bpifrance; reported mid-2026 talks to raise ~EUR 3B at ~EUR 20B were not yet closed as of July 2026
Thinking Machines Lab 2025 Private $12B (Jul 2025) Late-2025 talks to raise at a ~$50B valuation collapsed by January 2026 without a deal, leaving it at its $12B seed valuation
MiniMax 2021 Public $11.4B (Jan 2026) Publicly listed on the Hong Kong Stock Exchange since January 2026; post-IPO market cap moves daily and is not tracked here beyond the listing event
Z.ai (Zhipu) 2019 Public $7B (Jan 2026) Backers pre-IPO included Alibaba, Tencent, Meituan, Xiaomi and Saudi Aramco's Prosperity7; post-IPO market cap is set by the market, not funding rounds
Alibaba (Qwen) 2023 Public parent n/a Qwen team sits inside Alibaba Cloud (first Tongyi Qianwen release April 2023); funded by Alibaba, which pledged a multi-year AI/cloud capex program of tens of billions of dollars; no standalone valuation
Amazon (AGI / Nova) 2023 Public parent n/a Nova is developed by Amazon's in-house AGI team and funded from Amazon.com's (Nasdaq: AMZN) balance sheet; no standalone valuation exists for the division
Baidu 2000 Public parent n/a Baidu is itself the public parent (Nasdaq: BIDU, HKEX: 9888); Ernie is funded from Baidu's own balance sheet with no standalone divisional valuation
ByteDance (Seed) 2023 Private n/a Seed is a research org inside parent ByteDance, not a separately funded entity; ByteDance itself was estimated at roughly $300B in employee tender-offer activity as of November 2024 (see en.wikipedia.org/wiki/ByteDance), with no standalone valuation for the model team
Google DeepMind 2010 Public parent n/a Merged from DeepMind (founded 2010, acquired 2014) and Google Brain in 2023; funded from Alphabet's balance sheet, no standalone valuation exists, and Alphabet's AI-driven capex has run in the tens of billions per year
Meta AI (Superintelligence Labs) 2013 Public parent n/a FAIR founded 2013 under Yann LeCun; Meta's 2025 pivot included a ~$14.3B stake in Scale AI to recruit Alexandr Wang and an aggressive talent raid on rival labs; division has no standalone valuation
NVIDIA 1993 Public n/a NVIDIA is a directly public company (Nasdaq: NVDA) rather than a lab housed inside a larger parent; it surpassed a $4 trillion market cap in mid-2025 on AI accelerator demand, and Nemotron is its own open-model line rather than a division's side project
Tencent (Hunyuan) 2023 Public parent n/a Hunyuan sits inside Tencent Holdings (HKEX: 0700), funded from Tencent's own cloud and R&D budget; no standalone valuation exists for the model team

All 89 tracked models, most recent release first.

Modalities at launch: T text, I image, A audio, V video. Gray chips are inputs, green chips are outputs.

Weights: open means weights are downloadable, whether under a permissive license (Apache 2.0, MIT) or a usage-restricted one (Llama, Gemma); hover a badge for the license nuance. Closed means API or product access only. Badges show launch state; "open since" notes weights released later.

Status: active means currently offered as a first-line model. legacy means superseded by newer models but still served. deprecated means shutdown announced, still accessible until its retirement date. retired means no longer accessible.

Benchmark scores are as claimed in each vendor's launch materials; the changing mix of benchmarks across rows reflects how evaluations saturate and get replaced.

Data is hand-curated from vendor announcements. Each row's sources live in the repository dataset.

Showing 89 of 89 models

Weights Modalities Launch benchmarks Notable
2026-07-24 Claude Opus 5 Anthropic closed text and image in, text out 1M $5/$25 active Humanity’s Last Exam 56.3% · OSWorld 2.0 70.6% · BrowseComp 90.8% State-of-the-art on coding and knowledge work (Frontier-Bench, GDPval-AA) at a third of Fable 5's price; the launch post led with charts rather than published score tables.
2026-07-21 Gemini 3.6 Flash Google closed text and image and audio and video in, text out 1M $1.50/$7.50 active MLE-Bench 63.9% · DeepSWE 49% Efficiency-focused successor to 3.5 Flash, using up to 17% fewer output tokens; announced alongside 3.5 Flash-Lite and the security-focused 3.5 Flash Cyber, with 3.5 Pro still in testing.
2026-07-19 Qwen3.8-Max Alibaba open text in, text out n/a n/a active none published 2.4-trillion-parameter flagship previewed with open weights promised, days after Moonshot AI's competing Kimi K3.
2026-07-16 Grok 4.5 xAI closed text and image in, text out 500k $2/$6 active SWE-bench Pro 64.7% · Terminal-Bench 2.1 83.3% Positioned as Opus-class at a fraction of the price.
2026-07-15 Inkling Thinking Machines Lab open text and image and audio and video in, text out 1M n/a active none published Mira Murati's lab's first model: a 975B-parameter MoE (about 41B active) trained on 45T multimodal tokens, released under Apache 2.0 and positioned as a customization base for the Tinker platform rather than a frontier flagship.
2026-07-09 GPT-5.6 OpenAI closed text and image in, text out 1M $5/$30 active GPQA Diamond 94.6% · Terminal-Bench 2.1 88.8% · SWE-bench Pro 64.6% First release split into three durable capability tiers, Sol (flagship), Terra, and Luna, each advancing on its own cadence; Sol Ultra multi-agent mode tops Terminal-Bench 2.1 at 91.9%.
2026-07-02 Hy3 Tencent open text in, text out 256k n/a active none published 295B MoE (21B active, 3.8B MTP) under Apache 2.0; in Tencent's own blind evaluation across 270 experts, Hy3 scored 2.67/4 versus GLM-5.1's 2.51/4, with the largest edge in frontend development, data and storage, and CI/CD tasks.
2026-06-30 Claude Sonnet 5 Anthropic closed text and image in, text out 1M $2/$10 active Humanity’s Last Exam 34.6% (no tools) · SWE-bench Verified 72.7% · SWE-bench Pro 63.2% Mid-tier model pitched as running agent workloads that recently required flagship models, at introductory discount pricing.
2026-06-09 Claude Fable 5 Anthropic closed text and image in, text out 1M $10/$50 active SWE-bench Pro 80.3% · Terminal-Bench 2.1 88.0% First Mythos-class model above the Opus tier, built for days-long autonomous tasks; briefly suspended under export controls days after launch.
2026-06-02 MiniMax-M3 MiniMax open text and image and video in, text out 1M n/a active none published 428B MoE (about 23B active) with MiniMax Sparse Attention; natively multimodal, released under a MiniMax community license.
2026-05-28 Claude Opus 4.8 Anthropic closed text and image in, text out 1M $5/$25 legacy none published Honesty-focused upgrade about four times less likely to let flaws in its own code pass unremarked, plus a $10/$50 fast mode and a parallel-subagent preview in Claude Code.
2026-05-19 Gemini 3.5 Flash Google closed text and image and audio and video in, text out 1M $1.50/$9 legacy none published I/O 2026 Flash generation succeeded by 3.6 Flash two months later.
2026-05-18 Qwen3.7-Max Alibaba closed text in, text out n/a n/a legacy none published Proprietary flagship variant, with the Qwen3.7-Plus tier following a month later.
2026-04-24 DeepSeek-V4 DeepSeek open text in, text out 1M n/a active none published V4-Pro and V4-Flash preview with 1M context and open weights.
2026-04-23 GPT-5.5 OpenAI closed text and image in, text out 1M $5/$30 active Terminal-Bench 2.0 82.7% · ARC-AGI-2 85.0% · FrontierMath Tiers 1-3 51.7% OpenAI's current frontier model, pushing 1M context and long-horizon agentic work.
2026-04-16 Claude Opus 4.7 Anthropic closed text and image in, text out 1M $5/$25 legacy CursorBench 70% · BigLaw Bench (Harvey) 90.9% (high effort) Better advanced software engineering with a new xhigh effort level and higher-resolution vision, though a new tokenizer maps the same text to roughly 1 to 1.35 times as many tokens.
2026-04-15 Qwen3.6 Alibaba open text in, text out 131k n/a active none published Apache-licensed open-weight release the same month as the proprietary Qwen3.6-Plus.
2026-04-08 Muse Spark Meta closed text and image in, text out n/a n/a active Humanity’s Last Exam 58% (Contemplating) · FrontierScience Research 38% (Contemplating) First model from Meta Superintelligence Labs and a pivot away from the open-weight Llama strategy; consumer-only at launch with a multi-agent Contemplating mode.
2026-04-02 Gemma 4 Google open text and image in, text out 256k n/a active MMLU-Pro 85.2% (31B) · LiveCodeBench v6 80.0% (31B) · AIME 2026 89.2% (31B, no tools) Dropped the restrictive Gemma license for Apache 2.0; four variants (E2B to 31B dense) distilled from Gemini 3, with the 31B ranking third among open models on the Arena text leaderboard.
2026-03-05 GPT-5.4 OpenAI closed text and image in, text out 1M $2.50/$15 legacy none published Thinking and Pro at launch with mini and nano following two weeks later; the mainline 5.3 number was skipped, with only the Codex coding variant shipping under it.
2026-02-17 Claude Sonnet 4.6 Anthropic closed text and image in, text out 1M $3/$15 legacy none published Full upgrade of Sonnet's coding, computer use, long-context reasoning, agent planning, knowledge work, and design skills, with a 1M-token context window in beta at unchanged $3/$15 pricing.
2026-02-16 Qwen3.5 Alibaba open text in, text out 131k n/a legacy none published 397B-A17B open-weight release alongside the proprietary Qwen3.5-Plus, able to operate desktop and mobile applications.
2026-02-14 Doubao Seed 2.0 ByteDance closed text and image in, text out n/a n/a active none published Pro variant positioned against GPT-5.2 and Gemini 3 Pro for long-chain reasoning and agentic tasks; ByteDance said it leads on multiple multimodal, math, and coding benchmarks at roughly a tenth of competitors' token pricing.
2026-02-11 GLM-5 Z.ai (Zhipu) open text in, text out 200k n/a active SWE-bench Verified 77.8% · GPQA Diamond 86.0% Frontier release scaling from GLM-4.5's 355B to 744B total parameters (40B active) with DeepSeek Sparse Attention, closing the gap with frontier closed models.
2026-02-05 Claude Opus 4.6 Anthropic closed text and image in, text out 1M $5/$25 legacy SWE-bench Verified 80.9% · Humanity’s Last Exam 53.0% (with tools) · BrowseComp 86.8% (multi-agent harness) Brought Opus to a 1M-token context in beta at unchanged pricing, with premium long-context rates above 200k tokens.
2025-12-15 Nemotron 3 NVIDIA open text in, text out 1M n/a active none published NVIDIA's flagship open family in Nano, Super, and Ultra sizes built on a hybrid Mamba-attention MoE with a 1M-token context; the 550B/55B-active Ultra followed in June 2026 as NVIDIA's largest open-weight model; NVIDIA Nemotron Open Model License.
2025-12-11 GPT-5.2 OpenAI closed text and image in, text out 400k $1.75/$14 legacy none published The Code Red response to Gemini 3 Pro, shipped in Instant, Thinking, and Pro variants.
2025-12-02 Mistral Large 3 Mistral AI open text and image in, text out 256k $0.50/$1.50 active none published 675B MoE flagship family under Apache 2.0.
2025-12-02 Amazon Nova 2 Amazon closed text and image and video and audio in, text and image out 1M n/a active none published Second-generation family (Lite, Pro, Omni, Sonic) with reasoning and 1M context; launched alongside Nova Forge custom training and Nova Act agents.
2025-12-01 DeepSeek-V3.2 DeepSeek open text in, text out 128k $0.28/$0.42 active AIME 2025 93.1% · GPQA Diamond 82.4% · SWE-bench Verified 73.1% Sparse-attention flagship with integrated reasoning-mode tool calls, plus a Speciale variant hitting olympiad gold-level results.
2025-11-24 Claude Opus 4.5 Anthropic closed text and image in, text out 200k $5/$25 active SWE-bench Verified 80.9% Cut Opus pricing by two-thirds while topping SWE-bench Verified, ending Opus's premium-only positioning.
2025-11-18 Gemini 3 Pro Google closed text and image and audio and video in, text out 1M $2/$12 active GPQA Diamond 91.9% · Humanity’s Last Exam 37.5% (no tools) · ARC-AGI-2 31.1% Day-one launch across Search, the Gemini app, and the API, with record LMArena scores and generative UI output.
2025-11-17 Grok 4.1 xAI closed text and image in, text out 256k n/a legacy none published Vendor-announced flagship default across xAI surfaces, launched at the top of LMArena.
2025-11-13 Ernie 5.0 Baidu closed text and image and audio and video in, text and image and video out n/a n/a active none published 2.4T-parameter natively omni-modal flagship unveiled at Baidu World 2025, jointly modeling text, images, audio, and video in a single autoregressive framework.
2025-11-12 GPT-5.1 OpenAI closed text and image in, text out 400k $1.25/$10 active SWE-bench Verified 76.3% Instant/Thinking split refined GPT-5's router with adaptive reasoning and a warmer default persona.
2025-10-22 MiniMax-M2 MiniMax open text in, text out 205k $0.30/$1.20 legacy none published Modified-MIT-licensed MoE tuned for coding and agents at a fraction of frontier pricing; became a default open agentic model.
2025-10-15 Claude Haiku 4.5 Anthropic closed text and image in, text out 200k $1/$5 active SWE-bench Verified 73.3% · Terminal-Bench 41.75% (32K thinking) Small-tier model matching five-month-old frontier (Sonnet 4) coding performance at a third of the cost.
2025-09-30 GLM-4.6 Z.ai (Zhipu) open text in, text out 200k n/a legacy none published MIT-licensed successor to GLM-4.5 with 200k context, expanded from GLM-4.5's 128K.
2025-09-29 Claude Sonnet 4.5 Anthropic closed text and image in, text out 200k $3/$15 active SWE-bench Verified 77.2% · OSWorld 61.4% Billed as the best coding model on release, with 30-plus-hour autonomous agent runs.
2025-09-23 Qwen3-Max Alibaba closed text in, text out 262k n/a legacy none published Proprietary trillion-parameter-plus flagship, trained on about 36 trillion tokens, available only via API.
2025-08-21 DeepSeek-V3.1 DeepSeek open text in, text out 128k n/a legacy none published Hybrid thinking and non-thinking modes billed as the start of the agent era.
2025-08-18 Nemotron Nano 2 NVIDIA open text in, text out 131k n/a legacy none published First fully NVIDIA-pretrained generation after the Llama derivatives: a 9B hybrid Mamba-2/Transformer reasoning model (12B base variant also released) with up to 6x the throughput of same-size Transformers, shipped with most of its pretraining corpus opened as the Nemotron Pretraining Dataset; NVIDIA Open Model License.
2025-08-07 GPT-5 OpenAI closed text and image in, text out 400k $1.25/$10 legacy SWE-bench Verified 74.9% · AIME 2025 94.6% (no tools) · GPQA Diamond 88.4% (GPT-5 pro) Unified router-based system merging fast and reasoning modes, at aggressively commoditized pricing.
2025-08-05 gpt-oss-120b OpenAI open text in, text out 128k n/a active GPQA Diamond 80.1% · MMLU-Pro 80.8% OpenAI's first open-weight language models since GPT-2, released under Apache 2.0 alongside the 21B-parameter gpt-oss-20b; the 117B MoE runs on a single 80GB GPU and lands near o4-mini on reasoning.
2025-08-05 Claude Opus 4.1 Anthropic closed text and image in, text out 200k $15/$75 legacy SWE-bench Verified 74.5% Incremental Opus upgrade focused on agentic coding, released two days before GPT-5.
2025-07-28 GLM-4.5 Z.ai (Zhipu) open text in, text out 128k $0.60/$2.20 active AIME 2024 91.0% · SWE-bench Verified 64.2% · GPQA 79.1% Agent-focused open-weight MoE undercutting even DeepSeek on price, cementing the Chinese open-model surge of mid-2025.
2025-07-11 Kimi K2 Moonshot AI open text in, text out 128k $0.60/$2.50 active SWE-bench Verified 65.8% · LiveCodeBench v6 53.7% · AIME 2025 49.5% 1T-parameter MoE under a modified MIT license, the strongest open agentic model at release.
2025-07-09 Grok 4 xAI closed text and image in, text out 256k $3/$15 active Humanity’s Last Exam 25.4% (no tools) · GPQA 87.5% · ARC-AGI-2 15.9% Reasoning-first flagship with a multi-agent Heavy tier, briefly topping several frontier benchmarks.
2025-06-16 MiniMax-M1 MiniMax open text in, text out 1M n/a legacy none published First open-weight large-scale hybrid-attention reasoning model (456B MoE, 45.9B active), Apache 2.0, 1M context.
2025-05-22 Claude 4 (Opus 4 / Sonnet 4) Anthropic closed text and image in, text out 200k $15/$75 legacy SWE-bench Verified 72.5% (Opus 4) · Terminal-bench 43.2% (Opus 4) · GPQA Diamond 79.6% (Opus 4) Opus returned as the coding and agent flagship; first Anthropic release under ASL-3 safeguards.
2025-04-28 Qwen3 Alibaba open text in, text out 131k n/a active AIME 2024 85.7 (235B-A22B) · AIME 2025 81.5 (235B-A22B) · LiveCodeBench 70.7 (235B-A22B) Hybrid thinking/non-thinking MoE family under Apache 2.0 that led open-model leaderboards through 2025.
2025-04-16 o3 OpenAI closed text and image in, text out 200k $10/$40 legacy AIME 2024 91.6% · GPQA Diamond 83.3% · SWE-bench Verified 69.1% Flagship reasoning model with agentic tool use; its later price cut reset reasoning-model economics.
2025-04-14 GPT-4.1 OpenAI closed text and image in, text out 1M $2/$8 legacy SWE-bench Verified 54.6% · MMLU 90.2% · GPQA Diamond 66.3% API-first workhorse family that brought a 1M-token context window to OpenAI's lineup.
2025-04-07 Llama-3.1-Nemotron-Ultra-253B NVIDIA open text in, text out 131k n/a legacy GPQA (reasoning on) 76.0% · AIME 2025 (reasoning on) 72.5% · LiveCodeBench (reasoning on) 66.3% Flagship of the Llama Nemotron reasoning family (Nano 8B and Super 49B debuted at GTC in March 2025): a 253B model derived from Llama-3.1-405B via Neural Architecture Search with toggleable reasoning; dual NVIDIA Open Model License plus Llama 3.1 Community License.
2025-04-05 Llama 4 (Scout/Maverick) Meta open text and image in, text out 10M n/a active MMLU-Pro 80.5% (Maverick) · GPQA Diamond 69.8% (Maverick) · LiveCodeBench 43.4% (Maverick) Meta's MoE, natively multimodal generation with a claimed 10M-token context on Scout; a contested launch that cooled Meta's open-weights momentum.
2025-03-25 Gemini 2.5 Pro Google closed text and image and audio and video in, text out 1M $1.25/$10 active GPQA Diamond 84.0% · AIME 2025 86.7% · SWE-bench Verified 63.8% Thinking-by-default flagship that put Google back at the top of leaderboards in 2025.
2025-03-16 Ernie 4.5 Baidu open open since Jun 2025 text and image in, text out n/a n/a legacy none published Multimodal flagship launched closed, then fully open-sourced under Apache 2.0, the biggest Chinese open release between DeepSeek and the summer 2025 wave.
2025-03-12 Gemma 3 Google open text and image in, text out 128k n/a legacy LMArena Elo 1339 (27B) · MMLU-Pro 67.5% (27B) · GPQA Diamond 42.4% (27B) Multimodal open-weight family (1B to 27B) under the use-restricted Gemma license; the 27B ranked top-10 on LMArena at launch while fitting on a single GPU.
2025-02-27 GPT-4.5 OpenAI closed text and image in, text out 128k $75/$150 retired GPQA Diamond 71.4% · AIME 2024 36.7% · SWE-bench Verified 38.0% OpenAI's largest pretraining-scaling bet, priced so high it was retired from the API within months.
2025-02-24 Claude 3.7 Sonnet Anthropic closed text and image in, text out 200k $3/$15 legacy SWE-bench Verified 62.3% · GPQA Diamond 78.2% (extended thinking) · TAU-bench (retail) 81.2% First hybrid reasoning model, blending instant replies and extended thinking in one model; shipped with Claude Code.
2025-02-17 Grok 3 xAI closed text in, text out 131k $3/$15 legacy AIME 2025 93.3% (Think, cons@64) · GPQA Diamond 84.6% (Think) · LiveCodeBench 79.4% (Think) Trained on the 200k-GPU Colossus cluster; introduced Think mode and DeepSearch. Image input never shipped for grok-3 in the API.
2025-01-22 Doubao 1.5 Pro ByteDance closed text and image in, text out n/a n/a legacy none published Sparse-MoE flagship matching GPT-4o-class benchmarks at a fraction of the cost; powered Doubao, China's largest AI assistant.
2025-01-20 DeepSeek-R1 DeepSeek open text in, text out 128k $0.55/$2.19 legacy AIME 2024 79.8% · GPQA Diamond 71.5% · SWE-bench Verified 49.2% MIT-licensed o1-class reasoning model whose release triggered a global market shock and a wave of RL replication.
2024-12-26 DeepSeek-V3 DeepSeek open text in, text out 128k $0.27/$1.10 legacy MMLU 88.5% · GPQA Diamond 59.1% · MATH-500 90.2% GPT-4o-class open weights reportedly trained for under $6M, upending assumptions about frontier training cost.
2024-12-11 Gemini 2.0 Flash Google closed text and image and audio and video in, text out 1M n/a legacy MMLU-Pro 76.4% · GPQA Diamond 62.1% · Natural2Code 92.9% Opened Google's agentic era with native tool use; image and audio output were announced at launch but limited to early-access partners.
2024-12-03 Amazon Nova Amazon closed text and image and video in, text out 300k n/a legacy none published Amazon's first in-house frontier family (Micro, Lite, Pro, Premier), launched at re:Invent 2024 with an emphasis on price-performance; Premier remained in training at launch and shipped the following spring.
2024-11-05 Hunyuan-Large Tencent open text in, text out 256k n/a legacy none published 389B-total, 52B-active MoE, the largest open MoE of its day; Tencent's entry into the open-weights race.
2024-09-19 Qwen2.5 Alibaba open text in, text out 131k n/a legacy MMLU 85+ · HumanEval 85+ · MATH 80+ Broad Apache-2.0 family (0.5B to 72B) that became the default base for open fine-tunes and distillations.
2024-09-12 o1 OpenAI closed text in, text out 128k $15/$60 deprecated AIME 2024 83.3% (cons@64) · GPQA Diamond 78.0% · Codeforces 89th percentile First mainstream reasoning model, introducing chain-of-thought test-time compute as a product.
2024-08-13 Grok-2 xAI open open since Aug 2025 text and image in, text out n/a n/a legacy none published Beta launch on X Premium and Premium+, alongside the smaller Grok-2 mini; the full weights were released on Hugging Face about a year later under a restricted xAI Community License.
2024-07-23 Llama 3.1 405B Meta open text in, text out 128k n/a legacy MMLU 88.6% · HumanEval 89.0% · GSM8K 96.8% First open-weight model credibly matching closed frontier models on benchmarks.
2024-06-20 Claude 3.5 Sonnet Anthropic closed text and image in, text out 200k $3/$15 retired GPQA Diamond 59.4% · HumanEval 92.0% · GSM8K 96.4% Beat Opus at a fifth of the price and became the default coding model of 2024.
2024-06-14 Nemotron-4 340B NVIDIA open text in, text out 4k n/a legacy Arena Hard 54.2 · MT-Bench (GPT-4-Turbo judge) 8.22 NVIDIA's first frontier-scale open release: a 340B dense family (Base, Instruct, Reward) under the NVIDIA Open Model License, trained on 9T tokens and positioned as a synthetic-data-generation engine; over 98% of the Instruct model's alignment data was itself synthetic.
2024-05-13 GPT-4o OpenAI closed text and image in, text out 128k $5/$15 legacy MMLU 88.7% · GPQA 53.6% · HumanEval 90.2% Natively multimodal omni model; the launch demoed real-time voice, but the API shipped with text and image in, text out.
2024-05-06 DeepSeek-V2 DeepSeek open text in, text out 128k $0.14/$0.28 retired MMLU 77.8 (Chat RL) · HumanEval 81.1 (Chat RL) · GSM8K 92.2 (Chat RL) Efficient MoE with multi-head latent attention that started China's LLM price war.
2024-04-18 Llama 3 (8B/70B) Meta open text in, text out 8k n/a legacy MMLU 82.0% (70B) · HumanEval 81.7% (70B) · GSM8K 93.0% (70B) Closed most of the open-vs-closed quality gap at the 70B scale.
2024-03-04 Claude 3 (Haiku/Sonnet/Opus) Anthropic closed text and image in, text out 200k $15/$75 retired MMLU 86.8% (Opus) · GPQA Diamond 50.4% (Opus) · HumanEval 84.9% (Opus) First release to beat GPT-4 on headline benchmarks and the origin of the three-tier Haiku/Sonnet/Opus naming.
2024-02-26 Mistral Large Mistral AI closed text in, text out 32k $8/$24 legacy MMLU 81.2% · HumanEval 45.1% Mistral's flagship API bet, launched alongside the Microsoft Azure partnership.
2024-02-15 Gemini 1.5 Pro Google closed text and image and audio and video in, text out 1M n/a retired MMLU 81.9% · GSM8K 91.7% · HumanEval 71.9% Made million-token context real, an order of magnitude beyond anything shipping at the time.
2023-12-06 Gemini 1.0 (Ultra/Pro/Nano) Google closed text and image and video in, text out 32k n/a retired MMLU 90.0% (Ultra, CoT@32) · GSM8K 94.4% (Ultra) · HumanEval 74.4% (Ultra) Google's first natively multimodal frontier family, unifying DeepMind and Brain efforts; audio understanding was demoed but not in the launch API.
2023-11-03 Grok-1 xAI open open since Mar 2024 text in, text out 8k n/a retired MMLU 73.0% · GSM8K 62.9% · HumanEval 63.2% xAI's debut, later notable for the Apache-2.0 release of its 314B MoE weights.
2023-09-27 Mistral 7B Mistral AI open text in, text out 8k n/a legacy MMLU 60.1% · HumanEval 30.5% · GSM8K 52.2% (8-shot) Apache-2.0 model that beat Llama 2 13B and proved small European labs could compete.
2023-07-18 Llama 2 Meta open text in, text out 4k n/a legacy MMLU 68.9% (70B) · GSM8K 56.8% (70B) · HumanEval 29.9% (70B) First commercially usable open-weight frontier-adjacent model; seeded the modern open ecosystem.
2023-07-11 Claude 2 Anthropic closed text in, text out 100k $11.02/$32.68 retired HumanEval (Codex P@1) 71.2% · GSM8K 88.0% · Bar exam (MBE) 76.5% Made 100k-token context mainstream when rivals were at 8-32k.
2023-05-10 PaLM 2 Google closed text in, text out 8k n/a retired none published Google's I/O 2023 answer to GPT-4 that powered Bard before the Gemini era.
2023-03-16 Ernie Bot Baidu closed text in, text out n/a n/a retired none published The first major Chinese answer to ChatGPT; Ernie 4.0 (Oct 2023) claimed GPT-4 parity domestically.
2023-03-14 GPT-4 OpenAI closed text in, text out 8k $30/$60 retired MMLU 86.4% · HumanEval 67.0% · GSM8K 92.0% Announced as multimodal, but image input stayed gated at launch; defined the state of the art for over a year.
2023-03-14 Claude 1 Anthropic closed text in, text out 9k n/a retired none published Anthropic's first commercial model and the debut of Constitutional AI training in production.
2022-11-30 GPT-3.5 / ChatGPT OpenAI closed text in, text out 4k n/a retired none published The ChatGPT launch that took LLMs mainstream and kicked off the current AI race.

Sources and current references

Where to check current pricing and model docs directly with each vendor, since both move faster than a hand-curated dataset can track.

Company Current pricing Docs / models
OpenAI Pricing Docs
Anthropic Pricing Docs
Google DeepMind Pricing Docs
Meta AI (Superintelligence Labs) Pricing Docs
Mistral AI Pricing Docs
DeepSeek Pricing Docs
xAI Pricing Docs
Alibaba (Qwen) Pricing Docs
Moonshot AI Pricing Docs
Z.ai (Zhipu) Pricing Docs
Thinking Machines Lab Pricing Docs
NVIDIA Pricing Docs
MiniMax Pricing Docs
ByteDance (Seed) Pricing Docs
Baidu Pricing Docs
Tencent (Hunyuan) Pricing Docs
Amazon (AGI / Nova) Pricing Docs

Every model row in the underlying dataset carries its own source URLs (vendor announcements and documentation); the dataset is available via Export above.

Benchmark scores and leaderboards move; for live comparisons check the vendors' model pages above.