AI Model Release Tracker
Last updated August 5, 2026
This tracks how major AI models looked the day they launched: context window, price, modalities, weights availability, and launch benchmark claims, plus the dated changes that followed. Most trackers overwrite that history with current specs. This one keeps the launch-state snapshot alongside it, hand-curated from vendor announcements.
89
Models tracked
17
Companies
6.3/yr
Avg releases, US big four (2026 pace)
$2.00
Median input price, active models (per 1M tokens)
16 days
US frontier release cadence, past year
Release timeline by company
One row per company, ordered by its first release. Each dot marks a launch date. The "2025" and "2026 pace" columns on the right show release counts for all releases, regardless of the granularity toggle below.
- United States
- China
- France
- filled = closed weights
- outline = open weights today
Export
| Company | Model | Release date | Weights |
|---|---|---|---|
| OpenAI | GPT-3.5 / ChatGPT | 2022-11-30 | closed |
| OpenAI | GPT-4 | 2023-03-14 | closed |
| OpenAI | GPT-4o | 2024-05-13 | closed |
| OpenAI | o1 | 2024-09-12 | closed |
| OpenAI | GPT-4.5 | 2025-02-27 | closed |
| OpenAI | GPT-4.1 | 2025-04-14 | closed |
| OpenAI | o3 | 2025-04-16 | closed |
| OpenAI | gpt-oss-120b | 2025-08-05 | open |
| OpenAI | GPT-5 | 2025-08-07 | closed |
| OpenAI | GPT-5.1 | 2025-11-12 | closed |
| OpenAI | GPT-5.2 | 2025-12-11 | closed |
| OpenAI | GPT-5.4 | 2026-03-05 | closed |
| OpenAI | GPT-5.5 | 2026-04-23 | closed |
| OpenAI | GPT-5.6 | 2026-07-09 | closed |
| Anthropic | Claude 1 | 2023-03-14 | closed |
| Anthropic | Claude 2 | 2023-07-11 | closed |
| Anthropic | Claude 3 (Haiku/Sonnet/Opus) | 2024-03-04 | closed |
| Anthropic | Claude 3.5 Sonnet | 2024-06-20 | closed |
| Anthropic | Claude 3.7 Sonnet | 2025-02-24 | closed |
| Anthropic | Claude 4 (Opus 4 / Sonnet 4) | 2025-05-22 | closed |
| Anthropic | Claude Opus 4.1 | 2025-08-05 | closed |
| Anthropic | Claude Sonnet 4.5 | 2025-09-29 | closed |
| Anthropic | Claude Haiku 4.5 | 2025-10-15 | closed |
| Anthropic | Claude Opus 4.5 | 2025-11-24 | closed |
| Anthropic | Claude Opus 4.6 | 2026-02-05 | closed |
| Anthropic | Claude Sonnet 4.6 | 2026-02-17 | closed |
| Anthropic | Claude Opus 4.7 | 2026-04-16 | closed |
| Anthropic | Claude Opus 4.8 | 2026-05-28 | closed |
| Anthropic | Claude Fable 5 | 2026-06-09 | closed |
| Anthropic | Claude Sonnet 5 | 2026-06-30 | closed |
| Anthropic | Claude Opus 5 | 2026-07-24 | closed |
| PaLM 2 | 2023-05-10 | closed | |
| Gemini 1.0 (Ultra/Pro/Nano) | 2023-12-06 | closed | |
| Gemini 1.5 Pro | 2024-02-15 | closed | |
| Gemini 2.0 Flash | 2024-12-11 | closed | |
| Gemma 3 | 2025-03-12 | open | |
| Gemini 2.5 Pro | 2025-03-25 | closed | |
| Gemini 3 Pro | 2025-11-18 | closed | |
| Gemma 4 | 2026-04-02 | open | |
| Gemini 3.5 Flash | 2026-05-19 | closed | |
| Gemini 3.6 Flash | 2026-07-21 | closed | |
| Meta | Llama 2 | 2023-07-18 | open |
| Meta | Llama 3 (8B/70B) | 2024-04-18 | open |
| Meta | Llama 3.1 405B | 2024-07-23 | open |
| Meta | Llama 4 (Scout/Maverick) | 2025-04-05 | open |
| Meta | Muse Spark | 2026-04-08 | closed |
| Mistral AI | Mistral 7B | 2023-09-27 | open |
| Mistral AI | Mistral Large | 2024-02-26 | closed |
| Mistral AI | Mistral Large 3 | 2025-12-02 | open |
| DeepSeek | DeepSeek-V2 | 2024-05-06 | open |
| DeepSeek | DeepSeek-V3 | 2024-12-26 | open |
| DeepSeek | DeepSeek-R1 | 2025-01-20 | open |
| DeepSeek | DeepSeek-V3.1 | 2025-08-21 | open |
| DeepSeek | DeepSeek-V3.2 | 2025-12-01 | open |
| DeepSeek | DeepSeek-V4 | 2026-04-24 | open |
| xAI | Grok-1 | 2023-11-03 | open |
| xAI | Grok-2 | 2024-08-13 | open |
| xAI | Grok 3 | 2025-02-17 | closed |
| xAI | Grok 4 | 2025-07-09 | closed |
| xAI | Grok 4.1 | 2025-11-17 | closed |
| xAI | Grok 4.5 | 2026-07-16 | closed |
| Alibaba | Qwen2.5 | 2024-09-19 | open |
| Alibaba | Qwen3 | 2025-04-28 | open |
| Alibaba | Qwen3-Max | 2025-09-23 | closed |
| Alibaba | Qwen3.5 | 2026-02-16 | open |
| Alibaba | Qwen3.6 | 2026-04-15 | open |
| Alibaba | Qwen3.7-Max | 2026-05-18 | closed |
| Alibaba | Qwen3.8-Max | 2026-07-19 | open |
| Moonshot AI | Kimi K2 | 2025-07-11 | open |
| Z.ai (Zhipu) | GLM-4.5 | 2025-07-28 | open |
| Z.ai (Zhipu) | GLM-4.6 | 2025-09-30 | open |
| Z.ai (Zhipu) | GLM-5 | 2026-02-11 | open |
| Thinking Machines Lab | Inkling | 2026-07-15 | open |
| NVIDIA | Nemotron-4 340B | 2024-06-14 | open |
| NVIDIA | Llama-3.1-Nemotron-Ultra-253B | 2025-04-07 | open |
| NVIDIA | Nemotron Nano 2 | 2025-08-18 | open |
| NVIDIA | Nemotron 3 | 2025-12-15 | open |
| MiniMax | MiniMax-M1 | 2025-06-16 | open |
| MiniMax | MiniMax-M2 | 2025-10-22 | open |
| MiniMax | MiniMax-M3 | 2026-06-02 | open |
| ByteDance | Doubao 1.5 Pro | 2025-01-22 | closed |
| ByteDance | Doubao Seed 2.0 | 2026-02-14 | closed |
| Baidu | Ernie Bot | 2023-03-16 | closed |
| Baidu | Ernie 4.5 | 2025-03-16 | open |
| Baidu | Ernie 5.0 | 2025-11-13 | closed |
| Tencent | Hunyuan-Large | 2024-11-05 | open |
| Tencent | Hy3 | 2026-07-02 | open |
| Amazon | Amazon Nova | 2024-12-03 | closed |
| Amazon | Amazon Nova 2 | 2025-12-02 | closed |
Frontier benchmark scores by company
The running best launch score for each company with two or more scoring releases on the selected benchmark, as claimed in vendor launch materials, linear scale from 0 to 100. Solid segments run from one release to the next; dotted segments show the frontier being held, with no new release raising it, through the chart's right edge (or, for a superseded benchmark version, up to the version break). These are vendor launch claims, not a third-party re-run, so cross-vendor comparability is imperfect. SWE-bench and Terminal-Bench each cover two benchmark versions; the dotted vertical divider marks where vendor reporting switched from one to the other. Terminal-Bench and Humanity's Last Exam carry the 2026 frontier, since SWE-bench Verified fell out of vendor launch reporting in 2026.
- Anthropic
- DeepSeek
- Moonshot AI
- OpenAI
- Z.ai (Zhipu)
- xAI
Export
Single-score models shown as points without a trend line: SWE-bench Verified: Kimi K2, Gemini 2.5 Pro; SWE-bench Pro: Grok 4.5 .
| Company | Model | Release date | Version | SWE-bench score at launch |
|---|---|---|---|---|
| Anthropic | Claude 3.7 Sonnet | 2025-02-24 | SWE-bench Verified | 62.3% |
| Anthropic | Claude 4 (Opus 4 / Sonnet 4) | 2025-05-22 | SWE-bench Verified | 72.5% |
| Anthropic | Claude Opus 4.1 | 2025-08-05 | SWE-bench Verified | 74.5% |
| Anthropic | Claude Sonnet 4.5 | 2025-09-29 | SWE-bench Verified | 77.2% |
| Anthropic | Claude Opus 4.5 | 2025-11-24 | SWE-bench Verified | 80.9% |
| Z.ai (Zhipu) | GLM-4.5 | 2025-07-28 | SWE-bench Verified | 64.2% |
| Z.ai (Zhipu) | GLM-5 | 2026-02-11 | SWE-bench Verified | 77.8% |
| OpenAI | GPT-4.5 | 2025-02-27 | SWE-bench Verified | 38% |
| OpenAI | GPT-4.1 | 2025-04-14 | SWE-bench Verified | 54.6% |
| OpenAI | o3 | 2025-04-16 | SWE-bench Verified | 69.1% |
| OpenAI | GPT-5 | 2025-08-07 | SWE-bench Verified | 74.9% |
| OpenAI | GPT-5.1 | 2025-11-12 | SWE-bench Verified | 76.3% |
| DeepSeek | DeepSeek-R1 | 2025-01-20 | SWE-bench Verified | 49.2% |
| DeepSeek | DeepSeek-V3.2 | 2025-12-01 | SWE-bench Verified | 73.1% |
| Moonshot AI | Kimi K2 | 2025-07-11 | SWE-bench Verified | 65.8% |
| Gemini 2.5 Pro | 2025-03-25 | SWE-bench Verified | 63.8% | |
| Anthropic | Claude Fable 5 | 2026-06-09 | SWE-bench Pro | 80.3% |
| OpenAI | GPT-5.5 | 2026-04-23 | SWE-bench Pro | 58.6% |
| OpenAI | GPT-5.6 | 2026-07-09 | SWE-bench Pro | 64.6% |
| xAI | Grok 4.5 | 2026-07-16 | SWE-bench Pro | 64.7% |
- Anthropic
- DeepSeek
- Meta
- NVIDIA
- OpenAI
- Z.ai (Zhipu)
- xAI
Export
Single-score models shown as points without a trend line: Llama-3.1-Nemotron-Ultra-253B, Llama 4 (Scout/Maverick) .
| Company | Model | Release date | GPQA Diamond score at launch |
|---|---|---|---|
| OpenAI | GPT-4o | 2024-05-13 | 53.6% |
| OpenAI | o1 | 2024-09-12 | 78% |
| OpenAI | o3 | 2025-04-16 | 83.3% |
| OpenAI | GPT-5 | 2025-08-07 | 88.4% |
| OpenAI | GPT-5.6 | 2026-07-09 | 94.6% |
| Gemini 2.0 Flash | 2024-12-11 | 62.1% | |
| Gemini 2.5 Pro | 2025-03-25 | 84% | |
| Gemini 3 Pro | 2025-11-18 | 91.9% | |
| xAI | Grok 3 | 2025-02-17 | 84.6% |
| xAI | Grok 4 | 2025-07-09 | 87.5% |
| Z.ai (Zhipu) | GLM-4.5 | 2025-07-28 | 79.1% |
| Z.ai (Zhipu) | GLM-5 | 2026-02-11 | 86% |
| DeepSeek | DeepSeek-V3 | 2024-12-26 | 59.1% |
| DeepSeek | DeepSeek-R1 | 2025-01-20 | 71.5% |
| DeepSeek | DeepSeek-V3.2 | 2025-12-01 | 82.4% |
| Anthropic | Claude 3 (Haiku/Sonnet/Opus) | 2024-03-04 | 50.4% |
| Anthropic | Claude 3.5 Sonnet | 2024-06-20 | 59.4% |
| Anthropic | Claude 3.7 Sonnet | 2025-02-24 | 78.2% |
| Anthropic | Claude 4 (Opus 4 / Sonnet 4) | 2025-05-22 | 79.6% |
| NVIDIA | Llama-3.1-Nemotron-Ultra-253B | 2025-04-07 | 76% |
| Meta | Llama 4 (Scout/Maverick) | 2025-04-05 | 69.8% |
- Anthropic
- OpenAI
- xAI
Export
Single-score models shown as points without a trend line: Terminal-Bench 2.0: GPT-5.5; Terminal-Bench 2.1: GPT-5.6, Claude Fable 5, Grok 4.5 .
| Company | Model | Release date | Version | Terminal-Bench score at launch |
|---|---|---|---|---|
| Anthropic | Claude 4 (Opus 4 / Sonnet 4) | 2025-05-22 | Terminal-Bench 2.0 | 43.2% |
| OpenAI | GPT-5.5 | 2026-04-23 | Terminal-Bench 2.0 | 82.7% |
| OpenAI | GPT-5.6 | 2026-07-09 | Terminal-Bench 2.1 | 88.8% |
| Anthropic | Claude Fable 5 | 2026-06-09 | Terminal-Bench 2.1 | 88% |
| xAI | Grok 4.5 | 2026-07-16 | Terminal-Bench 2.1 | 83.3% |
- Anthropic
- Meta
- xAI
Export
Single-score models shown as points without a trend line: Muse Spark, Gemini 3 Pro, Grok 4 .
| Company | Model | Release date | Humanity's Last Exam score at launch |
|---|---|---|---|
| Anthropic | Claude Opus 4.6 | 2026-02-05 | 53% |
| Anthropic | Claude Opus 5 | 2026-07-24 | 56.3% |
| Meta | Muse Spark | 2026-04-08 | 58% |
| Gemini 3 Pro | 2025-11-18 | 37.5% | |
| xAI | Grok 4 | 2025-07-09 | 25.4% |
Capability and industry milestones
(11)
A short list of releases and industry moments that changed what was available at launch, drawn from the dataset's notable lines.
- Nov 2022 GPT-3.5 / ChatGPT launches, taking LLMs mainstream and kicking off the current AI race.
- Mar 2023 GPT-4 ships and defines the state of the art for over a year; its announced image input stays gated until GPT-4V that fall.
- Jul 2023 Llama 2 becomes the first commercially usable open-weight frontier-adjacent model, seeding the modern open ecosystem.
- Feb 2024 Gemini 1.5 Pro makes million-token context real, an order of magnitude beyond anything shipping at the time.
- Sep 2024 o1 introduces chain-of-thought test-time compute as a product, the first mainstream reasoning model.
- Oct 2024 Claude 3.5 Sonnet ships computer use in beta alongside an upgraded model.
- Jan 2025 DeepSeek-R1 releases as an MIT-licensed o1-class reasoning model, triggering a global market shock and a wave of RL replication.
- Feb 2025 Claude 3.7 Sonnet ships as the first hybrid reasoning model, blending instant replies and extended thinking, alongside Claude Code.
- Apr 2025 GPT-4.1 brings a 1M-token context window to OpenAI’s API-first lineup.
- Sep 2025 Claude Sonnet 4.5 ships with 30-plus-hour autonomous agent runs, marking agentic coding going mainstream.
- Jul 2026 Thirty-four companies and organizations, including OpenAI, Meta, Microsoft, Mistral, NVIDIA, and Hugging Face, sign a joint letter arguing open-weight models are essential to American AI leadership; Anthropic is notably absent.
The companies behind the models
Reported valuations for the private labs in this dataset, from their public funding announcements. End labels double as the legend.
Export
Valuations as reported in each round's press coverage or company announcement at the time of the raise; see Sources and references below for links to each company.
| Company | Date | Valuation | Round | Detail |
|---|---|---|---|---|
| Anthropic | 2024-03-27 | $18.4B | Amazon investment (final tranche) | Completed Amazon's $4B strategic investment announced Sep 2023 |
| Anthropic | 2025-03-03 | $61.5B | Series E | Lightspeed-led round at $61.5B post-money |
| Anthropic | 2025-09-02 | $183B | Series F | ICONIQ-led, co-led by Fidelity and Lightspeed |
| Anthropic | 2026-02-12 | $380B | Series G | Led by GIC and Coatue; co-led by D. E. Shaw Ventures, Dragoneer, Founders Fund, ICONIQ, MGX |
| Anthropic | 2026-05-28 | $965B | Series H | Led by Altimeter, Dragoneer, Greenoaks, Sequoia; includes ~$15B previously committed hyperscaler capital; overtook OpenAI as most valuable AI startup |
| OpenAI | 2023-04-28 | $29B | Tender offer | Employee share sale led by Thrive, Sequoia, a16z after the ChatGPT breakout |
| OpenAI | 2024-10-02 | $157B | Venture round | Thrive-led round with Microsoft, Nvidia, SoftBank participating |
| OpenAI | 2025-03-31 | $300B | SoftBank-led round | Largest private tech raise to date; tied to Stargate infrastructure buildout |
| OpenAI | 2025-10-02 | $500B | Secondary share sale | Employee tender at $500B made OpenAI the most valuable private company; followed the for-profit restructuring |
| OpenAI | 2026-03-31 | $852B | Strategic round | $122B committed capital; Amazon (~$50B), Nvidia (~$30B), SoftBank (~$30B) anchoring |
| xAI | 2024-05-26 | $24B | Series B | Valor, Vy Capital, Sequoia, a16z among backers |
| xAI | 2024-12-23 | $50B | Series C | Included Nvidia and AMD as strategic investors |
| xAI | 2025-03-28 | $80B | xAI-X merger | All-stock acquisition of X (Twitter) valuing xAI at $80B and X at $33B |
| xAI | 2025-09-30 | $200B | Venture round | Reported ~$10B raise around the Grok 4 era at ~$200B |
| xAI | 2026-01-06 | $230B | Series E | Nvidia- and Cisco-backed; Valor, StepStone, Fidelity, QIA, MGX, Baron participating |
| Moonshot AI | 2024-02-21 | $2.5B | Series B | Alibaba-led round, among the largest for a Chinese AI startup at the time |
| Moonshot AI | 2024-08-05 | $3.3B | Extension | Tencent and Gaorong joined at $3.3B |
| Moonshot AI | 2026-05-07 | $20B | Venture round | Led by Meituan's Long-Z Investments with Tsinghua Capital, China Mobile, CPE Yuanfeng; China's top-funded LLM startup |
| Mistral AI | 2023-12-11 | $2B | Series A | a16z-led round (~EUR 385M) six months after founding |
| Mistral AI | 2024-06-11 | $6.2B | Series B | General Catalyst-led ~EUR 600M at ~EUR 5.8B valuation |
| Mistral AI | 2025-09-09 | $13.7B | Series C | ASML-led EUR 1.7B (ASML took ~11% for EUR 1.3B) at EUR 11.7B post-money; Nvidia, a16z, General Catalyst, Bpifrance participating |
| Z.ai (Zhipu) | 2024-12-17 | $2.8B | Venture round | ~CNY 3B round with state-affiliated and corporate backers |
| Z.ai (Zhipu) | 2026-01-08 | $7B | Hong Kong IPO | First LLM company to list publicly (HKEX 2513); raised ~HK$4.17B at ~US$7B; shares later surged to a ~$93B market cap by July 2026 |
Valuations are reported figures from private funding rounds as covered in the press; public parent companies (Google, Meta, Alibaba) are excluded rather than assigned artificial division valuations.
Single-round companies not plotted: DeepSeek, MiniMax, Thinking Machines Lab.
| Founded | Ownership | Notable | ||
|---|---|---|---|---|
| Anthropic | 2021 | Private | $965B (May 2026) | Founded by ex-OpenAI researchers (Amodei siblings); major backing from both Amazon and Google; widely expected to IPO after the Series H |
| OpenAI | 2015 | Private | $852B (Mar 2026) | Confidentially filed a draft S-1 in June 2026; reportedly targeting a ~$1T IPO valuation |
| xAI | 2023 | Private | $230B (Jan 2026) | Elon Musk's lab; merged with X Corp in March 2025, so valuations after that date are for the combined entity |
| DeepSeek | 2023 | Private | $55B (Jun 2026) | Bankrolled solely by founder Liang Wenfeng's High-Flyer hedge fund until June 2026, no outside investors for its first three years; July 2026 reports of a follow-on at ~$74B ahead of a planned IPO |
| Moonshot AI | 2023 | Private | $20B (May 2026) | Kimi K2 made it the breakout open-weight lab of 2025; reports of a follow-on near $30B+ and a Hong Kong IPO were unconfirmed as of July 2026 |
| Mistral AI | 2023 | Private | $13.7B (Sep 2025) | Backed by ASML and the French state's Bpifrance; reported mid-2026 talks to raise ~EUR 3B at ~EUR 20B were not yet closed as of July 2026 |
| Thinking Machines Lab | 2025 | Private | $12B (Jul 2025) | Late-2025 talks to raise at a ~$50B valuation collapsed by January 2026 without a deal, leaving it at its $12B seed valuation |
| MiniMax | 2021 | Public | $11.4B (Jan 2026) | Publicly listed on the Hong Kong Stock Exchange since January 2026; post-IPO market cap moves daily and is not tracked here beyond the listing event |
| Z.ai (Zhipu) | 2019 | Public | $7B (Jan 2026) | Backers pre-IPO included Alibaba, Tencent, Meituan, Xiaomi and Saudi Aramco's Prosperity7; post-IPO market cap is set by the market, not funding rounds |
| Alibaba (Qwen) | 2023 | Public parent | n/a | Qwen team sits inside Alibaba Cloud (first Tongyi Qianwen release April 2023); funded by Alibaba, which pledged a multi-year AI/cloud capex program of tens of billions of dollars; no standalone valuation |
| Amazon (AGI / Nova) | 2023 | Public parent | n/a | Nova is developed by Amazon's in-house AGI team and funded from Amazon.com's (Nasdaq: AMZN) balance sheet; no standalone valuation exists for the division |
| Baidu | 2000 | Public parent | n/a | Baidu is itself the public parent (Nasdaq: BIDU, HKEX: 9888); Ernie is funded from Baidu's own balance sheet with no standalone divisional valuation |
| ByteDance (Seed) | 2023 | Private | n/a | Seed is a research org inside parent ByteDance, not a separately funded entity; ByteDance itself was estimated at roughly $300B in employee tender-offer activity as of November 2024 (see en.wikipedia.org/wiki/ByteDance), with no standalone valuation for the model team |
| Google DeepMind | 2010 | Public parent | n/a | Merged from DeepMind (founded 2010, acquired 2014) and Google Brain in 2023; funded from Alphabet's balance sheet, no standalone valuation exists, and Alphabet's AI-driven capex has run in the tens of billions per year |
| Meta AI (Superintelligence Labs) | 2013 | Public parent | n/a | FAIR founded 2013 under Yann LeCun; Meta's 2025 pivot included a ~$14.3B stake in Scale AI to recruit Alexandr Wang and an aggressive talent raid on rival labs; division has no standalone valuation |
| NVIDIA | 1993 | Public | n/a | NVIDIA is a directly public company (Nasdaq: NVDA) rather than a lab housed inside a larger parent; it surpassed a $4 trillion market cap in mid-2025 on AI accelerator demand, and Nemotron is its own open-model line rather than a division's side project |
| Tencent (Hunyuan) | 2023 | Public parent | n/a | Hunyuan sits inside Tencent Holdings (HKEX: 0700), funded from Tencent's own cloud and R&D budget; no standalone valuation exists for the model team |
Full release table
Export
All 89 tracked models, most recent release first.
Modalities at launch: T text, I image, A audio, V video. Gray chips are inputs, green chips are outputs.
Weights: open means weights are downloadable, whether under a permissive license (Apache 2.0, MIT) or a usage-restricted one (Llama, Gemma); hover a badge for the license nuance. Closed means API or product access only. Badges show launch state; "open since" notes weights released later.
Status: active means currently offered as a first-line model. legacy means superseded by newer models but still served. deprecated means shutdown announced, still accessible until its retirement date. retired means no longer accessible.
Benchmark scores are as claimed in each vendor's launch materials; the changing mix of benchmarks across rows reflects how evaluations saturate and get replaced.
Data is hand-curated from vendor announcements. Each row's sources live in the repository dataset.
Showing 89 of 89 models
| Weights | Modalities | Launch benchmarks | Notable | ||||||
|---|---|---|---|---|---|---|---|---|---|
| 2026-07-24 | Claude Opus 5 | Anthropic | closed | text and image in, text out | 1M | $5/$25 | active | Humanity’s Last Exam 56.3% · OSWorld 2.0 70.6% · BrowseComp 90.8% | State-of-the-art on coding and knowledge work (Frontier-Bench, GDPval-AA) at a third of Fable 5's price; the launch post led with charts rather than published score tables. |
| 2026-07-21 | Gemini 3.6 Flash | closed | text and image and audio and video in, text out | 1M | $1.50/$7.50 | active | MLE-Bench 63.9% · DeepSWE 49% | Efficiency-focused successor to 3.5 Flash, using up to 17% fewer output tokens; announced alongside 3.5 Flash-Lite and the security-focused 3.5 Flash Cyber, with 3.5 Pro still in testing. | |
| 2026-07-19 | Qwen3.8-Max | Alibaba | open | text in, text out | n/a | n/a | active | none published | 2.4-trillion-parameter flagship previewed with open weights promised, days after Moonshot AI's competing Kimi K3. |
| 2026-07-16 | Grok 4.5 | xAI | closed | text and image in, text out | 500k | $2/$6 | active | SWE-bench Pro 64.7% · Terminal-Bench 2.1 83.3% | Positioned as Opus-class at a fraction of the price. |
| 2026-07-15 | Inkling | Thinking Machines Lab | open | text and image and audio and video in, text out | 1M | n/a | active | none published | Mira Murati's lab's first model: a 975B-parameter MoE (about 41B active) trained on 45T multimodal tokens, released under Apache 2.0 and positioned as a customization base for the Tinker platform rather than a frontier flagship. |
| 2026-07-09 | GPT-5.6 | OpenAI | closed | text and image in, text out | 1M | $5/$30 | active | GPQA Diamond 94.6% · Terminal-Bench 2.1 88.8% · SWE-bench Pro 64.6% | First release split into three durable capability tiers, Sol (flagship), Terra, and Luna, each advancing on its own cadence; Sol Ultra multi-agent mode tops Terminal-Bench 2.1 at 91.9%. |
| 2026-07-02 | Hy3 | Tencent | open | text in, text out | 256k | n/a | active | none published | 295B MoE (21B active, 3.8B MTP) under Apache 2.0; in Tencent's own blind evaluation across 270 experts, Hy3 scored 2.67/4 versus GLM-5.1's 2.51/4, with the largest edge in frontend development, data and storage, and CI/CD tasks. |
| 2026-06-30 | Claude Sonnet 5 | Anthropic | closed | text and image in, text out | 1M | $2/$10 | active | Humanity’s Last Exam 34.6% (no tools) · SWE-bench Verified 72.7% · SWE-bench Pro 63.2% | Mid-tier model pitched as running agent workloads that recently required flagship models, at introductory discount pricing. |
| 2026-06-09 | Claude Fable 5 | Anthropic | closed | text and image in, text out | 1M | $10/$50 | active | SWE-bench Pro 80.3% · Terminal-Bench 2.1 88.0% | First Mythos-class model above the Opus tier, built for days-long autonomous tasks; briefly suspended under export controls days after launch. |
| 2026-06-02 | MiniMax-M3 | MiniMax | open | text and image and video in, text out | 1M | n/a | active | none published | 428B MoE (about 23B active) with MiniMax Sparse Attention; natively multimodal, released under a MiniMax community license. |
| 2026-05-28 | Claude Opus 4.8 | Anthropic | closed | text and image in, text out | 1M | $5/$25 | legacy | none published | Honesty-focused upgrade about four times less likely to let flaws in its own code pass unremarked, plus a $10/$50 fast mode and a parallel-subagent preview in Claude Code. |
| 2026-05-19 | Gemini 3.5 Flash | closed | text and image and audio and video in, text out | 1M | $1.50/$9 | legacy | none published | I/O 2026 Flash generation succeeded by 3.6 Flash two months later. | |
| 2026-05-18 | Qwen3.7-Max | Alibaba | closed | text in, text out | n/a | n/a | legacy | none published | Proprietary flagship variant, with the Qwen3.7-Plus tier following a month later. |
| 2026-04-24 | DeepSeek-V4 | DeepSeek | open | text in, text out | 1M | n/a | active | none published | V4-Pro and V4-Flash preview with 1M context and open weights. |
| 2026-04-23 | GPT-5.5 | OpenAI | closed | text and image in, text out | 1M | $5/$30 | active | Terminal-Bench 2.0 82.7% · ARC-AGI-2 85.0% · FrontierMath Tiers 1-3 51.7% | OpenAI's current frontier model, pushing 1M context and long-horizon agentic work. |
| 2026-04-16 | Claude Opus 4.7 | Anthropic | closed | text and image in, text out | 1M | $5/$25 | legacy | CursorBench 70% · BigLaw Bench (Harvey) 90.9% (high effort) | Better advanced software engineering with a new xhigh effort level and higher-resolution vision, though a new tokenizer maps the same text to roughly 1 to 1.35 times as many tokens. |
| 2026-04-15 | Qwen3.6 | Alibaba | open | text in, text out | 131k | n/a | active | none published | Apache-licensed open-weight release the same month as the proprietary Qwen3.6-Plus. |
| 2026-04-08 | Muse Spark | Meta | closed | text and image in, text out | n/a | n/a | active | Humanity’s Last Exam 58% (Contemplating) · FrontierScience Research 38% (Contemplating) | First model from Meta Superintelligence Labs and a pivot away from the open-weight Llama strategy; consumer-only at launch with a multi-agent Contemplating mode. |
| 2026-04-02 | Gemma 4 | open | text and image in, text out | 256k | n/a | active | MMLU-Pro 85.2% (31B) · LiveCodeBench v6 80.0% (31B) · AIME 2026 89.2% (31B, no tools) | Dropped the restrictive Gemma license for Apache 2.0; four variants (E2B to 31B dense) distilled from Gemini 3, with the 31B ranking third among open models on the Arena text leaderboard. | |
| 2026-03-05 | GPT-5.4 | OpenAI | closed | text and image in, text out | 1M | $2.50/$15 | legacy | none published | Thinking and Pro at launch with mini and nano following two weeks later; the mainline 5.3 number was skipped, with only the Codex coding variant shipping under it. |
| 2026-02-17 | Claude Sonnet 4.6 | Anthropic | closed | text and image in, text out | 1M | $3/$15 | legacy | none published | Full upgrade of Sonnet's coding, computer use, long-context reasoning, agent planning, knowledge work, and design skills, with a 1M-token context window in beta at unchanged $3/$15 pricing. |
| 2026-02-16 | Qwen3.5 | Alibaba | open | text in, text out | 131k | n/a | legacy | none published | 397B-A17B open-weight release alongside the proprietary Qwen3.5-Plus, able to operate desktop and mobile applications. |
| 2026-02-14 | Doubao Seed 2.0 | ByteDance | closed | text and image in, text out | n/a | n/a | active | none published | Pro variant positioned against GPT-5.2 and Gemini 3 Pro for long-chain reasoning and agentic tasks; ByteDance said it leads on multiple multimodal, math, and coding benchmarks at roughly a tenth of competitors' token pricing. |
| 2026-02-11 | GLM-5 | Z.ai (Zhipu) | open | text in, text out | 200k | n/a | active | SWE-bench Verified 77.8% · GPQA Diamond 86.0% | Frontier release scaling from GLM-4.5's 355B to 744B total parameters (40B active) with DeepSeek Sparse Attention, closing the gap with frontier closed models. |
| 2026-02-05 | Claude Opus 4.6 | Anthropic | closed | text and image in, text out | 1M | $5/$25 | legacy | SWE-bench Verified 80.9% · Humanity’s Last Exam 53.0% (with tools) · BrowseComp 86.8% (multi-agent harness) | Brought Opus to a 1M-token context in beta at unchanged pricing, with premium long-context rates above 200k tokens. |
| 2025-12-15 | Nemotron 3 | NVIDIA | open | text in, text out | 1M | n/a | active | none published | NVIDIA's flagship open family in Nano, Super, and Ultra sizes built on a hybrid Mamba-attention MoE with a 1M-token context; the 550B/55B-active Ultra followed in June 2026 as NVIDIA's largest open-weight model; NVIDIA Nemotron Open Model License. |
| 2025-12-11 | GPT-5.2 | OpenAI | closed | text and image in, text out | 400k | $1.75/$14 | legacy | none published | The Code Red response to Gemini 3 Pro, shipped in Instant, Thinking, and Pro variants. |
| 2025-12-02 | Mistral Large 3 | Mistral AI | open | text and image in, text out | 256k | $0.50/$1.50 | active | none published | 675B MoE flagship family under Apache 2.0. |
| 2025-12-02 | Amazon Nova 2 | Amazon | closed | text and image and video and audio in, text and image out | 1M | n/a | active | none published | Second-generation family (Lite, Pro, Omni, Sonic) with reasoning and 1M context; launched alongside Nova Forge custom training and Nova Act agents. |
| 2025-12-01 | DeepSeek-V3.2 | DeepSeek | open | text in, text out | 128k | $0.28/$0.42 | active | AIME 2025 93.1% · GPQA Diamond 82.4% · SWE-bench Verified 73.1% | Sparse-attention flagship with integrated reasoning-mode tool calls, plus a Speciale variant hitting olympiad gold-level results. |
| 2025-11-24 | Claude Opus 4.5 | Anthropic | closed | text and image in, text out | 200k | $5/$25 | active | SWE-bench Verified 80.9% | Cut Opus pricing by two-thirds while topping SWE-bench Verified, ending Opus's premium-only positioning. |
| 2025-11-18 | Gemini 3 Pro | closed | text and image and audio and video in, text out | 1M | $2/$12 | active | GPQA Diamond 91.9% · Humanity’s Last Exam 37.5% (no tools) · ARC-AGI-2 31.1% | Day-one launch across Search, the Gemini app, and the API, with record LMArena scores and generative UI output. | |
| 2025-11-17 | Grok 4.1 | xAI | closed | text and image in, text out | 256k | n/a | legacy | none published | Vendor-announced flagship default across xAI surfaces, launched at the top of LMArena. |
| 2025-11-13 | Ernie 5.0 | Baidu | closed | text and image and audio and video in, text and image and video out | n/a | n/a | active | none published | 2.4T-parameter natively omni-modal flagship unveiled at Baidu World 2025, jointly modeling text, images, audio, and video in a single autoregressive framework. |
| 2025-11-12 | GPT-5.1 | OpenAI | closed | text and image in, text out | 400k | $1.25/$10 | active | SWE-bench Verified 76.3% | Instant/Thinking split refined GPT-5's router with adaptive reasoning and a warmer default persona. |
| 2025-10-22 | MiniMax-M2 | MiniMax | open | text in, text out | 205k | $0.30/$1.20 | legacy | none published | Modified-MIT-licensed MoE tuned for coding and agents at a fraction of frontier pricing; became a default open agentic model. |
| 2025-10-15 | Claude Haiku 4.5 | Anthropic | closed | text and image in, text out | 200k | $1/$5 | active | SWE-bench Verified 73.3% · Terminal-Bench 41.75% (32K thinking) | Small-tier model matching five-month-old frontier (Sonnet 4) coding performance at a third of the cost. |
| 2025-09-30 | GLM-4.6 | Z.ai (Zhipu) | open | text in, text out | 200k | n/a | legacy | none published | MIT-licensed successor to GLM-4.5 with 200k context, expanded from GLM-4.5's 128K. |
| 2025-09-29 | Claude Sonnet 4.5 | Anthropic | closed | text and image in, text out | 200k | $3/$15 | active | SWE-bench Verified 77.2% · OSWorld 61.4% | Billed as the best coding model on release, with 30-plus-hour autonomous agent runs. |
| 2025-09-23 | Qwen3-Max | Alibaba | closed | text in, text out | 262k | n/a | legacy | none published | Proprietary trillion-parameter-plus flagship, trained on about 36 trillion tokens, available only via API. |
| 2025-08-21 | DeepSeek-V3.1 | DeepSeek | open | text in, text out | 128k | n/a | legacy | none published | Hybrid thinking and non-thinking modes billed as the start of the agent era. |
| 2025-08-18 | Nemotron Nano 2 | NVIDIA | open | text in, text out | 131k | n/a | legacy | none published | First fully NVIDIA-pretrained generation after the Llama derivatives: a 9B hybrid Mamba-2/Transformer reasoning model (12B base variant also released) with up to 6x the throughput of same-size Transformers, shipped with most of its pretraining corpus opened as the Nemotron Pretraining Dataset; NVIDIA Open Model License. |
| 2025-08-07 | GPT-5 | OpenAI | closed | text and image in, text out | 400k | $1.25/$10 | legacy | SWE-bench Verified 74.9% · AIME 2025 94.6% (no tools) · GPQA Diamond 88.4% (GPT-5 pro) | Unified router-based system merging fast and reasoning modes, at aggressively commoditized pricing. |
| 2025-08-05 | gpt-oss-120b | OpenAI | open | text in, text out | 128k | n/a | active | GPQA Diamond 80.1% · MMLU-Pro 80.8% | OpenAI's first open-weight language models since GPT-2, released under Apache 2.0 alongside the 21B-parameter gpt-oss-20b; the 117B MoE runs on a single 80GB GPU and lands near o4-mini on reasoning. |
| 2025-08-05 | Claude Opus 4.1 | Anthropic | closed | text and image in, text out | 200k | $15/$75 | legacy | SWE-bench Verified 74.5% | Incremental Opus upgrade focused on agentic coding, released two days before GPT-5. |
| 2025-07-28 | GLM-4.5 | Z.ai (Zhipu) | open | text in, text out | 128k | $0.60/$2.20 | active | AIME 2024 91.0% · SWE-bench Verified 64.2% · GPQA 79.1% | Agent-focused open-weight MoE undercutting even DeepSeek on price, cementing the Chinese open-model surge of mid-2025. |
| 2025-07-11 | Kimi K2 | Moonshot AI | open | text in, text out | 128k | $0.60/$2.50 | active | SWE-bench Verified 65.8% · LiveCodeBench v6 53.7% · AIME 2025 49.5% | 1T-parameter MoE under a modified MIT license, the strongest open agentic model at release. |
| 2025-07-09 | Grok 4 | xAI | closed | text and image in, text out | 256k | $3/$15 | active | Humanity’s Last Exam 25.4% (no tools) · GPQA 87.5% · ARC-AGI-2 15.9% | Reasoning-first flagship with a multi-agent Heavy tier, briefly topping several frontier benchmarks. |
| 2025-06-16 | MiniMax-M1 | MiniMax | open | text in, text out | 1M | n/a | legacy | none published | First open-weight large-scale hybrid-attention reasoning model (456B MoE, 45.9B active), Apache 2.0, 1M context. |
| 2025-05-22 | Claude 4 (Opus 4 / Sonnet 4) | Anthropic | closed | text and image in, text out | 200k | $15/$75 | legacy | SWE-bench Verified 72.5% (Opus 4) · Terminal-bench 43.2% (Opus 4) · GPQA Diamond 79.6% (Opus 4) | Opus returned as the coding and agent flagship; first Anthropic release under ASL-3 safeguards. |
| 2025-04-28 | Qwen3 | Alibaba | open | text in, text out | 131k | n/a | active | AIME 2024 85.7 (235B-A22B) · AIME 2025 81.5 (235B-A22B) · LiveCodeBench 70.7 (235B-A22B) | Hybrid thinking/non-thinking MoE family under Apache 2.0 that led open-model leaderboards through 2025. |
| 2025-04-16 | o3 | OpenAI | closed | text and image in, text out | 200k | $10/$40 | legacy | AIME 2024 91.6% · GPQA Diamond 83.3% · SWE-bench Verified 69.1% | Flagship reasoning model with agentic tool use; its later price cut reset reasoning-model economics. |
| 2025-04-14 | GPT-4.1 | OpenAI | closed | text and image in, text out | 1M | $2/$8 | legacy | SWE-bench Verified 54.6% · MMLU 90.2% · GPQA Diamond 66.3% | API-first workhorse family that brought a 1M-token context window to OpenAI's lineup. |
| 2025-04-07 | Llama-3.1-Nemotron-Ultra-253B | NVIDIA | open | text in, text out | 131k | n/a | legacy | GPQA (reasoning on) 76.0% · AIME 2025 (reasoning on) 72.5% · LiveCodeBench (reasoning on) 66.3% | Flagship of the Llama Nemotron reasoning family (Nano 8B and Super 49B debuted at GTC in March 2025): a 253B model derived from Llama-3.1-405B via Neural Architecture Search with toggleable reasoning; dual NVIDIA Open Model License plus Llama 3.1 Community License. |
| 2025-04-05 | Llama 4 (Scout/Maverick) | Meta | open | text and image in, text out | 10M | n/a | active | MMLU-Pro 80.5% (Maverick) · GPQA Diamond 69.8% (Maverick) · LiveCodeBench 43.4% (Maverick) | Meta's MoE, natively multimodal generation with a claimed 10M-token context on Scout; a contested launch that cooled Meta's open-weights momentum. |
| 2025-03-25 | Gemini 2.5 Pro | closed | text and image and audio and video in, text out | 1M | $1.25/$10 | active | GPQA Diamond 84.0% · AIME 2025 86.7% · SWE-bench Verified 63.8% | Thinking-by-default flagship that put Google back at the top of leaderboards in 2025. | |
| 2025-03-16 | Ernie 4.5 | Baidu | open open since Jun 2025 | text and image in, text out | n/a | n/a | legacy | none published | Multimodal flagship launched closed, then fully open-sourced under Apache 2.0, the biggest Chinese open release between DeepSeek and the summer 2025 wave. |
| 2025-03-12 | Gemma 3 | open | text and image in, text out | 128k | n/a | legacy | LMArena Elo 1339 (27B) · MMLU-Pro 67.5% (27B) · GPQA Diamond 42.4% (27B) | Multimodal open-weight family (1B to 27B) under the use-restricted Gemma license; the 27B ranked top-10 on LMArena at launch while fitting on a single GPU. | |
| 2025-02-27 | GPT-4.5 | OpenAI | closed | text and image in, text out | 128k | $75/$150 | retired | GPQA Diamond 71.4% · AIME 2024 36.7% · SWE-bench Verified 38.0% | OpenAI's largest pretraining-scaling bet, priced so high it was retired from the API within months. |
| 2025-02-24 | Claude 3.7 Sonnet | Anthropic | closed | text and image in, text out | 200k | $3/$15 | legacy | SWE-bench Verified 62.3% · GPQA Diamond 78.2% (extended thinking) · TAU-bench (retail) 81.2% | First hybrid reasoning model, blending instant replies and extended thinking in one model; shipped with Claude Code. |
| 2025-02-17 | Grok 3 | xAI | closed | text in, text out | 131k | $3/$15 | legacy | AIME 2025 93.3% (Think, cons@64) · GPQA Diamond 84.6% (Think) · LiveCodeBench 79.4% (Think) | Trained on the 200k-GPU Colossus cluster; introduced Think mode and DeepSearch. Image input never shipped for grok-3 in the API. |
| 2025-01-22 | Doubao 1.5 Pro | ByteDance | closed | text and image in, text out | n/a | n/a | legacy | none published | Sparse-MoE flagship matching GPT-4o-class benchmarks at a fraction of the cost; powered Doubao, China's largest AI assistant. |
| 2025-01-20 | DeepSeek-R1 | DeepSeek | open | text in, text out | 128k | $0.55/$2.19 | legacy | AIME 2024 79.8% · GPQA Diamond 71.5% · SWE-bench Verified 49.2% | MIT-licensed o1-class reasoning model whose release triggered a global market shock and a wave of RL replication. |
| 2024-12-26 | DeepSeek-V3 | DeepSeek | open | text in, text out | 128k | $0.27/$1.10 | legacy | MMLU 88.5% · GPQA Diamond 59.1% · MATH-500 90.2% | GPT-4o-class open weights reportedly trained for under $6M, upending assumptions about frontier training cost. |
| 2024-12-11 | Gemini 2.0 Flash | closed | text and image and audio and video in, text out | 1M | n/a | legacy | MMLU-Pro 76.4% · GPQA Diamond 62.1% · Natural2Code 92.9% | Opened Google's agentic era with native tool use; image and audio output were announced at launch but limited to early-access partners. | |
| 2024-12-03 | Amazon Nova | Amazon | closed | text and image and video in, text out | 300k | n/a | legacy | none published | Amazon's first in-house frontier family (Micro, Lite, Pro, Premier), launched at re:Invent 2024 with an emphasis on price-performance; Premier remained in training at launch and shipped the following spring. |
| 2024-11-05 | Hunyuan-Large | Tencent | open | text in, text out | 256k | n/a | legacy | none published | 389B-total, 52B-active MoE, the largest open MoE of its day; Tencent's entry into the open-weights race. |
| 2024-09-19 | Qwen2.5 | Alibaba | open | text in, text out | 131k | n/a | legacy | MMLU 85+ · HumanEval 85+ · MATH 80+ | Broad Apache-2.0 family (0.5B to 72B) that became the default base for open fine-tunes and distillations. |
| 2024-09-12 | o1 | OpenAI | closed | text in, text out | 128k | $15/$60 | deprecated | AIME 2024 83.3% (cons@64) · GPQA Diamond 78.0% · Codeforces 89th percentile | First mainstream reasoning model, introducing chain-of-thought test-time compute as a product. |
| 2024-08-13 | Grok-2 | xAI | open open since Aug 2025 | text and image in, text out | n/a | n/a | legacy | none published | Beta launch on X Premium and Premium+, alongside the smaller Grok-2 mini; the full weights were released on Hugging Face about a year later under a restricted xAI Community License. |
| 2024-07-23 | Llama 3.1 405B | Meta | open | text in, text out | 128k | n/a | legacy | MMLU 88.6% · HumanEval 89.0% · GSM8K 96.8% | First open-weight model credibly matching closed frontier models on benchmarks. |
| 2024-06-20 | Claude 3.5 Sonnet | Anthropic | closed | text and image in, text out | 200k | $3/$15 | retired | GPQA Diamond 59.4% · HumanEval 92.0% · GSM8K 96.4% | Beat Opus at a fifth of the price and became the default coding model of 2024. |
| 2024-06-14 | Nemotron-4 340B | NVIDIA | open | text in, text out | 4k | n/a | legacy | Arena Hard 54.2 · MT-Bench (GPT-4-Turbo judge) 8.22 | NVIDIA's first frontier-scale open release: a 340B dense family (Base, Instruct, Reward) under the NVIDIA Open Model License, trained on 9T tokens and positioned as a synthetic-data-generation engine; over 98% of the Instruct model's alignment data was itself synthetic. |
| 2024-05-13 | GPT-4o | OpenAI | closed | text and image in, text out | 128k | $5/$15 | legacy | MMLU 88.7% · GPQA 53.6% · HumanEval 90.2% | Natively multimodal omni model; the launch demoed real-time voice, but the API shipped with text and image in, text out. |
| 2024-05-06 | DeepSeek-V2 | DeepSeek | open | text in, text out | 128k | $0.14/$0.28 | retired | MMLU 77.8 (Chat RL) · HumanEval 81.1 (Chat RL) · GSM8K 92.2 (Chat RL) | Efficient MoE with multi-head latent attention that started China's LLM price war. |
| 2024-04-18 | Llama 3 (8B/70B) | Meta | open | text in, text out | 8k | n/a | legacy | MMLU 82.0% (70B) · HumanEval 81.7% (70B) · GSM8K 93.0% (70B) | Closed most of the open-vs-closed quality gap at the 70B scale. |
| 2024-03-04 | Claude 3 (Haiku/Sonnet/Opus) | Anthropic | closed | text and image in, text out | 200k | $15/$75 | retired | MMLU 86.8% (Opus) · GPQA Diamond 50.4% (Opus) · HumanEval 84.9% (Opus) | First release to beat GPT-4 on headline benchmarks and the origin of the three-tier Haiku/Sonnet/Opus naming. |
| 2024-02-26 | Mistral Large | Mistral AI | closed | text in, text out | 32k | $8/$24 | legacy | MMLU 81.2% · HumanEval 45.1% | Mistral's flagship API bet, launched alongside the Microsoft Azure partnership. |
| 2024-02-15 | Gemini 1.5 Pro | closed | text and image and audio and video in, text out | 1M | n/a | retired | MMLU 81.9% · GSM8K 91.7% · HumanEval 71.9% | Made million-token context real, an order of magnitude beyond anything shipping at the time. | |
| 2023-12-06 | Gemini 1.0 (Ultra/Pro/Nano) | closed | text and image and video in, text out | 32k | n/a | retired | MMLU 90.0% (Ultra, CoT@32) · GSM8K 94.4% (Ultra) · HumanEval 74.4% (Ultra) | Google's first natively multimodal frontier family, unifying DeepMind and Brain efforts; audio understanding was demoed but not in the launch API. | |
| 2023-11-03 | Grok-1 | xAI | open open since Mar 2024 | text in, text out | 8k | n/a | retired | MMLU 73.0% · GSM8K 62.9% · HumanEval 63.2% | xAI's debut, later notable for the Apache-2.0 release of its 314B MoE weights. |
| 2023-09-27 | Mistral 7B | Mistral AI | open | text in, text out | 8k | n/a | legacy | MMLU 60.1% · HumanEval 30.5% · GSM8K 52.2% (8-shot) | Apache-2.0 model that beat Llama 2 13B and proved small European labs could compete. |
| 2023-07-18 | Llama 2 | Meta | open | text in, text out | 4k | n/a | legacy | MMLU 68.9% (70B) · GSM8K 56.8% (70B) · HumanEval 29.9% (70B) | First commercially usable open-weight frontier-adjacent model; seeded the modern open ecosystem. |
| 2023-07-11 | Claude 2 | Anthropic | closed | text in, text out | 100k | $11.02/$32.68 | retired | HumanEval (Codex P@1) 71.2% · GSM8K 88.0% · Bar exam (MBE) 76.5% | Made 100k-token context mainstream when rivals were at 8-32k. |
| 2023-05-10 | PaLM 2 | closed | text in, text out | 8k | n/a | retired | none published | Google's I/O 2023 answer to GPT-4 that powered Bard before the Gemini era. | |
| 2023-03-16 | Ernie Bot | Baidu | closed | text in, text out | n/a | n/a | retired | none published | The first major Chinese answer to ChatGPT; Ernie 4.0 (Oct 2023) claimed GPT-4 parity domestically. |
| 2023-03-14 | GPT-4 | OpenAI | closed | text in, text out | 8k | $30/$60 | retired | MMLU 86.4% · HumanEval 67.0% · GSM8K 92.0% | Announced as multimodal, but image input stayed gated at launch; defined the state of the art for over a year. |
| 2023-03-14 | Claude 1 | Anthropic | closed | text in, text out | 9k | n/a | retired | none published | Anthropic's first commercial model and the debut of Constitutional AI training in production. |
| 2022-11-30 | GPT-3.5 / ChatGPT | OpenAI | closed | text in, text out | 4k | n/a | retired | none published | The ChatGPT launch that took LLMs mainstream and kicked off the current AI race. |
Sources and current references
Where to check current pricing and model docs directly with each vendor, since both move faster than a hand-curated dataset can track.
| Company | Current pricing | Docs / models |
|---|---|---|
| OpenAI | Pricing | Docs |
| Anthropic | Pricing | Docs |
| Google DeepMind | Pricing | Docs |
| Meta AI (Superintelligence Labs) | Pricing | Docs |
| Mistral AI | Pricing | Docs |
| DeepSeek | Pricing | Docs |
| xAI | Pricing | Docs |
| Alibaba (Qwen) | Pricing | Docs |
| Moonshot AI | Pricing | Docs |
| Z.ai (Zhipu) | Pricing | Docs |
| Thinking Machines Lab | Pricing | Docs |
| NVIDIA | Pricing | Docs |
| MiniMax | Pricing | Docs |
| ByteDance (Seed) | Pricing | Docs |
| Baidu | Pricing | Docs |
| Tencent (Hunyuan) | Pricing | Docs |
| Amazon (AGI / Nova) | Pricing | Docs |
Every model row in the underlying dataset carries its own source URLs (vendor announcements and documentation); the dataset is available via Export above.
Benchmark scores and leaderboards move; for live comparisons check the vendors' model pages above.