Subscribe
The Evolution of LLMs
Projects The Evolution of LLMs

Interactive LLM timeline

The Evolution of LLMs

Trace the evolution of large language models, with sourced releases and versioned benchmark evidence checked 2026-09-09.

Explore the evolution of language models, from the Transformer to reasoning systems and agents. The interactive timeline and story view separate release chronology from the dated benchmark comparison below.

The current comparison

Checked 2026-09-09. AA Intelligence Index v4.3: GLM-5.3 scores 44.9, versus Claude Fable 5.1 at 53.4. The general reasoning gap is 8.5 points within the reviewed cohort. AA Coding Agent Index v1.4: Kimi K3 scores 62.6, versus Claude Fable 5.1 at 70.4. The coding systems gap is 7.8 points within the reviewed cohort. AutomationBench-AA v1.0.6 objective score: GLM-5.3 scores 62.2, versus GPT-6 Astra at 68.5. The workflow automation gap is 6.3 percentage points within the reviewed cohort.

Benchmark methodology, current general results and coding harness results identify the evidence behind this snapshot. Each number measures a specific model/configuration within its benchmark. Calculations use unrounded values. The reviewed cohort is a subset of the publisher’s results, not a claim to cover every model or self-hosted setup.

What changed in this sweep

  • Claude Fable 5.1 — Generally available for long-running coding and knowledge work. Safety fallback routes some tasks to Opus models, so the benchmark configuration must travel with the score.
  • Claude Mythos 5.1 — Available to vetted organisations through trusted-access programmes. This restricted cyber and biology model is tracked separately from the generally available ceiling.
  • GPT-6 Astra — OpenAI launched Astra for computer use, coding, science and professional work with a phased rollout. Access varies by product and account. Independent results are distinct from the vendor's headline claims.
  • Gemini 3.8 Flash — Google's new Flash model targets longer software and enterprise workflows. Its hosted distribution and speed-oriented positioning are separate from open-weight availability.
  • Muse Spark 1.3 — The xhigh variant is available through Meta's API and Muse Code. The max variant was a limited partner preview at launch. Keep the two reasoning configurations distinct.
  • GLM-5.3 — The earlier Coding Plan launch now has a downloadable checkpoint. The current custom licence must be read separately from GLM-5.2's MIT terms. It now leads the measured open general cohort.
  • GLM-5.3-Flash — Downloadable multimodal Flash weights are verified in this sweep. MIT licensing and a lower cost per benchmark task make this a distinct deployment option from the larger custom-licensed GLM-5.3.
  • Qwen3.8 2.4T A95B — The downloadable text checkpoint is distinct from hosted Qwen3.8-Max, whose additional vision, tools and serving features must not be attributed automatically to the weights.
  • Qwen3.8 27B — A smaller downloadable Qwen3.8 model broadens the practical deployment floor. The observed date records this verification, rather than inferring a public launch from repository creation.
  • Qwen3.8-Flash-Next — An experimental Qwen4 architecture preview combines sparse attention, gated residuals and n-gram embeddings. Its 125B language model activates 6B parameters, with another 51B embedding and 4B multi-token-prediction parameters.
  • DeepSeek V4 Pro 0813 — A matching downloadable 0813 checkpoint is now verified. This supersedes the August snapshot's hosted-only classification without altering that historical observation.
  • DeepSeek V4 Flash Vision Exp — An experimental downloadable vision branch is verified. Its availability is recorded without assigning an unverified comparable general or coding score.
  • MiniCPM5-2B — OpenBMB's compact local model targets tool use, coding and reasoning. It has about 2.52B total parameters, around 1.98B outside embeddings, and native 131,072-token context. Vendor averages are not an AA Intelligence Index score.
  • Agnes 2.5 Pro Beta — A hosted beta for multimodal reasoning and agentic work. The August launch evaluation used an older Intelligence Index, so its original 49-point result is not inserted into the September v4.3 comparison.
  • Gemini 3.8 Flash Cyber — A specialised model for vulnerability detection and patching, offered through the Fairwind programme. Its restricted security access is distinct from ordinary Gemini 3.8 Flash.

How to read the numbers

Each tab is a separate measurement. General reasoning uses Intelligence Index v4.3. Coding systems uses Coding Agent Index v1.4, which measures a model inside a named harness, not the model alone. Workflow automation uses AutomationBench-AA v1.0.6 objective score, not the older broad Agentic Index. Workflow scores are percentages of objectives completed under the benchmark rules, not whole-task success rates. Compare only within a tab and version. No percentage of capability retained is inferred from index ratios. An open-weight model being tested through a hosted service does not prove identical self-hosted performance. Scores retain source precision and display one decimal. The gap is calculated from unrounded values. Current general results and the dated September 7 workflow dataset are separate measurements; their GLM-5.3 general values differ, so the current general table is used consistently. Earlier records without versioned evidence are retained as unverified history and do not produce validated gaps.

The floor means downloadable capability, not a promise that the model will fit on a laptop or that its licence is unrestricted. Restricted Mythos access is excluded from the ordinary proprietary ceiling. Missing comparable scores remain unscored.