Subscribe
The AI Era
Projects The AI Era

Monthly AI timeline

The AI Era

A public month-by-month timeline of AI breakthroughs, model launches, open-weight releases, media models, and agent shifts from GPT-3 through August 18, 2026.

This timeline is a month-by-month record of the AI era rather than a loose collection of milestones. The August 18 refresh carries it through Muse Glimmer 30B, Grok 4.6, DeepSeek V4 Pro GA, Gemini 3.7 Flash, and GLM-5.3.

Use it when you want sequence, not just significance. A lot of AI discussion collapses events together. This page shows how model launches, product shifts, infrastructure moves, and agent patterns accumulated over time.

What is included

The emphasis is on the modern period from GPT-3 onward: major model releases, reasoning shifts, agent tooling, open-weight frontier moves, speech and multimodal infrastructure, and the transition from chat interfaces to more capable operating layers.

Current snapshot

July’s most important sequence took ten days to resolve. Moonshot launched Kimi K3 as an API model on July 16, Anthropic released Claude Opus 5 on July 24, and Moonshot then made K3’s full checkpoint downloadable. That last step matters: a promise became an artifact, and K3 moved into the actual open-weight floor.

The current independent scores put Claude Opus 5 at 63 Intelligence and Kimi K3 at 60. Coding Agent Index v1.3 has Claude Code with Opus 5 and Codex with GPT-5.6 Sol tied at 67, six points above Kimi Code with K3. The Agentic Index gap is about five points, 55 to 50.

The next sequence is about deployment shape as much as raw score. Muse Glimmer 30B makes a useful multimodal agent local and Apache-2.0. Grok 4.6 and Gemini 3.7 Flash deepen the proprietary agent tier. DeepSeek V4 Pro 0813 is a hosted GA update rather than a new public checkpoint, while GLM-5.3 ships to Coding Plan users before its general API or public weights.

Artificial Analysis also moved to Intelligence Index v4.1.1 and Coding Agent Index v1.3. The timeline records that as an event because benchmark changes alter the ruler, not just the ranking. Older snapshots remain historical rather than being rewritten to look artificially continuous.

How to read it

Treat it as a chronology first and a commentary surface second. It works best alongside the broader architectural pieces on the site, because the question is not only what launched, but what each launch changed in the way people work with AI.