Subscribe
Floor / Ceiling
Projects Floor / Ceiling

Open AI baseline tracker

Floor / Ceiling

A public tracker for the moving floor of open-weight AI against the proprietary frontier ceiling, refreshed through August 18, 2026.

This tracker compares two moving AI baselines: the proprietary ceiling and the open-weight floor.

The ceiling is the strongest capability available through closed systems. The floor is the strongest model whose weights you could still run if those proprietary systems disappeared tomorrow.

What the floor means

The floor is not the average open model. It is the strongest downloadable fallback for a domain such as coding, agentic work, or general reasoning.

That makes two checks essential. The model must be independently measured, and the actual weights must exist. An announced release is not a durable floor. A downloadable checkpoint with a restrictive licence is real open weight, but not the same thing as permissively licensed open source.

Current snapshot

The August 18 snapshot keeps Kimi K3’s downloadable checkpoint as the measured floor. It is a 2.8T-total / 104B-active multimodal MoE with a 1M-token context window. Its custom Kimi K3 licence allows broad use but adds restrictions for some large-scale commercial deployments, which is why the tracker labels it open weight rather than simply open source.

Artificial Analysis v4.1.1 scores Kimi K3 at 60 Intelligence. Claude Opus 5 sets the proprietary general ceiling at 63 and the Agentic Index ceiling at 55, while K3 reaches about 50 Agentic. On Coding Agent Index v1.3, Claude Code with Opus 5 (xhigh) and Codex with GPT-5.6 Sol (max) tie at 67; Kimi Code with K3 scores 61. The current strict gaps are therefore three points for general intelligence, five for agentic work, and six for coding.

Those numbers begin a new August snapshot rather than revising July. Artificial Analysis changed the Intelligence Index to v4.1.1 and the Coding Agent Index to v1.3, including explicit coding-harness effects. The new scores are the best current comparison, but they are not perfectly continuous with the older series.

What else changed

DeepSeek V4 Flash 0731 is the other important open release. It has MIT-licensed weights, 284B total / 13B active parameters, 1M context, and a current Artificial Analysis Intelligence score of 52. It does not displace K3, but its roughly $0.14 input and $0.28 output pricing per million tokens makes the cost curve nearly as important as the benchmark result.

The newest cohort does not change the gap. Grok 4.6 scores 61 Intelligence and Gemini 3.7 Flash scores 56, both below Opus 5. Muse Glimmer 30B scores 35 but matters as an Apache-2.0 local agent that can fit in under 20GB when quantised. DeepSeek V4 Pro 0813 is a hosted GA revision without a matching updated public checkpoint. GLM-5.3 is live for Coding Plan users, but its general API, public weights, and comparable independent score are still pending.

The conclusion is narrower and stronger than “open has caught up”. The strict general gap is now very small. The coding and agentic gaps are still visible, and the proprietary systems retain advantages in reliability, harness integration, and long-running execution. But a surprisingly large share of frontier capability is now downloadable and durable.

Why track it monthly

AI progress is easier to understand when the floor and ceiling are separated. Some months the ceiling jumps. Other months the floor catches up. The long-run question is not which lab is briefly ahead, but how much capability remains available even without the frontier APIs.