Interactive guide
The Intelligence Stack
This guide follows AI from chatbots through tool use and agents to orchestration across applications. The model layer changes quickly, while tools, memory, verification and workflow design determine what a system can actually finish.
The current comparison
Checked 2026-09-09. AA Intelligence Index v4.3: GLM-5.3 scores 44.9, versus Claude Fable 5.1 at 53.4. The general reasoning gap is 8.5 points within the reviewed cohort. AA Coding Agent Index v1.4: Kimi K3 scores 62.6, versus Claude Fable 5.1 at 70.4. The coding systems gap is 7.8 points within the reviewed cohort. AutomationBench-AA v1.0.6 objective score: GLM-5.3 scores 62.2, versus GPT-6 Astra at 68.5. The workflow automation gap is 6.3 percentage points within the reviewed cohort.
Benchmark methodology, current general results and coding harness results identify the evidence behind this snapshot. Each number measures a specific model/configuration within its benchmark. Calculations use unrounded values. The reviewed cohort is a subset of the publisher’s results, not a claim to cover every model or self-hosted setup.
What changed in this sweep
Astra and Fable 5.1 extend work across tools and applications, while downloadable GLM and Qwen releases widen the choices for operating your own model layer. Orchestration, recovery, context management and verification remain separate system responsibilities. The demo uses this dated snapshot and is explicitly labelled as illustrative. Explore model releases and sources.
How to read the numbers
Each tab is a separate measurement. General reasoning uses Intelligence Index v4.3. Coding systems uses Coding Agent Index v1.4, which measures a model inside a named harness, not the model alone. Workflow automation uses AutomationBench-AA v1.0.6 objective score, not the older broad Agentic Index. Workflow scores are percentages of objectives completed under the benchmark rules, not whole-task success rates. Compare only within a tab and version. No percentage of capability retained is inferred from index ratios. An open-weight model being tested through a hosted service does not prove identical self-hosted performance. Scores retain source precision and display one decimal. The gap is calculated from unrounded values. Current general results and the dated September 7 workflow dataset are separate measurements; their GLM-5.3 general values differ, so the current general table is used consistently. Earlier records without versioned evidence are retained as unverified history and do not produce validated gaps.
The floor means downloadable capability, not a promise that the model will fit on a laptop or that its licence is unrestricted. Restricted Mythos access is excluded from the ordinary proprietary ceiling. Missing comparable scores remain unscored.