Skip to content
Dossier — AI

Artificial General Intelligence

Last updated September 6, 2026

Overview

Artificial general intelligence — AGI — is the shorthand for machine systems whose capabilities match or exceed human performance across most economically and intellectually valuable tasks. It is at once a research goal, a marketing term, a policy category and a philosophical question. This dossier tracks what the phrase is being used to mean, what the evidence supports, and how the institutions around it are forming.

Why it matters

Whether or not any particular system deserves the label, the trajectory of capability is real, and it is already reshaping labour markets, research, security and governance. The AGI question is a proxy for a practical one: how much should we prepare for systems that can do most of what we can do, and how soon?

Current state

Frontier models now perform at or above human expert level on a widening set of benchmarks, including professional examinations, competition mathematics and software engineering tasks, while remaining brittle in ways that humans are not. There is no agreed definition of AGI and no accepted test. Labs differ sharply in their forecasts. Evaluation is shifting from task performance toward autonomy, reliability and safety-relevant behaviour.

Key players
Frontier AI laboratories
Train and deploy the most capable models; publish capability and safety evaluations.
Academic alignment and interpretability groups
Study how models reason and how their behaviour can be verified.
National AI safety institutes
Government bodies evaluating frontier models and drafting standards.
Open-weight communities
Release and improve models whose weights are freely downloadable, diffusing capability.
Timeline
  1. 2017

    The transformer architecture is published

    The design underlying nearly every subsequent large model.

  2. 2020

    Scaling laws formalised

    Predictable improvement with compute, data and parameters becomes a research programme.

  3. 2022

    Conversational assistants reach the public

    Large models become a mass-market product and a policy issue.

  4. 2023–2024

    Governments establish AI safety institutes

    Frontier evaluation becomes a state function.

  5. 2025–2026

    Reasoning and agentic systems

    Models that plan, use tools and act over many steps enter deployment.

Important developments
  • Recent

    Reasoning models show emergent planning in evaluation

    See The Signal.

  • Recent

    Multilateral call for shared frontier standards

    Voluntary reporting on evaluations and incidents.

Core technology

Large transformer-based models trained on internet-scale data, refined with human and AI feedback, extended with tool use, retrieval and long-horizon planning. Interpretability tools examine internal features; evaluation suites test capability, honesty and dangerous-capability thresholds.

Key papers
  • Attention Is All You Need, NeurIPS (2017)The transformer.
  • Scaling Laws for Neural Language Models, arXiv (2020)
  • Language Models (Mostly) Know What They Know, arXiv (2022)Calibration in large models.
  • Concrete Problems in AI Safety, arXiv (2016)
Open questions
  1. 01Is there a meaningful capability threshold, or is 'AGI' a gradient with no natural line?
  2. 02Do current architectures generalise to genuinely novel problems, or interpolate across vast training data?
  3. 03Can trustworthiness — calibration, honesty, corrigibility — be verified rather than merely observed?
  4. 04What governance is possible when capability diffuses through open weights?
What changed

The centre of gravity has moved from 'can it answer?' to 'can it act?'. Agentic systems that plan and use tools have made autonomy, not accuracy, the frontier question.

What to watch

Evaluation standards for autonomous behaviour; interpretability results that move from research to requirement; whether open-weight releases keep pace with frontier systems; the first binding international standards.

LAST UPDATED SEPTEMBER 6, 2026 · EDITORIAL DEMONSTRATION CONTENT