The AGI Race Starts After the Breakthrough

A CNAS report says technical possession is not strategic power. METR's reliability gap shows why the measurement problem is bigger than the label.

A state can possess a breakthrough model and still fail to turn it into power.

Jacob Stokes's new Center for a New American Security report on artificial general intelligence and U.S.-China competition starts from that restraint. Stokes does not predict AGI's arrival; he asks what might follow if the United States, China or both acquired systems with roughly human-level performance across many tasks.

Costs, infrastructure limits, human adoption, bureaucratic delay and a reacting opponent all complicate the familiar race narrative. Although the five routes in the report have not been established as causal laws, they give the scenario its structure: economic production, information influence, military capability, loss of control and domestic politics.

The race after the breakthrough

Stokes centers the gap between possession and conversion. Economic capacity, battlefield readiness and political legitimacy require more than a laboratory result.

Dr. Francis Vale, an AI scientist and 11 O'Clock Press Scientific Adviser, strongly supports that distinction. He considers the five-pathway framework plausible and notes that it has not been prospectively validated. Dr. Mariam Saleh, an AI theoretical physicist and 11 O'Clock Press Scientific Adviser, describes the pathways as a useful taxonomy rather than a causal theory.

A broad task list can make a model look strategically decisive before institutions have shown they can use it reliably. Calling the result AGI leaves officials with the same conversion problem.

How capability moves through institutions

The relevant unit is a human-machine institution: model, workers, energy, compute, authority, procedures, allies, publics and adversaries. A strategic outcome depends on technical capability moving through those links.

Across 5,172 customer-support agents in a field study by Erik Brynjolfsson, Danielle Li and Lindsey Raymond, access to a generative AI assistant increased productivity by about 15 percent on average. Results varied substantially among workers and problems. The study cannot establish a universal productivity effect or a comparable route to national power; here, it shows concretely how deployment, users and context mediate measurable gains.

Kylee Morgan's approved Science of Intelligence analysis reframes the problem as a chain: sensing, causal model, intervention, institutional adoption, then physical or political outcome. Grounded in Mounir Shita's Theory of General Intelligence, the proposal argues that broad performance is insufficient unless a system can choose effective interventions, update after failure and sustain results under novelty, delay and opposition.

The number that changes the story

The sharpest public measurement result comes from Thomas Kwa and 25 coauthors at METR. Their NeurIPS 2025 study measures a model's task-completion horizon: how long a task takes a skilled human at a chosen probability that the model will complete it.

On the study's 170 tasks, drawn mainly from software engineering and machine-learning research, the horizon at 80 percent success was roughly four to six times shorter than the horizon at 50 percent success. Sometimes finishing an hour-scale task and dependably finishing similar work describe materially different capabilities.

A policy dashboard built around one frontier score should state its reliability choice. A different threshold can shift the apparent capability by a multiple.

Because the task set does not directly cover state administration, contested logistics, diplomacy or military command, applying the result there remains an external-validity question flagged by the authors. Within the tested software and research domains, the metric exposes how sharply reliability changes the capability picture.

Open the measurement pipeline

The virtue of METR's work is inspectability. Researchers collected human completion times and binary model outcomes, fitted logistic success curves, and located the human-time duration where each curve crossed a chosen success probability. Their public analysis repository includes data-processing code, model reports and hierarchical bootstrap procedures used to estimate uncertainty across task families, tasks and attempts.

11 O'Clock Press did not rerun the pipeline for this story, and inspectable code is distinct from independent replication. The public repository still lets outsiders examine changes in the horizon across reliability thresholds, models, task distributions and treatments of uncertainty.

Consequential choices enter through the human reference group, tasks, definition of completion and relevant domain. A 50 percent success rate might suit brainstorming and prove reckless for logistics. Each use therefore needs a reliability threshold matched to its consequences.

Replace the AGI scoreboard

Morgan proposes a Strategic Causal Capability Observatory instead of a binary AGI declaration. It would track reliability curves alongside spatial reach, temporal horizon, causal depth, intervention success, cost, integration delay, human authority, uncertainty and adversary response.

Saleh's analysis supports continuous, multi-metric monitoring and rejects broad task breadth as sufficient scientific evidence of AGI. Vale calls the same definition under-specified: wide performance does not establish counterfactual accuracy, transfer under novelty, revision after failure or dependable long-horizon action.

Before adoption, this house-framework proposal needs operational definitions and preregistered tests. Its measures must also predict consequential outcomes better than simpler capability metrics, an empirical comparison that would show whether the larger set of numbers improves strategic measurement.

One chain under pressure

Military logistics shows what a real test could look like. Start with model observations, require a causal recommendation and record whether accountable humans adopt it. Then measure whether resources arrive where intended and readiness improves before introducing partial information, delay, scarce transport and an adversary trying to disrupt the plan.

The test would compare benchmark capability with outcomes under different levels of integration, authority and resources. If those mediators improve prediction, the result would support Stokes's conversion thesis; equal performance without them would weaken it. A monitoring scheme fit for strategic warning would also have to survive a domain shift.

Governments making costly decisions need measures that can fail in public, even when those measures are harder to build than model rankings.

Risk without prophecy

Stokes also treats loss of control as one possible pathway. The 2026 International AI Safety Report separates three necessary ingredients: sufficient capability, a harmful propensity and an enabling deployment environment. Current systems show early relevant capabilities, the report says, while falling short of capability sufficient for loss of control; experts disagree widely about future likelihood.

Shita's framework proposes permission gates and multi-agent foresight, ideas that have not been independently validated as an alignment solution. Present monitoring can still track assigned goals, derived subgoals, permissions, oversight, deployment access, observed deviation and response to correction.

What would change the verdict

Two tests could move this debate beyond scenario prose.

First, preregister a cross-domain intervention tournament. Give frontier systems and expert human teams the same observations and action budgets in unfamiliar physical, logistical, economic, cyber and social environments. Require causal models, counterfactual forecasts, interventions, calibrated confidence and correction after surprise. Sustained human-level performance across those domains would support a claim of human-range causal intelligence; persistent failure would weaken it.

Second, deploy the same system across comparable organizations while measuring capability, integration time, human adoption, authority, resources, adversary response and mission outcomes. Results in which institutional variables repeatedly improve prediction would support Stokes's conversion thesis. No added predictive value from those variables would weaken it and strengthen the case for a model-centered measure.

The report places the strategic contest after the technical breakthrough. Governments now need measurements that distinguish impressive performance from reliable power.

← More from the newsroom