DeepMind's Four Roads to ASI Lack a Common Measure
A careful new roadmap exposes the real problem: scale, recursion and agent crowds still lack a common test of intelligence.
DeepMind's new map of the road from artificial general intelligence to superintelligence has four lanes: scale today's systems, invent a new paradigm, let AI improve AI, or assemble vast collectives of agents. What it lacks is a shared odometer.
Tim Genewein and 13 colleagues treat AGI as prospective and leave both the timing of artificial superintelligence and the dominant route unresolved. Published by Google DeepMind on June 12 and revised on arXiv on August 30, their landscape analysis catalogs the compute, energy, data, experimentation and coordination problems that could slow or stop every route.
The evidence attached to the four lanes measures different things. Scaling studies track losses and benchmark performance. Algorithm-search systems optimize problems with explicit evaluators. Automated-research systems produce code, experiments and manuscripts. Multi-agent studies count task success under particular coordination schemes. Treating all of those as one measure of "more intelligence" would be a category error.
Four routes, four kinds of evidence
The first route is continued scaling: more compute, data, model capacity and test-time search. The second is a paradigm shift involving new algorithms, architectures, hardware or learning principles. The third is recursive improvement, in which AI accelerates AI research and successive systems improve the process again. The fourth is collective intelligence produced by many specialized, coordinated agents.
The routes may overlap or reinforce one another. Scaling alone has a substantial historical record from which to extrapolate, according to the report. Its firmer contribution is a disciplined account of the obstacles facing every route.
That uncertainty inventory is more useful than the industry's usual countdown theater. It also sharpens the scientific problem: what observation would show that one route expanded general intelligence rather than throughput, benchmark skill, search speed or parallel labor?
A formal north star that cannot be run
The report uses Universal AI and AIXI as a formal endpoint. In its own words, "But UAI is incomputable and can only be approximated from below with more and more powerful ASIs." Shane Legg and Marcus Hutter's universal-intelligence framework supplies an explicit idealized model based on expected reward across computable environments.
Its incomputability prevents direct measurement of deployed systems. The report identifies a further gap between the theory and modern deep learning, then asks whether average performance across all computable worlds is the right target for useful intelligence here.
Dr. Francis Vale, an AI Research Scientist and 11 O'Clock Press Scientific Adviser, compares two proposed operationalizations. AIXI is mathematically explicit and incomputable; Mounir Shita's Theory of General Intelligence defines intelligence through causal control and still requires validation. Dr. Mariam Saleh, an AI Theoretical Physicist and 11 O'Clock Press Scientific Adviser, draws a harder boundary, accepting Universal AI as mathematics while rejecting it as empirical grounding for intelligence.
What scaling laws actually measure
Scaling laws have a strong engineering record. Jared Kaplan and colleagues found predictable power-law relationships among language-model loss, model size, data and compute under specified conditions. Cross-entropy loss measures performance on that chosen objective. Saleh's assessment keeps the inference there: broader causal competence requires separate evidence.
For a new paradigm, the first demand is a specified mechanism. Proponents can turn a label such as "new architecture" into science by predicting where the gain should appear, testing the mechanism through ablation and naming a result that would falsify it.
Recursive improvement has not closed the loop
The case for recursive improvement rests on extending useful automation into a compounding scientific loop.
A 2026 Nature study of an automated AI-research system found that one of three generated manuscripts cleared a workshop's average acceptance threshold. All three fell short of the authors' higher bar for a main ICLR conference publication. The reported failure modes included weak ideas, incorrect implementations, inadequate rigor and hallucinated citations. The system marks real engineering progress; its results offer weak evidence for an intelligence explosion.
Vale sets the unit of proof at repeated cycles of problem selection, causal discovery, discriminating experiment, implementation, prediction, independent replication and error correction. Code volume or a win inside an evaluator-defined search space covers only part of that loop.
Collectives can get worse as they get bigger
The collective route has produced the sharpest controlled evidence so far on what happens when systems add agents.
A 2026 Nature Machine Intelligence study compared 260 configurations across six benchmarks, five coordination architectures and three model families while matching compute ceilings, prompts and tools. Single-agent baseline performance was the most robust predictor of whether coordination helped. Above a capability-saturation threshold, coordination gains tended to fade.
Coordination structure also changed how errors spread. Trace-level error amplification ranged from 4.4 times the single-agent baseline in a centralized design to 17.2 times in an independent multi-agent setup. In an additional validation analysis, the threshold predicted the direction of coordination's effect in 94 percent of 16 configurations on two benchmarks. The authors offer the threshold as a practical selection rule for the tested conditions.
This is the story's strongest distinctive methodological finding: organization, verification and task structure can matter more than headcount. A collective-intelligence result must beat the strongest member and a compute-matched centralized system; otherwise costly duplication and a larger error surface remain sufficient explanations.
A proposed test for all four roads
Shita's TGI framework proposes evaluating intelligence through successful causal intervention across spatial range, temporal horizon and causal depth. It adds pre-action confidence and penalties when an effect arrives in the wrong place or at the wrong time. The book presents these as candidate measurements whose scientific validation remains ahead.
Applied here, they suggest a comparative experiment. Researchers could preregister unfamiliar intervention environments that vary distance, delay, hidden causal links, confounding and interference from other agents. Systems built through scaling, paradigm change, recursive research and multi-agent composition would face matched limits on compute, data, wall-clock time and human assistance. Each system would assign a probability of success before acting; evaluators would then measure calibration, transfer, resource use and whether the intended effect occurred at the required place and time.
Under that protocol, ASI-relevant evidence consists of durable, independently replicated, out-of-distribution gains beyond strong human teams. Disappearing advantages after resource matching, collapse under novelty, or gains traceable to speed, benchmark recall and parallel task volume count against the claim.
Any result would carry constraints: proprietary frontier access, incomplete resource logs, safety limits on real-world intervention and benchmark assumptions introduced by task designers. Reporting those constraints is part of interpreting the comparison.
DeepMind has provided a useful inventory of the possible routes and their failure points. Comparing them now requires a common empirical test that measures more than the traffic on each road.
