A One-Year AGI Forecast Needs an Agreed Finish Line

Selected AI benchmarks show rapid progress alongside deep weaknesses elsewhere. Calibrating an AGI date requires a fixed event, system and test.

OSWorld performance climbed from about 12% to 66.3% in a year, putting AI agents within six percentage points of that benchmark's human baseline, according to Stanford's 2026 AI Index. An advanced Gemini system produced five correct solutions to the 2025 International Mathematical Olympiad problems and received a gold-level score.

Elsewhere, the same broad category of frontier AI looks far less capable. Stanford reports 50.6% accuracy for the leading model on ClockBench, compared with 90.1% for people. ARC-AGI-3 places agents in unfamiliar interactive environments and requires them to explore, infer a goal and plan from the environment alone. People solved all calibrated environments; frontier systems scored below 1% as of March 2026.

The percentages have scales specific to their systems, constructs and protocols. Read side by side, the results trace a jagged capability profile rather than a timetable for artificial general intelligence.

Ben Goertzel, chief executive of SingularityNET and chair of the Artificial General Intelligence Society, told IBM Think that progress in coding, research automation and groups of AI agents had raised his expectations. “But it increased my probability weight that we could get an AGI breakthrough within the next nine or 12 months.”

Goertzel described a change in probability and acknowledged uncertainty. Assessing that forecast requires a defined event, an account of how the evidence changes its likelihood and an outcome that would count against it. IBM's article describes a field still seeking a standard definition and accepted AGI test.

What the benchmark record establishes

The OSWorld and mathematics results are substantial engineering achievements, each bounded by its task. Google DeepMind noted that IMO coordinators graded the submitted solutions; validation covered the answers alone, leaving the model and production process outside its scope. ClockBench tests visual time reading; ARC-AGI-3 tests adaptation when rules and goals are initially hidden. Evidence for generality has to extend beyond any one result.

Dr. Mariam Saleh, an AI scientist and 11 O'Clock Press Scientific Adviser, calls the record a jagged capability profile. Her assessment identifies the missing validated mapping from benchmark slopes or increased use of agent groups to an operational AGI event by August 2027. Fellow Scientific Adviser and AI scientist Dr. Francis Vale points to the unfixed endpoint and capability set that one unchanged system would have to pass.

A reproducible route from current public evidence to the nine-to-twelve-month inference is missing. This finding concerns the forecast's scientific calibration; an AGI breakthrough could still occur within the window.

A test designed before the deadline could narrow the gap. Before running it, researchers would record the system configuration, permitted tools, human assistance and scoring rules. One unchanged system would then face every declared threshold on private, contamination-resistant tasks, matched to human baselines, that cover novel problem solving, continual causal learning, cross-domain transfer and long-horizon planning. This design prevents replacement of a failed task after its result is known.

From proposed ingredients to controlled tests

Goertzel presents continual learning, persistent self-models, predictive coding and neural-symbolic systems as possible ingredients of general intelligence. The IBM report gives one concrete technical boundary: his team's predictive coding work in transformers reaches roughly GPT-2 scale, with current frontier-model scale still beyond the reported result. The article stops short of linking a primary experiment that would let outsiders inspect the setup.

A matched ablation would give systems the same nonstationary tasks and resources while allowing only one to revise an online causal model and maintain explicit self-monitoring; the control would have those functions removed or frozen. Pre-registered measures could track transfer, intervention accuracy, recovery after environmental change and resource use. The mechanism would gain support from consistent improvement on held-out tasks and lose support after a null or reversed result. Establishing the necessity or sufficiency of any proposed ingredient awaits such evidence.

Goertzel's proposed “MIT test” would ask an AI-controlled robot to navigate campus, take classes and exams, work with professors and produce an original thesis under human rules. Completing it would demand embodiment, communication, planning and research. The result would bundle contributions from prior training, tools, academic field, institutional accommodations, human assistance, hardware and examiner judgment.

François Chollet's skill-acquisition framework supplies a more focused question for that scenario: how much prior information and experience does the system need to acquire a new capability? Measuring that efficiency could reveal whether a successful doctorate reflected broad adaptation or extensive scaffolding.

Testing the house framework too

Mounir Shita's Science of Intelligence framework proposes spatial range, temporal horizon and causal depth as dimensions of intelligence. Applied to the forecast, they would test the breadth of environments in which a system can act, the length of time its plans survive disruption and the number of linked causes it can model and control.

TGI originates in the house theory, and scientific consensus around it has yet to form. Evaluating it calls for independent calibration, comparisons with rival frameworks and declared failure conditions, including a test of whether its dimensions predict transfer beyond the tasks used to define them. Failure on that test would weaken TGI as a measuring framework even if the system showed other capabilities.

AI capability is advancing quickly across some tasks and faltering on others. Judging whether that uneven progress supports a one-year window requires the forecast to specify its event, system, tests and failure conditions in advance.

← More from the newsroom