What the AGI Debate Gets Wrong: A Sober Look at Where AI Actually Is in 2026

Everyone is arguing about whether AGI is here. Almost nobody agrees on what it is. Here's what the 2026 data actually shows about AI's real capabilities.

By Ubedulla · 6 min read
What the AGI Debate Gets Wrong: A Sober Look at Where AI Actually Is in 2026

If you want to start a fight between two AI researchers in 2026, ask them whether AGI has arrived. Sam Altman talks about slipping past human-level intelligence toward superintelligence. Demis Hassabis says current systems are "nowhere near" human-level. Yann LeCun says large language models will never get there at all. These are three of the most informed people on the planet about frontier AI, and they cannot agree on where we are — because the agi debate has quietly stopped being a scientific question and become a definitional one.

That matters more than it sounds. The people building the most powerful AI systems in history genuinely disagree about what they are building, how far along they are, and what "done" looks like. And while the headlines fixate on whether some threshold has been crossed, the actually measurable story — what these systems can and cannot do, task by task, hour by hour — gets far less attention than it deserves.

So let's skip the vibes and look at what the numbers say.

The AGI debate is mostly a fight about definitions

There has never been a consensus definition of artificial general intelligence. Melanie Mitchell laid this out in Science: the term has meant everything from "human-level performance on most cognitive tasks" to "AI that can do any economically valuable work" to something closer to consciousness, depending on who is speaking and, often, what they are selling. OpenAI's charter definition — highly autonomous systems that outperform humans at most economically valuable work — is not the same thing DeepMind means, which is not what LeCun means, which is not what a philosopher means.

This is why the loudest disagreements evaporate under inspection. When one camp says AGI is here and another says it's a decade away, they are frequently describing the same systems and the same evidence. The dispute is over where to draw a line that was never drawn in the first place. A debate where the terms are undefined isn't a debate — it's a branding exercise.

What the people building it actually said

At Davos in January, the disagreement was on full display, and it was instructive. According to Fortune's coverage, Hassabis put roughly 50% odds on AGI within a decade and said it may require "one or two more breakthroughs." Dario Amodei predicted AI would replace the work of software developers within a year and reach Nobel-level science within two. LeCun argued LLMs are structurally incapable of humanlike intelligence and that "language is easy" compared to modeling the physical world.

Notice what's happening: Hassabis is describing a scientific milestone, Amodei an economic one, LeCun an architectural one. All three can be simultaneously right about their own claim and the headline still reads "AI leaders clash on AGI." Meanwhile the timelines keep compressing — Hassabis, who predicted 2030–2035 last year, has since narrowed his window to 2029–30. Whether that reflects new evidence or competitive pressure is left as an exercise for the reader.

What the measurements say instead

Here's the part of the story that doesn't fit on a conference stage. The old benchmarks are dead — frontier models now cluster above 90% on MMLU and similar exams, making score differences statistically meaningless at the top. But newer, harder measurements paint a much more specific picture.

The most useful one comes from METR, which tracks how long a task an AI agent can complete autonomously. In its January 2026 Time Horizon update, the leading model — Claude Opus 4.5 — could complete software tasks that take humans about 320 minutes, at a 50% success rate. That horizon has been doubling roughly every seven months since 2019, and faster recently. It is a genuinely steep curve, and if it holds, agents handling multi-day projects are years, not decades, away.

But read the fine print, because it cuts both ways:

  • 50% success is the headline number. For most real work you need 95% or better, and the reliable-completion horizon is far shorter than five hours.
  • The confidence intervals are enormous. METR's own estimate for Opus 4.5 spans 170 to 729 minutes — the researchers are honest that the measurement is noisy.
  • The tasks are mostly software. Coding is where models are strongest and where training data is richest; the curve says little about, say, running a negotiation or a lab.
  • Hard exams still humble everything. Frontier models score around a third on Humanity's Last Exam, where human domain experts average about 90%.

The jagged frontier is the real story

Put those numbers together and you get something no single word — certainly not "AGI" — captures: systems that are superhuman in some directions and preschool-level in others. A model can draft a working compiler patch and then confidently misread a train timetable. Anyone who uses these tools daily knows the shape of this already — genuine, compounding usefulness punctuated by failures a human would never make.

Hassabis's list of missing pieces is a decent map of the jagged edges: learning from a handful of examples, continuous learning after deployment, durable long-term memory, and reasoning that doesn't degrade on genuinely novel problems. None of these are solved. All of them are being worked on. Whether closing them requires "one or two breakthroughs" or a decade of grind is exactly the thing nobody knows — and exactly the thing confident timelines paper over.

The most useful question in 2026 is not whether AI is "generally intelligent." It's which specific tasks a system can do, at what reliability, for how long, and with how much human oversight. On that question — unlike the AGI question — the data is finally getting good.

Better questions than "is it AGI yet?"

The binary framing fails because nothing about this technology is binary. Capabilities arrive unevenly, diffuse slowly, and matter differently across domains. A more honest scoreboard looks like this: How long is the autonomous task horizon at 95% reliability? Which job tasks — not jobs — have actually been automated at scale? Where do error rates still make deployment irresponsible? What happens to the METR curve when tasks leave the software domain?

These questions have answers you can check. "Is it AGI?" does not, and the incentives around it are terrible: labs benefit from proximity to AGI when raising money and distance from it when facing regulation. The word does real work in contracts and policy — OpenAI's relationship with Microsoft famously hinged on it — which is precisely why it will keep being stretched.

None of this is an argument that the stakes are low. If METR's curve holds even approximately, the next few years get strange fast, and Amodei's labor-market warnings deserve engagement rather than eye-rolls. It's an argument for keeping your eye on the measurements instead of the milestone. The milestone will be declared, disputed, and re-declared. The measurements will just keep moving.

FAQ

Has AGI already been achieved?

By some definitions, arguably; by most, no. Frontier models exceed average human performance on many exams and text tasks, but they still lack continuous learning, robust long-term memory, and reliable autonomy over long tasks — gaps that leaders like Demis Hassabis say put current systems "nowhere near" human-level. Any confident yes or no is really a claim about definitions, not capabilities.

When do most experts think AGI will arrive?

There is no consensus, but the center of gravity among frontier-lab leaders has moved to the late 2020s and early 2030s. Hassabis has narrowed his estimate to roughly 2029–30, Altman suggests even sooner, while skeptics like Yann LeCun argue current LLM architectures will never get there without fundamentally new approaches. Treat all timelines as arguments, not forecasts.

Why does the AGI debate matter for ordinary users?

Mostly, it doesn't — what matters is the capability curve underneath it. Autonomous task horizons doubling every several months affect your job and your tools regardless of what label anyone applies. Judging tools by what they reliably do today will serve you better than tracking whether someone declares the finish line crossed.

About the author

Ubedulla

Founder & Editor

Founder and editor of The Bot Post, covering AI news and technology.

Related Articles