My working definition of artificial general intelligence used to be quite simple.

I should be able to say, "Go and do that," and the AI should be able to understand the outcome, work out the route, use the tools, recover when something goes wrong and come back with the work done.

That still feels like an important threshold. It is also incomplete.

A system can work autonomously for a long time and remain narrow. A chess engine can search brilliantly without becoming an accountant. A coding agent can work across a repository without being able to run a hospital, negotiate a contract or learn an unfamiliar physical task.

What I was describing was agency. AGI needs agency, but agency is not the same thing as general intelligence.

AGI is not one magic switch. It is the point at which breadth, performance, adaptation, agency and reliability become general enough to change what intelligence means as an economic resource.

That is my synthesis after looking at the history, the competing definitions and the public evidence available on 27 August 2026.

It is not a definition everyone will accept. That is rather the point. There is no single agreed AGI finish line.

AGI is a moving target

We often talk about AGI as though a laboratory will ring a bell one morning and announce that it has arrived.

The history of artificial intelligence suggests something messier.

In 1950, Alan Turing asked whether machines could think and replaced that slippery question with an operational test based on conversation. The 1956 Dartmouth summer project helped establish artificial intelligence as a field. Researchers pursued programmes that could reason, solve problems, use language and learn.

Many tasks once treated as signs of intelligence have since become ordinary software. Computers beat world champions at chess. They recognise speech, translate languages, generate images and answer difficult examination questions. Each time a capability becomes reliable, people are tempted to say it was never really intelligence.

The modern phrase artificial general intelligence became more established through the 2007 book Artificial General Intelligence and the first AGI conference in 2008. It was useful because it separated broad, adaptable intelligence from systems built for a specific task.

But it did not settle what should count.

PeriodMilestoneWhy it matters
1950Turing publishes Computing Machinery and Intelligence.It turns a philosophical question into observable behaviour, while never claiming that conversation measures every form of intelligence.
1956The Dartmouth project establishes the name artificial intelligence.The founding ambition is broad: learning, language, abstraction and problem-solving, not one narrow application.
2007 to 2008The book Artificial General Intelligence is published, followed by the first AGI conference.A research community forms around general, adaptable machine intelligence as distinct from task-specific AI.
2010sDeep learning delivers large gains in perception, language and games.Capability expands, but systems still depend heavily on their training distribution and narrow objective.
2020sFoundation models become multimodal, tool-using and increasingly agentic.One model can perform many tasks. The debate shifts from whether breadth exists to whether it is adaptable and dependable enough to be general.

The three definitions that matter most

There are dozens of definitions of intelligence and nearly as many ways to define AGI. For practical decisions, I think they fall into three useful families.

1. Economic AGI

OpenAI's Charter defines AGI as "highly autonomous systems that outperform humans at most economically valuable work".

This definition is clear about impact. It does not require a machine to feel, possess a body or reproduce every feature of a human mind. If it can autonomously outperform people across most valuable work, the economic change has arrived.

The problem is measurement. What is "most" work? What is "outperform"? Is a system general if it needs different harnesses, tools and human reviewers for every occupation? Does it count if it is technically capable but too expensive or unreliable to use?

2. Capability AGI

Google DeepMind's Levels of AGI framework separates breadth from depth. A system can be narrow or general, and within either category it can be emerging, competent, expert, virtuoso or superhuman. The paper also treats autonomy and deployment risk as related but separate questions.

This is useful because it stops one spectacular benchmark becoming proof of general intelligence. Being superhuman at mathematics does not automatically make a system general. Being mediocre at many tasks does not make it useful.

The difficulty is choosing a sufficiently broad and honest set of tasks. Benchmarks can saturate. Tests can leak into training data. A score can hide brittleness, cost and the work humans did around the model.

3. Adaptive-learning AGI

Shane Legg and Marcus Hutter described intelligence in terms of an agent's ability to achieve goals across a wide range of environments. François Chollet later argued that intelligence should be measured by the efficiency with which a system acquires new skills, not merely by the skills it already possesses.

This definition asks the hardest question: can the system face something genuinely unfamiliar, form the right abstractions and learn without requiring another enormous training run?

A model may contain an astonishing stock of human knowledge and still be poor at learning a genuinely novel rule from sparse evidence. If AGI means adaptable intelligence rather than a very large library, that gap matters.

My five-part AGI test

I do not think one definition is sufficient on its own. For an organisation deciding whether a system is becoming generally intelligent, I would ask five questions.

  1. Breadth: Can it work across many substantially different cognitive domains, not just variations of one benchmark?
  2. Performance: Can it do those tasks at a useful human level, with the quality, speed and cost included?
  3. Adaptation: Can it learn unfamiliar tasks and recover from novelty rather than merely retrieve a familiar pattern?
  4. Agency: Can it plan, use tools, act and continue towards an outcome with proportionate supervision?
  5. Reliability: Can it do all of that dependably enough that people can responsibly delegate consequential work?

The dimensions multiply rather than add.

A broad system that is unreliable is an impressive assistant. A reliable system that only does one task is excellent automation. A powerful agent that cannot adapt outside familiar patterns is a capable worker in a bounded environment.

AGI is the claim that these qualities have become general together.

Who is closest?

The honest answer is that no public evidence lets us declare a winner, and no public system has demonstrated all five properties.

There are at least three reasons to be cautious. Developers choose which results to publish. Benchmark comparisons often use different harnesses, budgets and scoring methods. And the public sees products, not the full internal frontier.

So the chart below is not a percentage-to-AGI score. It is a dated map of the evidence each company has made public.

A dated qualitative evidence matrix for public frontier AI systems from OpenAI, Anthropic, Google DeepMind and SpaceXAI. All show strong breadth and partial agency, while novel adaptation and dependable long-horizon reliability remain incomplete or unclear.
Public evidence checked on 27 August 2026. Strong means the developer has published substantial evidence, not that the capability is complete or independently proven.
Public systemBreadth and performanceAgencyWhat still prevents an AGI conclusionEvidence boundary
OpenAI GPT-5.6 and CodexStrong public results across coding, knowledge work, browsing, science and computer use.Strong tool use and multi-agent coordination inside an increasingly capable harness.Public evidence still shows dependence on explicit goals, tools, environments, checks and human authority. Novel learning and dependable whole-job performance are not demonstrated.Most capability comparisons are vendor-published. Product task duration is not the same as reliable autonomous completion.
Anthropic Claude Fable 5, Mythos 5 and Opus 5Strong public claims across software, professional work, research, vision and cybersecurity.Strong long-context and coding-agent behaviour, with explicit work on longer-running agents.Adaptation to genuinely unfamiliar tasks and reliable completion across whole occupations remain unproven. Safeguards and access tiers also change what users can observe.Many headline results are Anthropic evaluations. External maintainers and experts still validate consequential findings.
Google DeepMind Gemini, Deep Think and Project AstraStrong multimodal, scientific and general reasoning evidence across a broad research portfolio.Project Astra demonstrates a persistent multimodal assistant direction; Gemini supports tool-using agents.A universal assistant is not automatically an autonomous worker. Public evidence does not show one system reliably learning and completing most economically valuable work.Project Astra belongs to Google DeepMind, not OpenAI, and is a research programme whose capabilities feed products such as Gemini Live.
SpaceXAI Grok 4.6 and Grok BuildVendor evidence places Grok at the public frontier for reasoning, knowledge work and coding.Designed for long-running agents, workflows and parallel execution.There is not enough independent public evidence of broad novel adaptation or reliability across consequential real work.The current evidence is recent and largely vendor-reported. Long-running execution remains different from general intelligence.

Other important systems should not be ignored. Meta, Mistral, Alibaba/Qwen, Moonshot/Kimi and open-source communities contribute substantial models and research. But a four-row table is already at risk of implying a league table that the evidence cannot support.

The conclusion that holds up is simpler: the public frontier is a cluster, capabilities are jagged, and the missing evidence is remarkably consistent.

We have breadth. We have increasingly deep performance. We have useful agency.

We do not yet have public proof of robust, efficient adaptation and dependable long-horizon performance across most real work.

Why a long-running agent is not AGI

This matters to the argument I made in The Coworker Is Here. The Employee Is Next.

If I can give an agent a release goal and it works for two days, that is a large operational change. It may alter how I organise a company. It does not prove that the underlying model is generally intelligent.

The harness may be preserving context, choosing tools, splitting the task, running tests, retrying failures, asking a stronger model for help and stopping at an authority boundary.

That does not make the achievement fake. Human intelligence also depends on notebooks, institutions, colleagues, tools and culture.

It does mean we should judge the whole system rather than pretending the model alone did everything.

METR's time-horizon work is useful here. It estimates the human-expert duration of software tasks at which an agent reaches a given probability of success. METR explicitly warns that this is not the length of time an AI can work independently. A 50% time horizon is also not a sensible delegation threshold for a safety-critical task.

The International AI Safety Report 2026 reaches a similar conclusion. Current general-purpose systems are impressive across well-specified tasks, but their performance remains jagged. They still struggle with long-term planning, coherent strategy, unexpected obstacles and reliable automation of many whole jobs.

We are getting better workers. The evidence does not yet establish a general employee.

What about Astra?

I originally associated the name Astra with OpenAI. That was wrong, and it is worth correcting clearly.

Project Astra is Google DeepMind's research towards a universal AI assistant. It is designed to understand what a person sees and hears, remember context, use tools and converse naturally across devices. Some of its capabilities are being incorporated into Gemini Live.

Astra is relevant because it moves intelligence out of a text box and into an ongoing relationship with a person and their environment.

But Astra is not public evidence of an agent independently working for days, and it is not an announced OpenAI route to AGI.

A future system could combine Astra-like perception and presence with Codex-like work, deep adaptation, long-term memory and dependable execution. That would move further through my five gates.

Even then, the name on the product would not answer the question. The evidence would.

The Bugatti becomes a bicycle

Suppose we reach a system that can do most valuable cognitive work at a high level.

At first, it may be like driving a Bugatti to the shops.

It will work, but the compute, energy, specialised chips, engineering and verification may make it absurdly expensive for everyday use. The first organisations able to deploy it widely will be governments, laboratories and companies with capital.

Then the engineering curve begins. Models become smaller. Chips improve. Inference is optimised. Work moves to cheaper capability tiers. Harnesses stop repeating themselves. Competition lowers prices.

The Stanford AI Index reported that the inference cost of a system performing around GPT-3.5 level fell by more than 280 times between November 2022 and October 2024. That does not prove AGI will follow the same curve. It shows why today's price should not be treated as a permanent boundary.

The Bugatti may become a bicycle: widely available, ordinary and woven into daily life.

There is an important catch. Yesterday's frontier becomes cheap while tomorrow's frontier remains expensive. Access can broaden at one level while power concentrates at the next.

That is why I keep returning to the cost of a successful AI task. Intelligence may become abundant in technical terms without becoming equally available in economic or political terms.

What AGI would mean for society

I do not think the most useful answer is "everyone loses their job" or "nothing really changes". Both are too easy.

Work becomes a question of task design

The IMF estimated that about 40% of global employment is exposed to AI, rising to around 60% in advanced economies. Exposure is not the same as replacement. Some tasks are automated, some are improved and some occupations are reorganised.

AGI would push that reorganisation much further. The scarce skill may move from producing every artefact to choosing outcomes, designing constraints, judging evidence and taking responsibility.

That makes competence in the loop more important, not less.

Small organisations gain leverage

A founder, charity, school or local business could gain access to capabilities that previously required a large professional team. The ability to research, analyse, build, test and communicate could spread much more widely.

That is the optimistic case: intelligence levels the playing field.

Capital and infrastructure gain power

If the best systems require enormous compute, the organisations controlling chips, data centres, energy, models and distribution can gain extraordinary influence.

The relevant question is not only who invented AGI. It is who can afford to run it, who decides what it may do and who can inspect the evidence.

Science may accelerate

AI systems are already assisting mathematics, biology, materials research and software security. A system that could form good hypotheses, learn a new field, design tests and work reliably with human researchers could compress parts of discovery.

But scientific fluency is not scientific truth. Claims still need experiments, reproducible evidence and challenge.

Governance has to become operational

Principles are not enough when systems can act. AGI-era governance needs permissions, logs, budgets, evaluations, competent review, appeal, shutdown routes and named accountability.

Capability can be broad. Permission should remain specific.

How will we know?

I would not accept a declaration based on one exam, one extraordinary demonstration or one company's marketing page.

I would want:

  • independent evaluations across genuinely different domains
  • novel tasks that were not foreseeable from public training material
  • full accounting for tools, scaffolding, human help, time and cost
  • success rates suitable for the consequence of the work
  • evidence from deployment, not only a benchmark
  • clear failure cases and the ability to recover from them
  • repeated results across laboratories and environments

I would also resist turning AGI into a binary distraction.

A system does not need to meet everyone's philosophical definition of AGI before it changes a company, an occupation or a country. The threshold for economic disruption may arrive before the threshold for scientific agreement.

My answer

What is AGI?

It is not merely a chatbot that knows a lot. It is not merely a model that wins a benchmark. It is not merely an agent that can keep going.

It is a system that combines broad capability, high performance, efficient adaptation, useful agency and dependable execution across enough of the world that intelligence stops behaving like a specialist tool.

Have we reached it?

Not on the public evidence I can verify as of 27 August 2026.

Are we building systems that contain more of its ingredients?

Absolutely.

And perhaps the most important question is no longer simply, "When will AGI arrive?"

Who will be able to use how much intelligence, at what cost, with what data, permissions and accountability?

That is where the definition stops being academic.

Sources and notes

This is a dated evidence review, not a declaration about private systems. Product capabilities change quickly. Vendor benchmark claims are labelled as vendor evidence, and no single benchmark is treated as proof of AGI.