What Is Artificial General Intelligence? Separating Capability from Promise
A guide to separating definition, evidence, system, and promise when someone claims an AI is general.
Re-edited on July 30, 2026, this article begins with a precaution: “artificial general intelligence” does not currently name a particular machine or a goal with a universal test. It names an aspiration—systems able to operate across many tasks and contexts—whose definitions change with the comparison, allowed resources, and required autonomy. An impressive demonstration is therefore not enough to announce AGI.
The distinction matters outside laboratories. A company may market an assistant that writes, codes, and analyzes images as “general,” even if it fails on a new instruction, depends on external tools, or needs constant supervision. Readers need to separate four layers: definition, evidence, deployed system, and promise. That discipline remains useful as models and product names change.
Define before measuring
Any useful definition must make it possible to say what observation would contradict it. In their collection of definitions of intelligence, Shane Legg and Marcus Hutter identified a common core: the ability to achieve goals across a wide range of environments. It is a starting point, not a certificate. The relevant breadth, efficiency, goals, and comparison with people remain open.
François Chollet proposed in On the Measure of Intelligence that measurement should focus on the efficiency with which a system acquires new skills using experience and prior knowledge. This corrects a common confusion: accumulating learned answers is not the same as adapting. When an exam resembles training, it may measure data coverage or engineering rather than generalization.
The Levels of AGI framework separates performance, generality, and autonomy, and proposes levels instead of a magical boundary. Its value is that it forces people to finish the sentence. General across which task set? Compared with which human percentile? With assistance, tools, and retries? Two teams can use “AGI” for incompatible goals while appearing to disagree about one object.
Generality must also be separated from anthropomorphism. Fluent conversation, first-person language, or descriptions of emotion are interaction properties. They do not demonstrate stable memory, a causal world model, intention, or human understanding. An operational definition lists observable behavior; it does not turn a user’s impression into evidence about an internal state.
Broad capability can remain fragile
Current systems can cover domains that once required separate programs. That is a real expansion, but “many tasks” does not mean “any task” or “learns like a person.” The GPT-3 paper evaluated few-shot learning across numerous language datasets and documented uneven results as well as limitations. The proper evidence is that task profile, not the conclusion that the model understood the world generally.
The contrast with extraordinary but narrow systems helps. AlphaGo combined policy and value networks, tree search, human games, and self-play for Go. Defeating elite players demonstrated a powerful solution inside an environment with defined rules, actions, and objectives. It did not show spontaneous transfer to contracts, driving, or medicine. The magnitude of a result and the breadth of a capability are separate axes.
Fragility appears when something the exam held fixed changes: question format, language, a rule, the case distribution, a tool, or the cost of error. A system may answer one hundred tests well yet fail to decide when it does not know. Claims about generality should measure adaptation to held-out tasks, required demonstrations, computing cost, variation across runs, and recovery from mistakes.
The evaluated unit must be explicit. Sometimes a result comes from a model; sometimes it comes from a system with search, a calculator, memory, filters, instructions, and a person selecting outputs. That composition may be extremely useful. Attributing everything to the model, however, prevents reproduction and hides dependencies. The honest question is what each component did and what happens when it is absent.
Promising techniques are not a staircase to AGI
Deep networks, recurrent memory, self-supervision, reinforcement learning, search, and transfer solve different problems. A recurrent architecture handles sequences; a self-supervised task constructs signal from data; reinforcement optimizes decisions against a reward; search explores alternatives; transfer reuses representations or parameters. No label by itself implies general reasoning, safety, or autonomy.
Nor do the techniques form an inevitable historical sequence. An advance may improve planning while reducing robustness, extend context while increasing cost, or master a benchmark by exploiting a shortcut. Integrating several methods creates a new system that must be evaluated as such. Saying “it combines X and Y, so it is closer to AGI” replaces a test with an imagined direction.
To audit a claim, build a matrix. Put familiar tasks, novel tasks, adversarial variations, and consequential actions in the rows. Put success, adaptation data, time, tools, human intervention, and error cost in the columns. An empty cell does not prove inability, but it limits the conclusion. An aggregate score should never erase which capabilities were observed and which were inferred.
Benefit and risk depend on use, not the label
Promises that a future AGI will solve energy, disease, or climate change mix a hypothetical capability with outcomes that depend on data, institutions, resources, and political choices. A system may generate hypotheses without access to experiments, optimize an electrical grid without building infrastructure, or propose a molecule without establishing clinical safety. Benefits must be specified as a verifiable chain of decisions, not a list of human problems.
The same applies to risk. Before extreme scenarios come observable failures: discrimination, data leakage, inappropriate automation, provider dependence, cyber misuse, and loss of operational control. Work on evaluating dangerous capabilities proposes investigating whether general-purpose models could facilitate manipulation, cyberattacks, resource acquisition, or other harm. An evaluation does not prove that harm will occur; it identifies capabilities that need controls and monitoring.
The NIST AI Risk Management Framework organizes work into govern, map, measure, and manage. That sequence avoids waiting for the philosophy of AGI to be settled before acting. Teams can document context, affected people, limits, metrics, oversight, and incident response for a real system. Risk management begins with the deployment that exists, not with a laboratory’s most ambitious label.
Regulation can also operate without a single answer about AGI. The European Union’s AI Act establishes duties according to uses and, for general-purpose models, according to responsibilities and possible systemic risks. It suggests a sound method: asking who provides a system, who deploys it, what it is used for, and which impact it can cause is often more productive than debating whether the system “is general.”
A protocol for reading the next announcement
First, copy the exact definition. If it contains only words such as “human,” “reasons,” or “general,” ask for tasks, comparison population, and allowed resources. Second, find the primary evidence: technical report, system card, data, and protocol. Third, separate measured results from selected demonstrations and forecasts. Fourth, identify the evaluated unit—model or system—and the human intervention.
Then look for missing tests: held-out tasks, distribution shifts, calibration, ability to abstain, independent replication, and costs. Ask whether training may have contaminated the test and whether failures, rather than just averages, were published. Finally, translate “important” into a specific decision: purchase, deployment, regulation, research, or observation. Each decision requires a different evidence threshold.
AGI can remain a valuable scientific and social question without becoming a countdown. Progress is better described as a profile of capabilities, limits, and resources than as distance from a goal with no shared scale. The transferable skill is separating definition, evidence, system, and promise. With those four layers, a reader can recognize a real advance without granting it capabilities that nobody measured.
This article was produced with artificial intelligence under human editorial oversight.