Weak AI, Strong AI, and AGI: Three Questions Not to Confuse
Weak versus strong AI does not separate rules from machine learning. A five-axis framework for evaluating capabilities without inflating labels.
Re-edited on July 30, 2026, this article corrects a misleading classification: “weak AI” does not necessarily mean a simple rule-based program, and “strong AI” is not the name for systems that use machine learning. Those terms arose in a philosophical debate about minds and understanding. Mixing them with the breadth of a task or the technique being used collapses three different questions into one label.
To evaluate a system, it helps to separate what it can do, how it learns, and what claim is being made about it. A neural model can be highly capable while remaining confined to a domain. A hand-coded program can solve a complex task. Producing answers on many subjects does not, by itself, demonstrate human understanding, consciousness, or general capability.
“Weak” and “strong” described a claim about minds
John Searle drew the distinction in his 1980 paper Minds, Brains, and Programs. Under the view he called strong AI, an appropriately programmed computer would not merely simulate a mind: it would have cognitive states, and the program would explain cognition. Under the weak reading, the computer is a powerful tool for studying and testing hypotheses about minds. The difference was therefore not “simple algorithm versus advanced algorithm.”
Searle’s well-known Chinese room argument was intended to separate correct symbol manipulation from understanding what the symbols mean. Responses to the argument remain disputed; the practical point here is to use the vocabulary accurately. Calling a speech recognizer, control system, or data-trained model “strong AI” claims much more than performance on a task. There is no accepted operational test that turns a high score into evidence of a mind or consciousness.
It is equally inaccurate to say that weak AI merely follows predefined instructions. A system can learn millions of parameters from data, adapt its outputs to new inputs, and still be a specialized tool. “Learned” does not mean “general”; “complex” does not mean “strong”; “fluent” does not mean “conscious.”
Specialized AI and general intelligence are another axis
The specialized/general distinction asks about the scope of capabilities, not whether a mind exists. Specialized AI is designed and evaluated within a bounded set of tasks and conditions. Artificial general intelligence, or AGI, would refer to a much broader capacity to acquire and transfer skills across domains. The boundary is disputed and depends on which tasks, resources, and degrees of transfer a definition requires.
AlphaGo illustrates the difference. The Nature paper on the system documented a combination of policy and value networks with tree search that defeated the European Go champion 5–0. It was an extraordinary achievement inside a game with defined rules, actions, and objectives. Superiority in that setting does not imply that the same system can drive a car, diagnose a disease, or learn another game without redesign.
Apparent breadth is not sufficient either. A language interface may accept requests about coding, history, and cooking, yet the number of topics does not reveal how it responds to a genuinely new task, how many examples it needs, or which failures carry across domains. In On the Measure of Intelligence, François Chollet proposes measuring the efficiency with which a system acquires new skills while accounting for experience and prior knowledge. It is a specific proposal, not the final definition of intelligence, but it exposes what a fixed demonstration conceals: how much work the system completed before the test.
There is no ladder from rules to consciousness
A useful classification needs several axes. The first is scope: one task, a family of tasks, or transfer to new domains. The second is adaptation: fixed rules, learned parameters, updates during use, or later supervised learning. The third is autonomy: recommending, deciding, or acting, and under what supervision. The fourth is the environment: digital, physical, stable, or changing. The fifth is consequence: what happens when the system is wrong.
Those axes do not form a single ladder. A statistical filter may learn from data and have little authority to act. A rule-based controller may operate machinery directly and require strict assurances. A chatbot may cover many subjects but depend on human review before any action with real effects. The architecture—trees, neural networks, or generative models—explains part of the mechanism; it does not determine scope, autonomy, or risk by itself.
Performance must also be separated from generality. Beating people on a narrow test establishes a result under the conditions of that test. A broader claim requires unseen sets, different tasks, controlled variations, comparable costs, and a protocol that prevents adapting the system to the exam. A product name or an impressive conversation is not a substitute for that evidence.
How institutions define actual systems
Current operational definitions avoid settling the philosophical issue first. The OECD’s updated definition, approved in November 2023, describes a machine-based system that infers, from the input it receives and for explicit or implicit objectives, how to generate predictions, content, recommendations, or decisions that can influence physical or virtual environments. The memorandum also explains that systems vary in autonomy and adaptiveness after deployment.
The OECD pairs that definition with a classification framework examining people and planet, economic context, data and input, model and task, and output and impact. It does not put every system into a metaphysical box; it requires a description of context and effects. Two technically similar models may need different controls if one recommends films and the other influences access to employment.
The European Union AI Act likewise defines an AI system in Article 3 through its machine-based operation, varying levels of autonomy and adaptiveness, and inference of outputs from inputs. It then structures obligations around particular uses and risks. A legal definition serves a regulatory purpose; it does not certify that a system thinks or settle the debate over AGI.
A method for reading any claim about AI
When faced with “this AI is more advanced,” the first question is: on what task? Identify the input, output, conditions, and metric. The second is: compared with what? The reference may be an earlier version, another tool, human specialists, or a very weak baseline. The third is: what was kept outside the test? A result on repeated data, one language, or a simulated environment does not support a universal conclusion.
Then separate learning from operation. Were the parameters trained once, or do they change during use? Does the system merely propose an action or execute it? Can it abstain? Who reviews it? Finally, examine the consequence of error: a wrong label on a photograph does not carry the same cost as a clinical recommendation or employment decision. That description supports appropriate metrics, supervision, and routes for challenge without depending on a grand label.
“Weak,” “strong,” “specialized,” and “general” can be useful when their meanings are defined, but they should not function as certificates. A useful claim also states the system version, test date, allowed tools, comparison, and known failures, so another reader can determine whether the evidence still describes the deployed product. The transferable skill is to decompose any classification into five testable questions: scope, adaptation, autonomy, evidence, and consequence. That makes it possible to recognize a real improvement without turning it, through a verbal leap, into a claim about general intelligence or a mind.
This article was produced with artificial intelligence under human editorial oversight.