IA 360
General Artificial Intelligence (AGI)

What It Means for AI to Have Common Sense

Common sense is not a single function. This guide shows how to test what knowledge, exceptions, and grounding an AI system actually demonstrates.

Admin IA360 4 min read AI-generated Leer en español
What It Means for AI to Have Common Sense

This guide, revised on July 30, 2026, starts with an ordinary scene: someone says, “I put the ice cream in the cupboard,” and we expect a listener to notice the problem without first receiving a lesson in thermodynamics. An AI system may complete the sentence fluently and still fail to recognize that the action conflicts with the goal of preserving the food. The gap between producing a plausible answer and handling the world’s tacit expectations is the problem of commonsense reasoning.

It is neither one faculty nor a test of AGI

John McCarthy set out an early ambition in “Programs with Common Sense”: a program should deduce a sufficiently wide class of immediate consequences from what it is told and what it already knows. The definition remains useful because it moves the question from “Does it look intelligent?” to “Which consequences can it infer, and under what conditions?”

But “common sense” bundles different problems together. It includes physical knowledge—open containers can spill liquid—temporal regularities, causal relations, social expectations, intentions, local norms, and implicit language. Solving a pronoun question does not show that a system understands an object’s weight, and recognizing a social convention does not establish that it can plan around an exception. Nor does either result, by itself, make a model artificial general intelligence. This is a family of capabilities that must be separated before it can be measured.

Tacit knowledge: what a request leaves unsaid

Human instructions omit almost everything we take for granted. “Put the groceries in the fridge” does not specify that sealed cans may go in a cupboard, ice cream belongs in the freezer, or a broken bottle should not be handled like an intact one. A useful system has to supply some of this context without fabricating decisive details.

One approach is to represent relations explicitly. ConceptNet 5.5, for example, connects words and phrases with labeled edges and combines knowledge from expert resources, crowdsourcing, and games. Such a graph may supply the fact that an object serves a certain purpose or is typically found in a certain place. Storing a relation, however, does not settle when to apply it, which exception defeats it, or whether it remains valid in another language and culture.

Default reasoning must allow revision

Common sense works with provisional conclusions. If someone leaves a glass on a table, it is reasonable to assume it will remain there; if we later learn that the table was overturned, that inference must be withdrawn. This is default, or non-monotonic, reasoning: adding information can require abandoning an earlier conclusion. Classical deductive logic, by contrast, does not withdraw a valid consequence when compatible premises are added.

The distinction provides a practical test. Asking what usually happens is not enough. Introduce a relevant exception and see whether the system updates its answer, identifies the assumption that no longer holds, and preserves unrelated facts. A fluent answer that retains the first conclusion after the change reveals brittleness, even if the normal case was correct.

Language needs grounding beyond text

Text contains many regularities about the world, but reading descriptions of experience is not the same as having that experience. “Experience Grounds Language” argues that communication rests on shared physical and social experience and organizes grounding into layers that extend from interaction with the environment to culture. This distinction sharpens the reading of product claims: more language data may improve text prediction without demonstrating perception, action, or shared experience.

Grounding also changes evaluation. A purely verbal test is not enough to examine physical causality. CLEVRER uses controlled videos of collisions and descriptive, explanatory, predictive, and counterfactual questions. Its scenes are deliberately simple so that temporal and causal relations can be isolated. Passing it would not establish “complete common sense,” either, but it provides more specific evidence than a convincing conversation.

Benchmarks sample narrow slices of the problem

The Winograd Schema Challenge uses pairs of sentences in which resolving a reference requires background knowledge rather than an obvious syntactic cue. Its design made one capability visible; it did not create a universal intelligence scale. After models began scoring highly on variants of the task, the authors of WinoGrande built a larger, adversarial dataset to reduce spurious bias. That sequence teaches an enduring lesson: scores may rise because the intended capability improved, because training resembles the test, or because the dataset exposes shortcuts.

Other tests isolate different domains. CommonsenseQA created 12,247 questions from ConceptNet relations and human-authored distractors. Social IQa assembled 38,000 questions about motivations, reactions, and likely consequences in social situations. Those figures describe the size and method of the studies; they are not percentages of “common sense acquired.” A multiple-choice test also measures selecting among supplied options, not necessarily generating an explanation, acting in an environment, or recognizing uncertainty.

How to audit a commonsense claim

When a vendor says that a model “now reasons with common sense,” ask for five specifics. First, which domain was tested—physical, causal, temporal, social, or linguistic? Second, does the test require retrieving knowledge, applying a rule, or revising a conclusion after an exception? Third, are the situations genuinely new relative to training, and were shortcuts controlled? Fourth, is the output a selected answer, a checkable explanation, or an action? Fifth, what happens when information is missing or several answers are plausible?

Then look for an error matrix, not merely an average. A system can post a high mean score while failing systematically when a negation changes, a rare exception appears, or a proper name is replaced. Equivalent variants of a problem also matter: rephrase the question, reverse the roles, or replace objects with others that share the relevant properties. If the answer changes without a material reason, the evidence points to surface sensitivity rather than a stable rule.

A minimum test battery anyone can reproduce

\n

Even without access to training data, three families of paired prompts can be prepared. In the first, keep the question and change an irrelevant detail, such as a name: the conclusion should remain stable. In the second, alter the property that supports the answer—the container is no longer sealed, for example: the conclusion should change, and the system should identify the new fact. In the third, remove necessary information: a cautious answer should express uncertainty or request that information instead of confidently filling the gap.

\n

The evaluation becomes more useful when each variation is labeled in advance: which one should preserve the answer, which one should change it, and why. Score correctness, explanation, and consistency separately. This prevents a correct option supported by a contradictory explanation from receiving full credit. Repeating the pairs in a different order can also expose contamination from recent context. This is not a certificate of general intelligence; it is a small control that turns a demonstration into evidence that can be falsified.

\n

What counts as progress

A defensible advance combines at least four signals: coverage across several domains, resistance to paraphrase, coherent revision when facts change, and traceable assumptions. Combining statistical models with symbolic tools, memory, or perception may help, but an architecture’s label is not evidence for those behaviors. A recurrent network, a Transformer, or a neuro-symbolic system does not possess common sense by definition.

The reader’s transferable skill is precise: break every “common sense” claim into domain, assumption, exception, grounding, and test. This audit does not settle the AGI research program, but it prevents a striking demonstration from being mistaken for a general competence and turns a vague label into questions that can actually be checked.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close