IA 360
Current Affairs

Grok 4 and Heavy: reading a benchmark without mistaking it for reliability

xAI introduces Grok 4 with search and tools, reserving Grok 4 Heavy for a new subscription tier. Its scores are useful only when model, tools, compute and protocol are separated.

4 min read AI-generated Leer en español
Grok 4 and Heavy: reading a benchmark without mistaking it for reliability

On July 9, 2025, xAI introduced Grok 4 and Grok 4 Heavy. The main model includes native tool use, web search and X search; the Heavy variant spends more compute at answer time to consider several hypotheses in parallel. Grok 4 became available to SuperGrok and Premium+ subscribers and through the API, while Heavy was tied to a new SuperGrok Heavy tier.

The announcement is built around striking scores, but a reasoning number is neither a general intelligence label nor a reliability guarantee. Interpretation requires separating four pieces: which set was evaluated, which tools the system received, how much compute it used and who verified the result. That method remains useful when model names and leaderboard leaders change.

Model, tools and compute are different things

xAI attributes the advance to reinforcement learning at scale on its 200,000-GPU Colossus cluster. The company says it expanded verifiable training data beyond mathematics and coding and used more than an order of magnitude more compute than its earlier reasoning training. These are vendor statements about the process, not a reproducible recipe: the announcement does not disclose architecture, corpus, exact budget or enough detail to repeat the run.

Grok 4 was trained to decide when to use a code interpreter, browse the web or search X. That changes the unit under evaluation. An answer can depend on the planning model, the retrieval system, the pages available and executed code. A result “with tools” measures that full system. It should not be compared directly with one obtained without browsing or Python.

Heavy adds what xAI calls parallel test-time compute, described as considering multiple hypotheses at once. The page depicts several instances working, but does not publish a protocol in which independent agents debate, compare every result and select one under a stated rule. The careful wording is that it uses more parallel compute during inference. More attempts can increase the chance of finding a solution; they also raise time and cost and do not guarantee truth when every attempt begins from the same false premise.

What Humanity's Last Exam measures

Humanity's Last Exam, or HLE, was created as a set of closed, verifiable questions across mathematics, science, the humanities and other areas of expert knowledge. Its initial version contained 3,000 short-answer and multiple-choice questions, some multimodal. The authors sought a test difficult to solve through quick Internet retrieval and suitable for automated grading.

xAI says Grok 4 Heavy reached 50.7% on HLE's text-only subset with tools. That denominator contains several decisions: multimodal questions are excluded, Python and Internet access are allowed, and the first result is scored through pass@1. The number does not describe the bare model or the full collection. Nor does it say how long Heavy took or how much inference resource it spent per question.

HLE's value lies in delimiting verifiable expert knowledge under a protocol. Its authors did not design it to measure the entire assistant experience: a strong score does not establish instruction following, calibration, privacy, resistance to manipulation or consistency over a long conversation. Solving known-answer questions is not the same as conducting autonomous research, framing the right problem or recognizing missing evidence.

A valid comparison keeps the dataset version and conditions fixed. “No tools,” “with tools,” the full set and the text-only subset are different tracks. When two labs publish scores without identical search access, time limits and inference budgets, the difference may reflect scaffolding as much as the model. Before repeating a ranking, write one complete sentence containing the system, set, variant, tools and metric.

What ARC-AGI-2 adds — and what it does not

xAI also reports 15.9% for Grok 4 on ARC-AGI-2. The ARC-AGI-2 technical paper explains that each task presents a few input-output grid pairs. The system must infer an unstated rule and apply it to a new grid. No historical knowledge or specialist vocabulary is required; the benchmark targets rule composition, context-dependent application and symbols defined inside the problem itself.

The test was designed to resist memorization and exhaustive search. Its authors validated every task with people and found that test takers solved 66% of the tasks they attempted on average; every task was solved by at least two people in two attempts or fewer. In May 2025, published baseline models remained below 5%, a band the authors considered too close to noise for consistent interpretation.

A 15.9% result would therefore be a meaningful jump under a comparable protocol, but xAI's announcement is the source for this run. The page does not link an independent report detailing configuration, attempts, cost or output files. ARC-AGI also distinguishes public, semi-private and private sets. Without the exact set and policy, the figure should be attributed to xAI rather than stated as a universally certified property.

HLE and ARC-AGI-2 do not measure the same thing. The first requires broad expert knowledge on closed questions; the second minimizes prior knowledge and demands inference of novel visual transformations. Strength on both supplies complementary signals. It does not produce an average called “intelligence” or automatically predict contract drafting, customer service or news verification.

What the launch material did not provide

The announcement describes training, tools, demonstrations and results, but on July 9 it did not attach a model card or safety report covering bias, misuse, jailbreaks, hallucinations and mitigations. Nor did it publish domain-level error rates or confidence intervals for every comparison. That absence does not prove that no internal tests existed; it means readers could not audit them from the presented material.

The page introduces SuperGrok Heavy but does not record a monthly price for that tier. A commercial figure should therefore not anchor a headline without a locatable primary source. Even with a published price, “the most expensive plan” would require defining competitors, region, taxes, monthly versus annual billing and included features. A shelf-price comparison ages worse than the technical explanation without those denominators.

API availability also needs details that one launch sentence cannot answer: rate limits, retention, regions, latency, token pricing, administrative controls and version stability. The page mentions a 256,000-token context window and text-and-image understanding, but a large window does not guarantee equal attention to every input or faithful compliance with every instruction in a long response.

A purchasing protocol more useful than the ranking

An organization can begin with ten or twenty representative, verifiable tasks: questions grounded in internal documents, recent searches, calculations, code and cases where the correct answer is “there is not enough evidence.” It should run Grok 4 and, when available, Heavy with the same prompt, tools and time limit. Beyond accuracy, record correct citations, persuasive errors, latency, consumption and variation across repeats.

Measuring Heavy's value means comparing the increase in success with the increase in resources. Solving two more cases while multiplying time or cost may be worthwhile in high-value research and not in routine support. If several hypotheses converge on one false answer, parallelization has not created independence. Tests should include conflicting sources, ambiguity and missing data, not only clean problems with known solutions.

A separate safety track is also required: prompt attacks, sensitive data, harmful content, overconfidence and human intervention. An excellent academic benchmark does not offset a severe operational failure, just as a robust policy does not make a weak model useful. They are simultaneous requirements, not interchangeable points in an overall grade.

On July 9, Grok 4 showed a clear commitment to tools and inference compute. The durable lesson is not memorizing who led a table, but reconstructing the test behind a percentage. When set, variant, tools, metric, budget and verifier are named, a benchmark becomes evidence. When they are missing, it remains a marketing number that cannot rigorously support a real decision.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close