What is GPT-4?
The GPT-4 technical report, published by OpenAI on 15 March 2023, states in its scope section that it “contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar.” What the vendor does say, where the circulating figures come from, and how to read any product whose maker expressly declares what it withholds.
On 15 March 2023, OpenAI published the "GPT-4 Technical Report." It runs to hundreds of pages, carries almost three hundred signatories, and contains one sentence worth reading in full before anything else about this model:
"Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar."
A technical report stating, in its own scope section, that it withholds the technical details. That is neither a contradiction nor an oversight: it is an explained decision. But it completely changes what can be asserted about this model, and it is why almost everything circulating about "how GPT-4 works inside" comes from nowhere at all.
What the vendor does say
The list is short and worth knowing exactly, because it is everything that exists at source.
It is "a large-scale, multimodal model which can accept image and text inputs and produce text outputs." It is "a Transformer-style model pre-trained to predict the next token in a document." It was trained "using both publicly available data (such as internet data) and data licensed from third-party providers." And it was then "fine-tuned using Reinforcement Learning from Human Feedback (RLHF)."
On results the abstract is equally concrete: the model "exhibits human-level performance on various professional and academic benchmarks, including passing a simulated bar exam with a score around the top 10% of test takers." And it adds a less-quoted and rather more interesting point: they developed methods that behave predictably across scales, which let them "accurately predict some aspects of GPT-4's performance based on models trained with no more than 1/1,000th the compute of GPT-4."
That is all. No parameters, no training data, no hardware, no cost.
What nobody knows, however often they read it
Here is the usable part. Every time you meet a figure for GPT-4's size — the trillions of parameters, the specific internal architecture, the number of experts, the exact training cost — that figure cannot have come from the vendor, because the vendor wrote down that it does not say.
It can come from three places, worth telling apart: a reasoned technical estimate by third parties, which is legitimate when presented as an estimate; an unconfirmed leak; or nowhere at all, which is the commonest case — somebody said it, somebody else repeated it, and by the third repetition it was travelling without quotation marks.
The capability you take away is this: when a vendor expressly states it does not disclose something, any confident claim about that thing is, at best, somebody else's estimate. And what to look for is not whether the number sounds plausible, but who measured it and how. If the answer is "it is said," there is no answer.
It reaches far past this model. It applies to a processor's performance where the tests are unpublished, to a product's formula, to a platform's ranking algorithm. A vendor's silence does not create knowledge in the press, however much it can look that way.
The finding in the report almost nobody reports
Among everything the document does reveal there is one result that went nearly unnoticed and is probably the most important they published. It sits in the abstract: they could "accurately predict some aspects of GPT-4's performance based on models trained with no more than 1/1,000th the compute of GPT-4."
Translated: they trained versions a thousand times smaller and cheaper, measured how those did, and from that calculated in advance how the large model would behave before spending the money to build it.
That sounds like a laboratory detail and it is a change of kind. For years, training a large model was a bet: you spent a fortune and saw what emerged. Being able to predict the outcome from an experiment a thousand times cheaper turns that bet into a budgetable project, with a forecast defensible to whoever signs the cheque. It is the difference between alchemy and engineering, and it explains a good deal of why large models have proliferated since: it stopped being a lottery.
With a candour worth underlining, because it is in the sentence itself: they say "some aspects," not all. Certain capabilities still appear discontinuously and a small model does not anticipate them. That "some" is the word separating the actual claim from the headline anyone would have written.
And there is a second capability, twin to the first: in technical documents, the quantifiers are the content. "Some aspects," "most benchmarks we tested," "around the top 10%." Anyone reading past those hedges is reading a different document from the one published.
Why it is called technical if it withholds the technique
Because it documents something else, and that part is substantial: what the model can do, where it fails, and which risks were measured before release. Capabilities, limitations and safety properties — as the document presents itself — rather than engineering.
The argument they give is worth understanding even if you reject it, because it will recur with every large model released. They cite two reasons: the "competitive landscape" — publishing the recipe hands it to whoever is behind — and the "safety implications" of anyone being able to rebuild a system of that power. The text adds a commitment: they declare themselves "committed to independent auditing" of their technologies and say they plan to make further technical details available to third parties who can advise them.
One may think the first reason weighs more than the second, and there are people who argue exactly that. But that is a legitimate debate about publication policy, and it is quite different from pretending the data are available.
What to do with this in practice
If you are evaluating this model for real use, the consequence is concrete: you cannot audit the system, only its behaviour. There are no weights to inspect and no training data to review, so the only available evidence is empirical — testing it against your own cases, yours, not the brochure's.
And there the report itself offers a useful hint about reading its results. Passing a bar exam in the top 10% is a real figure, verifiable in the document; what it does not say, and cannot say, is how the model will behave with your firm's filings. An exam measures what an exam measures. The distance between "performs at human level on this test" and "is useful to me for this" has to be walked yourself, and nobody walks it for you.
Where to go next, with no middlemen
The full report is public and free. Two parts repay going to directly: the half-page scope section, where the sentence quoted above sits; and limitations, where the authors themselves enumerate what the model still gets wrong. It is the section no summary carries and the most useful of the lot.
The capability you take from this is reading first what a vendor declares it will not tell you — because that gap is exactly the space where figures with no origin grow.
This article was produced with artificial intelligence under human editorial oversight.