IA 360
OpenAI Codex

GPT-4.5: OpenAI's Largest Model Scales Pre-Training

OpenAI releases GPT-4.5, its largest and most knowledgeable model yet. It promises fewer hallucinations, while its tests show strengths distinct from reasoning models.

Admin IA360 7 min read AI-generated Leer en español
GPT-4.5: OpenAI's Largest Model Scales Pre-Training

OpenAI announced GPT-4.5 on Thursday, February 27, 2025, a research preview it described as its largest, best chat, and most knowledgeable model yet. The company presents it as an advance from scaling pre-training and post-training, not as a reasoning model. The announcement does not use “Orion” or say the model lies outside the frontier; neither label is an official characterization of the release.

What GPT-4.5 Is and Who Can Use It

GPT-4.5 first rolled out to ChatGPT Pro subscribers as a research preview, and developers on all paid API tiers received access that day. OpenAI's published schedule placed Plus and Team in the following week, followed by Enterprise and Edu one week later. That was the declared timetable, not simultaneous availability across every plan.

The company is emphatic on one point: GPT-4.5 is not meant to be a drop-in replacement for GPT-4o, the general-purpose workhorse that powers most of its API and ChatGPT. It supports file and image uploads and works with ChatGPT's canvas tool, but for now it lacks capabilities that GPT-4o already has, like realistic two-way voice mode.

A preview with explicit limits

The GPT-4.5 system card defines it as a research preview intended to study strengths and limitations. Its preparedness evaluation rates CBRN and persuasion risk as medium, and cybersecurity and model autonomy as low; these are categories in OpenAI's own framework, not an external certification. The company says its pre-deployment evaluations found no significant increase in safety risk versus existing models, a conclusion bounded by its own process and tests.

The framework also sets a decision rule: only models with post-mitigation risk at medium or below may be deployed, and only models at high or below may be developed further. This is an internal governance threshold, not a probability of harm. “Low” model autonomy does not mean zero risk or guarantee every response's behaviour.

The central limit lies in the design itself: GPT-4.5 scales unsupervised learning and does not generate a chain of thought before responding as o1 and o3-mini do. That gives it a different profile, not a general inability to solve problems or evidence that it “doesn't convince.”

For post-training, OpenAI says it combined supervised fine-tuning and reinforcement learning from human feedback with new scalable techniques using data derived from smaller models. It attributes improved steerability, nuance and natural conversation to this work. The source does not publish the full recipe or training set: “better understands intent” remains a conclusion from the manufacturer's evaluations, not direct access to the system's cognition.

Where It Excels and Where It Falls Short

GPT-4.5 was built by scaling compute and data during pre-training, together with architecture, optimization, and new supervision techniques. The technical report places it in the unsupervised-learning line of GPT-3.5 and GPT-4, while the o-series scales reasoning through reinforcement learning. OpenAI treats these as complementary axes.

This time, the pattern breaks. OpenAI says the larger size has given GPT-4.5 "a deeper world knowledge" and "higher emotional intelligence." The model responds in a warmer, more natural tone, holds its own on creative tasks like writing, and, according to the company, understands human intent better.

One strength measured by OpenAI is factual reliability. On SimpleQA, a set of short questions with answers intended to be incontrovertible, the company reports higher accuracy and a lower hallucination rate for GPT-4.5 than for GPT-4o, o1, and o3-mini. This is a result on that set under the manufacturer's methodology; it does not measure the truth of every response or justify comparing GPT-4.5 here with systems absent from the primary source.

Internal coding results are mixed. The announcement appendix places GPT-4.5 at 38.0% on SWE-bench Verified, versus 30.7% for GPT-4o and 61.0% for o3-mini high. On SWE-Lancer Diamond it reports 32.6%, versus 23.3% and 10.8%, respectively. OpenAI labels these figures as best internal performance, so they compare manufacturer-selected configurations and do not rank every model for every software task.

On the academic tests OpenAI does publish, GPT-4.5 scores 71.4% on GPQA and 36.7% on AIME 2024, while o3-mini high reaches 79.7% and 87.3%. On multilingual MMMLU, the order reverses: 85.1% for GPT-4.5 and 81.1% for o3-mini high. Capability depends on the task; “largest model” does not mean first on every metric, while “without reasoning” does not mean unable to reason in the everyday sense.

The company points to things benchmarks don't capture. In one informal test, it asked GPT-4.5, GPT-4o, and o3-mini to draw a unicorn in SVG, a graphics format based on mathematical formulas and code. GPT-4.5 was the only one that produced anything resembling a unicorn. In another test, given the prompt "I'm going through a tough time after failing a test," all three models gave useful responses, but GPT-4.5's was the most socially appropriate.

Cost: an acknowledged limit

OpenAI describes GPT-4.5 as very large, compute-intensive, more expensive than GPT-4o, and not a replacement for it. The company said it would evaluate whether to keep serving it in the API long term while balancing this capability against building future models. The linked announcement contains no historical price table supporting multiples of 15 and 30; it supports only the qualitative relationship “more expensive” and uncertainty about continued service.

What pre-training does—and does not—show

OpenAI says GPT-4.5 is "at the frontier of what's possible with unsupervised learning" and that each new order of magnitude of compute brings novel capabilities. The release does not establish that scaling is flattening or that pre-training has reached a ceiling: in OpenAI's tests, the model improves some tasks and trails o3-mini high on others.

The technical distinction is more useful than a verdict. Pre-training aims to expand world knowledge, pattern recognition and intuition; the o-series adds reasoning-time compute before answering STEM or logic tasks. OpenAI presents the axes as complementary and says a more knowledgeable base can support agents that reason and use tools. That is the manufacturer's product hypothesis, not a demonstrated law covering every model.

A Stepping Stone, Not a Summit

The announcement does not publish GPT-4.5's training cost, delays or internal expectations. It does document a worldwide preview for developers on every paid tier through Chat Completions, Assistants and Batch, with function calling, Structured Outputs, streaming, system messages and image input. In ChatGPT it supported search, files, images and canvas, but not Voice Mode, video or screen sharing.

The same model name did not imply the same surface: the API accepted image input while ChatGPT restricted other audiovisual modalities. Any capability claim should name the product, interface and date, not just the model.

Evaluating the release requires four separate questions: what each benchmark measures, which modality the product supports, who had access on that date and what cost the source publishes. “Largest model” answers only the declared scale; it does not establish universal superiority, deliberate reasoning, complete availability or commercial value.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close