IA 360
Current Affairs

OpenAI unveils GPT-4, a model that understands images

GPT-4 improves ChatGPT’s reasoning, delivers strong exam results and can analyze images as well as text. OpenAI is making it available today to ChatGPT Plus subscribers and is preparing API access.

5 min read AI-generated Leer en español
OpenAI unveils GPT-4, a model that understands images

On March 14, 2023, OpenAI introduced GPT-4, a model able to receive text and images and produce text. The announcement and technical report published evaluations, access conditions and limitations. “Multimodal” describes input and output types; it does not show that the system understands every image or that an academic score predicts professional work.

The launch matters for two reasons. First, it turns conversations with an AI into a more useful tool for tasks that require interpreting documents, charts or photographs. Second, it raises the bar for language models once again, just four months after ChatGPT brought the technology to the mainstream.

From chatbot to a system that can see

GPT-4 is a multimodal model: it can process more than one type of information. In its presentation, OpenAI showed how the system interprets an image and answers questions about it. For example, it can explain what a photograph contains, extract information from a chart or suggest which ingredients are missing based on an image.

It does not generate images: its output is still text. The novelty is that it can combine what it sees with a written instruction to produce a response.

This capability will not be immediately available to everyone. OpenAI is initially testing it with Be My Eyes, an app that connects blind and low-vision people with volunteers who describe their surroundings. The integration points to one of the clearest uses for conversational computer vision: describing a scene, reading labels or helping people understand an image without relying on someone else being available.

Better scores, with important caveats

OpenAI evaluated GPT-4 using standardized and professional exams. In a simulation of the U.S. bar exam, the model ranked in approximately the top 10% of participants, compared with GPT-3.5’s bottom 10%. It also reached high percentiles on the U.S. SAT and GRE. Source

These figures show a real improvement in the ability to follow instructions, handle long texts and solve problems with a familiar structure. But they do not mean the system has general human understanding or can practice a profession autonomously.

An exam measures responses within a limited format and uses relatively well-defined grading criteria. Real-world work involves incomplete information, legal consequences and the need to verify sources. GPT-4 can write a convincing explanation and still include a false detail or fabricate a reference.

OpenAI acknowledges that limitation. The company says GPT-4 is 40% more likely than GPT-3.5 to produce factually accurate responses in its internal adversarial truthfulness evaluations, but it still hallucinates: it presents incorrect claims with confidence. That is a meaningful improvement, not a guarantee of reliability. Source

More context for conversations and documents

The initial version of GPT-4 can handle up to 8,192 context tokens, a measure that includes the instruction as well as the text it receives and generates. OpenAI also offers a 32,768-token variant for cases involving much longer documents. Source

In practical terms, that increase makes it possible to analyze contracts, reports, code or long conversations without splitting them into as many fragments. For businesses, it opens the door to assistants that work with their own documentation; for users, it improves tasks such as summarizing study materials, reviewing drafts or preparing questions about a complex text.

Its usefulness will depend, however, on how the data entered into these services is protected. A model can process internal documents, but that makes it essential to review terms of use, access controls and the handling of confidential information before incorporating it into a workflow.

Available in ChatGPT Plus as the AI assistant race heats up

GPT-4 is available starting today to people who pay for ChatGPT Plus, OpenAI’s $20-a-month subscription, although usage limits apply. The company has also opened a waitlist for its API, the route that allows developers to integrate the model into their own applications. Source

Microsoft has also confirmed that the latest version of its Bing search engine already uses GPT-4. The news explains why Bing had displayed more advanced conversational and summarization capabilities than earlier models would have suggested, although it had also produced episodes of erratic responses during its first few public weeks.

OpenAI has not disclosed GPT-4’s parameter count or details about its architecture, training data or the computational cost of developing it. Parameters are internal values that the model adjusts during learning, but their number alone does not determine quality: the data and training matter just as much, as do the later mechanisms used to align responses with human instructions.

The presentation solidifies a new phase for AI assistants. Until now, the general public had discovered chatbots that could write. GPT-4 adds a layer of visual analysis and a greater ability to work with complex instructions. The immediate challenge will be determining whether that improvement holds up beyond exams and demonstrations, where a plausible answer is not enough and getting it right is essential.

Multimodal describes the interface

An image input may contain objects, small text, tables, diagrams or spatial relations. Each needs its own cases. Add cropping, low resolution and contradictory information, then require the region supporting an answer. Describing a photograph does not establish graph reading or radiograph interpretation.

At announcement time, visual capability was not generally open. That condition matters: selected demonstrations cannot measure failure rate, variation or real use. The record should separate shown capability, test access and commercial availability.

An exam is a sample with conditions

Percentiles and scores depend on exam version, instructions, answer selection and possible contamination. To transfer them into work, build real tasks with verifiable criteria and compare against the previous process. Strong academic performance may coexist with errors on local documents or tools.

The technical report acknowledges that GPT-4 still hallucinates. Relative improvement in a vendor evaluation does not make every answer true. Require sources, allow abstention and reserve human review for sensitive decisions. Reliability belongs to the system, not only the model.

Context does not mean perfect memory

A window indicates how much text fits, not how faithfully it is retrieved. Distribute facts across beginning, middle and end, introduce conflicting versions and request citations. Measure using the exact length and configuration; another context variant is another product.

Record tools and post-cutoff data too. If an application searches or calculates, do not attribute tool output to the base model. Separating weights, instructions, retrieval and execution prevents complete packages being compared as one intelligence.

The transferable skill is to turn a model announcement into a matrix of modality, access, benchmark, limit and own task. It survives the next version and prevents “better” from replacing a concrete question.

Evaluation needs negative cases

Questions the model can solve are not enough. Include unreadable images, impossible requests, incomplete documents and instructions conflicting with a rule. Score whether it recognises the limit, asks for context or invents. Useful improvement may mean better abstention even if fewer answers are produced.

Publish denominators by category and inspect a sample of errors. A high average may hide failure concentrated in another language, format or group. The matrix becomes a control only when it preserves that distribution rather than the headline figure alone.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close