IA 360
Current Affairs

Gemini 1.0: how to read a benchmark without confusing model and product

Gemini Ultra scored 90.04% on MMLU with a protocol using up to 32 reasoning samples. A guide to comparing variants, inference, availability and the actual product.

4 min read AI-generated Leer en español
Gemini 1.0: how to read a benchmark without confusing model and product

Google unveiled Gemini 1.0 on December 6, 2023, the AI model the company had spent months teasing as its direct answer to OpenAI's GPT-4. Sundar Pichai, CEO of Google and Alphabet, and Demis Hassabis, who leads Google DeepMind, called it "our most capable and general model yet," according to the company's official announcement.

The news isn't just about performance. Google is emphasizing that Gemini is natively multimodal: unlike other systems that bolt image, audio or video modules onto a text model, Gemini was trained from the ground up to process and reason across those formats together. It's a technical distinction the company views as central to its results.

Three sizes for three audiences

Gemini comes in three versions, each built for a different scenario:

  • Gemini Ultra, the most powerful model, built for complex tasks.
  • Gemini Pro, designed to scale across a wide range of tasks.
  • Gemini Nano, built to run directly on mobile devices without relying on the cloud.

Google has already moved on Pro: starting today, it powers the English-language version of Bard, the company's conversational assistant, replacing the model Bard ran on until now. Nano, meanwhile, is already running on the Pixel 8 Pro, enabling on-device features without sending data to external servers. Ultra, the flagship model, isn't available yet — Google says it will finish safety and trust testing before opening it up to developers and large customers in early 2024, when a more advanced version of Bard, called Bard Advanced, will also launch.

What the benchmarks actually say

The precise announcement says Ultra exceeded the state of the art on 30 of 32 academic benchmarks, not that it defeated GPT-4 thirty times under one protocol. The table mixes results published by other labs with measurements Google ran through APIs in November, and each row may use a different number of examples, instructions or sampling method. Thirty summarizes a heterogeneous table; it is not a league with the same referee and equipment for every entrant.

MMLU shows why the small print matters. The original paper defines 57 subjects and a multiple-choice test of knowledge and problem solving. The Gemini technical report gives Ultra 90.04%, but that result uses as many as 32 chains of thought and a mechanism that selects a consensus answer when it clears a validation threshold. In the same table’s uniform five-shot comparison, Ultra scores 83.7% while GPT-4 is listed at a reported 86.4%.

Both numbers are legitimate when their protocols travel with them; they mislead when swapped. The 90.04% demonstrates a particular inference system, not one spontaneous answer. It also explains the claim about exceeding expert performance: the report places that human baseline at 89.8%. Anyone comparing models should copy the version, dataset, number of examples, sampling method and selection rule. Without those five fields, a decimal has no stable meaning.

A demo is not the product

An edited demonstration can show a real capability while concealing the user experience. Ask which inputs the model received, whether they were still images or video, which text instructions were added, how many attempts were recorded, how long each response took and what was accelerated. The honest test is not to call editing fraudulent. It is to avoid turning a montage of possibilities into a measurement of latency, reliability or availability.

The alternative is reproducible: save ten image, audio and text sequences; define a correct answer; run each case three times; and record accuracy, omissions, latency and cost. If a model recognizes objects but loses temporal order, its multimodality has a concrete limit. If it needs auxiliary text that viewers never saw, that text belongs to the system. A video can inspire a test; it cannot replace one.

“Native” describes training, not quality

Google says Gemini was jointly pre-trained on text, images, audio and video, then fine-tuned with multimodal data. That is more specific than joining independent models at the end, but it does not guarantee equal accuracy in every modality or reliable links between them. A test should require cross-modal operations: locating visual evidence for a textual answer, following change over time, or connecting a sentence to part of an audio clip. Recognizing each format separately is not enough.

The report itself contains a less visible warning about contamination. Its authors searched the training corpus for evaluation data, declined to report LAMBADA after finding issues, and showed that only one hundred fine-tuning steps on extracts associated with HellaSwag changed its result sharply. That transparency does not invalidate the table. It shows that a benchmark can measure memory of its material as well as generalization. The older and more public the exam, the more valuable a held-out test becomes.

For an organization, evaluation must end in a useful output. Take twenty real cases, remove personal data, define which mistakes are severe and measure the whole task: quality, time, cost, repeatability and human review. Then compare Pro with the previous system through the same interface. If MMLU improves while extraction from your documents worsens, the benchmark is not “wrong”; it measures another axis. A vendor supplies a capability hypothesis. A buyer decides whether that capability solves the work.

One final caution: “multimodal” names input types, not automatic access to all of them in every product. A model family may process video while a particular interface accepts only text. Before buying or integrating it, inspect that surface’s contract: accepted formats, maximum size, region, language, retention and price. The family name does not replace the documentation for the actual access point.

Availability: model, product and device

There was no single Gemini experience on launch day. Bard received a tuned version of Pro in English; Ultra remained in testing and was promised for 2024; Nano was already running on Pixel 8 Pro. A score from Ultra therefore cannot be transferred to the Bard that someone could open on December 6. Evaluated model, deployed model and visible product are three different objects.

The Pixel announcement names two Nano uses: summarizing recordings and suggesting replies in Gboard. It also says local execution helps keep sensitive data on the phone and enables features without a network. “On device” does not guarantee that an entire application is private or works offline; it describes the component that actually runs there. Check the function, language, hardware and data path.

A record for auditing any launch

Before accepting that one model “beats” another, build a six-row record. First: exact version and date. Second: input and output modality. Third: actual availability, not a promise. Fourth: benchmark protocol. Fifth: extra inference resources, such as multiple samples or tools. Sixth: one repeated task of your own. This prevents Ultra from being compared with Bard, a demo with an API, or a 32-attempt result with one answer.

Gemini 1.0 matters because it combines text, images, audio and video in a family spanning data centers and phones. The transferable skill is not memorizing who won on December 6. It is preserving the qualifiers attached to every claim: variant, protocol, surface and date. When those fields accompany a number, the benchmark informs. When they disappear, the number merely decorates a campaign.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close