IA 360
Current Affairs

PaLM 2 was not the product: Google turned it into a layer

PaLM 2 powered Bard, Workspace and Search, but each integration was a different system. Evaluation must follow six layers to consequence.

5 min read AI-generated Leer en español
PaLM 2 was not the product: Google turned it into a layer

Google introduced PaLM 2 on May 10, 2023 and said it already powered more than 25 products and features. Bard, Workspace, Cloud and a search experiment would use it in different ways. The news was not only a new model: it was a decision to insert probabilistic capability into interfaces, data and permissions already used by millions.

That changes evaluation. A model benchmark alone cannot predict the experience in Gmail or Search. The final product adds context, data retrieval, instructions, filters, buttons, review and consequences. The transferable skill is following that chain and assigning each outcome to the correct layer.

What Google announced about the model

The official PaLM 2 introduction highlighted gains in multilingual work, reasoning and programming. Google said it had trained the model more heavily on text spanning over 100 languages, scientific material containing mathematical expressions and a large amount of publicly available source code.

The company offered four sizes, from smallest to largest: Gecko, Otter, Bison and Unicorn. It said Gecko was light enough to run on a phone even offline. The announcement did not disclose parameter counts. Names, intended uses and vendor claims were confirmable; model size, corpus and training cost could not be reconstructed.

“Trained on more than 100 languages” did not mean equal competence in all of them. Data share, script, task and criterion could vary. Multilingual evaluation must publish results by language and phenomenon—understanding, translation, factuality and mixed-language code. An average can conceal a serious drop in lower-resource languages.

From model to Bard

Google said Bard had moved to PaLM 2 and removed the waitlist, offering English access in more than 180 countries and territories. It added Japanese and Korean and announced future image features, code export and integrations with first- and third-party services.

PaLM 2 generated answers, but Bard defined the product: which instructions it received, which tools it could use, how it displayed sources, which data it retained and which action sat one click away. A vendor-reported reasoning gain did not guarantee every mathematical answer or code sample was correct.

Testing Bard requires tasks and consequences. For code: compile it, run tests, review dependencies and scan for vulnerabilities. For information: verify claims and links. For exports: show a preview and preserve the user’s decision. “Better model” is a hypothesis each workflow translates into its own metric.

Duet AI changed existing work

The Duet AI for Workspace introduction described help drafting in Gmail and Docs, generating images in Slides, organizing plans and classifying data in Sheets, and making backgrounds in Meet. Several features were in testing or announced for later.

On a blank page, an incorrect draft is visible. Inside a sheet with real data, a classification error can propagate into a decision. In email, an invented statement may reach a customer. The same text generation acquires different risk depending on data, recipient and reversibility.

The right experiment compares the full workflow with and without assistance. Measure minutes to an approved output, corrections, surviving errors, recipient satisfaction and review load. Counting words rewards volume; counting accepted deliverables measures utility.

Record provenance too. A document should distinguish human text, accepted suggestion, later editing and external source. That trail supports correction and learning. If everything merges into a file without history, neither worker nor organization can reconstruct how a datum appeared.

Search changes the contract with the source

Google presented Search Generative Experience as a Search Labs experiment for the United States, in English and by signup. Its interface would show a generated snapshot, links for deeper reading and follow-up questions. Google acknowledged known model limitations and restricted the query types where it would appear.

A conventional search engine ranks documents; a synthesis writes an answer. In the first case, users see titles and choose a source. In the second, they may accept a sentence before checking its origin. The interface, not only the model, changes the threshold of trust.

Search testing should preserve three units: synthesis accuracy, support for every claim and usefulness of links. Text can sound correct while citing pages that do not support it. It can also answer correctly while narrowing perspective by selecting one interpretation.

For sensitive topics, evaluate abstention, documented disagreement and access to originals. Time to a verifiable source matters as much as time to an answer. A synthesis that saves seconds but adds minutes of checking has moved work rather than removed it.

A six-layer chain

The first layer is model and version. The second is instructions and context. The third contains retrieval and tools. The fourth applies policy and permissions. The fifth is the interface that displays, hides or confirms. The sixth is the user and the consequence of a decision.

When a result fails, locate the layer. If retrieval missed the correct source, changing model style is not enough. If a user sent without review because “accept” dominated the interface, design participates in the failure. If one language performs worse, a global average cannot excuse it.

Every deployment needs a card: dated model, accessible data, tools, retention, controls, metrics and owner. The same PaLM 2 family could appear in Bard, a spreadsheet or a phone, but these were not the same systems. Comparing products by model name ignores most of what determines behavior.

An evaluation matrix for every layer

For the model, record accuracy, calibration, language, latency and cost on a fixed set. For context, vary length, order and distractors. For retrieval, measure whether the right document appears and supports each sentence. These are separate tests: an answer can fail while two layers work.

For tools and policy, test insufficient permissions, hostile input, limits and abstention. In the interface, observe whether readers find sources, distinguish draft from fact and understand what acceptance will do. Study users through real tasks, not only satisfaction questions.

The final row is operations: model updates, incidents, complaints, correction time and rollback. A favorable pilot does not ensure post-deployment stability. Keep metrics by segment and date so an average improvement cannot hide regressions.

This matrix prevents magical attribution. If prefilled fields save time, PaLM 2 does not deserve all the credit. If the interface hides a source, blaming only the model is insufficient. Layer-level measurement gives a concrete lever for improvement.

Adopt without buying the demonstration

First choose a narrow task with a baseline. Then test the real product, not a prepared conversation, using ordinary cases, extremes and missing data. Reviewers unaware of the version score output and harm under the same time budget.

Next calculate supervision cost and define an exit. A feature may remain a suggestion, require approval or stay disabled. Its level follows impact and reversibility, not the confidence of model prose.

PaLM 2 let Google deploy a common family across many products. Its durable lesson is that the model is one layer, not the system. Whoever maps the chain from weights to consequence can measure value, locate failures and demand evidence where a demonstration only displays an attractive answer.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close