IA 360
Current Affairs

Llama 2: “free” does not mean unlicensed or cost-free

Llama 2 opened weights and commercial use under a conditional license. Evaluation requires separating artifacts, rights, restrictions, costs and responsibility.

5 min read AI-generated Leer en español
Llama 2: “free” does not mean unlicensed or cost-free

Meta introduced Llama 2 on July 18, 2023 as a family of models available free of charge for research and commercial use. The joint announcement with Microsoft promised weights and starter code, local operation on Windows, and access through Azure, AWS and Hugging Face. But “free” described the license price for many users, not an absence of conditions, costs or responsibility.

The distinction mattered to any company comparing downloadable weights with an API. Downloading a model provides control over more of the system; it also makes the deploying team responsible for infrastructure, updates, testing and safeguards. The useful question was not whether Llama 2 was open in the abstract, but which artifacts were supplied, which rights the applicable text granted and how much an approved answer cost.

A family, not one assistant

The Llama 2 technical paper described pretrained models and Llama 2-Chat versions fine-tuned for dialogue, with released sizes of 7, 13 and 70 billion parameters. The base variant learned text continuation; the chat version added supervised fine-tuning and reinforcement learning from human preferences. Choosing between them changed both behavior and adaptation work.

The authors reported a pretraining corpus of two trillion tokens, 40% more tokens than Llama 1, and a context window doubled to 4,096 tokens. Those figures described construction. They did not guarantee that the model understood every document of that length or would beat another system on a particular task.

The paper said Llama 2-Chat outperformed open chat models on most of its tests and, in its human evaluations of helpfulness and safety, could be a substitute for some closed models. It also warned that human evaluation was noisy because it depended on the prompt set, subjective guidelines and individual raters. A vendor comparison is evidence worth reading, not universal certification.

Read rights as verbs

The Llama 2 Community License, dated the same July 18, granted a limited, worldwide, non-exclusive and royalty-free right to use, reproduce, distribute, copy, modify and create derivative works from the materials. That is a broad grant, but “limited” matters: the following clauses established conditions.

Anyone redistributing the materials or derivatives had to provide a copy of the agreement and preserve an attribution notice in a Notice file. Use had to comply with law and the incorporated Acceptable Use Policy. The materials or their results also could not be used to improve another large language model, except Llama 2 or its derivatives.

There was a specific enterprise condition. If the licensee’s and affiliates’ products or services exceeded 700 million monthly active users in the month before the release date, the entity had to request another license from Meta. It remained unauthorized until receiving one, and Meta could grant it at its sole discretion. This was neither an automatic fee for every large company nor a download limit; it was an exception defined by users, date and authorization.

The license supplied the materials and results “as is,” without warranties, and made the user responsible for deciding suitability and assuming risks of use or redistribution. “Free for commercial use” therefore did not mean “any use,” “no contract” or “Meta answers for the result.” A serious review turns the agreement into a verb table: run, modify, fine-tune, redistribute weights, offer a service and use outputs for training.

Five layers that must stay separate

The first layer is the weights: numerical files that enable inference but do not expose every training example. The second is code for loading, fine-tuning or serving them. It can have its own version, dependencies and license. The third is the model license; the fourth is the use policy incorporated into it; the fifth is the operating product, comprising hardware, data, filters, interface, logs and people.

A provider can open some layers and close others. Downloadable weights do not require disclosure of the exact corpus, the full annotation process or all infrastructure. Conversely, a closed API may publish a detailed system card and provide contractual assurances. The label “open” compresses different decisions and prevents comparison unless accompanied by an artifact list.

The list starts with answerable questions. Can complete weights be downloaded? Is the necessary code available? May users modify and redistribute? Are there restrictions by use, scale or training competitors? What information exists about data and evaluation? Each answer should point to a versioned document, not a marketing sentence.

Free to license, costly to operate

A royalty-free model still consumes compute. Its real bill includes GPUs or hosted service, memory, storage, transfer, availability, monitoring and engineering time. It also includes human evaluation: a cheap answer requiring extensive review can cost more per resolved case than a more expensive inference.

Measure cost per accepted output. If one hundred requests incur ten euros of infrastructure but only fifty meet the criterion, the basic cost is not ten cents per useful response but twenty, before review and correction. Apply the same calculation to every model size and alternative with the same task set.

The 70-billion-parameter variant might offer more capability than the 7-billion model on some problems, but it also required more memory and compute. Quantization could reduce resources while changing quality or speed. File size does not settle the choice; the boundary between quality, latency, cost and case requirements does.

Local control is not automatic privacy

Running the model on owned infrastructure lets an organization choose where prompts travel and how long they remain. That possibility becomes control only if the whole pipeline stays inside the chosen boundary. An interface might call external services for telemetry, filtering, search or logging; a library may download components; operators may retain traces.

Before promising privacy, a diagram should follow every datum from input to backup. Identify processors, regions, retention, permissions, encryption and deletion. Local weights remove one dependency, but they do not cleanse training data or authorize submission of personal information, secrets or third-party material.

Self-hosting also transfers patching, access control, isolation and incident response. An unauthenticated endpoint or vulnerable dependency can erase the benefit of avoiding an API. Control is a demonstrated operating capability, not a magical property of a downloaded file.

A reproducible adoption test

First freeze 50 to 100 real tasks, including normal, ambiguous and adversarial cases. For each, define what counts as correct, which error requires abstention and which harm is unacceptable. Prompts must be identical for Llama 2 and its alternative, as must the budget for context, tools and retries.

Then compare at least accuracy, hallucination, abstention, safety, latency, cost and review minutes. Reviewers should not know the model name. Report averages alongside failures by category: a system may improve general prose while worsening citations, calculations or the language that matters.

The test needs a version card: base or chat variant, size, weight hash, quantization, library, system prompt, sampling parameters, hardware and date. Without it, a silent update or configuration change can look like a model improvement. Preserve a rollback path in case a version breaks the workflow.

Finally, fill a matrix with five columns: artifact, right, restriction, cost and owner. It separates what Meta supplied from what the team must build, and makes it possible to evaluate a later license without repeating the whole argument. Llama 2 materially expanded access to capable models, but its durable lesson is more precise: judge an “open” model by reading files, license and operation separately.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close