Open Letter Calls for Six-Month Pause on AI More Powerful Than GPT-4
The Future of Life Institute is calling for a six-month halt to training systems more powerful than GPT-4. Signed by Elon Musk, Yoshua Bengio and more than 1,000 others, the letter demands safety rules and independent oversight.
On March 22, 2023, the Future of Life Institute published an open letter asking laboratories to pause for at least six months the training of systems more powerful than GPT-4. Its signature list changes over time and is not a stable denominator. The text itself supports auditing the proposal: scope, duration, conditions and measures to develop during a pause.
The request comes just one week after OpenAI launched GPT-4. It does not call for an end to all artificial intelligence research, but for a temporary halt to the race to train ever-larger language models while common safety measures are agreed upon.
The request: halt systems beyond GPT-4
GPT-4 is a large language model: a system capable of generating text, summarizing documents, writing code and answering questions based on vast amounts of data. OpenAI has also shown that it can work with images, although that capability is not broadly available to all users.
The letter asks labs to suspend training models more powerful than GPT-4 for six months. If that pause cannot be guaranteed voluntarily, the signatories propose government intervention.
The text does not dispute that AI can deliver benefits in areas such as education, science and productivity. Its warning focuses on the speed of deployment: It argues that companies are caught up in a competition to launch systems with capabilities that are not yet sufficiently understood or controlled.
Choosing GPT-4 as the cutoff matters. This is not an abstract call to regulate AI, but a response to a generation of models that is beginning to move these tools from technical demonstrations into mass-market products. Microsoft has already integrated OpenAI technology into Bing, while Google is accelerating the launch of Bard.
Safety before an indefinite moratorium
The requested six months are not intended to be a period of inactivity. The letter proposes using them to develop shared safety protocols that independent experts can audit. The measures it puts forward include systems for identifying AI-generated content, traceability mechanisms, external evaluations before deployment and public bodies with oversight powers.
The underlying idea is simple: Before increasing model capabilities, we should get better at measuring their risks. Language models can produce convincing but false answers, reproduce biases in their training data or enable large-scale disinformation campaigns. They can also automate parts of professional and educational work before companies, governments and workers have established clear rules for using them.
The debate is not new. Researchers such as Bengio have been warning for years that developing highly capable systems requires more research into alignment, the field focused on ensuring that an AI acts in accordance with human goals and constraints. What has changed now is the issue’s public scale: Tools that were recently confined to technology labs have reached millions of people within a few months.
A request that will be difficult to implement
The letter carries political and symbolic weight, but it does not create a legal obligation. A global moratorium would require coordination among companies with different commercial interests and countries competing for technological capabilities. It is also difficult to define precisely what it means for a model to be “more powerful” than GPT-4: Size is not the only factor, and capabilities can improve through data, training techniques, external tools or combinations of models.
There is also an obvious tension in the list of signatories. Musk co-founded OpenAI and now runs companies with direct interests in AI, while other signatories come from organizations that develop or fund the technology. That does not invalidate the discussion about safety, but it does require separating legitimate concerns from each actor’s business incentives.
The letter puts an uncomfortable question on the table: If companies can rapidly launch increasingly capable models, who decides when they are ready for the public? For now, the answer depends largely on the companies themselves. Pressure for regulators, independent researchers and evaluation bodies to get involved has just grown.
A pause needs a verifiable unit
“More powerful than GPT-4” does not define scope by itself. Capability may grow through size, data, tools, inference time or combinations. An operating agreement would need reference tests, included activities and treatment of improvements not requiring a new training run.
Training, evaluation, deployment and safety research must also be distinguished. Stopping one phase while another continues may be coherent, but it must be stated. Without run records and audit, a voluntary promise cannot be checked.
The calendar is not the outcome
A period has value only if it produces artefacts: shared protocols, external evaluations, incident reporting and release criteria. Each needs an owner, method and publication. “Working on safety” without deliverables turns a moratorium into symbolic waiting.
The letter proposes content provenance and supervisory bodies among other measures. They do not all depend on pausing training or solve the same risk. Map each tool to its threat: a label helps with origin but does not correct a flawed automated decision.
Signatures, arguments and authority are different layers
A famous name attracts attention but does not prove a premise. Read the list beside the text and possible declared interests. Conversely, a conflict does not automatically invalidate an argument: test mechanism and evidence.
The letter creates no legal obligation. Moving from request to rule requires authority, jurisdiction, definitions, supervision and consequences. That distinction prevents a private initiative being reported as a changed legal framework.
How to compare governance proposals
Build a table with risk, action, subject, duration, evidence and compliance route. Add which research may continue and what happens at the end. Then test evasion scenarios: training elsewhere, improvement through tools or third-party weight publication.
The transferable skill is to turn an AI-risk declaration into an auditable mechanism. Asking who must do what, for how long and how it is checked enables debate without confusing rhetorical urgency with effective design.
Ability to comply should be measured first
A small laboratory and one embedded in a large cloud do not have equal visibility over compute, providers and partners. The design should specify required records and who may inspect them without exposing unnecessary secrets. If only transparent organisations are auditable, a pause may punish precisely those that report.
International coordination adds different calendars and authorities. A voluntary agreement may adopt shared definitions while governments prepare rules, but it does not replace law. Naming the vehicle—commitment, contract, standard or regulation—clarifies what consequence follows non-compliance.
Prior evaluation reduces circular debate
Before the period, fix a battery of capabilities and risks, passage criteria and result publication. During a pause, test whether controls reduce failures. At the end, the decision should refer to that evidence; if elapsed time alone is enough, there is no technical exit condition.
Record opportunity costs without making them an automatic veto. Delayed beneficial research, displacement toward non-participants and concentration are mechanism risks. Compare them with risks of continuing without controls under explicit assumptions.
Debate improves when each side can state which observation would change its position. If no imaginable evidence counts, the discussion is no longer about a testable mechanism. The table preserves legitimate disagreement without pretending one figure can settle political decisions.
This article was produced with artificial intelligence under human editorial oversight.