John Schulman leaves OpenAI to join Anthropic
John Schulman, an OpenAI co-founder and key figure in the development of ChatGPT, is leaving the company to work at Anthropic. His departure comes during a period of upheaval in OpenAI’s leadership and safety teams.
On August 5, 2024, John Schulman said in his own statement that he was leaving OpenAI for Anthropic to focus on alignment and hands-on technical work. The post denies that a lack of OpenAI support caused the move; any more dramatic explanation would require additional evidence rather than mind-reading.
The move carries significance beyond a change of jobs. Schulman was part of OpenAI’s founding group in 2015 and has been one of the most prominent technical voices in work on the safety and training of language models. Anthropic, the creator of the Claude family of models, is also one of OpenAI’s main rivals in the race to develop advanced AI assistants.
From reinforcement learning research to ChatGPT
Before ChatGPT popularized conversational assistants, Schulman was already known in research circles for his contributions to reinforcement learning. The technique trains a system through rewards: the model tries out actions, receives signals about which ones are preferable, and adjusts its behavior.
At OpenAI, that work proved essential to developing reinforcement learning from human feedback, known as RLHF. The method uses human evaluations so that a model does more than generate plausible text—it produces responses that are more useful and better aligned with instructions. It was one of the building blocks that made it possible for ChatGPT to go from a technical demonstration to a product the general public could readily use.
Schulman said his decision was a personal one and that he wanted to return to more hands-on technical work, particularly alignment research. At companies developing increasingly capable models, alignment is no secondary concern: it seeks to reduce unwanted behavior, improve controllability and anticipate risks before systems reach millions of users.
Anthropic strengthens its safety profile
Anthropic was founded in 2021 by former OpenAI employees, including Dario Amodei and Daniela Amodei. From the outset, it has placed model safety at the center of its commercial and scientific strategy. Its “constitutional AI” approach, for example, seeks to guide Claude’s behavior through an explicit set of principles rather than relying solely on case-by-case human judgments.
Schulman’s arrival gives Anthropic a researcher with hands-on experience in the most delicate phase of developing an assistant: post-training. This is when a base model, trained on enormous quantities of text, is adapted to follow instructions, refuse dangerous requests and sustain a useful conversation.
The competition between the two companies is no longer limited to releasing models with better results on technical benchmarks. It also extends to winning the trust of businesses and government agencies that need to know how a system responds to ambiguous instructions, sensitive data or high-risk uses.
More changes at OpenAI
The departure comes during a tumultuous summer for OpenAI. In May, Ilya Sutskever, a co-founder and chief scientist, and Jan Leike, who co-led the superalignment team with him, left the company. Leike said at the time that safety had taken a back seat to product priorities within the company, a criticism OpenAI implicitly pushed back on by defending the importance of its safety work.
It also emerged Monday that Greg Brockman, the president and another co-founder, will take an extended leave of absence through the end of the year. Brockman has not announced that he is leaving OpenAI, but his temporary absence adds to a period of internal transition following the company’s governance crisis last November.
OpenAI retains a dominant position in the AI assistant market, backed by its partnership with Microsoft and ChatGPT’s enormous adoption. However, the departure of researchers who helped build its technical foundations shows that the fight for specialized safety talent is as intense as the competition for customers, computing capacity and more powerful models.
For Anthropic, hiring Schulman strengthens its credibility in an area it has made a defining part of its identity. For OpenAI, the challenge will be to show that it can retain that critical expertise while accelerating the development of new products and keeping the risks posed by its systems under control.
A departure does not establish the state of a programme
The personal statement establishes destination and stated motives, but not the budget, staff or authority of the teams left behind. To assess alignment changes at OpenAI or Anthropic, follow owners, publications, evaluations, vacancies and release processes. Counting resignations without measuring functions turns people into substitutes for institutional evidence.
Alignment is not one property either. The InstructGPT paper co-authored by Schulman explains that preference fine-tuning seeks helpful, honest and less harmful answers, but represents the preferences of particular evaluators and retains errors. Following instructions, resisting manipulation and remaining controllable while using tools are different tests.
How to measure post-training
An evaluation set should include clear instructions, rule conflicts, hostile external data and tasks where abstention is correct. Preserve the base version to identify change and check whether improving one behaviour degrades another. An average may hide a system becoming more agreeable and less truthful.
Who defines the reward matters too. Labelers, researchers, customers and affected people may disagree. Useful documentation describes the evaluator population, instructions and disagreements rather than vaguely invoking “human values”.
The transferable skill is to separate a researcher’s biography from the programme they join. For any appointment, ask which function they hold, which method they use, which test they publish and which decision they can stop. Talent matters, but a safety policy must survive talent moving between employers.
An organisation chart must translate into controls
A safety team may research without deciding releases, advise without veto power, or own a mandatory evaluation. The three arrangements produce similar headlines and different consequences. Useful documentation identifies whom the team reports to, which artefacts it must approve and what happens when it disagrees with product or leadership.
Incentives are observable too. If recognition comes only from increasing capabilities or accelerating a release, safety depends on individual heroism. Stable budgets, published criteria, audits and an escalation path turn intent into process. No structure removes conflict, but it makes conflict visible and correctable.
Evaluations need institutional memory
Every version should preserve failure cases, mitigation changes and regressions. When someone leaves, that record lets a team repeat tests and understand why a control exists. Without it, successors inherit conclusions without evidence or repeat incidents already studied.
To follow an organisation, build a table of functions and tests rather than a scoreboard of famous names. Update it only from documents, publications or declared responsibilities. An appointment may matter; its effect appears later in methods, authority and verifiable results.
Public communication deserves the same discipline. “Working on alignment” can cover fundamental research, behavioural training, evaluation, product safety or policy. Asking for the concrete object prevents one person from standing in for an entire field. Honest follow-up may conclude that evidence of an effect is not yet sufficient; that uncertainty is better than turning a job move into a technical diagnosis. It also separates a declared priority from resources, authority and observable outcomes while preserving the date of every claim.
This article was produced with artificial intelligence under human editorial oversight.