A man who helped invent RLHF now says RLHF may be the problem. Diogo Almeida — co-author of the InstructGPT, ChatGPT and GPT-4 papers, and a core inventor of the human-feedback reinforcement-learning paradigm — broke two years of silence after leaving OpenAI and, on September 15, released his own model, Jev, with a training method called RLCD. His public criticism cuts to a specific failure: today's AI is incredible at assistance, but not yet at automation — and RLHF is precisely the training that taught models to hedge.
His argument targets a
[1][2]concrete failure mode. Ask a current model "option A or option B" and it lists five advantages, three risks, then adds "please weigh according to the actual situation." In a chat, that is decorum. In an automation pipeline, it is fatal: a chain of sub-agents waits for a decision and receives a PR disclaimer. Almeida's observation is that preference optimization penalizes visible uncertainty more than confident mistakes — models learn to phrase things well rather than to get things done. He argues for a third training objective beyond human preference and verifiable rewards: calibrated decision-making.
Jev is the productized version of this critique. According to Almeida's launch materials, Jev claims 20-200x faster inference, 40-400x lower cost, and free output tokens, positioned as "frontier composable intelligence optimized for decision-making." The caveat needs stating immediately: these are all company claims, not independently verified; "free output" also implies a business model resting on downstream inference or subscriptions. What deserves attention is not the numbers but the direction — when an engineer who helped shape ChatGPT argues the industry bet on the wrong training objective for four years, that itself is a signal worth recording.
Set against this week's industry context, the critique forms a strange complement to the slowdown debate. Amodei wants the industry to slow down, worried models are too strong. Almeida criticizes models as too smooth — weak precisely in decision-making, not in language. The intersection is a judgment rarely stated directly: the core metrics the industry optimized (human preference, instruction following, conversational fluency) are systematically misaligned with what it actually needs (reliable decisions, calibrated uncertainty, task completion). If that judgment holds, training objectives themselves need rewriting, whatever the release cadence.
The stance should be stated: this piece does not endorse Jev, nor does it assume RLHF is doomed. Almeida's critique has an obvious weakness — the paradigm he now wants to overturn is the one that made ChatGPT possible, and readers should weigh motive and position themselves. But the distinction between assistance and automation as two different capabilities is worth more than any single model: it may be the most useful sentence about AI in 2026.
A last observation about the timing. Almeida surfaced on September 15, the same week labs called for a slowdown and a federal court received an antitrust suit over that very call. All three events share a premise: the direction of the industry is being contested at the level of objectives, not just speed. The slowdown camp argues about pace; Almeida argues about what the objective function should have been all along. When the inventor of the current objective says it was the wrong objective, the debate stops being about margins and starts being about foundations.
[1][2]