Tyler Romero, a machine learning engineer, published a full derivation this month of the training method behind today’s reasoning models: reward a correct final answer, and let the math figure out which of the model’s choices deserve credit. His post follows a single example, the model solving “What is 17 times 24?”, from the first token it guesses to the update that makes the right path more likely next time.
The stakes are practical. Labs training models to solve math problems, write working code, or pass verifiable tests (tasks with a checkable right answer) do not have a person grading every intermediate step. They only get a pass or fail on the finished answer. Romero’s post explains the trick that makes that sparse signal enough to train on: a technique called the policy gradient, first formalized by Ronald Williams in 1992 and now the shared ancestor of PPO and GRPO, the two reinforcement learning methods most commonly used to post-train large language models.
The core idea, stripped of notation, is this: you cannot compute the exact best direction to adjust a model’s billions of parameters, because that would require checking every possible response it could generate. Instead, you sample a batch of the model’s own attempts, score each one against the final answer, and nudge the model’s weights toward whatever it did on the attempts that succeeded. Romero calls this “supervised fine-tuning on your own samples, weighted by reward,” a plain description of what is otherwise dressed up in heavy math notation.
The explainer also covers why GRPO, the method DeepSeek popularized in its DeepSeekMath paper, became the default for many open reasoning models. PPO, the algorithm OpenAI used to build ChatGPT through human feedback, requires training a second, separate network just to estimate how good an answer is likely to be before comparing it to the reward. GRPO skips that network. It generates a group of attempts at the same problem and uses the group’s own average score as the baseline, which is cheaper to run and, according to Romero’s walkthrough, mathematically equivalent in expectation.
Romero closes with a warning that matters more to practitioners than the derivation itself. The math assumes the model generating the practice attempts and the model being updated are the exact same model. That assumption breaks in production. A separate, faster system, commonly vLLM or SGLang, produces the practice attempts, and it does not always match the training model’s math: it can run at a different numerical precision, or, in an asynchronous setup, it may still be a few training updates behind, generating from an older set of the model’s parameters, the weights, than the version currently being trained. That gap between the model that generated the data and the model being trained is where reinforcement learning runs quietly destabilize, and it is a more common failure mode than a bad reward function.
For any team running reinforcement learning on top of a large language model, that mismatch between the generating system and the training weights is the first place to check before assuming the reward signal itself is broken.
Reported from Tyler Romero’s own technical blog, tylerromero.com, published in September 2026.