A widely used trick for making chatbots answer faster stops paying for itself once a server gets busy, according to engineer Joe Barrow, because the trick never saved any work in the first place. Writing on his personal site on 8 October, Barrow argues that the speedup comes from putting idle hardware to use. When no hardware is idle, the same trick takes capacity away from everyone else’s requests.

Start with how a chatbot writes. Language models produce text in small pieces called tokens, roughly a word or a chunk of one. Normally the model writes one token, feeds it back in, and writes the next, so a 300-word answer means hundreds of separate trips through the model. Each trip forces the chip to pull the model’s billions of stored numbers out of memory, and for a large model that haul takes longer than the arithmetic that follows it.

The technique is called speculative decoding. A second, much smaller “draft” model guesses the next several tokens cheaply. The large model then checks the whole guess at once, reading its numbers from memory one time instead of once per token, and keeps every guess it would have written itself. Good guesses mean several tokens for the price of one memory haul.

Barrow’s contribution is an argument about what that costs. Checking five guessed tokens takes more arithmetic than writing one, and wrong guesses get thrown away, so the total work goes up. He calls the common belief that checking is cheaper than writing a misconception. In his telling, the gain comes from the chip spending time on arithmetic that it would otherwise have spent waiting on memory.

He illustrates it with a toy example in which the chip does its best work at five tokens per pass. A server could fill those five slots with five different users, each moving ahead one token, or with one user moving ahead five. The effort is about the same. When traffic is light, the second option is free speed, because the other four slots would have sat empty.

Busy servers are different. Providers batch many users together to keep expensive chips occupied, and the slots fill with real requests. Speculation then adds work to a chip already working flat out, and every rejected guess is pure waste that slows the other users down. Barrow notes that vLLM, the open-source software many teams use to serve models, ships a setting that switches speculation off above a chosen batch size, which he offers as evidence that the limit is real.

He also points to research that varies the guess length with load: guess deep when the server is quiet, guess shallow or skip guessing when it is full. He names dSpark as an example of that idea. The five-slot figure is a teaching device from a blog post, not a measured crossover for any real chip or model, and nothing here has been peer reviewed or tested against a vendor benchmark.

The practical default in the industry is to turn speculation on and leave it there. Barrow’s argument implies that default holds only while the chip has slack, and a quiet test machine always has slack. A team paying for inference should measure its own throughput with speculation on and off at peak load, and if the advantage vanishes at the busiest hour, the setting is costing money.

Joe Barrow, “Why is Speculative Decoding Fast?”, jbarrow.ai, published 8 October 2026.