OpenAI says the first chip it designed itself, an inference processor called Jalapeño, delivers 1.5 to 1.9 times the performance per watt of Nvidia’s GB200 and GB300 systems at peak throughput. The claim comes from OpenAI’s own presentation at the Hot Chips conference, where Richard Ho, who runs the company’s hardware effort, showed what he described as the first numbers taken off real, fabricated chips rather than simulations. More Than Moore, a semiconductor analysis outlet, laid out the figures on 30 September in a long interview with Ho.

Inference means running a finished model to answer requests, as opposed to training it. OpenAI also claims 1.7 to 3.6 times better latency than the Nvidia parts. Its largest figure, up to 104 times better, applies only at operating points where OpenAI says Nvidia hardware struggles to reach, so it is not a general speed multiple. Per the interview, each chip carries 216 GiB of next-generation HBM4 memory, a compute die and a separate input/output chiplet. It is rated at 700 watts peak, with sustained draw measured nearer 550 watts. A full system links 2,048 accelerators.

The more unusual choice is the design philosophy. Much of the industry is splitting inference across different chips for different stages of answering a prompt. OpenAI built one balanced part for the whole job. Ho argues that a data center committed to fixed ratios of specialized hardware carries a risk if the workload mix shifts. A chip that does everything can be moved around the fleet as demand changes. He concedes the generalist approach may carry a hidden cost, but says OpenAI cannot identify one yet.

The chip’s memory layout is also a bet. Nvidia’s GPUs share one pool of memory across their cores. Jalapeño ties each bank of memory to specific cores, so data rarely has to travel. Ho says OpenAI has not found an open model it cannot fit onto this design, though he allows that very long context windows could cause trouble. The company has not yet analyzed that case.

The comparison with Nvidia deserves a close look. Every figure above is a claim from the chip’s own designers, presented in OpenAI’s own talk and relayed in an interview with one of them. No outside lab has reproduced the results. Ho says the team had the silicon back only around mid-May and sprinted to get results in time for the conference, so the data set is narrow. Further benchmarking work is unlikely, he says, because the team has moved on to production.

OpenAI also leaned on its own models to build the chip. Ho says unmodified internal models, not fine-tuned ones, found a way to save more than 13 percent of die area in one case. The conference talk added that model-driven tuning lifted one attention routine from under one percent of the hardware’s theoretical ceiling to close to 90 percent, after about 40 hours of work. Ho still does not trust the models to sign off a design. Standard chip-design tools do that, because a model that is 99.99 percent correct cannot be sent to manufacturing. He also admits his team did not fully understand some of the model’s source-code changes at first, and ran full validation to make sure nothing else broke.

Broadcom is OpenAI’s design partner, and Celestica handles board and rack integration. Ho credits Broadcom with access to chip wafers and memory that a software company would struggle to secure alone. The B0 revision, a second stepping that was planned from the start, is now in qualification. The company is moving toward volume production, though the interview gives no shipment numbers or deployment dates.

Ho says OpenAI intends to push Nvidia hard while remaining one of its biggest customers. For anyone budgeting inference capacity, the useful number is not the 104-times headline but the first production price per token, which OpenAI has not published.

Reported by More Than Moore on 30 September 2026, in an interview with OpenAI’s Richard Ho conducted by Dr. Ian Cutress.