Prime Intellect, the lab that trains open models and sells the tooling to improve them, has begun selling access to the system it uses to run those models for its own work. The service is called Prime Inference. It offers pay-as-you-go endpoints for variable demand and reserved capacity for steady workloads, running on the company’s GPUs across several data centers.

The pitch rests on scale the company says it has already reached. In a post on its own blog, Prime Intellect says that before launch the system pushed close to a trillion tokens a day through the company’s internal work: coding agents that run for long stretches, reinforcement learning rollouts, model evaluations, and generated training data. It also says paying customers have run large deployments on it since January. Both figures are the company’s own. No outside party has audited the volume or the customer claim, and the post names no customers.

The first public model is GLM-5.3, which went live on OpenRouter, a marketplace where developers pick among model providers, on 22 September. Prime Intellect says that endpoint is among the fastest for that model, has a near-zero rate of failed tool calls, and has had 100 percent uptime since launch. That last number covers about ten days at the time of the post. A perfect record over under two weeks is a start, not a track record, and the post does not say how the speed ranking was measured.

The engineering section is where the post gets specific. Agent traffic is heavy on repeated context: the company describes a typical turn as adding about 6,000 new tokens to a prompt of roughly 140,000. Rereading all of that every time would be wasteful, so the system keeps the processed history in memory and routes each session back to workers that already hold it. It also runs the two halves of the job, digesting a long prompt and writing the reply, on separate groups of GPUs so a big incoming prompt does not stall someone else’s answer. In the company’s own tests, that split cut the slow-tail delay between generated tokens by nearly 40 percent.

Memory is the other constraint. Prime Intellect says it shrank the stored form of the model’s attention data to fit about 50 percent more cached text on each decoding machine, and it plans to contribute the custom code to FlashInfer, an open-source library for this work. The stack combines Nvidia’s Dynamo, vLLM, Mooncake, and FlashInfer, and the company says it developed it with Nvidia and Inferact, the company that works on vLLM.

One section deserves attention from anyone building coding agents. The company says that while serving GLM-5.3 it found the system sometimes dropped calls to tools the agent had never declared and reported a normal finish, leaving the agent with nothing to do and no error to handle. Other calls arrived with missing or wrongly typed arguments. Prime Intellect says it fixed this by constraining what the model can emit and by contributing code upstream to Dynamo. These are the quiet failures that make a hosted model feel flaky, and a vendor naming them is more informative than a vendor claiming a number.

The business logic is plain. A lab that needs huge volumes of inference to generate training data can spread the fixed cost of that fleet by renting out the slack, and the OpenAI-compatible API means a customer can switch by changing an address and a key. Prime Intellect says batch and asynchronous jobs at lower prices, plus dedicated deployments for fine-tuned models, are still on the roadmap. The post lists no prices.

For teams running agents on open models, the sensible test is a short bake-off: send your own tool-calling workload to this endpoint and your current provider, and compare error rates and cost per finished task before taking any of these numbers on trust.

Reported by Prime Intellect on 2 October 2026