A cloud computing company called Modal, working with database researchers at Carnegie Mellon University, built a system that processes more than a billion AI-generated tokens per minute on a single Nvidia H100 chip, more than ten times the speed of the open source software most companies use today. Modal published the results on its blog on September 24, 2026.
The system, named Quail, targets a narrow but growing job: using AI models to answer fuzzy questions inside a company’s database, the kind of query a data analyst would ask about, say, which drug side effects show up alongside heart problems. That is different from asking a chatbot to write SQL in plain English. Here the SQL itself calls the AI model, feeding it rows of data and asking it to filter, join, or classify them, a pattern called AI-SQL that Snowflake, Databricks, and Google’s BigQuery have each shipped their own version of.
Modal says Quail runs 1.84 times faster than vLLM, the inference software most AI infrastructure teams default to, averaged across the company’s own new benchmark for these database-style queries. On one especially demanding query involving several table joins, the gap widens to more than tenfold, and Modal says that translates to under six cents per billion tokens processed on its cloud.
Those are Modal’s own benchmark numbers, run on its own infrastructure, and the company has not published third-party verification.
The speed comes from treating the query planner and the inference engine as one system rather than two. A standard chatbot or coding-agent workload is unpredictable: the software has no idea what a user will type next, so it cannot plan ahead for the AI model’s memory cache (the technical term is a KV cache, the working memory a model builds up as it processes text). A database query is different. The full set of requests the query will generate is known in advance, so Quail’s planner can order and evict cache entries the way a chess player thinks several moves ahead, instead of reacting one message at a time. The team also built on an existing memory-sharing method (Hydragen-style cascade attention), and stripped out generation steps that database filtering does not need, since a yes-or-no filter check only requires the model to produce a single token, not a full written response.
This is not a story about a smarter model. It is a story about the cost of using the models already available, applied to a workload that so far has drawn less attention than chatbots and coding agents. Every large company running AI over its own data, insurance claims, medical records, customer complaints, is paying a GPU bill for exactly this kind of filtering and joining today, mostly on general-purpose inference software not built for it. Quail’s code and documentation are public, so any team running these workloads on Modal or elsewhere can test the claims directly rather than take the benchmark on faith.
For any team spending meaningful GPU budget on AI-driven data filtering or classification inside a database pipeline, Quail’s public benchmark is now the number to beat before renewing an inference contract built on general-purpose serving software.
Reported by Modal on its company blog on 24 September 2026.