Jaya Gupta, an AI investor who co-wrote the analysis with researchers Avanika Narayan and Jon Saad-Falcon, posted a lengthy argument on X on Wednesday claiming enterprises are overpaying for intelligence most of their workloads do not require. Her thesis: outside a narrow band of open-ended discovery problems, frontier models are routinely overqualified for the tasks companies actually run, while the industry’s own usage metrics reward burning more compute rather than solving problems cheaply. The claim cuts against a default assumption inside AI buying committees, that the newest, most capable model is always the safest purchase.

Gupta grants that frontier intelligence is real before arguing it is narrow. She points to an OpenAI internal model that reportedly disproved a longstanding conjecture tied to Erdős’s planar unit-distance geometry problem, work mathematician Tim Gowers said would have qualified for immediate publication in the Annals of Mathematics had a human written it. She also notes Anthropic says its Opus 4.6 model surfaced and confirmed over 500 serious software vulnerabilities. Those results matter for research-grade discovery work.

Her argument is that discovery work is a sliver of the economy. Citing Bureau of Labor Statistics figures, she notes jobs in the life, physical and social sciences make up under 1% of American employment, with an estimated 2,000 mathematicians, about 20,000 physicists, and roughly 37,000 people in computer and information research work nationally, against millions of nurses, claims adjusters and logistics coordinators whose work applies existing rules to specific situations rather than searching for new ones.

The concept she builds the argument around is a per-workload capability threshold. Below it, a model cannot do the job and more intelligence helps enormously. Above it, the constraint shifts to context, reliability and cost, and extra intelligence stops improving outcomes. As evidence that thresholds for common work are being crossed, she cites Qwen3.8-27B, an open-weight model with just 27 billion parameters, beating Claude’s Opus 4.6 Max on multiple coding and autonomous-agent benchmark suites.

Gupta also frames a structural conflict of interest. Model providers are paid on usage: how many tokens get consumed, how long sessions run, how many agents spin up. Enterprises are paid on outcomes: cost per resolved ticket, time to completion, escalation rates and customer retention. She cites OpenAI data showing enterprise reasoning-token usage grew roughly 320 times over the past year, alongside a PwC survey of 4,454 chief executives in which 56% said AI had yet to produce a meaningful financial return.

To argue routing already works in practice, not just in theory, Gupta cites her own prior research on a system called Minions, in which a large cloud model broke a complex task into pieces that smaller models running locally then carried out. A hybrid configuration kept 97.9% of the cloud model’s accuracy while cutting cost 5.7-fold, and a more aggressive setup held 87.9% accuracy at a 30.4-fold cost reduction. On that basis, she predicts routers and gateways become standard infrastructure, sorting each unit of work to the cheapest system still capable of the correct result.

Gupta’s figures are drawn entirely from her own published research and from secondhand citations of OpenAI, Anthropic, Amazon and PwC data rather than independent verification, so the specific multiples are best read as one advocate’s synthesis rather than an audited benchmark. The framing is nonetheless useful for budget-setting: for a CTO drafting next year’s inference spend, the practical move is to stop pricing AI usage in tokens per query and start pricing it in cost per successfully resolved task, then let a routing layer, not the model’s name, decide which system earns each call.

Posted by Jaya Gupta on X on August 19, 2026.