A company that clocks a faster single step in an AI deployment and books the gain as more overall output has likely made a forecasting error, according to researcher Abi Olvera. Olvera, speaking with Annelies Gamble in a post published on X, argues that most AI forecasts collapse the same way: they measure what a model can do to one task and assume the gain carries cleanly through the rest of the workflow.

Olvera served seven-plus years in the diplomatic corps at the State Department, working on crisis preparedness, national security, and emerging technology risk in postings that included Dakar and Cairo. She now writes the Positive Sum Substack, holds an advisory role at Golden Gate Institute, and is affiliated with the Institute for AI Policy and Strategy. Her position is that capability and impact are separate forecasting problems, and the industry keeps treating the first as a stand-in for the second. “It’s not something you can get from first principles,” she told Gamble.

Radiology is her clearest case study. Geoffrey Hinton made the point starkly in 2016, telling a Toronto AI conference, “people should stop training radiologists,” predicting deep learning would surpass them within five years. AI did take hold in reading scans: the FDA has since cleared hundreds of imaging tools built on it. Radiology staffing did not shrink, though. At Mayo Clinic, an early and heavy adopter of the technology, radiology headcount reportedly rose 55 percent since 2016, even as the department stood up a 40-person AI team running over 250 models.

Dr. Charles E. Kahn Jr., a radiology professor at the University of Pennsylvania’s Perelman School of Medicine, summed up why: “There’s been amazing progress, but these AI tools for the most part look for one thing.” Reading the image was never the whole job. Consulting with surgeons, writing reports, and weighing a scan against a patient’s history all remained. Once image interpretation sped up, the rest of the workflow needed more people to keep pace, not fewer.

That gap points to the test an operator should run before booking any AI-driven speedup as throughput: did the accelerated step sit on the workflow’s critical path, or was it running in parallel with a slower human, institutional, or physical constraint that now has to absorb more volume? A support team that halves ticket-drafting time has not halved resolution time if escalation, verification, or customer follow-up was already the binding constraint. Time saved upstream disappears into slack downstream instead of showing up as delivered output.

Olvera traces the blind spot to geography. Technologists cluster around capability because their own work runs close to end to end inside a chat window. Adoption looks different in institutional centers like Washington, she said: “People in DC are less likely to be using Claude Cowork or Codex. If people use it, a lot of times they might be using the chatbot version, which is great, but that doesn’t really unleash the parts of AI that are crazy surprising.” Neither vantage point sees the full workflow. That is why Olvera pushes for what she calls cross-framework research, pairing practitioners with technologists on the same task rather than publishing in parallel.

One example she cites: a controlled study gave novices either plain internet access or internet access paired with a frontier AI model, then tracked who could finish real wet-lab molecular biology tasks over eight weeks unsupervised. The AI group improved on some intermediate steps but was not meaningfully more likely to complete the full workflow, evidence that tacit hands-on skill, not information access, was the limiting factor.

Olvera extends the same logic to policy. Rather than writing rigid rules for a technology whose trajectory is still uncertain, she argues governments should build the capacity to notice and respond as evidence firms up: better reporting, better evaluations, more technical expertise inside institutions.

For operators, the fix is a due-diligence step, not a reason to slow deployment. Before a pilot’s task-level speedup lands in a budget line or a headcount plan, map the workflow’s actual bottleneck and confirm the AI is touching it rather than the step next to it. Olvera’s own framing applies directly: “Bottlenecks become more important when everything else gets automated.” A model that accelerates the easy majority of a job while leaving the hard remainder untouched will not move the throughput number the forecast promised.

Annelies Gamble published this analysis on X on August 11, 2026, drawing on an interview with researcher Abi Olvera.