Venture investor Alana Levin, writing on X with co-author Caleb Shack, argues that the agent economy’s real bottleneck sits past the model itself: it is the missing machinery for judging an agent’s ongoing output. That claim matters because most enterprise agent budgets today buy capability, a model that can complete the task, without buying the verification layer that would tell anyone whether the job was actually done well. Levin names three specific places where that verification machinery does not yet exist: tacit standards, self-compounding feedback, and liability when an agent’s output goes wrong.
Levin’s framework borrows hiring vocabulary, splitting an agent deployment into three stages:
- Screening: testing whether the agent can do the task.
- Onboarding: giving it access to internal tools and company-specific context.
- Performance review: judging whether the output is genuinely good and feeding correction back in.
Two of the three already have working infrastructure, she writes. Offline evals cover screening. Platforms such as Glean and Cognee cover onboarding by giving agents a shared substrate of company knowledge. Performance review, the stage that decides whether an agent’s output can be trusted, is the one nobody has built.
The first hard constraint is defining “good.” Early automation projects picked tasks with a binary pass or fail: did the job complete, yes or no. Levin argues the frontier is moving toward work judged against “a company’s tacit standards”, her term for the unwritten, experience-based judgment that lives inside an organization and was never codified into a rulebook. A support ticket resolved correctly is easy to grade. A strategy memo written in a founder’s voice, weighing tradeoffs the way that founder would, is not. Building an eval for that requires writing down judgment nobody has ever had to document, including for the humans who currently do the job.
The second constraint is how feedback compounds. Levin describes today’s online evals as a loop where a human records whether the agent got it right, grades the attempt, and files that verdict in a separate database somebody has to maintain. The agent itself does not retain it. A genuinely self-improving system would need feedback to accumulate inside the agent and compound across the organization without a human re-entering it each time. That raises a question Levin says has no settled answer. Corrections tuned to one company’s standards are valuable training signal, and nobody has agreed whether the customer or the vendor owns them. She points to Palantir’s term for this, “Sovereign AI,” as one framing already in circulation among operators.
The third constraint is liability. If an agent inside a company acts on faulty judgment, permissions, access controls, and legal exposure sit upstream of any technical fix. Levin credits her firm’s in-house counsel with keeping that question central to how she evaluates founders building autonomous systems.
AI Insiders covered OpenAI’s own enterprise research today, which found a widening gap between heavy and typical users of its products inside companies. Read together, the two pieces make the same point from opposite ends: consumption is not verification, and neither is capability alone.
The operator question worth asking before expanding any agent deployment is not whether the agent can do more of the job. It is who signs off that the work met the standard, and whether that judgment gets recorded anywhere the organization can learn from later. A company that buys agent capacity without building that verification layer has bought output it cannot grade.
Alana Levin, writing on X with Caleb Shack, published this analysis on August 12, 2026.