Andreessen Horowitz published new data on computer-use agents on August 10, showing the best model now clears 85 percent of tasks on OSWorld-Verified, a benchmark that scores an agent driving an actual Ubuntu, Windows, or macOS desktop. A year earlier the top score was 42 percent. Human testers average around 72 percent on the same benchmark, so the current leader, Claude Fable 5, now sits above the human baseline. All of those numbers trace back to a June 2026 leaderboard maintained at llm-stats.com.

That jump explains why computer-use agents moved from demo reels into production workflows over roughly the past eighteen months. One founder told the firm that models only became reliable enough to run unsupervised once Opus 4.6 shipped in February. But a16z’s own reporting makes clear the model was never really the constraint holding adoption back. It was everything wrapped around it: verification, retries, escalation paths, and the accumulated knowledge of how one specific company’s back office actually functions.

That distinction points to where this technology earns its keep. Consumer demos of computer-use agents tend to show them booking a restaurant table or filling an online cart, tasks performed on modern software that mostly already has an API and never needed a screen-clicking agent to begin with. The workflows a16z’s sources described running at scale look nothing like that: updating customer records in a CRM, signing into government portals and insurance systems, pulling records from regulatory databases, processing contracts, and closing tickets in systems like ServiceNow. None of that software was built with integration in mind. Some of it is decades old. Much of it will never get a clean API, because the vendor has no reason to build one or the system is too entrenched to replace. That missing interface is precisely what created the market: a business turns to a screen-clicking agent only once every cheaper integration path has already been closed off. Computer-use agents are, structurally, a tool for automating the parts of enterprise software the API economy skipped over.

The economics support that read. a16z estimates a computer-use agent costs $6 to $8 per hour of inference to run, with a wider range of $3 to $15 depending on how much of the workload gets cached as deterministic code instead of executed live, screenshot by screenshot, on a frontier model. That compares with roughly $10 an hour for offshore business-process outsourcing and $30 to $45 an hour fully loaded for a domestic back-office worker, a figure built from a Bureau of Labor Statistics wage benchmark for customer-service roles plus benefits and overhead. Two examples show the scale involved: a CPG data platform running 15 million to 20 million automated portal interactions monthly cut in half the engineering headcount it once needed for scraper upkeep, and a global systems integrator now runs 27 live workflows handling 1,500 to 2,100 IT tickets daily, aiming to redeploy 20 to 25 percent of the relevant staff elsewhere.

It is worth naming what a16z is here. This is a venture firm with capital already committed across the agent and computer-use category, publishing proprietary benchmark and cost figures that argue for its own thesis. The founder anecdotes come from companies inside the firm’s own network, not an independent audit, and the piece does not disclose how many of the cited operators are portfolio companies. The underlying OSWorld-Verified numbers are third-party and checkable against the leaderboard. The cost comparisons and adoption stories are not.

For operators sizing this category, the filter a16z lands on holds up regardless of who is making the argument: look for workflows with no existing API, high volume, repeatable steps, and a result a machine can confirm without a person checking. That is a narrower target than “agents that use computers.” It rules out most of the flashy consumer demos and points squarely at the unglamorous software still running enterprise back offices, the systems nobody ever built a plug for.

Andreessen Horowitz published this analysis on August 10, 2026.