OpenAI published two internal studies on August 13 measuring how its enterprise customers use its products, and the standout figure is a widening gap: measured per active user, firms in the top 10 percent of monthly usage now push 8.3 times the output tokens a typical firm does, against a 2.6 times gap back in January. Enterprise buyers deciding how aggressively to roll out agents will be pointed to this comparison. It is worth pausing on before accepting it as a productivity claim.
The two reports are “Enterprise Signals,” a survey of agentic activity across OpenAI’s business customers, and a companion working paper titled “How Organizations Use AI: Evidence from ChatGPT,” which tracks adoption across firms, roles, and seniority levels. OpenAI defines frontier firms as the top decile of monthly usage and typical firms as those sitting between the 45th and 55th percentiles.
Output tokens per user is a consumption metric, not a productivity metric, and OpenAI itself describes it only as “a proxy for depth of use.” A study built on that measure showing heavy users generate more output than light users is close to a tautology: it confirms that customers who buy more of a metered product consume more of it. The company is publishing research about its own customers, in the unit it charges them for.
The more useful number in the release describes what heavy users do differently, not how much they produce. As of June, OpenAI’s coding agent Codex generated 64 percent of combined Codex and ChatGPT output tokens among enterprise accounts, evidence that agentic, multi-step work is displacing single-turn assistance. Frontier firms also connect agents to company systems far more often: among weekly actives, Plugins reach 21 percent inside frontier firms against 9 percent elsewhere, and skills reach 19 percent against 3 percent. Internally, OpenAI says 95 percent of its own employees use Plugins weekly, a gap the company frames as headroom rather than a ceiling.
That capability gap is the actionable finding. Firms separating themselves from the pack are not simply issuing more prompts; they are wiring agents into CRMs, document stores, and repeatable workflows through Plugins and skills, which is what turns a chatbot into something that can complete a task rather than describe one.
Adoption of Codex is also spreading well beyond engineering. Since February, weekly active enterprise Codex users grew 108 times in legal, 41 times each in sales and recruiting, and 26 times in marketing, compared with 5 times in engineering, the function where agentic coding tools first took hold. Separately, the underlying data has junior staff out-messaging executives by 13 a week once six months have passed since adoption, the opposite of what many self-reported surveys on AI usage have found.
OpenAI’s lone customer example, Virgin Atlantic, says its engineering teams now refactor legacy code in 30 minutes rather than two weeks, and its product teams compress weeks of competitive research into hours using ChatGPT Work. Those figures come from the airline via OpenAI’s own release, with no independent benchmark or third-party audit attached, and Enterprise Signals draws its broader claims from a sample OpenAI describes as more than 10 million messages without disclosing the customer set behind it.
Operators evaluating an enterprise AI rollout in the next quarter should treat the 8.3x figure as marketing framing and focus instead on the capability gap: whether agents in their organization can already reach internal tools, data, and repeatable workflows through something like Plugins or skills, or whether usage is still confined to open-ended chat.
OpenAI published these findings in a company post titled “How enterprises put AI to work” on August 13, 2026.