Thomson Reuters just showed that a legacy company can build a competitive AI model without touching OpenAI or Anthropic’s API, provided it has the right raw material to feed it. The professional-information giant unveiled “Thomson,” a large language model built from Alibaba’s open-weight Qwen family rather than licensed from a frontier lab. Getting there cost roughly $40 million in engineering salaries and compute, spread across more than two years, per the company’s own disclosure reported by The Decoder.

That number reframes an earlier figure the company had circulated: $450,000, which only covers the compute for the latest training run itself. The bulk of the spend went into a longer pipeline. Thomson Reuters and Imperial College first adjusted Qwen3.5-397B for safety and political balance, an intermediate build nicknamed “Snowdon.” From there the company layered in training on its own archives, Westlaw, Practical Law, Checkpoint, and Reuters among them, followed by expert review and reinforcement learning inside its proprietary tools. Fewer than one in ten of the company’s eligible documents has been used for training so far, leaving substantial headroom. CTO Joel Hron frames the real output as a repeatable production pipeline rather than any single checkpoint, what he calls the “model factory.”

Independent numbers temper the marketing. On Stanford’s LegalBench, Thomson posts 0.823, behind both Gemini 3.1 Pro and GPT-5.5. On the Harvey Legal Agent Benchmark it lands just under Anthropic’s Opus 4.8. It does top the field on instruction following and on PrBench Legal, a harder test, before dropping noticeably on reasoning and code generation. Thomson also runs with test-time scaling enabled during these comparisons, while GPT-5.5 was evaluated without its reasoning mode switched on, a methodological tilt in Thomson’s favor.

The clearer evidence sits in the company’s internal Deep Research benchmark. Restricted to open web access, Thomson’s factual-accuracy score is 0.53 against GPT 5.4’s 0.65, a gap evaluation lead Andrew Bean admits keeps Thomson in the pack rather than out front. Hand both models the company’s proprietary archives and the scores converge: 0.83 for Thomson versus 0.82 for GPT 5.4. Because GPT 5.4 gains almost as much as Thomson does once it sees the same material, the proprietary content, not the custom training, appears to be doing most of the lifting.

That is the crux of the build-versus-rent decision now facing every data-rich incumbent watching OpenAI and Anthropic price frontier access. Thomson Reuters can make the economics work because it holds three things at once: content no competitor can license, a standing bench of full-time subject-matter experts to grade output, and workflows, document review chief among them, where correctness has an objective answer. Research chief Jonathan Schwartz describes the payoff as cumulative. Unlike a rented model, every expert correction becomes an asset the company keeps rather than value handed to an outside vendor.

Whether $40 million turns out to be a bargain or a costly experiment depends on what ships next. Thomson now powers the Tabular Analysis tool inside CoCounsel Legal, a narrow, repetitive task suited to a cheaper, smaller model, while the broader product still calls on multiple vendors. A lighter open-weight variant is headed to Hugging Face under noncommercial terms, and Hron says conversations with outside law firms about licensing access remain preliminary.

The takeaway for operators without Thomson Reuters’ archive or expert bench is blunt: the benchmarks show the model’s edge nearly vanished once a general-purpose competitor got the same proprietary data. Building an in-house model pays off only when a company already owns data a rival cannot buy and a way to measure whether the output is actually right. Absent both, a similar spend risks becoming a maintenance line item with no moat behind it.

Reported by The Decoder (byline Maximilian Schreiner), published August 24, 2026, drawing on Thomson Reuters’ own blog post disclosures.