Thomson Reuters bets $40M on owning its AI instead of renting from OpenAI or Anthropic
OpenAI — with “Thomson,” the professional information company is launching its first in-house language model, built on Alibaba’s Qwen. The model hits top marks when it can tap into the company’s own content and tools.
Thomson Reuters spent about $40 million on staff and computing power over more than two years, according to the company. The more widely touted figure of $450,000 covers only the final training run of the current version. Even the full sum leaves out the real capital: That’s decades of content from Westlaw, Practical Law, Checkpoint, and Reuters, plus the working hours of hundreds of domain experts. The foundation is Alibaba’s open Qwen, most recently Qwen3.5-397B, the company says. Working with Imperial College, Thomson Reuters first retrained the Chinese model for safety, ethics, and political neutrality. This intermediate version is called “Snowdon,” named after the mountain in Wales.Ad Then came pre-training on the company’s own content, post-training with domain experts, and agentic reinforcement learning inside the company’s own tool environments. So far, less than 10 percent of the available content has gone into training.Ad CTO Joel Hron says the company has “changed the open source starting point like probably close to a half dozen times already.” The bigger finding is “less the individual model and more the model factory we built,” adds research chief Jonathan Schwartz. The blog post claims Thomson ranks among the best models in the world. The company’s own numbers paint a more sober picture: On Stanford LegalBench, Thomson (0.823) trails Gemini 3.1 Pro and GPT-5.5. On the Harvey Legal Agent Benchmark it sits just behind Opus 4.8. It leads on instruction following and the tough PrBench Legal. On reasoning, and especially coding, it falls off sharply. The comparison is also skewed by method, as Thomson competes with test-time scaling, while GPT-5.5 runs without a reasoning mode.Ad In the company’s in-house Deep Research benchmark with web access alone, Thomson scores 0.53 on factual accuracy, while GPT 5.4 hits 0.65. Only with access to the company’s content does Thomson edge past GPT 5.4, 0.83 to 0.82. With web access, Thomson is “within the scope of the other models, but certainly not the leader yet,” admits evaluation lead Andrew Bean.