Thomson Reuters didn't build its model from scratch, and that's exactly the interesting part: it took an open-weight model and turned it, with $40 million and its proprietary legal archive, into something that competes with frontier models on the tasks its customers care about.
On August 24, 2026, Thomson Reuters announced "Thomson," its first internally developed frontier language model. The model combines the company's legal, tax, and news knowledge archive with a third-party open model base, instead of depending exclusively on an external LLM provider for the core of its AI product.
Thomson starts from an open-weight model — Qwen3.6-35B-A3B — and applies "continual training" on the company's own Westlaw, Practical Law, and tax and news content. The total reported investment was $40 million, covering both talent and compute — a fraction of what it costs to train a frontier model from scratch.
Thomson's first deployment is in Tabular Analysis, a high-volume document review capability within its CoCounsel Legal AI assistant. According to Thomson Reuters, the model trains and runs at a fraction of the cost of comparable frontier models, and stays entirely under the company's control and ownership.
If your company has a deep proprietary data archive in a specific domain (legal, tax, healthcare, engineering, whatever), this case is evidence that you no longer need to be OpenAI or Anthropic to have a competitive model of your own in your niche — but you do need the right data archive, not just budget. Before committing indefinitely to paying an external provider for tokens on your most critical use case, it's worth evaluating whether continual training on an open model, applied to your own archive, gives you better control and unit cost in the long run.
Carlos Montiel is an enterprise AI solutions architect. He implements LLMs, Agents, RAG, and orchestrators for companies across Guatemala and Latin America. Reach out for a consultation.
Contact Carlos Montiel