Thomson Reuters Built Its Own AI Model for $40 Million

"Start with a strong foundation, specialize it deeply for the work that matters, and you can build intelligence that is highly capable and entirely under your control."
On August 24, Thomson Reuters announced the launch of Thomson, the company's first proprietary large language model, trained on the legal, tax, and news content that makes up its core product portfolio and designed to operate under the governance standards its professional clients require.
Thomson Reuters said it spent $40 million to develop Thomson. The company built the model on an open-source foundation and trained it on content from Westlaw, Practical Law, Checkpoint, and Reuters, with a focus on legal and tax applications.
The company has used less than 10% of its total content library so far, meaning the training data available to the model has significant room to grow, according to the press release.
"Start with a strong foundation, specialize it deeply for the work that matters, and you can build intelligence that is highly capable, far more efficient and entirely under your control," said Joel Hron, Chief Technology Officer at Thomson Reuters.
Thomson Reuters said Thomson gives it greater control over the model and its development. The company plans to use Thomson across its legal and tax products, alongside models from other providers.
Thomson was designed to meet what Thomson Reuters calls Fiduciary-Grade standards, a term the company uses to describe a set of data handling commitments under which customer data is never used for training without explicit consent.
Thomson's first production deployment is Tabular Analysis within CoCounsel Legal. CoCounsel Legal will remain multi-model, deploying Thomson where it provides advantages in legal and tax-specific tasks while continuing to use other frontier models elsewhere.
CoCounsel Legal will continue to use multiple models. Thomson Reuters said it will use Thomson for tasks where the model provides advantages in legal and tax applications, while using other models for other tasks.
In an evaluation by Jonathan H. Choi of Washington University School of Law, Thomson performed comparably with ChatGPT and Claude on corporate tax questions. All three models answered the questions correctly, according to Thomson Reuters.
Thomson Reuters' own evaluations indicate the model demonstrates improvements in instruction following and navigating domain-specific content compared to its base open-source model.
A smaller version of Thomson will be available on Hugging Face for academic and non-commercial use. Thomson Reuters plans to extend the model across its broader legal and tax product portfolio beyond the initial CoCounsel deployment.
Key Takeaways
- Thomson Reuters invested $40 million to develop its proprietary AI model, Thomson.
- Train the model on specialized legal and tax content to meet professional client governance standards.
- Utilize only 10% of its total content library, indicating significant growth potential for the model.
- Build AI on an open-source foundation for greater control and efficiency in applications.
- Emphasize deep specialization to create highly capable AI tailored to specific industry needs.