Thomson Reuters has developed its first in-house language model, named 'Thomson,' designed for legal work. The model is built on Alibaba's Qwen and trained using the company's own data and domain experts. According to the company, Thomson performs best when it can access exclusive company content, such as internal tools and data. The model will initially be used for document review, with a smaller version set to be released under a non-commercial license. Source: thedecoder
Thomson Reuters invested $40 million over more than two years to develop the model, with additional costs covering decades of content from Westlaw, Practical Law, Checkpoint, and Reuters, as well as the labor of hundreds of domain experts. The model is based on Alibaba's Qwen3.5-397B, with initial retraining at Imperial College focused on safety, ethics, and political neutrality. This intermediate version was named 'Snowdon' after a mountain in Wales. Further training involved pre-training on the company's own content and post-training with domain experts, along with agentic reinforcement learning within the company's own tool environments. Source: thedecoder
The company's benchmarks show Thomson performs well in instruction following and the PrBench Legal test but lags behind models like GPT-5.5 and Opus 4.8 in reasoning and coding tasks. On the Stanford LegalBench, Thomson scored 0.823, trailing Gemini 3.1 Pro and GPT-5.5. With web access alone, Thomson scores 0.53 on factual accuracy, while GPT 5.4 scores 0.65. However, Thomson edges past GPT 5.4 when it can access internal content, achieving a score of 0.83. Source: thedecoder