You’ve decided to start roasting your own coffee. Not buying a fancier bag, actually roasting: green beans shipped in from a specific farm, a drum roaster taking up counter space, a probe thermometer, a logbook of first-crack times you’re now tracking like it’s a lab notebook. Your partner asks the obvious question over breakfast: the shop three blocks away sells a perfectly good bag for twelve dollars, why are you doing this? You mumble something about control over the roast profile, about beans nobody else has access to, about how you drink four cups a day so the math works out eventually. Some of that is true. Some of it is you liking the ritual. The honest version of that answer only survives if all three things are simultaneously true: control that actually changes the outcome, an ingredient nobody else can get, and enough volume to make the fixed cost worth it. Miss one and you’ve spent a Saturday and four hundred dollars making worse coffee than the shop sells.
Now, what does home-roasting coffee have to do with a $40 million line item in an enterprise AI budget? Everything, actually.
Most companies evaluating whether to build a proprietary AI model instead of renting one from OpenAI, Anthropic, or Google skip straight to the ambition and skip the audit. “We have unique data” gets used to justify a nine-figure training run without anyone checking whether that data is actually exclusive, whether anyone in the building can tell a good model output from a bad one at scale, or whether the volume of use even clears the fixed cost. The failure mode usually isn’t building the wrong model. It’s building a model at all when the underlying conditions were never there to begin with.
On August 24, 2026, Thomson Reuters launched Thomson, its first in-house large language model, after spending roughly $40 million and two years building it with Safe Sign Technologies, an AI safety research firm founded by lawyers and researchers out of Harvard and Cambridge, which TR acquired in 2024. Thomson starts from an open-source foundation, then trains on decades of Westlaw case law, Practical Law guidance, and Reuters news that no general-purpose model has access to. On the company’s own benchmarks across legal, tax, reasoning, and coding tasks, Thomson lands ahead of GPT-5.4 and Claude Sonnet 5, though it still trails Claude Opus 4.8 (TR’s own numbers, worth flagging, since nobody grades their own homework for free). It ships first inside CoCounsel Legal’s Tabular Analysis feature, where it becomes the default model, ahead of a broader rollout across TR’s legal and tax products.
Here’s the part that actually teaches the lesson. Thomson Reuters didn’t rip out its frontier-model relationships to get there. CoCounsel Legal’s next generation runs on Anthropic’s Claude Agent SDK, and the two companies expanded their partnership the same year Thomson launched. TR uses its own model where its own model wins (fiduciary-grade legal drafting grounded in content nobody else owns) and reaches for Claude where a general frontier model is simply the better tool for the job.
Building your own model isn’t a replacement decision, it’s an addition you only make where you clear all three bars, and you keep renting everywhere else.
| Build (proprietary model) | Buy (frontier API) | |
|---|---|---|
| Cost structure | High fixed cost, cheaper per request at scale | Low fixed cost, scales with usage |
| Data advantage | Wins only with data no competitor can access | Neutral, same model for everyone |
| Quality control | Requires in-house experts to grade output | Vendor owns quality, you audit outputs |
| Time to value | Months to years before it beats renting | Immediate |
| Right for | High-volume, narrow, defensible domain | Everything else |
You’re probably thinking: fine, but that’s Thomson Reuters, they had two years, a $40 million budget, and an acquired AI safety lab to work with. Most companies don’t. That’s exactly the point, and it’s why TR itself is careful about where it applies this. Less than 10% of the company’s proprietary legal content has gone into training Thomson so far, which means the current model is a floor, not a ceiling, and the company is treating “build” as a targeted bet on the narrowest slice of its business where the three conditions are strongest (Westlaw and Practical Law content), not a company-wide replacement strategy. In my experience, the leaders who get this decision wrong aren’t the ones who never consider building. They’re the ones who treat “we have data” as sufficient justification without ever asking whether anyone can objectively measure if the resulting model is actually better.
Before your next AI budget conversation, ask your team one question: which of the three conditions (data nobody else has, people who can objectively grade the output, and enough volume to amortize the cost) are we actually missing, and are we willing to say that out loud before we approve the spend?