It eventually stopped making sense because of inference costs. Running something internal with 30% GPU utilization is just too cost inefficient compared to using an API. Idk how Reuters will manage to solve this fundamental problem.
It eventually stopped making sense because of inference costs. Running something internal with 30% GPU utilization is just too cost inefficient compared to using an API. Idk how Reuters will manage to solve this fundamental problem.
That's actually easy, because you can solve it through doing nothing and simply declaring that optimizing for lowest cost is not the main goal.
The fundamental-ness of that problem is entirely man-made and thus can easily be declared void as long as you have the cash to back that up.
Which might be a winning strategy in a world where everyone else is not doing that. Plus that your knowledge stays in-house, etc.
At the Thomson Reuters family of companies (technically then: Refinitiv Ltd. sold to LSEG), the first foundational model (in the sense of "trained entirely from scratch") was trained already in 2018 (i.e., pre-ChatGPT); it would even have been earlier, but the electricity wires and fuses in the rented 5 Canada Sq, Canary Wharf office had to be replaced first at the time to deal with the current needed to serve the GPUs.
Unfortunately for reuters tho, they dont really have a choice. A lot of their data moat is not necessary live data as in linkedin, and the only way they can keep that moat is by doing this. I guess that justifies any cost.
It won't match Anthropic or OpenAI, but it could be economic?
There is no such thing as "secure infra" hosted by someone else.
This might still be fine, depending on your threat model, of course, but if your weights absolutely must never leave the confines of your org, you cannot use any shared hosting provider, because they just offer legal coverage of incidents. But if your moat is your knowledge, legal doesn't matter as much as the knowledge being suddenly unmoated.