Yep, this is the fundamental issue. It's a 35BA3B model and they probably finetuned it and evalled it in one bursty week on an 8xH100 rental just fine. But long term inference is always going to be easier in an API.
Unfortunately for reuters tho, they dont really have a choice. A lot of their data moat is not necessary live data as in linkedin, and the only way they can keep that moat is by doing this. I guess that justifies any cost.