let's say that they ask a single question for each session they get. they are immediately doubling the compute they need in processing and then post-processing the same session twice.
nothing trivial about it. not saying they cannot feed "their own LLM" saying it isn't trivial especially at scale.
if you do not trust me try it without the "at scale" part.