2) I do not believe the "billions" number is OpenAI/Anthropic's training costs. I suspect it includes business expenses (including the big $$ to the guy who came up with "frontier model") and infrastructure, etc. That includes the data-center costs for running the cloud. And the "R&D" expenses which includes who-knows-what. And the settlement payment for data access. Etc. Why are the cost breakdowns not available to the public? Not because of thoughtfulness, altruism, care for the human race, but because of business plans.
3) "Piggybacking off OpenAI/Anthropic" - Scraping the web is "piggybacking" too, and so is buying existing data or even paying for new data. The "L" in LLM stands for "language" which is our common heritage.
But what does this have to do with anything anyway? People argue that truly Free (FOSS) LLMs couldn't be be developed because of costs, but I do not think that that is obvious. This used to be the argument against Linux and Wikipedia.