Should be fun.
Edit: clarification
Should be fun.
Edit: clarification
It's about the same valuation as bun, lol.
Given that US legal system is precedent-base that... changes things.
Just to give an idea of the scale of it:
Let's say a modern SOTA LLM has 1T params and is therefore trained on 100T tokens
1000 tokens of text = 750 words of prose, which may take 15 min to 3hr to write (Gemini's estimate)
1000 tokens of code = 50-70 lines of code, which may take 15min to 5hr to write
We just want a rough estimate of the value of this, so let's say that 1000 tokens took 1hr of human labor to generate at an average wage of $50/hr
So, if 1000 tokens cost $50 of human labor, then that 100T of training data cost $5T.
So, the value of what the AI companies took from society might better be estimated in the trillions of dollars, not billions.
And of course what they are doing with all this data is building generative AI, so it's not just the value of what they took, but more importantly the future opportunity they are stealing from everyone by replacing human labor with their automaton who's profits they intend to keep for themselves.
Anthropic's actions seem performative. Others have already speculated on the likely audience(s).
As cited in a peer comment here[0]:
In June 2025, Judge William Alsup of the U.S. District
Court for the Northern District of California ruled on
summary judgment that using books without permission to
train AI was fair use if they were acquired legally, but he
denied Anthropic’s request for summary judgment related to
piracy—finding that the piracy was not fair use.[1]
Of note in the judge's finding; "the piracy was not fair use".0 - https://news.ycombinator.com/item?id=48667411
1 - https://authorsguild.org/advocacy/artificial-intelligence/wh...