The biggest technical hurdle to sharing the work among interested parties is the web only authenticates the pipe, not the content.
The biggest technical hurdle to sharing the work among interested parties is the web only authenticates the pipe, not the content.
"Our goal is to democratize the data so that everyone, not just big companies, can do high-quality research and analysis."
Because they share it openly including with those doing AI, they wind up on "AI crawler" lists, which are increasingly used by blocking tools that just "use the AI list", by people who don't like AI, or, quite ironically, people who are trying to prevent the excess traffic that poorly mannered AI crawlers cause. (Common Crawl's crawler is well mannered, uses good user-agent, respects robots.txt including crawl-delay, etc)
No, it's really not, as most of the people who actually spend the time and effort to produce that content did not consent to it being used to train AI.
> copyright & capitalism
That's a really disingenuous way to say "the creators of that data didn't consent to training or commercial use and I want to steal their effort".
To clarify: the creators of the majority of online content haven't consented to their content being used to build AI models for any company or organization. For US-based "creators", that includes both domestic companies like Anthropic, OpenAI, Google, and foreign companies like ByteDance.