Who exactly is running all these scrapers? There are, what, maybe 15 major AI labs, if that?
And none of them are smart enough to realize they could just `git clone` all the content and use it offline?
And none of them are smart enough to realize they could just `git clone` all the content and use it offline?
That is super interesting, thank you!
> At that price point, its actually very affordable to many thousands of organizations to get their own copy.
I'm still confused as to who is actually doing it though! Maybe it's affordable to scrape and store, but training a competitive AI model is going to cost much more, right?