Common Crawl
commoncrawl.org
commoncrawl.org
What if you want to make a simple link previewer? An abstract crawler for scientific articles? Most websites are behind cloudflare which will block/captcha you, but happily whitelist only google & major social sites. Tha answer is measures that bring the web back to basics, not this over-SEOed bot infested ecosystem. FANGS succeeded in sucking out all the information of the web, but they suck at creating protocols that are interoperable (even twitter now needs its own tags!).
Incidentally, maybe the next search engine should use a push-system , where websites will ping to it whenever they have updates. If the engine has unique features, it might be actually seen as a measure to reduce the loads from bots.
Cloudflare's protection is to guard against traffic spikes and automated malicious attacks; Google and social sites are allow-listed because they're trustable entities with well-defined online presences and a vested interest (generally) in not breaking sites. "How do we do away with the need to put up an anti-automated-traffic screen plus whitelist" is in the same category of problem as "How do we change the modern web to address automated malicious attacks?"
This simply isn't true. By my reading, Cloudflare implies any automated traffic that doesn't openly advertise itself as such is bad. They certainly seem to be happy to assist you in blocking it regardless of whether it was actually malicious or actually causing traffic problems. (https://blog.cloudflare.com/super-bot-fight-mode/)
More reliable services for those willing to de-pssudonymize and declare themselves in an auditable fashion.
In theory someone could do something similar with terms, or you could first use URLs to filter the text size you download into Elastic or Solr and do your own custom search that way.
The indexes are really neat though, I highly recommend playing with them.
So setting up a new search engine that way would require going to every site and convincing them to notify you of changes. Wouldn't that be even more limited than the current Cloudflare whitelist system? At least there's some chance you can get around the whitelist system.
I mean, I get it, then sites can send to any number of indexers, but let's be honest, like you say, any new search engine has to get sites to push data out to them. That's just not. going. to. happen.
I see a lot of people subscribe to the idea of this being the feeder to alternative search engines.
I'd guess part of the problem with doing things this way is the 'crawl priority' of what the search engine thinks are the next best pages to crawl, it's totally out of their hands or at least, they'd still need to crawl on top of the Common Crawl data.
The recent UK CMA report into monopolies in online advertising estimated Google's index to be around 500-600 billion pages in size and Bing's to be 100-200 billion pages in size [0]. Of course, what you define as a 'page' is subjective given URL canonicals and page similarity.
At the very least, the common crawl gets around crawl rate limiting problems by being one massive download.
Would be interesting to know if there's an appreciable % of site owners blocking it, though going on past data (there is some data in the UK CMA about this also), it's not a huge issue.
[0] https://assets.publishing.service.gov.uk/media/5efc57ed3a6f4... (page 89)
Ask HN: What would be the fastest way to grep Common Crawl? - https://news.ycombinator.com/item?id=22214474 - Feb 2020 (7 comments)
Using Common Crawl to play Family Feud - https://news.ycombinator.com/item?id=16543851 - March 2018 (4 comments)
Web image size prediction for efficient focused image crawling - https://news.ycombinator.com/item?id=10107819 - Aug 2015 (5 comments)
102TB of New Crawl Data Available - https://news.ycombinator.com/item?id=6811754 - Nov 2013 (37 comments)
SwiftKey’s Head Data Scientist on the Value of Common Crawl’s Open Data [video] - https://news.ycombinator.com/item?id=6214874 - Aug 2013 (2 comments)
A Look Inside Our 210TB 2012 Web Corpus - https://news.ycombinator.com/item?id=6208603 - Aug 2013 (36 comments)
Blekko donates search data to Common Crawl - https://news.ycombinator.com/item?id=4933149 - Dec 2012 (36 comments)
Common Crawl - https://news.ycombinator.com/item?id=3690974 - March 2012 (5 comments)
CommonCrawl: an open repository of web crawl data that is universally accessible - https://news.ycombinator.com/item?id=3346125 - Dec 2011 (8 comments)
Tokenising the english text of 30TB common crawl - https://news.ycombinator.com/item?id=3342543 - Dec 2011 (7 comments)
Free 5 Billion Page Web Index Now Available from Common Crawl Foundation - https://news.ycombinator.com/item?id=3209690 - Nov 2011 (39 comments)
As an aside, it always jars me when a site hijacks default browser scrolling functionality. In my experience, making it as fast as possible is a _far_ better use of dev resources than figuring out how to make scrolling unique (no matter what the marketing department says).
I assume you are saying this because this site does it. What do you mean? I can't see any difference from normal scrolling functionality on there.
They don't approach the scale of what Google crawls, they state as much. Nor do they do it on the same timeline as Google. This is really nice for research or kick starting a project but this isn't a long term viable solution for alternative search engines. Between breadth, depth, timeline/speed, priority, and information captured it falls well short.
[1] https://blog.sia.tech/announcing-skynet-premium-plans-faster...
Of course Sia's Skynet are package deals right now and I guess they're currently bootstrapping the network with users. Filecoin has no operational storage yet. Storj quotes 10$/Terabyte/month [1] so that will come out expensive.
1. https://www.storj.io/blog/2019/11/announcing-pioneer-2-and-t...
In fact someone could easily setup some IPFS nodes that fetch the data from the current host if requested over IPFS.[1] This way people could access it via IPFS and provide an alternate mirror of the data.
[1] https://github.com/ipfs/go-ipfs/blob/master/docs/experimenta...
The main benefits here would be
- Even if the source is unavailable there may be other copies on IPFS which would be transparently used.
- There may be some performance benefits in rare cases.
- If you are accessing this on a bunch of machines your IPFS gateways would handle downloading the source once, then automatically using the local copy from inside your network.
The maindownside is that if Amazon is donating their resources why bother with IPFS?
Storing a dataset of this size and making it available online is not inexpensive. Amazon has generously donated their services to handle both of these tasks; it would be foolish to turn them down.
EDIT: the state of the crawls are summarized at https://commoncrawl.github.io/cc-crawl-statistics/.
https://commoncrawl.github.io/cc-crawl-statistics/plots/craw...
Amazon makes plenty of money from people using AWS to processes the data.
The Internet archive seems more like a library with exceptions applying, but common crawl seems to advertise also many other purposes that go beyond archiving publically relevant content.
Would this be possible in Europe, too? My feeling is that US legislation different here. Do you have to actively claim copyright in the US or enforce technically e.g. via DRM? Anyone can use anything without a license as long nobody finds out?
1) Whether it is legal for Common Crawl to collect this data in the first place? Presumably they don't have explicit permission from every site they have crawled?
and
2) Whether it is legal to use this data to train ML algorithms for commercial use (e.g. GPT-3)? They provide a bunch of examples of this on their site but the terms of use are kind of vague about this.
https://en.wikipedia.org/wiki/Feist_Publications,_Inc.,_v._R....
https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
Presumably this is because they lack the money to do so. Have they attempted to estimate how much it would cost per year to crawl the web as aggressively and comprehensively as Google does? I've checked their site and didn't find anything like that.
If they came up with a number, say $2 or $10 billion per year, it might actually be possible to gather enough donations to dethrone Google.
A lot of Google competitors would love to see them dethroned. And it would be a huge win for virtually everyone else too. There's no one in the world that wants Google to maintain their web search monopoly indefinitely.
https://www.google.com/search?q=%22whiskey+checker+pickle+bo...
That doesn’t work with common crawl.
Particularly their "Year in Search" entries, like: https://trends.google.com/trends/yis/2020/US/