Processing 40 TB of code from ~10M projects with a server and Go for $100 (2019)
boyter.org
boyter.org
https://news.ycombinator.com/item?id=21121735 (80 comments)
Also, according to the page, only 59 million of all files are named “Makefile”.
I suspect that the file language recognition has major false positives for the Makefile language.
Anyone knows if it is possible to download similar data set for youtube and reddit? I have ideas for search engine based on it, but I don't want to write/maintain scraper scripts.