I ran a HEAD request against them all to sum up the total file size, and it's 2.67TB total.
Here's a Datasette Lite URL that lets you explore the size metadata about those files: https://lite.datasette.io/?json=https://gist.github.com/simo...
And a SQL query that shows the breakdown across the different sources:
https://lite.datasette.io/?json=https://gist.github.com/simo...
Sizes here are in GB:
common_crawl 1341.6166818914935
c4 806.7667234372348
github 212.1786002581939
wikipedia 111.89125544670969
book 100.43162744678557
arxiv 87.35323827341199
stackexchange 74.54870238155127
Common Crawl is in there a few times - they have the following folders: common_crawl/2020-05 198 files
common_crawl/2021-04 176 files
common_crawl/2023-06 175 files
common_crawl/2022-05 157 files
common_crawl/2019-30 153 files
And then C4 as well, which is "a colossal, cleaned version of Common Crawl's web crawl corpus. It was based on Common Crawl dataset": https://paperswithcode.com/dataset/c4