Given that there are plenty of existing, open-source crawling engines out there, I don't see how this decision is really accomplishing anything. Concretely, Apache Nutch[1] can crawl at "web scale" and is apparently the crawler used by Common Crawl.
There’s a more general issue here, which is this: who gets to crawl the web?
This, to me, is the most interesting issue raised by this article. In principle, there's no particular reason that, say, Google, has to dominate search. If somebody clever comes up with a better ranking algorithm, or some other cool innovation, they should be able to knock Google off their perch the same way Google displaced Altavista. BUT... that's only true if anybody can crawl the web in the first place... OR something like Common Crawl reaches parity with the Google's of the world, in both volume and frequency of crawled data.
The first scenario is definitely questionable. Sure, you can plug the Googlebot user agent string into your crawler, but plenty of sites are smart enough to look at other factors and will reject your requests anyway. (I know, I used to work for a company that specialized in blocking bots, crawlers, etc.)
It really is a bit of a catch-22. Site owners legitimately want to keep bad crawlers/bots from A. consuming excessive resources, and B. stealing content, from their sites. But too much of this will lock us into a search oligopoly that isn't good for anybody (except maybe Google shareholders).