CFD was merely used as an example of something that does scale well. I'm not sure it was the best example, since CFD isn't very common. But basically you have a volume mesh and each cell iterates on the Navier-Stokes equation. So if you have N processor cores, you break the mesh in N pieces, each of which get processed in parallel. Doubling the number of cores allows you process double the amount in the same time, minus communication loses (each section of the mesh needs to communicate the results on its boundary to its neighbors).
I don't fully understand the graph, but it looks like his point is that Alpha Go Zero uses 1e5 times as many resources than AlexNet, but does not produce anywhere near 10,000 times better results. We saw that with CFDt 1e5 more cores resulted in 1e5 better results (= scales). The assertion is that DL's results are much less than 1e5 better, hence it does not scale.
Basically the argument is:
1. CFD produces N times better results given N times more resources [this is implied, requires a knowledge of CFD]. That is, f(ax) = a f(x). Or, f(ax) = 1 a * f(x).
2. Empirically, we see that DL has used 1e5 more resources, but is not producing 1e5 times better results. [No quantitative analysis of how much better the results are is given]
3. Since DL has f(a * x) = b * a * f(x), where b < 1, DL does not scale. [Presumably b << 1 but the article did not give any specific results]
This isn't a very rigorous argument and the article left out half the argument, but it is suggestive.