Path-breaking Papers About Image Classification
blog.paralleldots.com
blog.paralleldots.com
> By Moore’s law, we will reach computing power of human brain by 2025 and all of the humanity by 2050.
Their graph does show exponential growth, but the data points cut off at the year 2000. Not surprising, given that Moore's law has reached its end in the last decade. ML improvements now depend upon better algorithms to make them more parallel, and the economies of scale which make more parallel computation units available. I don't think we're anywhere near that exponential graph, however, and we'll keep getting further from it.
Perhaps quantum computing will become a widespread reality and blow the field open, but I'm not holding my breath that it will happen in the next few decades.
Considering machine learning is all on GPU's and TPU's now, I think this is still a fair assessment.
What we are seeing is an increase in die sizes; more parallel cores. Parallel cores still require parallel algorithms, so I stand by my earlier statement.
This is partly why we see slowing scaling in performance and why you cannot just throw cores at the problem.
Think of a transistor as having two regions, separated by a channel. When the transistor is on, charge carriers flow through the channel between the regions. When off, charge carriers do not flow. But when we move to smaller and smaller length scales, the channel is so small that charge carriers will tunnel through and reach the other region. How will you distinguish on/off behavior now?
("3d" transistors built on their side don't really count)
And for storage the main issue isn't how big it is, but how cheap it is to manufacture.
GPUs and multicore CPUs do not perform the same work as the single core CPUs that dominated before ~2003, when Moore's Law first slowed. By expanding the measure of speedup to include GPU/multicore, arguments like yours require not only a change in hardware but in benchmark code as well.
Longstanding general app benchmarks like SPEC emphasized everyday tasks that rarely benefit from parallelism, appropriately revealing the general ineffectiveness of adding GPUs or multicore to everyday apps. The only fair way to assess the impact of GPU/multicore is to continue using the same benchmark code as when mono-core CPUs reigned. When you do that, the value of adding GPU/multicore essentially disappears and the speedup of Moore's Law duly fades (again ~2003).
Thus until users begin to run deep learning code on their computer's GPU, AI-code won't speed up exponentially, nor continue with future GPU advancement. The hardware basis driving Kurzweil's Singularity has truly run out of steam.
I think the graph was originally produced for Ray Kurzweil's 1999 book "The Age of Spiritual Machines".
Why is this post about a specific machine-learning task even referencing Kurzweil, anyway?
It states "exponential decline in top 5 error rate", the decline looks more like diminishing returns to me, especially if you push the 2017 data point out to where it should be (they've omitted 2016).
It's nice that the error rate is low, but the caption appears to oversell it.
This graph reminds me of a very closely related one I saw in a talk a few years ago [1]. It was showing decline in voice recognition error rates over time, with a highlighted band for "human performance".
The speaker, Roger Moore (the academic, not the actor, and not the Moore with the law), pointed out that this line, while encouraging, hid two important points.
1) For linear improvement, exponentially more training data was needed. 2) No insight into how living beings solve the same task.
These aren't necessarily fatal flaws, but they're worth remembering.
Also, now that the top-5 error rate been brought down considerably, what is the next benchmark for the research community to beat? A new dataset, top-1 error rate on Imagenet?
And which framework would you recommend to code these in?
Google since brought up their Neural Architecture search that can automatically design network, which I think is way ahead of rest of the competitors here.
ResNet: https://github.com/KaimingHe/deep-residual-networks Wide ResNets: https://github.com/szagoruyko/wide-residual-networks ResNeXt: https://github.com/facebookresearch/ResNeXt DenseNet: https://github.com/liuzhuang13/DenseNet