Karpathy: “Computer Vision: we are far away” (2012)
karpathy.github.io
karpathy.github.io
> I’ve seen some arguments that all we need is lots more data from images, video, maybe text and run some clever learning algorithm: maybe a better objective function, run SGD, maybe anneal the step size, use adagrad, or slap an L1 here and there and everything will just pop out. If we only had a few more tricks up our sleeves! But to me, examples like this illustrate that we are missing many crucial pieces of the puzzle and that a central problem will be as much about obtaining the right training data in the right form to support these inferences as it will be about making them.
Really echoes his answer about how much data they gather (his answer: it's not about how much data we gather, it's about which data).