The irresponsibility is in seeing the existence of a technique to solve a more complex version of a toy problem (e.g. find the location in this photo that exactly matches pattern x), and inferring that the same technique, given more power, will exhibit super-human behaviour.
It's a reasonable claim that, in the task as described, the technique outperforms humans but that's about as exciting a claim as saying how much quicker the latest supercomputer is at arithmetic than a human. The point being that human object recognition isn't about labelling a scene with nouns, but somehow instinctively knowing the relevant objects for a situation and, if required, the appropriate situation-specific noun.
I therefore request the final sentence of the article be rewritten as "Or put another way, it may only be a matter of time, funding and motivation before your smartphone (with equivalent computing resources of the likes of Google Inc. and University of Tokyo) can ascribe one or more nouns from a set of size of order 1000 to regions of a photo, given that huge amounts of pre-processing and man hours have been dedicated to the pre-processing of that exact set of nouns to create a training set, and that you take the photo in similar lighting conditions as the training set, don't apply any filters, and that the objects referred to by the appropriate nouns is neither small nor thin, better than you.".