Not it is not. Learning new object classes from a single image or a few images is very hard. See
http://www.sciencemag.org/content/350/6266/1332.short
Machine translation is a joke.
Put any comment on this page through Google translate to another language and back to English and see what you get.
I did a small part of yours. Hardly human-level for just a small simple sentence.
> But even if you do not make this assumption, identifying the object involves spitting the distance from the performance of the human level.