I think we do. We see a building we've never seen before and we know it's a building because it has certain features that we use to classify it as a building. The examples aren't scarce.
I also think a good indicator of us doing it is the use of "y" and "ish" and "sort".
As for sthlm's point 2:
>2. The software can't recognize a feather if it's never seen a feather like that. It's not a sentient being.
This is Asimo in 2009: