I think the biggest problem is this swiping thing. Imagine Apple had released the first iPhone with a mouse plugged to it. It's a pure betrayal of the original Glass vision, where all user => machine communication goes through voice.
The other problem is that apps cannot (at least for now) change the language model, so Glass will always be in either "search" or "dictation" mode.
Wouldn't the hardest HCI problem be a direct cortical interface?