* Sending all microphone input to a remote server would be really bad for battery life and data usage.
* It still works even in airplane mode.
* On ChromeOS (which isn't Android, but may share the same code) their privacy whitepaper [1] says:
If you opt-in to the feature, Chrome OS will listen for
you to say "Ok Google" and then send the audio of the
next thing you say, plus a few seconds before, to
Google. Detection of the phrase "Ok Google" is performed
locally on your computer, and the audio is only sent to
Google after it detects "Ok Google"."
[1] https://www.google.com/chrome/browser/privacy/whitepaper.htm...Those keyword identification algorithms can be implemented at the hardware level. I believe this is done for specialized devices like Echo and Kinect, not sure how phones do it.
It still works even in airplane mode.
I agree that they probably only send speech recorded after the activation words. However, I don't arguments think your arguments hold up. Speech data compresses exceptionally well and generally requires lower sampling rates and/or resolutions. So, it's feasible to compress the speech signal (possibly hardware-assisted), buffer it, and send it to the server at some regular interval or when the user wakes the phone.
Not if you transcribe it first, and upload compressed text.