The speech recognition is performed in the cloud, not locally, so under the current architecture they do indeed need to upload the audio capture following the wakeword.
One of the other things I used them for frequently before disabling it on all my devices and removing all of the networked mics from my house was things like "remind me to $THING at 21 tonight".