Where does it process its speech and with what?
You can then use any language to create skills to make it perform actions.
Take a look at how to build a connected speaker with Snips: https://medium.com/snips-ai/building-a-voice-controlled-home...
Looks like it can all be done offline + OSS except for Speech-To-Text. But I think the key idea is even if it's using cloud for processing, you can at least trust it isn't listening when it shouldn't be (as long as always-on, wake word detection is offline).