Not like either except for the speech component. Yours can be done using the speech open src sft but the accuracy will be bad unless you buy or make a good acoustic model. Very painful.
Yes, but the acoustic model (or models if you account for male/female and different accents) can be improved with all the input that the service gets over time. The bad accuracy can be fixed by human intelligence using Amazon Mechanical Turk.