Here's some examples where bandwidth and latency wins with speech:
1. "Play here comes the sun" vs. opening spotify, waiting, clicking the search box, typing here comes the sun, pressing enter, waiting, scanning the page and clicking the right song.
2. "Send email to John asking him if he would like to Play golf" vs. opening Gmail, waiting, clicking compose, start typing john, click the right email, tab to subject... etc.
There are cases where keyboard and mouse input is better... e.g. editing text, graphics production and editing, etc.. But certainly not in "almost all tasks" as you say. I think speech is the 3rd big computer interface that complements the mouse and keyboard and will make computers more productive and convenient for everyone regardless if you have a disability.
Which John? Which of that John's contact points you have saved?
..and why don't you have the keyboard shortcuts for those actions committed to muscle memory by now?
Even shortcuts (which peer comments are relying upon) aren't all that fast - they require additional selection movement with the keyboard or mouse before they can be used.
Copy two words
Select line
Paste before word
etc.
Opening apps is ever simpler: "open spotify". Compare the complexity and time required to say those two words against moving your hand to the mouse, moving the mouse to a 100x100 pixel target, and clicking twice within 100ms. Even compare it against using "Cmd-Space Spotify".
It'd require a learning period, but so does - for example - teaching the concept of the mouse to someone who's only ever used a tablet.
EDIT: And I'll copy this from another of my posts - getting good voice control won't take our keyboards and mice away from us.
Vs properly enunciating "Kah-Pee f-i-l-e-1 to f-i-l-e-2"
enunciating: 'copy snake-geary-street-financial-report snake-divisadero-street-financial-report'
versus typing: 'cp gearyStreetFinancialReport divisaderoStreetFinancialReport'
If you're trying to exactly replicate something designed (and named) for text input, you're absolutely right, but I thought we were talking about hypothetical designed-for-voice systems.
'cp g-[TAB] divisaderoStreetFinancialReport'
I'd expect that to be an advantage of voice stuff; that you can go fast in new kinds of large scope contexts, maybe even whole-machine context. A system designed from the ground up could exploit that in interesting ways.
typing and shell help is always going to be faster than speaking
`c g-[TAB] g-[TAB]` then replace the couple characters at the front with 'divisadero'
there's no way you can do that faster speaking