Ironically, I think only technical people would even want to do something like that. The less technical you get, the more high level (and ambitious) you need to go.
You can see this a lot in game dev questions. Beginner questions will be "How do I make an MMORPG?" and the more advanced questions will be "how do I return x from y" or whatever, and then it scales between the two ends of the spectrum.
As a more simple problem space, building programs from UML charts was one of Java's promise, and it failed miserably, not because the technology was lacking, but because it's just a damn hard problem.
As of now ee have nothing approaching "non technical people to be able to instantly create things" if the "things" you want are useful in any way.
You have to read code and debug it that is inevitable, you can't say that there will not be any bugs if you use voice instead of writing.
I'd encourage anyone who hasn't tried it to give Copilot a try. There really has never been anything like it in my memory, and while I totally agree there have been dozens of efforts to allow non-technical people to generate code, I think Copilot may be on to something very special.
The crazy thing is... this probably will work.
In 20 years. But, it probably will work.
There is absolutely no reason you cannot use a neural network to transcribe appropriately phrased requirements into an AST.
But now, given how magical this thing is, it opens many doors to what's possible with no code.
I never really believe in anything no code before apart from Excel and RAD.
But basic tasks are going to get accessible to a lot of people sooner than I expected
That part is unlikely to change much, if it all.
Trained professionals will always outperform the latest fellow dragged off the street.
You free part of the industry from needing professionals for some simple tasks, which is enough to empower users with a lot of possibilities, and focus pros on where they are really needed.
"Hey phone, next time mum send me a text about voting, you can send back a 'ok boomer'?" or "hey phone, can you setup a webpage that list my tiktok videos up to last year?"
Most people are not going to hire a pro for that, but we could end up with a AI general enough to be able to do that for users. It's ok if the result is not extensible, maintainable or modular.
1. How do you check the output of the voice to code step? If you need as much expertise as you do now to actually review the code, then the voice to code step is just a layer that adds confusion
2. How would debugging work? Again, would you still need to be able to understand the code? Same issue.
3. What if you have to pause and think? This will affect how the voice to code interface interprets your speech.
4. How would you make a precise edit to your source audio using a voice interface?
5. How would you make changes which touch multiple components across the project? How would you coordinate this?
6. Precisely defining interfaces between components and using correct references to specific symbols is very difficult to do in natural speech, which typically uses context to resolve ambiguous references. The language you would be using would still have to resemble the strictness of a programming language even when spoken, but you have replaced a reliable checkable channel (input through keyboard, transfer as-as to text buffer, feedback from visual view of source) with an unreliable channel (input through microphone, transfer through complex signal processing and multiple neural network language models, through multiple representations, where you have to be able to check multiple representations for feedback about the structure of your program (initial speech-to-text step, text to source))
Voice to code into Scratch, plus 20 years of hard product development effort.
I'm not saying it'll be easy or work the first time. The Apple Newton failed miserably the first time.
I'm saying touch screens still worked out amazingly well two decades later.
If we can say anything about programming today, it's that programming today with 20 years of advancements is remarkably different to 20 years ago.
I fully expect the next 20 years to be similarly fruitful.
Before by switching from binary to assembly to higher level languages to frameworks/libraries, you generally reduce amount of "code" being written after each step, with voice programming this seems to be the opposite.