For example, say you want to loop from 1 to 10, print each number, and then print "done". The spoken-code equivalent might be something like "from i equals 1 to 10, print i, then print 'done'". When you translate that into code, it could reasonably end up as either of the two following snippets:
for(i = 0; i < 10; i++) { print(i) } print("done")
or
for(i = 0; i < 10; i++) { print(i) print("done") }
The worst part about designing a grammar for spoken language is figuring out intuitive natural language to substitute for the "silent" parts of code -- braces, semicolons, whitespace, caret position, etc. Obviously Python doesn't have those symbols specifically, but the concepts still exist -- if you start defining a function (with "create function [function_name]" in your code), what extra language do you have/need to say "okay, we're done defining this function and ready to go back to writing code wherever we were before defining this function".
Say you line up "create function foo", "set variable x 5", and "set variable y 8" voice commands. Will y get set to 8 within that function you created? If so, how do you signify the end of the function? How do you know where your caret jumps back to after finishing a function?
This is a neat project and I think it has a lot of potential for use if done correctly, but it's also insanely difficult to get something both powerful and intuitive when you're lossily translating spoken code to executable code because you're toeing a balance between saying something natural and saying literally just every character you would otherwise write.
In case it's helpful, here's my now-seven-years-old attempt and proof-of-concept in Javascript: https://github.com/drusepth/voice2code