While the flashy part is the “speech synthesis”, the science breakthrough is actually better frames as an machine learning problem.
Imagine you record someone moving their hand across a canvas. The hand movement becomes the input, the drawing the output. The ML problem they solved is to reconstruct the output (or at least something that resembles it) from the input. In the study’s case, that’s efferent motor signals.
There is a long history of mapping these kind of signals in the sensorimotor homunculus to the respective muscles they control downstream (and some really cool stuff like prosthetic limbs can already be controlled with it), but speaking is a notoriously hard motor task and requires a lot of muscles to work in unison in very precise ways. When you implant these multi-electrode arrays, you get a few hundred more or less random single neurons, astrocytes, and local field potentials from the nether in between. Nonsense, noisy data. Being able to map this back to the result they produce in the body is technically as complex as astonishing!