I always wondered about doing this with hardware. Surely there must be some other mechanism for accurately creating sounds then a flat magnetic speaker.
You could stick one (or two) transducers to a guitar or piano and then mic it up, and get the sound of the body of that instrument and parasympathetic resonance of any undamped strings.
The body is quite easily modelled with an impulse response, but parasympathetic resonance in DSP is still not great
Speakers, themselves, have an absurdly broad range of nonlinearity, even within a single speaker, the way it behaves at 10% power and the way it behaves at 90% power can be drastically different. That makes them fairly difficult to model. Impulse responses are the soup de jour, and do a good job, but, they only really model a slice of what a speaker can do, ie, a given amplitude.