The shape of the filters, the smoothing between the filters and the synthesis section, and (on some models) the patchability all create a very different result.
The reason the best analog vocoders are so expensive is because the filter for each band is much more complex than a plain old bandpass filter, with a much higher component count. Typically there's a flatter passband and a steeper slope than you'd expect.
You can do digital convolution with thousands of bins and it sounds nothing like analog vocoding. It's much cleaner, doesn't have those lovely harmonically spaced filter resonances, and creates sounds that can feel more acoustic than electronic.
Oh, it also might be of interest that the IVL algorithm isn’t FFT-based. I think their harmonizers sound better than the rest, so maybe FFT isn’t the best way to go.
Has anyone got more details?
If you look at the code of (phone/voice) codecs GSM/Speex/Opus you can see that you can estimate the spectral envelope (or the configuration of a physical tube model for the vocal tract) in time domain with linear prediction coefficients (LPC).
And it is simple, e.g. the often used Levinson-Durbin algorithm is just 22 lines of C code. It is an interesting exercise to build your own vocoder from scratch that fits in a single screen page.
Many of the code snippets I have seen (which likely have already processed your voice) are just translations of the Fortran code of the book "Linear Prediction of Speech" by Markel and Gray (1976).