I don't think it is discussing encoding time in the article, it says "features are extracted in chunks of 40ms". My reading is that its breaking down the speech into 40ms chunks, compressing it, and sending that.
But since the buffer size has to be 40ms then so the minimum latency is 40ms
Sure latency ends up being 40 ms but that's a function of needing to wait to send the encoded data + network headers at 6 Kbps not a function of the encoder being slow holding everything up.