Yes and no. Perceptually, you don't really need to model everything to get convincing speech sounds. Most of the realism actually comes from performance, and not the mathematical model.
In a way, lips and mouth are accounted for here, but in a more abstract away. The KL model approximates the vocal tract as a series of cylindrical tubes with varying diameters. Segments of the tubes actually correspond to things like the tongue and mouth somewhat. In this model there is a really neat tongue control that manipulates these segments. It's quite expressive!
This model is a 1d waveguide, so it doesn't account for things like the curvature of the tract. More modern vocal modelling techniques include implementing a 2-dimensional waveguide, which does allow for this control.