I think the two stances are consistent, and the drawing of shapes actually enforces the "draw what you see".
Breaking an object down into its constituent shapes, is based on what you see, rather than the image in your mind. It turns out the mind is really bad at retaining the shape of something, even when we're looking at it.
We "know" what a bicycle looks like, and our mind often skirts over the detail. When asked to draw a bicycle from memory, people are terrible at it[1]. Drawing from memory is like a slightly different task, but when people draw "what they know is there", they look at a scene and remember the objects in it, then look down at the paper and draw them from memory. Not a memory of what the scene looked like, but what it was. Each object in the scene becomes the butchered, bike form memory version of itself.
If you're building the scene out of basic shapes, then you first have to look at what the scene looks like, rather than what it is.
There are some shortcuts and techniques which might seem to go against the purist version of "draw what you see". Knowing the ratios between the eyes, mouth, nose etc. But actually these are based on previous practice in "drawing what you see" (or received practice) and are still taking you away from drawing a "face" and instead drawing a set of shapes - ie "what you see".
[1] https://www.wired.com/2016/04/can-draw-bikes-memory-definite...