I'm guessing probably, it's just sending a screenshot of the screen right after the voice input finished (there are easy ways to recognize pauses) and sending that to the multimodal version of gpt-4, the one which is able to work with image data.
ChatGPT properly recognizes context of the task: that's the typical newly created document in the 3d software Blender. Since it starts out with a box, the user wanted to shape it into a sphere. ChatGPT provides him with a list of operations: change the selection mode to vertices, select them all and apply a bevel function, which in effect, will cause a lousy spherelike object to be created.
Two, because what you mentioned; there are a couple of other ways (like your primitive, but also NURBS lathe of a half circle, even the box with a lot of smoothing steps, subdivided icosahedron with smoothed faces)
The app used is Blender, a 3d modeling application.
But hey your negative and simple comment could also be true who knows.