I actually tried that a while back, giving 3.5-turbo a multishot prompt that consisted of distance readings for ahead, left, right and back in an array, as extracted from lidar data, then giving it movement instructions. It performed rather terribly.
You've very much correct that their spatial awareness is terrible. Something as simple as drive forward, then back, turn left, etc. works just fine and they can generally translate it to a specified message format reasonably reliably, but give them something more complex to execute, like drive a robot in a square pattern (an example answer would be go forward, turn right, go forward, turn right, etc.) they start to generate nonsense.
I also tested it with the 30B WizardLM at the time which performed almost as well in terms of message format but had even worse awareness.
Part of the problem is that the training data contains next to no examples that would teach it how 3D space works. I considered making a dataset of driving a robot around with human movement commands and then logging the aggregated sensor data and commands for fine tuning so the prompt format would be pre-learned, but I'm not entirely sure how much it would help.