Agree, it's not easy. Learning the basics, for example projecting a rectangle with 3d coordinates to 2d coordinates, then feeding the 2d coordinates into a NN and ask for the (depth) third dimension. Can you teach the NN a perspective transform? Can you rotate the rectangle and recognize rotation. Can you add other rectangles to the scene and detect each? Can you add color and lighting to infer more properties and get better results? Shine some more info on the problem ;)
These are like unit tests of AI (basic shapes and transforms) and I agree physical reckoning is at the top, one of the big tests that is a capstone and something beautiful to behold in nature (eg. sports). Maybe the a virtual soccer game at the end?
From my lidar experience, I wanted to reach for a model rather than deal with noisy sensor data. I want to generate the output (3d world) with my model, then the NN learns the inverse (eg. the scene graph used to generate the scene).
I enjoy thinking about this stuff, though it really makes my head spiral sometimes when I relate it to my own reality. It's easy to feel like you're losing touch.