You really cannot expect audio processing to yield color information.
Beyond that, you are correct that the 3D shapes themselves cannot be derived perfectly accurately (see my other post)
Beyond that, you are correct that the 3D shapes themselves cannot be derived perfectly accurately (see my other post)