If you think about a sharp left turn with a wall on the right side of the road, there is a "stationary object" (the wall) in front of the vehicle throughout the turn that will be detected by radar and cameras. In order to navigate a situation like this, an autonomous vehicle has to have additional decision-making in place to effectively override "stationary object in front of me = stop" logic. But once you do that, it gets tricky to determine when you should ignore the "stationary object" and when you shouldn't.
On top of this, Tesla opted not to include LIDAR as a sensor in its vehicles, so it does not get a 3D point cloud to help identify what an object is and what to do with it. Instead, there's a high reliance on the cameras and radar, but these have some limitations. For example, it's difficult for a camera to distinguish between two objects of similar brightness and color, so for example if there's a white semi-trailer in front of a white wall 100ft away, the camera might not "see" the trailer. This can create situations where sensors are reporting conflicting data, which is bad - for example in the trailer situation, the camera would be saying "there's nothing in front of me until a wall 100ft away", but the radar would be saying "no there's something right in front of me". Or the radar might not "see" an obstacle that is a few feet off the ground while the camera is saying "there's something in front of me". Resolving these conflicts is hard.