You're basically describing one way AV companies already deal with issues. I'll describe a typical pipeline based on broad public information:
When an autonomous vehicle encounters a situation, it will initially attempt to resolve it automatically. Sometimes that fails, so it will fall back to asking a human for guidance. Depending on the company and the vehicle there may or may not be a local safety driver available as a secondary backup.
In either case, that situation will be flagged. The amount of detail varies depending on whether there's a human taking notes. One major difference from Tesla here that all of the high resolution cameras, the LIDARs, the radars, etc will be recorded and made available, alongside the full logging data for the ride. Depending on the car and the software it's running, stack tracing data may also be available.
Some of the consumers of this data collect these incidents into buckets to identify issues. This is used to determine things like rollout success (any particular company will have multiple "fleets" running the equivalent of canary/beta/stable) and for finding edge cases. Once one is identified (say through a news story), they'll search for similar events. From there, they can generate simulated test cases based on real incidents or do things more manually. There are multiple kinds of test cases as well. All subsequent updates will run against those test cases. Updates will be deployed to the vehicles when they're not on road. The exact frequency varies, but it's typically higher than Tesla's update cadence.