- The major caveat is that Waymo is not being directly compared against sober humans driving lawfully.[1] The reason why this caveat is so important is that technology which makes it impossible for humans to exceed a posted speed limit might be overall much safer than replacing human drivers with autonomous drivers. Uber isn't more dangerous than Waymo because humans are incompetent, it's because humans obey orders from impatient drivers and Waymo currently does not. This is a UI choice, not an AI advancement.
- More specifically, lawful driving is an important caveat because Tesla Autopilot had two different settings for driving unlawfully, according to the users' own sense of personal risk. An AV manufacturer who advertises "AI-assisted speeding" will almost certainly find a lot of customers, even if it's under the table. People don't speed and run red lights because they're too stupid to understand why it's dangerous: they do it because they're reckless and selfish. AI won't stop that, only regulation will.
- Another caveat is that Waymo was trained on human-dominated streets. Waymo being safer in a sea of human vehicles does not actually translate to Waymo being safer in a sea of Waymos. I think this is a low-probability risk but it's hardly a simple question: I believe Waymo has had issues where several AVs occupied the same street after an event and blocked traffic because they couldn't decide what to do - they were waiting on each other to behave like a human. But again, the risk seems like gridlock, not property damage or injury.
- And a minor but still important caveat is that SF and Phoenix have modern linear grids which have been mapped to death by AV manufacturers. As a Boston resident I am still holding my breath about their performance here :)
[1] Not because of anything insidious, it's just a granularity that both the analysis and the data struggle to capture.