Only the safety tokens are generated on your device, the action triggering scan happens in the cloud (just like all the others) and then it goes to human review, so if it's a hash collision it'd be caught there when they review the images, and they can only review the images that matched CSAM.
I honestly can't find the uproar here. Google devices can face match photos offline... so they are applying a neural net (scanning) ON THE DEVICE! How is that not worse than what apple do?