I don't want to sound like I'm disagreeing with you, but I want to raise some important points. Collecting too much data is definitely a bad thing. But there are things to consider:
> You are choosing to create that risk by storing any data beyond what is specifically necessary.
That's an interesting point and it begs that question of what is "necessary". Does that mean a company should distill its product down the the absolute minimum viable product without collecting any information beyond what is necessary to do its job? Nobody should use Google Analytics because it's not necessary to run a website? "Necessary" isn't black and white.
Pickups and dropoffs are very troublesome parts of an Uber ride (and definitely can't be avoided...the riders need to get in and out of the car without being run over). Uber currently has no way of improving that, without data at least. So what data is "necessary" to make that experience better? How can Uber, or any for-profit company, improve their product if they have no idea what's wrong with it? What kind of data can be collected that falls into the bucket of "necessary"?
> You could expunge user data when it is no longer needed.
Other than when a user leaves the service, how do you know when data is no longer "needed"? How can trends be measured over time if old data is expunged? How can fraud trends be detected if historical payment data is not preserved? How can Uber choose the best restaurants for UberEATS if order history is dropped over time?
When a user leaves a service for a time and returns, do they return to find a completely empty account? Would you expect this from Gmail, or Facebook?
At what point do you draw the line where you say, "I'm sure we'll never need this again."?
> You could find ways to store only aggregate data instead per-user values.
This assumes the data is homogeneous, which it's not. Different people use Uber (and any other service, for that matter) in different ways. Lumping everything together makes it impossible to discern patterns that are substantial but do not represent a majority of your user base: Some folks only ride to and from work. Some folks only ride to and from bars. Some folks only ever use Uber to go to the airport. How do you identify and optimize for these use cases?
It also calls into question how you bucket your data. Location data in California is not relevant to location data in New York. Location data in Rochester is not relevant to location data in New York City. Location data in Brooklyn is not very relevant to location data in Manhattan. At what point do you stop scrubbing your data into an aggregate form? How do you avoid scrubbing too much data away such that you've painted yourself into a corner when you try to solve a problem in the future? How do you avoid scrubbing too little, such that even when the UUIDs are stripped away you can use a big ol' hadoop job against pickup and dropoff locations to figure out which trips belong to which user?
It's worth asking also how much value you can provide to the user if you do all of these things. When I go home for the holidays once a year, I take Uber quite a lot. Should Uber forget my favorite destinations every year because the data gets expunged? Should it only show common destinations because my trip history has been anonymized?
Anyway, this is meant to point out that it's not simply a matter of keeping things as minimal as possible. It's hard to add value while also protecting users, even when the users' best interests are in mind.