The data is used for (or can be used for) a number of practical and important purposes. These are some examples of what we do as well as what we could do with such data:
- Fighting fraud
- Improving "suggested pickup" locations
- Improvements around POOL (we'd rather suggest a location that's easy to pick up from than one that's unsafe)
- Figuring out what side of the street you end up on (we can do a better job of choosing a route if we know you're able to get out of the car safely)
- Optimizing pickups and dropoffs around events or certain times of the day or week (Caltrain riders: imagine if we suggested pickups at 5th and Townsend instead of 4th and King during rush hour. Much easier to find your driver and prevents congestion)
- Dropping riders off closer to the "correct" entrance to their building
- Dropping riders off in a location that's easier for the driver to depart from without making the rider walk too far (e.g., instead of getting trapped in a weird road or parking lot)
- Analyze where riders ask drivers to make stops mid-trip and make better experiences around that.
As far as trust and safety goes, we can avoid charging you if your driver was looking for you at the wrong location and gave up. We can better see where you were actually dropped off at, if your driver lies and leaves the app running after you've left the car. Leave your phone in the driver's car, and we can potentially do more than just put you in touch with them.
Individual users' data is very closely guarded internally. It's immensely difficult to look at user data without specific access. Overwhelmingly, this data is queried in aggregate and fed into machine learning systems. The risk of abuse is exceptionally low.
But so you're saying it's possible? Somebody with high enough access at Uber might be able to stalk an ex-girlfriend and follow her home? Getting into trouble for it of course, but it wouldn't matter after a murder-suicide? Statistically improbable I know, but that's the sort of thing that people worry about
People with access to hardware can probably do anything if they really want to. People at the bank can remove all your money in the account by going directly into the database as well.
But that's the same story with that one engineer at Google that has SSH access to the server with the keys to the storage cluster for Gmail. Or the person that writes queries for the Hive cluster that manages Cortana queries at Microsoft. Or the person that writes the API layer that sits in front of the order history service at Amazon.
There's a million ways an engineer at any big company could abuse their powers. But we put safeguards in place, we watch each other, and we do our best to serve our users as best we can. I say this in a general sense, not specifically to Uber. It comes down to trust.
Yeah, I suppose this is true. Thanks for sharing your thoughts, I appreciate it.
This is why it's important to minimize the attack surface. Yes, there is always going to be risk when you store tracking data. You are choosing to create that risk by storing any data beyond what is specifically necessary. Uses such as "Optimizing pickups and dropoffs" are optional, which you more or less admit when you said it was something "we could do".
Alternatively, you could prioritize the safety of your customers by minimizing the risks they acquire when they use your service. The data collected for the ride could be minimized. You could expunge user data when it is no longer needed. You could find ways to store only aggregate data instead per-user values.
> we put safeguards in place
You also made design and business decisions that created the need for those safeguards. You are protecting your users from a problem that you created.
edit:
> we do our best to serve our users as best we can
You could be indemnifying you users for any damages that derive from the data you are choosing to store.
> You are choosing to create that risk by storing any data beyond what is specifically necessary.
That's an interesting point and it begs that question of what is "necessary". Does that mean a company should distill its product down the the absolute minimum viable product without collecting any information beyond what is necessary to do its job? Nobody should use Google Analytics because it's not necessary to run a website? "Necessary" isn't black and white.
Pickups and dropoffs are very troublesome parts of an Uber ride (and definitely can't be avoided...the riders need to get in and out of the car without being run over). Uber currently has no way of improving that, without data at least. So what data is "necessary" to make that experience better? How can Uber, or any for-profit company, improve their product if they have no idea what's wrong with it? What kind of data can be collected that falls into the bucket of "necessary"?
> You could expunge user data when it is no longer needed.
Other than when a user leaves the service, how do you know when data is no longer "needed"? How can trends be measured over time if old data is expunged? How can fraud trends be detected if historical payment data is not preserved? How can Uber choose the best restaurants for UberEATS if order history is dropped over time?
When a user leaves a service for a time and returns, do they return to find a completely empty account? Would you expect this from Gmail, or Facebook?
At what point do you draw the line where you say, "I'm sure we'll never need this again."?
> You could find ways to store only aggregate data instead per-user values.
This assumes the data is homogeneous, which it's not. Different people use Uber (and any other service, for that matter) in different ways. Lumping everything together makes it impossible to discern patterns that are substantial but do not represent a majority of your user base: Some folks only ride to and from work. Some folks only ride to and from bars. Some folks only ever use Uber to go to the airport. How do you identify and optimize for these use cases?
It also calls into question how you bucket your data. Location data in California is not relevant to location data in New York. Location data in Rochester is not relevant to location data in New York City. Location data in Brooklyn is not very relevant to location data in Manhattan. At what point do you stop scrubbing your data into an aggregate form? How do you avoid scrubbing too much data away such that you've painted yourself into a corner when you try to solve a problem in the future? How do you avoid scrubbing too little, such that even when the UUIDs are stripped away you can use a big ol' hadoop job against pickup and dropoff locations to figure out which trips belong to which user?
It's worth asking also how much value you can provide to the user if you do all of these things. When I go home for the holidays once a year, I take Uber quite a lot. Should Uber forget my favorite destinations every year because the data gets expunged? Should it only show common destinations because my trip history has been anonymized?
Anyway, this is meant to point out that it's not simply a matter of keeping things as minimal as possible. It's hard to add value while also protecting users, even when the users' best interests are in mind.
I think people could be picked up and dropped off just fine prior to this update so it's not something fundamentally broken. Could things be improved? Sure but the benefits need to be weighed vs. the tradeoff.
Given the many responses that you have gotten so far arguing against your stance, it might be helpful to consider why people are having such a reaction. We have valid concerns and they should be addressed.
Having said all that, it's still creepy. And the examples you gave (Google/Gmail, Microsoft/Cortana) are almost certainly also instances where data collected is being systematically used in what many would consider creepy ways.
For me, I go to and from home, the gym, and work. Every day. If someone was looking at that data, I'd be creeped out, but they'd learn almost nothing about me. If someone saw my credit card statements, I'd be mortified. And yet, almost everyone forgets to fill out and mail back that little slip you get with a new credit card that says "hey, actually I don't want you to sell all of my data to everyone on the planet."
Ha, this sounds exactly like "I have nothing to hide, so I'm okay with the NSA gathering all of my data".
Part of the issue is that location has a lot more "physicality" tied to it. One can't easily change where they live, work etc. whereas it's much easier to switch credit cards.
It certainly is; it's also something I've never said
so.. uh.. (eyebrow wag) what else do you use it for that you didn't enumerate?
The risk of abuse is exceptionally low.
Except, of course, events where you and I disagree about what abuse is. Like tracking journalists around for shits and giggles.
[1] https://en.m.wikipedia.org/wiki/National_security_letter
[2] http://www.buzzfeed.com/johanabhuiyan/uber-is-investigating-...
Have we gotten an NSL? I don't know, but I'd guess we probably have. It would be foolish to think otherwise.
That is completely irrelevant. The powers that be can siphon off the data as it comes in. The danger is in receiving excess data in the first place, making yourself a more enticing target for collection activities.
Uber also only uses strong TLS for connections, so "the powers that be" will have a heck of a time taking a real heap of gigabits (terabits?) per second of encrypted data and doing anything particularly useful with it.
Does uber offer a password/bank information storage service?
With such great security and unhackable storage, those are the next logical upgrade and would make a really great one-stop service for your customers.
The data volume isn't very big. Let's assume Uber records locations for 20 million users and pings each user every 5 minutes. We're talking about 20 million users * 12 events per hour * 24 hours * 16 bytes (my generous assumption of how much it would cost to store a user id, timestamp, and precise long/lat data) = ~100 GB a day or 3 TB a month. And 20 million / (5 * 60) = 67k events per second.
I could probably build out the infrastructure to process and store a year's worth of user location data for 20k a year on AWS.
What I'm getting at is that it's pretty straightforward and easy for Uber to maintain a database of every user's locations at all times, even if they aren't active users. This kind of data is valuable. Assuming Uber isn't already doing it, all it takes is one mid-level product manager to initiate a small project to kickstart it.
All of the points that you've listed here would work just as well with anonymized data. Obviously if users are being picked up or dropped off at their house (for instance) then making every journey truly anonymous would be impossible without damaging the data, but you could still strip out explicitly identifying metadata. Are you able to tell us what sort of measures are taken in this regard?
I don't work on this stuff myself, so I can't say.
I call bullshit. This coming from the same guys who were mapping one night stands using ride data? What were you guys studying there. Fornicating habits of young adults in large metropolitans?
http://www.whosdrivingyou.org/blog/ubers-deleted-rides-of-gl...
I almost forgot about the time you guys were tracking journalists.
Uber has a reputation and history of being a "shady" company.
I'm not going to try to defend our reputation. But it's worth saying that things are locked down _pretty damn tight_ around sensitive data. I've worked in enterprise file storage in the past, and the internal security at Uber is far better (relatively speaking), and continues to mature.
EDIT: BTW, did you hear from your support that there are users who aren't pleased with the new feature?
Someone well above your pay grade probably should. As far as I can tell, Uber's business model is psychopathic levels of regulatory and psychological arbitrage.
Commenting, to save this little gem for the inevitable time when Uber gets hacked. Again.
Which means absolutely nothing as an assurance. Even if that's the case today, in 5 or 10 years the company could go down, or change CEO and be all about exploiting the data, or selling it to advertisers or whatever. Or it could just be a hack that releases millions of ride information (it has happened to the best of web services).
That's the problem when you store data, making whether you have some "pretty damn tight locks" in place irrelevant.
But in all seriousness, I'd imagine the execs need to go through the same process an engineer would.
You see. You imagine. You believe. You are told. You do not know. And even if it were so today it would not have to be that way tomorrow. So even todays security isn't enough of a reassurance.
The only secure data is data never collected in the first place. And until the friggin disruptive startups start to recognize this I will try to not support them in making my data more insecure.
One is where we believe, based on your words (or some random person on the Internet who claims to be an employee), Uber employees are restricted from accessing customer data...you know, like LOVINT in the NSA. [1]
The second is where we believe, based on your words (and associated caveats), that Uber does not have a firehose feeding all this information to some other entity that has a much larger capacity for "machine learning", has far lesser actual oversight than what you claim to be in place in Uber, and uses that data for many things we don't know the impact of. This point might sound like I'm talking only about the NSA and trust in the U.S. government, but keep in mind that Uber operates in many countries, and all of their governments have an interest in gaining such data and using it for their own purposes.
Post Snowden, neither of these sides look harmless. I'm talking about the world society as a whole, not just about what one person may consider ("nothing to hide") or what negative experiences that some people may never go through in life.
[1]: https://www.washingtonpost.com/news/the-switch/wp/2013/08/24...
"Uber said it protects you from spying. Security sources say otherwise"
https://www.revealnews.org/article/uber-said-it-protects-you...
I'd ask bastawhiz to comment given his comments above.
But there's no oversight whatsoever. We should just take Uber's word for that -- despite the company's public track record of repeated privacy violations in the past?
And even if Uber itself has suddenly converted into an unimpeachable guardian of personal privacy, it could always be hacked (and probably already has, just like the NSA and everyone else.) The hackers are not likely to share the same moral compunctions about privacy.
I also noticed you said that it's queried in aggregate, but that is distinct from it living in aggregate. Does this mean that you actually do keep my individual GPS data and it could be looked up later?
I am still unconvinced that this is necessary, even with those (potential) improvements. I am relatively okay with allowing Uber access to my location outside of the app during the period between me requesting a ride and being picked up, and of course I'm okay with allowing access when I'm physically in the car. But after I get dropped off is a much harder thing to accept. Most of the things you suggested were for before I get picked up, not after. In fact, the only things you suggested were not on that bullet list but in the second to last paragraph ("we can avoid charging you if your driver was looking for you at the wrong location and gave up"; "We can better see where you were actually dropped off at, if your driver lies and leaves the app running after you've left the car").
I understand that Uber itself is trying to be a better company and is trying to move beyond its past life as a cavalier company that suggested looking up information about journalists who write negative things about the company, but it's a really hard image to shake and understandably many people are distrustful of a company whose executives were okay with this. I could very much see a member of Uber's counsel requesting GPS data on some user who filed a lawsuit and giving it to a private investigator (considering that Uber did in fact request information about and that a PI to do investigation into a person who filed a lawsuit against Kalanick).
I don't know specific programs that are in place or upcoming, not my jam.
> Does this mean that you actually do keep my individual GPS data and it could be looked up later?
Trip location data is attached to users. You can look it up yourself on the Uber website (riders.uber.com). With regard to data collected from rider devices, I don't know the technical details but I assume it's not living in aggregate.
At the end of the day, we're an engineering- and data-driven company. Pickups and dropoffs are overwhelmingly the most painful parts of the average Uber ride, and we know startlingly little about them: that's what this initiative is about. We want to make things better for our users, and overwhelmingly, things like this do lead to improvements. Whether you trust what we do with the five minutes of data is up to you. I'm not here to speak about our character, though I certainly wouldn't be around if I thought Uber is a bad company.
If Uber was really serious about user privacy, this would be a no brainer feature to offer.
It's also the case that the record of your trip isn't just data. It's the receipt for an actual thing that happened. A transaction between three or more parties takes place. That data needs to be preserved for any number of the reasons, including government compliance. Even just handling credit card disputes makes it invaluable.
A receipt is the piece of paper or electronic record showing how much your trip cost.
Show me an example of someone using trip data as a primary form for handling CC dispute. Please also back up your statement that trip data has to be kept forever for government compliance.
Lastly, what if my driver and I both want to delete trip data? Per your argument, is there an existing option to delete trip data then? I didn't think so.
A receipt is a proof of transaction and completely separate from whatever data you collected about your customers, your statement makes little sense. That it facilitates disputes doesn't change the nature of data privacy, it is just convenient for you and the bank.
Also, why isn't there an option to delete previous trip data?
Uber has your data even if you delete your account. I'm not sure if they can be convinced into deleting it, but they don't do that by default.Your iOS app effectively turns useless when choosing "Never". You can't do anything other than booking a cab. No access to account information, ride information, anything else.
----
Yeah, I get what it is when you say:
Improving "suggested pickup" locations
If your app watches me 24x7, then it'd be able to predict pickup locations better. Good job. It can't get creepier.https://techcrunch.com/2015/07/08/uber-suggested-pickup-poin...
All it does is help you pick spots near your pickup location that are more easily accessible to drivers.
For now. You haven't mentioned how long this data is kept, which means I can assume "forever". This means that a future change of policy, an acquisition, or a merger can expose this data to a variety of third parties (not to mention the obvious possible exposure to governments).
The reasons you have listed for gathering the data are not sufficient to offset the future risk to me. Moreover why should I do unpaid data acquisition for Uber while bearing all the risk in case the data is exposed? If you want higher detail location information, do what Google did and send out cars with sensors.
It's not just about the locations, but about the interactions between the rider(s), driver, the location, and the time. A location might be difficult to pull over at during rush hour, but nearly empty later on. Drivers might be unable to get very close to an event venue. Certain streets might be difficult to drop passengers off safely on, based on traffic conditions (though we wouldn't know that). And of course, there are many applications around customer support that couldn't be solved otherwise.
It is impossible to take your word for it. After so, so many companies have been caught intentionally or unintentionally mishandling user data, there is no reason to believe that Uber is the exception.
An extraordinary claim demands extraordinary proof.
I'm not asking you to. If you'd like to come in and interview for a position and eventually see for yourself, though, I'm happy to accept your résumé. :)
Your company is grossly overstepping privacy bounds, and your excuse making post is largely irrelevant. And it's wholly reasonable to be distrustful of a place where execs threaten to blackmail journalists and remain employed. It'd be ridiculous to trust Uber.
Giving people choice is good.
I can understand if you gave the "Always" option in addition to the "While Using" option with this explanation, but strong arming your users to give an all or nothing choice is bad.
I also doubt that Uber as a company, and Uber engineers individually, hadn't realized that already.
I'm going to deny permission, and call Uber from Apple maps. No reason for you to track me from M-F if I only use Uber Sa-Su.
They care because they are data driven and not personally culpable for their invasions of privacy. Why not collect everything you can without getting more permissions?
Personally, over never seen any evidence that Uber actually knows what it's doing except making a shit-ton of money. Even the self-drivIng car strikes me more as a sideffext of having a bunch of money they needed to spend and it won the rock-paper-scissors contest against rockets.
Doubly so for the "ME TOO!" Doordash / Postmates competitor.
http://www.nakedcapitalism.com/2016/11/can-uber-ever-deliver...
http://www.nakedcapitalism.com/2016/12/can-uber-ever-deliver...
Well... I still think Uber is drunk on money. Just VC money.
And I'm sure that if Uber does implode, Travis will still come out a hundred-millionare and a reputation of being a business genius, even though he never actually turned the growth into a successful sustainable business. Kind of like Sean Parker and Napster.
https://ftalphaville.ft.com/2016/12/01/2180647/the-taxi-unic...
Why not both? Why not collect all the information you possibly can, whether or not you immediately see obvious value to it, and then find ways to justify it later?
Because it's non-trivial and it costs a lot of money to do that, in terms of engineers' salaries, storage, etc.
Well they _have_ raised $15,000,000,000 in cash: http://www.nytimes.com/2016/06/21/business/dealbook/why-uber...