1 billion bytes = 1GB,
1 billion KB = 1TB,
1 billion MB = 1PB,
10 billion * 10MB = 100PB.
Which again leaves me wondering why Uber needs 100PB.
And it's not like Big Tech has ever misled people on this subject.
Remember, Uber also stores application positioning telemetry as well, so it isn't just the number of vehicles.
That said, 100kB does seem high. 10kB or so would seem more appropriate, as a very rough ballpark.
I had a friend head off in 2011 to work with Hailo in London, and designed much of the data ingest & storage architecture there. That was built on Cassandra, and some of the numbers were staggering, even for a (initially) single-city, < 2,000 driver system.
You're not just tracking start and end point of successful rides -- you're tracking users when their app is open (both drivers and users move around before and after initiating a ride, and you need to share that data between parties), you're also tracking empty vehicles so you can identify over / under serviced areas ... doubtless lots of other real time stuff you want to know. But the big thing is you want to then have that data available for mining, improving your service, regulatory obligations, and so on.
Most of what you describe has little meaning or use at a few months old. Sure you might want to collect detailed awareness of individual drivers over say a month, generic traffic patterns over six months, and little or nothing beyond.
Once the data has outlived that direct, primary, purpose you dispose of it or it's a famous, excessive, data breach waiting to happen. There is absolutely nothing wrong with only keeping billing data beyond say six months max.
The really invasive pattern-matches happen best retroactively at that scale.
Age of the data does not typically impact many such analyses.
There's a huge spectrum between your 'at most 30d' claim in another reply, and 'all data in perpetuity'.
30d is way too short. Perpetuity is probably the idealistic intent, and depending on storage costs over time this may be achievable.
> I suppose it demonstrates the difference in US and European perspectives on data.
How so? I described a the case of one UK-based company, and TFA is about one US-based company, both doing similar things.
> Most of what you describe has little meaning or use at a few months old.
Clearly this is not the case, as evinced by the actions of the people designing, building, operating, and maintaining these types of systems.