https://www.theverge.com/2022/1/28/22906071/apple-1-8-billio...
1.8B active devices / 300k nodes = (just) 6k devices per Cassandra node
https://www.theverge.com/2022/1/28/22906071/apple-1-8-billio...
1.8B active devices / 300k nodes = (just) 6k devices per Cassandra node
There is everything from Weather to Siri to Store Purchases etc.
And companies will syndicate data sets to different teams for performance and security reasons ie. lots of duplication.
Of course. That is not the point here.
Breaking monoliths into service boundaries yields easier ownership, maintenance, migration, and resilience.
One "tiny" company with a few verticals can be comprised of thousands of microservices, each handling their own dedicated objective. Authentication, reverse proxy, API gateway, SMS, email, customer list, marketing email gateway, CMS for marketers on product X, feature flags, transaction histories, GDPR compliance handling, billing intelligence, various risk models, offline ML risk enrichment, etc. etc. Each will have its own data needs and replication / availability needs.
This Apple number might seem crazy, but I'm not phased by it. I can picture it.
It's a sad and very inefficient picture though. Apple does not need this this much data processing. It's a grotesque amount per device. My most positive plausible interpretation is that maybe they're just wasting insane amounts of energy doing lots and lots doing of stupid analytics, as one tends to do.
See also the natural stochastic gradient ascent that produced our crazy complicated metabolic pathways (and all of biology).
Again. It is not just for device data.
There are backend services which your device interacts with e.g. Maps, Siri, Weather.
Apple is using less than one TB per server…
But when you see the 1000s of clusters it starts to make sense. They probably have a Cassandra cluster as their default storage for any use case and each one probably requires at least 3 nodes. They’re keeping the blast radius small of any issue while being super redundant. It probably grew organically instead of any central capacity management
Half a TB per node, which during regular compaction can double. And if you went over, your CPU and disk spent so much time on overhead such as JVM garbage collection that your compaction processes backlog, your node goes slower and slower, your disk eventually fills up, and it falls over. Later things got better and you could use bigger nodes if you knew what you were doing and didn't trip over any of the hidden bottlenecks in your workload. Maybe even fixed in the last few versions of Cassandra 3x and 4.0.
Past 100k servers you start needing really intense automation just to keep the fleet up with enough spares.
If you’ve got say 10k servers it’s much more manageable
The fun thing is Cassandra was born at FB but they don’t run any Cassandra clusters there anymore. You can use lots of cheap boxes but at some point the failure rate of using soo many boxes ends up killing the savings and the teams.
Isn't Intragram mostly Cassandra?
https://instagram-engineering.com/open-sourcing-a-10x-reduct...
Zippydb is honestly one of the best parts of fb infra. It let you select levels of consistency vs latency
How is that different from Cassandra's Tunable consistency model?
https://cassandra.apache.org/doc/4.1/cassandra/architecture/...
I’ve seen this type of design pattern implemented successfully with a variety of extremely large databases.
These could be K8s nodes for example, that don't make full utilization of the underlying VM, which would completely make sense at APPL scale.
For some background to other HN users out there, "virtual" nodes refer to logical nodes in distributed database software where # of virtual nodes >= physical nodes. This means if the size of data passes a certain threshold, physical nodes can redistribute the virtual nodes, reducing the amount of data shuffled across physical nodes (as opposed to a naive hash function that mods a key by a fixed number and requires all nodes to reshuffle data when a new node is added).
Apple does have a separate opt-in “Research” program to facilitate this kind of thing.
No it's not.
There is no evidence whatsoever that Apple is doing anything other than facilitating ads through their News, App Store, Maps etc properties.
Their revenue will be an insignificant fraction of Facebook, Google etc.
But unless they’re doing some voodoo magic then yes the data is leaving is your device in some form. Hence why I can view every heartbeat (aggregated by the minute) since I put on the original Apple Watch in 2015 despite changing all my devices and only restoring from iCloud. Indeed I expect it’s just part of your iCloud data storage.
Actually I just launched the health app now and I can see the app explicitly asks if you want to allow sharing your data with apple, so if you say ‘yes’ then you’re not only storing but allowing apple to query your data (minus PII).
Unlike other data in iCloud, if you lose your devices you lose your HealthKit data. This is not true for photos or emails, for example - which you keep if you lose your devices.