Apple Reportedly Storing over 8M Terabytes of iCloud Data on Google Servers
macrumors.com
macrumors.com
The point is that there are layers of security, and by moving the data outside of their physical control apple has given up one of those layers of security.
It might matter, yes. A company run by a guy like Eric Schmidt is a lot more likely to play nice with the US government when it comes to privacy compared to a company run by a guy like Cook, who from the outside seems obsessed with user-privacy (as long as China isn't directly involved).
Google is a huge company, the idea that it would set out to do something that everyone involved in would know is directly breaking the law (rather than doing something thats a grey area, or they know is legal but 'unethical' or could become a PR issue, or destroy trust in them and destroy their product) is fairly unbelievable.
This comes up again and again with stories about Google and (particularly) AWS cloud computing. I hope for better on HN!
There’s also mainstream reporting that Amazon employees used retail sales data to launch competing products: https://www.wsj.com/articles/amazon-scooped-up-data-from-its...
Maybe the first story is fake and the second is real… but they both point to a “win at all costs” company culture where policies might be violated, even if it threatens trust in the platform and a PR problem when exposed.
Hacking computers, networks, or services of your competitors, even when running in you data center is just bad business.
You're reading that a lot differently than I am.
The quoted post says:
>AWS proactively looked at traction of products hosted on its platform, built competing products, and then scraped & targeted customer list of those hosted products
None of that reads, to me, as them having had to use confidential data to do any of these things.
You can identify many organisations that are running on AWS without knowing anything about AWS accounts - blog posts, IP space, public code, social media comments from staff, linkedin and all sorts of other places will often reveal that.
Scraping/finding customer lists can be done using research, too. I've spoken to Account Manager-types at places and they've often used various tools that scrape other public resources to identify customers of competing services.
Swap out company names, and it's effectively what I've seen from a bunch of companies, without it delving into anything unethical.
Data is still data.
> Google is a huge company, the idea that it would set out to do something that everyone involved in would know is directly breaking the law (rather than doing something thats a grey area, or they know is legal but 'unethical' or could become a PR issue, or destroy trust in them and destroy their product) is fairly unbelievable.
They do this in Europe by not complying with GDPR. ( and they are not the only one)
> This comes up again and again with stories about Google and (particularly) AWS cloud computing. I hope for better on HN!
Google employs the same standard issue tech person you see here on HN.
And for the "encrypted data at rest" scenario we're talking about here, where symmetric encryption suffices, quantum crypto makes no sense what-so-ever.
This is extremely unlikely.
> What we know is that, extrapolating compute speed from the past decades and even assuming quantum computers become useable in practice, the best algorithms we currently have cannot be brute-forced within the next 50 years.
Quantum computers only offer a quadratic speedup against symmetric ciphers.
AES 256 will survive much longer than the next 50 years against brute force attacks.
It will either be broken spectacularly, using theoretical methods entirely inconceivable today, or live on – brute force is of no concern at all due to the amounts of energy and matter required to perform it against 256 bit keys.
From what I understand it simply can't be broken by brute force because simply iterating through every possible value of a 256 bit key would require more energy than there is in the universe, and that's without actually trying any of the combinations, just simply having a computer do a i++ through all possible values.
I'm not sure if quantum computing helps here in any way , someone else would need to chime in here with details.
Theoretically a quantum computer can brute-force AES-256 using 2^128 sequential steps using Grover's algorithm (i.e. a quadratic advantage over a classical computer). Parallelization diminishes the advantage, e.g. if you're limited to 2^64 sequential steps, you get a 2^64 speedup over classical, for a cost of 2^192 which is still ridiculously large.
Thus quantum computing is not a relevant threat for AES-256 or most other 256-bit symmetric crypto.
It's my understanding that when encryption gets "broken", it usually refers to something other than a simple brute force attack. Like, something that would make it so you don't need to run as many iterations or whatever.
I assume this because a brute force attack is something that is always possible from day 1, whereas an encryption scheme being broken is something that happens some time afterwards.
As a counter-example, DES would count as "unbroken" under your definition. The EFF built a machine in 1998 for under $250,000 that could crack a DES key in under 24 hours. I don't know what that would cost today, but I wouldn't be surprised if a couple GPUs could get you the same thing today.
The difference is whether such an attack has even a vanishing chance of succeeding. For AES, the hardware just isn't anywhere close to that. Afaik, there isn't anything that could even hypothetically threaten to make brute force attacks on AES feasible on the table today.
DES is both weak and broken, but it could be either without the other.
AES-256 was broken in 2011.[1] While only four times faster than brute force and thus not a practical attack, it suggests that compromise is possible. The Snowden documents indicated that the NSA was working on breaking AES-256. It seems unlikely they would waste effort on a task they considered impossible. Whatever they achieve will be achievable by others eventually.
On top of that, no implementation is perfect. Bugs are discovered in cryptographic APIs on a regular basis. Even if your API is perfect, the application calling the API can have bugs that allow compromise.
[1] https://web.archive.org/web/20120905154705/http://research.m...
Inch by inch "don't be evil" has been replaced with "maximize profit".
I have no doubt that if in 20 years decrypting "historical" Apple user data for "training purposes" is legal and will make Google leadership more money, they will pressure ethically flexible engineers to do it for them.
Assume anything profitable that is legally defensible somewhere in the world will be done by every surveillance capitalism company at some point.
Like, suddenly everything becomes plaintext.
Yes. If you want your data protected, it always matters where you store it.
Well, while there is truth to that, it isn't the whole story. There is a time value to all information that must be factored in. If nothing else, think of one overbroad* classification example: Battle plans, SECRET; Intelligence; TOP SECRET.
* By which I mean there are subtleties and exceptions too numerous to go into here, but the example remains largely illustrative.
The higher the classification, the higher the long term time value, the greater robustness required in one's controls.
In this case, information about/for large numbers of private individuals, today's strong enough symmetric encryption may be strong enough for quite some time.
Or it might not be. I'd love to see a detailed risk assessment....
Where are the keys though? If they ever end up on the same platform it doesn't matter.
They do of course have access to the ciphertext and access to traffic patterns, in crypto threat models traffic analysis tells the adversary a wealth of information.
It's not hard to imagine something business or government related in place of this of course. And do this analysis in aggregate and follow many people at once.
Apple could and for all we know possibly has implemented countermeasures for many of these cases, eg to make it hard to distinquish users from the mass of ciphertext.
What about this is practical?
Anyway, there are many scenarios that come to mind for knowing their IO sizes. Apps probably have IO fingerprints. Or you could send the set of suspected users differently sized files to probe them. Etc.
I asked for practical examples. I don’t need an in-depth report to see that this doesn’t qualify.
Most of this crypto stuff is completely impractical risk, especially compared to some phishing emails.
When it comes to communication. iCloud is file storage. Data at rest encryption.
You mean user devices may connect directly to Google storage? Did you observe it connecting to IPs in Google owned ASNs?
This should be pretty visible to Google, the rest of the traffic is handled better.
It just means Google may provide access to metadata outside of Apple’s control. Those metadata could be useful to do classification of anomalies on the basis of pattern of life analysis, or similar.
Are they making heavy use of public key cryptography? If so how? When I send a message to you, do I encrypt it using your public key? What about group messages? Does each conversation get its own key pair?
Also it’s interesting they decided to directly hit up google cloud… you’d think they would wrap it so at minimum they could tweak the underlying infrastructure without requiring every client to update.
They don’t: public key cryptography is not initially used.
The sender generates a random AES-256 key, applies it in CTR mode and uploads the encrypted blob to GCS.
Every receiving device gets a message with the key, the URI, and the SHA-1 of the blob. These messages are encrypted as usual and sent via APNS (<n>-courier.push.apple.com:5223)
> you’d think they would wrap it so at minimum they could tweak the underlying infrastructure without requiring every client to update
Apple does this: two other endpoints are *.blobstore.apple.com and the Chinese Guizhou-Cloud Big Data.
In my logs blobstore is used less than 1% of the time.
What?
I would think Apple is smart enough to mix storage blobs, so one blob is not one user. Plus all requests come from Apple datacenters, not user devices.
Edit: Not really trying to be facetious, but come on. Apple using a lot of storage at a lot of places isn't too surprising
If it was exabytes I doubt I would have looked.
It's also way beyond what we use in other units. Gigalitre is the biggest SI prefix I have seen outside of bytes.
But gigalitre? Surely, calling them cubic kilometres (edit: should be hm³, see followup comments) would be preferable?
He was also a high-level administrator who had a super-power that few administrators have: he understood science as well as people.
And, to top it all off, the book he coauthored with Johnson and Fleming is widely regarded as the foundational document in the field of modern Oceanography.
There is so much to say about Harald Sverdrup that students quickly become comfortable with the occasional use of "Sv" as an abbreviation for 10^6m^3/s.
my take : it's because numbers with lots of digits sell clicks easier -- up until scientific notation is needed, and then at that point the general readership can't fathom the number and generally doesn't care.
in other words, I bet 8,000,000 terabytes produces more clicks than '8 exabytes' and '6.4 x 10^19 bits' would have.
If I had been born a new-age journalist i'd have gone with 8,000,000,000,000 megabytes -- but only if that fit in the link/URL slug and headline for maximum click-bait exposure.
Speak for yourself, I've been trying to popularise "1Mm" for years. Admittedly not with a huge amount of success.
> or Sun to Earth 150 M km instead of 150 Gm
Surely "1 AU" is the preferred form for this one.
The point would be to put it in a familiar unit, even though '150 million' isn't easy to visualize.
It's roughly the same thing as describing a distant potentially-habitable planet as being "1.5 times the size of Earth" instead of giving its radius in metres. It's just that AUs wrap that "n times the distance between the sun and the earth" into a nice standardised unit.
You'd probably do the same thing that you would do if they asked what a km was: keep changing the units until you get one that they recognise, and then now they can understand the original unit you used.
(Although in this specific case given that they're asking about the definition of the unit you'd probably provide that context without needing further prompting in the first place, similar to if they asked about the circumference of the earth and you gave an answer in metres)
Everybody knows it’s 4.848×10^12 attoprsecs.
But also, I still have the Earth–Moon distance internalized as 385000 km, not 385 Mm, since that is what I learned as a kid.
But stating the Sun–Earth distance as 1 AU is just a tautology, restating the definition, and hence void of information. I suspect you’re jesting. ;-)
Yeah, I mainly only use it in chats (where I can immediately explain it if needed) or in person (where the actual pronunciation of "megametres" makes it pretty obvious). "Mm" and "Gm" being somewhat difficult to google due to "mm" and "gm" being different SI units makes it untenable to just expect people to work it out on their own.
> But stating the Sun–Earth distance as 1 AU is just a tautology, restating the definition, and hence void of information. I suspect you’re jesting. ;-)
To be clear, I don't mean that you'd answer "1 AU" if someone asked you how far away the sun was - obviously that's completely worthless information! The discussion was just on what units and prefixes are used for various magnitudes of measurement, and the AU is a fantastic unit to use for things that would otherwise be measured in hundreds of gigametres. For example, if you had a table containing the distances between various celestial bodies, it would be very convenient to have "Sun <-> Earth: 1 AU" in there along with things like "Sun <-> Mars: 1.5 AU".
And don't forget cubic centimeters (or cc) for engines. Anything from 2000 cc down to 50 cc is very common east of the Atlantic Ocean.
Most people don't know what an exabyte is. That's why you frame the amount of data in an amount of data (terabyte) that people know.
I do find it interesting that Apple could be one of Google's biggest Cloud clients though, that's very surprising.
It’s literally innuendo and conspiracy thinking to suggest they are.
8TB hard drive is about 200 dollars nowadays, so this amount of data is about 200m dollars worthy of retail hard drives.
Ofc the calculation is super inaccurate, which doesn't take redundancies into consideration, and discounted prices for someone like Google to purchase hardware, plus the discount for Apple as a big customer. But had the scale be comparable, in which case, Apple is paying Google billions per year to handle the data storage for itself, doesn't sound quite news worthy IMO, pretty price efficient even.
Edit: Previously stated 1.6B -> 200m
External hard drives can be obtained on sale for $15-16/TB. They can then be shucked to get internal drives. I suspect bulk buyers can get pricing equal or cheaper than this, rather than the retail of $25/TB
Edit: I googled it an apparently it’s an economy of scale thing. People are more likely to buy external hard drives, much more mass market thing, Costco and Walmart sell externals, therefore downward pressure on prices compared to internal drives.
Exactly, that’s a spec :)
Surely it didn't cost GOOG $6B upfront worth of infrastructure to get $300M of annual revenue from Apple. Yes, there's economies of scale, but I think you overstate the case.
Short answer: The cost of just the hard drives to replace an $89MM S3 storage bill was about $45 million dollars. That's not including bandwidth, racks, datacenter space, administration etc etc.
At that scale its less about sourcing the hardrives (yes its a problem, but now youre a big customer so you can ask for specialised things, and a 50% discount)
The big problem is the hosting, finding and building a number of datacentres local to where your customers are. Powering them, and connecting in a reliable way.
But even more, you need to make a storage interface that doesn't suck. I am currently working with a homegrown S3 "like" interface. It is a massive pain in the arse as it doesn't scale[1], has an entirely new set of words to describe standard things (trying to overwrite a file? it doesn't tell you that, it just says "predicate failed")
[1] large files transfer very fast, faster than s3. However all the tools are written single threaded. Add to this that each operation takes at least .7-1.5 seconds, shit gets slow very quickly. There is little to no documentation, the API is odd.
in short, much as it annoys me to say this: for general purpose storage, buy over build.
Not to mention the random access performance of spiny disks ain’t too great either.
Plus, it’s a no brainer leveraging Google and AWS. Their global footprint and expertise alone is worth it. Also, $300 million a year is a drop in the bucket to Apple’s bottom line. They probably make that alone from the millions of adapters they sell each year.
They also use AWS, which is over $350m a year. That contract ends in a year or two.
Afaik, they don't run on azure anymore.
It's not hard to believe they're spending $1B+ on external cloud partners.
At $57B net revenue ($274B), a high margin space such as the cloud services could make sense. Although if they're not selling it to the public it would only account for 1-2% net revenue.
Apple likes to learn from companies.. Soon there will be an Apple cloud using low energy ARM (M1X) servers. Let's call it the "Digital Apple Tree". Maybe a small reference to the DAT recorder, as Apple likes music. Of course this last part is just made up ;-)
why?
I'm paying Apple $120/year for 2TB of iCloud storage. I might not be doing the math right but at $300 million/year for 8 million TB, that's $75/year in google expense for the $120 I'm paying Apple.
There’s long lead times for data center level expansion.
So that leaves personal photos and videos. Apple has about 1 billion active users, so a back of the napkin calculation suggests that each user is storing 8GB of data. That actually sounds quite low. That's 1600 5MB photos without even talking about the crazy storage that personal videos would use.
The article talks about how much Apple stores on Google's servers, but doesn't say how much they store on their own servers or with other third parties. Maybe they have multiple times as much storage on their own servers.
Seems like a bargain, even at that scale.
Note: Apple charges $1/month for 50 GB of storage. That's ~84% margin Apple is making on their cloud storage service. Though granted they do give everyone 5 GB of storage for free, so the math will be lower (and they will still be generating a profit so long as 1 out of 5 are paying for iCloud storage).
Don’t know if they still are, but they are not only using, but deeply embedded in GCP.
On my phone I keep all my personal stuff in git with the remote set to my desktop (it has a public IP) over ssh. My photo stream is just rsynced occasionally over ssh.
Not having to deal with this is arguably even more convenient provided your OS vendor doesn't intentionally get in the way.
Also, I would suggest not port forwarding ssh directly to an internal machine. Use a VPN and then ssh.
The absolute state of things.
What Apple and Google are selling are different products. A fruit salad is different to what the fruit store sells and receives a higher price.
I ended up just plugging the iPhone into the computer and exporting the photos that way.
I think we could see full device backups to other cloud services later. I guess the problem is they would want to provide a specific api for 3rd parties to implement rather than trying to drop a 256GB zip file on google drive with no differential backups.
You can set an iOS geofence (without sharing your location with it) so it automatically syncs each time you come home.
No affiliation; happy user.
To test this theory, simply ask an average iPhone or Android user what encryption is.
What you are saying here is the equivalent to me buying a plane ticket, but also having the option to fly the plane as well.
If you want to manage your own data - use something like a PinePhone - the barrier to entry is high enough that people who are capable of managing their own data securely can use the device and achieve the data sovereignty outcome they require.
That said, I believe Apple should build and manage their own infrastructure. They have the resources and capabilities to do so. The longer Apple stick’s with GCP the larger the inertia of moving away and the less leverage they have when it comes to negotiating to maintain their high standard of privacy commitments.
Most normal usera may not be able to secure things effectively and they can still use the existing iCloud / Google Play infrastructure. That doesn't mean that other users, who want to manage their own data for one reason or another, shouldn't get the opportunity.
It's far more likely that building and maintaining this feature is not worth the development time for Apple and Google product teams at this point, since the possible market is small
By that I mean, storage in iCloud is optional. You can back up Macs locally pretty easily with Time Machine and you can back up iPhones to Macs locally as well (which then gets backed up via Time Machine). You can encrypt Time Machine and iPhone backups too if you want.
Is it as easy as using iCloud for backups, photo syncing, etc.? No, but that’s because iCloud is not just dumb storage, it’s a hosted software application (several, really).
edit: I guess my point is that different people have different computer skills and I think it's great that the tools now exist so people on the lower end of that spectrum can sort this out themselves without having to ask for help.
I've moved everything to the cloud. I have nothing running at home except networking gear, and a small "server" that pulls nightly backups from the clouds to a local USB drive.
In theory i could probably do without the local server if looking at Apple/Microsoft/Google data redundancy (Microsoft is multi geo, i can't figure out what Apple is).
Sadly i need to guarantee that some random account closure doesn't remove all my data, so the backup server stays for now. The way cloud prices are going, it will only be a question of months/years before it's cheaper/easier to just utilize two cloud services, one for main storage and one for backup storage, and with projects like the data transfer project [1], you don't even need to download them first.
Its a million drives. given the failure rate that I had when I was looking after 5 pb,(about 8k hdds) you'd be looking at at least 100-400 failed drives a week
But how do you detect that, how do you schedule replacements? whats your hamming factor for redundancy, how do you optimise your storage? for speed, power, geospatial, or redundancy?
what about capacity management, the lead time on growth would be large.
I was a sys admin for a large VFX place, so did a lot of storage admin, but unless it was a core business function, I'd buy over build that shit, for that scale.
It’s encrypted on your iPhone, but once it leaves your iPhone it becomes unencrypted once stored in iCloud.
This is also why they can comply with law enforcement requests.
Now, someone may argue that “only Apple can unlock it” if indeed there is any protection at all.
https://fortune.com/2020/01/21/apple-icloud-encryption-law-e...
https://www.reuters.com/article/us-apple-fbi-icloud-exclusiv...