Magic Pocket: Dropbox’s exabyte-scale blob storage system
infoq.com
infoq.com
Optimizing Magic Pocket for cold storage - https://news.ycombinator.com/item?id=19841887 - May 2019 (13 comments)
Dropbox Extending Magic Pocket with SMR Drive Deployment - https://news.ycombinator.com/item?id=17300661 - June 2018 (1 comment)
Inside the Magic Pocket - https://news.ycombinator.com/item?id=11645536 - May 2016 (29 comments)
Scaling to exabytes and beyond - https://news.ycombinator.com/item?id=11283064 - March 2016 (6 comments)
Dropbox’s Exodus from the Amazon Cloud - https://news.ycombinator.com/item?id=11282948 - March 2016 (240 comments)
[0] https://dropbox.tech/infrastructure/inside-the-magic-pocket
The system can run on any HDDs, but primarily runs on Shingled Magnetic Recording disks.
SMR is just one of the latest HDD technologies. Wiki tells me that is is about 10 years old now. I cannot believe the implementation of the storage hardware is going to have any affect upon how the software service runs. What am I missing? This sounds like a bit of "tech mumbo jumbo" on the lead-in. To be clear: I am not doubting the impressiveness of this system!The conclusion is also pithy. I like it.
Managing Magic Pocket, four key lessons have helped us maintain the system:
Protect and verify
Okay to move slow at scale
Keep things simple
Prepare for the worst The higher density of SMR drives, combined with its random-read nature, fills a niche between the sequential-access tape storage and the random-access conventional hard drive storage. They are suited to storing data that are unlikely to be modified, but need to be read from any point efficiently. One example of the use case is Dropbox's Magic Storage system, which runs the on-disk extents in an append-only way. Device-managed SMR disks have also been marketed as "Archive HDDs" due to this property.
Cool! I was probably wrong in my previous post: The hardware implementation may truly affect the software.It's an HDD technology that has to be combined with very specific type of logical volume management and protection workloads, and is very cheap at the tradeoff of these limitations.
To build a large, fault tolerant storage system on top of this "just another HDD technology" as you call it, is not something you can do at home with your normal tools. They did it, and they saved lots of cost because of that.
Do you ever walk into a meeting with an engineer SME talking about a topic which you had to look up on wiki, and because you don't understand what he's saying you dismiss it as "tech mumbo jumbo?" You will be very successful as middle-management at a large corporation.
Where can I get SMR drives that are a good deal? And what percent price improvement are you seeing?
When I look at the drives I can easily buy, the bigger models are all CMR and the smaller models keep having SMR snuck in without any notable price reduction.
I think of SMR like Optane. A bunch of consumer-level tech guys w/o a real job got all lit up and started arguing about it online, but its target market was not the "NewEgg website" market, it was the "million dollar storage array from IBM" market.
Remember, there are many workloads that are "write this on disk and keep it forever" - even historical data in a data warehouse showing soda sales by season.
I honestly haven't bought any storage or compute in 25 years - my personal crap is old "enterprise/business line" junk from work.
Where you, as a consumer, can get them - don't. I don't think they're meant to be used by consumers. You're gonna have a bad time unless you write software specific to this weird type of drive. What you might find out there is some kind of a storage 'consumer NAS in a box' product that has tiered storage inside, with these drives and other ones. But stay away from drives on their own w/o that special software in front.
Not the dimms as much but those had horrible prices per gig, and SMR is supposedly a way to save money.
> Where you, as a consumer, can get them - don't. I don't think they're meant to be used by consumers. You're gonna have a bad time unless you write software specific to this weird type of drive. What you might find out there is some kind of a storage 'consumer NAS in a box' product that has tiered storage inside, with these drives and other ones. But stay away from drives on their own w/o that special software in front.
That would be a reasonable statement if they weren't selling SMR drives to customers without even labeling them as such. Just raw drives that sometimes have the performance go to hell. If you wanted fast you should have gotten an SSD, I guess.
It's not that I can't easily get an SMR drive, it's that I can't get the bigger models and I can't get the price savings.
We use a custom disk format along with libzbc, but libzbd now provides many advantages, which we are looking to adopt. I did want the QCon talk to have some super straight-to-the-point conclusions and these, I believe, have saved us the most since I have been on the team. Largely due to the sheer scale of managing such a system.
Super interested to understand this better but having a hard time following the article. Is the hash index replicated across regions/zones, or are clients responsible for routing thier requests to the correct zone?
> The Lustre® file system is an open-source, parallel file system that supports many requirements of leadership class HPC simulation environments. Whether you’re a member of our diverse development community or considering the Lustre file system as a parallel file system solution, these pages offer a wealth of resources and support to meet your needs.
Hadn't heard of it.
MP was designed for multiple exabytes across many data centers and geographic regions. A different system design (or just using S3) would be more appropriate for use at smaller scales.
I hate the phrase "Our system has over twelve 9s of durability." Amazon was the first motherfucker to claim this, but the other cloud storage folks are also culpable, but at least they mostly had the modesty to add some weasel words like "designed for" and didn't just straight up claim there was less than a 1 in a trillion chance of a durability failure.
You don't have twelve 9s of durability. Your collection of copies of data on the hard drives do, assuming they exist in a vacuum and nothing bad happens to them except the normal sorts of things that cause hard drive failures that are nice and completely independent. But it completely ignores all other sources of problem, and those are so many orders of magnitude more common that you might as well claim "God-given, perfect durability" because it'd be just as accurate.
the tldr is that those numbers of 9s are just table stakes. no system should ever lose data due to routine disk failures. so then, as you mention, there's another whole art to mitigating those other sources of problems.
[1] https://www.facebook.com/atscaleevents/videos/17416916227706...
(a bunch of the early MP folks work at Convex now)
As far as I know Magic Pocket has had 100% durability, but that's obviously beside the point.
Should we trust these numbers though? Of course not, because the secret truth is that adherence to theoretical durability estimates is missing the point. They tell you how likely you are to lose data due to routine disk failure, but routine disk failure is easy to model for and protect against. If you lose data due to routine disk failure you’re probably doing something wrong."
https://medium.com/@jamesacowling/how-many-nines-is-my-stora...
That seems... really bad, for the Core Core Service that Everything Depends On.
note, I worked on S3 2015-2017
So yeah you need more storage-optimized server racks and all the associated manpower and maintenance, you also need them to be distributed across different datacenters and zones which of course also impacts latency and your ability to provide some appearance of consistency, then you also need the same distribution for stateless services serving the data.
On and on and on and you might be nearly doubling cost to get to an extra 9 that almost all of your customers won't care about.
I worked on DigitalOcean's object storage for about a year not long ago. Makes sense that those of us who have been in this space would be interested in this article haha.
Note: The consent decree, the actual legal obligation of the Bell system, actually specified the percentage of times someone would pick up the phone and not get a dial tone or operator. It was surprisingly high -- IIRC something like 2%, which is why you can see people toggling the hook in old movies. The five minutes per decade constraint I know because some of my old customers made digital phone switches and they had to provide that SLA to their phone company customers, or else not get the order.
Stepping back a bit, it seems like the FCC would be in charge of establishing rules like this but that consent decree I found (which might be the wrong one), is with the house Antitrust committee. (After my own Googling failed I also asked GPT-4 and it isn't aware of this either)
However the 5 min figure comes from my customers like DSC (R.I.P), Ericsson and Nokia building POTS switches. These guys were deadly serious, like the folks who made spacecraft and medical devices, plus the Ericsson and Nokia folks were nice too.
In DSC’s case they were so paranoid that they paid us a massive amount to maintain a special tool chain for just for them. It was frozen in time (no upgrades) and when they reported a bug and we sent them an updated tool chain they diffed the binaries and made sure that every delta was due to the big fix and nothing else (that the dev hadn’t snuck in some other patch for some reason)! They did some other headstands with their hardware and software, but in the end it didn’t save them.
It’s cool that the consent decree is online. That arrangement with the Bell system was very clever, though it led to a lot of weird anomalies and distortions but I think it did end up with a better phone system than the PTT model. It’s also been a better model than what’s happened with the power and water utilities.
The FCC back then was a better regulator for phone customers than the (now obsolete) ICC had been.
We had one 18 minute issue in the ten years of use. It was the best facility we ever used.
These companies weren't idiots: they had replication for their databases and leased redundant transmission service because they knew this kind of thing could happen. The problem is you'd buy transmission from two providers...whose fiber turned out to be in the same conduit, or even both had rented bandwidth on the same fiber.
At least people are smarter these days, and with widespread cloud service there are fewer people who need to keep track of this stuff.
The equipment was trusted to not fail in it's lifetime.
Telco basically set the bar for "5 9s."
You're not wrong, this is really really bad, especially for Dropbox, storage is their business so I expected way better.
These stats are no different to S3 at all. All of this engineering and moving away from AWS and for so few gain in availability.
I was initially excited when they moved away from AWS and expected industry leading higher availability when they moved away, but I was wrong.
This is disappointing for Dropbox which this is their main business, storing files without any minor hiccups or outages for years.
I would expect Dropbox, a file storage company that proudly invests heavily in tech and infrastructure to achieve a better availability than what they already were on AWS (99.99%)
In terms of availability the change is pretty much 0 and as a business / enterprise customer I might as well choose a different service with similar or higher 9s or (if my needs are complex) choose S3.
https://aws.amazon.com/blogs/publicsector/achieving-five-nin...
Here is a service that has managed to achieve 5x9s of availability:
Guaranteed availability is a bet they're willing to make, a gamble they've been on top of so far, a risk that, should something fail, they will pay out on according to their SLA.
I would have thought that given Dropbox's engineering talent, they would have designed a system that would account for 5x9s and even making that guarantee for enterprise or mission critical customers.
Can't even find their SLAs anywhere for these customers, so I presume that Dropbox doesn't care about them.
Guess I was wrong and this is just disappointing and made the move not worth it.
What makes you think a gain in availability matters or is necessarily a motivation for the project?
If they can achieve the same availability at far lower cost, it’s a win for them, which is why they would (and did) do it.
This isn't a win for enterprise / business / mission critical customers. Governments and public services cannot use this at all.
> Emergency response systems is 99.999% or “five nines” – or about five minutes and 15 seconds of downtime per year.
https://aws.amazon.com/blogs/publicsector/achieving-five-nin...
I see hospitals will be using AWS and other reliable hosts to use S3 rather than Dropbox.
One can have a custom built or an experienced IT team using AWS and achieve the same or even better availability using AWS and other providers that care about critical availability, if architected correctly.
https://docs.aws.amazon.com/wellarchitected/latest/reliabili...
Storage is our business and we target an even worse 99.95% availability[1].
Availability has a cost. That cost is complexity.
We would very much prefer to have boring outages more often than have fascinating outages very rarely.
It seems that Dropbox's system is complex enough but still isn't as designed to be highly available (5x9s) as a well architected system on AWS using S3.
Yet It is surprising a public billion dollar company like Dropbox cannot go beyond this for their mission critical customers (govt agencies, healthcare, military, etc) willing to spend hundreds of millions.
So yes, I expected way better.