Using Amazon Snowmobile to transmit 100PB of satellite images to AWS
wired.com
wired.com
"Never underestimate the bandwidth of a station wagon full of tapes hurtling down the highway." —Tanenbaum, Andrew S. (1989). Computer Networks. New Jersey: Prentice-Hall. p. 57. ISBN 0-13-166836-6.
Not really built in if it's out of band.
The rag-tag team of ex-spies have to band together with the younger techy team to hijack a Snowmobile and modify data on the fly, barreling down the highway.
I just got you unlimited international calling time
Sadly, Burt Reynolds isn't available to drive (he needs assistance to get around these days)
"Snowmobile is protected by 24/7 video surveillance and alarm monitoring, GPS tracking, and may be escorted by a security vehicle during transit." According to the FAQ.
By the time we're outside, the GPS is now in my truck, the driver and data of the real truck are under my control.
(The most unrealistic part of this scenario is that someone would have their main data center in Manhattan.)
I saw a few students nearly have a breakdown over that one.
I must stress this was a single question on a single first-year exam. Not representative of the whole course in the slightest.
You want bytes/seconds (i.e. quantity of data per unit time).
You have:
- Distance per unit time (km/h)
- Payloads per unit distance (# of flights per km)
- Quantity of data per payload (bytes per flight)
> The Snowmobile comes with a removable connector rack that needs to be mounted on one of your data center racks where it can be connected directly to your high-speed network backbone. The connector racks provides multiple 40Gb/s interfaces that can transfer up to 1 Tb/s in aggregate.[1]
so there is a bit of setup and teardown, but they're not plugging disks in one-by-oen
If I were designing it I'd expose a private S3 interface, or maybe NFS, to the customer from the truck, store the incoming bits on-disk in the form that S3 expects, then just copy raw blocks from disk and merge metadata databases on the AWS side so you're not doing two passes through the S3 API. I can't imagine they're doing something much different than that. Not every customer would be in a position to handle iSCSI or FC, for example. You could unload the disks at the destination, too, but that seems less efficient and more disruptive to their lifetime.
> When your Snowmobile is on site, AWS personnel will work with your team to connect a removable, high-speed network switch from Snowmobile to your local network and you can begin your high-speed data transfer from any number of sources within your data center to the Snowmobile.
At the maximal throughput of 1 Tbps, 100 PB of data would still take almost 10 days to copy.
It may very well be “mmapped” onto S3/Glacier API on the AWS side, but a copy is still necessary on the customer’s side.
Would have to re-watch https://www.youtube.com/watch?v=z8-DeSuKf9I to be sure.
And 100 PB is still a lot of data. Facebook in its IPO filing said their entire media library was at the time 100 PB. Backblaze’s total backup size was 150 PB in 2015, just 2 years ago. I’m sure S3 is at exabyte scale now, but even so, 100 PB is not a drop in anyone’s bucket.
> And 100 PB is still a lot of data. Facebook in its IPO filing said their entire media library was at the time 100 PB. Backblaze’s total backup size was 150 PB in 2015, just 2 years ago. I’m sure S3 is at exabyte scale now, but even so, 100 PB is not a drop in anyone’s bucket.
I'm ex-AWS, and used to work on a team closely involved with both S3 and Import/Export (the team that produced snowmobile, though that expansion launched after I left) I'm well aware of the kinds of capacity scale S3 operates at, as well as some other storage teams. Remember that backblaze, for all its scale, operates out of a small number of datacentres, where AWS has operations right across the world.
Being ex-AWS also why I'm familiar with the way AWS Security thinks, and I can easily picture how they'd hit the roof if S3 was "ported" on to the device.
That's before we even begin to tackle the concept of just how hard it would be to port S3 in to a small scale platform as the Snowmobile. Don't forget, you're not just talking about all the things that make up S3, but all the things that make up Amazon infrastructure as a whole. One of the reasons Amazon is able to push out so many services is because they've got a mature and well established ecosystem behind the scenes that is designed to operate at scale.
You'd be talking year(s) of effort to port S3 into a snowmobile, at best, vs. a matter of a few months to port and fully test a layer on top of the device. When it all boils down to it, the S3 API is pretty simple. Why on earth would you choose the most technically complicated way to approach the problem?
Do notice the context, however. We’re approaching this problem from the perspective of trying to determine the network bandwidth of a portable datacenter driving down the highway.
Assuming that container is really just a giant USB harddrive then it’s 10 days for copy at client’s site, 8 hours for the drive, and then another 10 days to assimilate the data back at AWS DC, or more if pushing it to 2 other DCs is slower.
Perhaps there’s some way to alleviate some of those concerns by shipping minimal S3-API server in the Snowmobile, then deploy actual production code when it’s assimilated back home. Perhaps the Snowmobile could become its own region temporarily.
Still, you’re right that is a lot of moving parts for something that you’d want to avoid… Except in this extremely important hypothetical, obviously.
Also, while it is probably true that S3 has way more storage distributed across the globe, I think the Snowmobile still follows the Sat Nav not DNS & BGP ;) I mean it delivers to a single DC, so that’s what needs to be considered.
Actually 100 PB is a lot on disk, but it’s even more over the network. They’ll probably want to push it during off-peaks, but should probably consider to just keep driving the truck to the next AZ.
http://www.seattletimes.com/seattle-news/crime/judge-blocks-...
I couldn't find the discussion but if I remember correctly most satellites shut down over a big chunk of the Indian Ocean because there's nothing there.
Oh, wait. Darnit.
Instead of a movie or few episodes, consider physical media that could store a year or decade of TV shows and movies.
What is more likely to occur in the next decade: increasing density and decreasing costs of storage media, OR, increased speeds at the local ISP level with decreasing costs?
It consists of a terabyte of data bringing together the latest music, Hollywood movies, TV series, mobile phone apps, magazines and even a classifieds section similar to Gumtree or Craigslist.
Every week, unidentified curators compile a selection of content and deliver it via a complex network of hundreds of distributors who, much like old-fashioned newspaper delivery boys, bring the Paquete to the door of its subscribers."
I'm also confronting the possibility that I'm old, and a lot of you whippersnappers have already forgotten about DVDs, CDs, eight tracks, etc. (Though are strangely very tuned in to vinyl.)
So yeah: Please get off my lawn. I'm going back to my rocking chair to listen to recaps of Google I/O streamed directly into my hearing aid.
I grew up with 5.25" floppy disks.
But I realized that maybe I'm asking the whole question wrong. We could actually be faster than those track using even CAT-5 today, it's just that the cost of the installation is prohibitive.
So I guess the proper form of the question would be when if ever the cost of transmitting using electrons/photons will be lower.
There's a giant nuclear fireball blasting half the planet at a time with EM, everything is radiating EM (even on some of the communication bands!), most of the things will diffract or reflect EM, and we don't know of a way to isolate EM using only EM.
So atoms seem to have a better asymptotic isolation in known physics, and thus conceivably might always be faster due to potential noise in the respective channels.
You tell the computer to copy, say, 200TB from Europe to Asia. It could automatically determine that shipping is faster, and so fill up a portable NAS, then use an API to request a transport drone, which then takes it to some ship/airplane/train and the reverse happens at the other point. The whole human involvement could be watching the progress bar.
I guess that's the point of the article the data we humans can now create is so enormous it's impossible to move it other than in bulk physically.