Where's my petabyte disk drive?
bit-player.org
bit-player.org
Here is a graph I got from Dave Anderson [director of strategic planning at Seagate] years ago. It shows that what looks like a smooth Kryder's Law curve is actually the superposition of a series of S-curves, one for each successive technology generation. Naturally, because the easy transitions get done first, the cost of each successive transition increases, perhaps even exponentially. Since margins are constrained and so, these days, are volumes, to generate a return on the investment in each transition requires that the technology be kept in the market longer. The longer interval between transitions translates to a lower Kryder rate.
http://4.bp.blogspot.com/-bkuDDrBpcZE/TpMsLTEspsI/AAAAAAAAA9...
0. https://en.wikipedia.org/wiki/Mark_Kryder
1. http://www.lockss.org/locksswp/wp-content/uploads/2012/09/un...
2. http://blog.dshr.org/2012/10/storage-will-be-lot-less-free-t...
and in turn, demand for storage is slowed by lack of need for storage.
The real question should be: Where's my 4K video? My 10-bit-channel images? My lossless audio?
Yes they are here but not nearly as common as content delivered in ancient formats from over a decade ago. For example go on a wallpapers subreddit, the vast majority is still in 8-bit 1920x1080 JPEG. Video is still mostly 1080. The leading music services still deliver in 2-channel lossy formats. And this is 2016.
Most of this I suppose is because of the general slowness of the internet and the usage caps in many areas of the world.
So maybe improve internet service -> create more detailed content and let people save it -> people will want to demand more storage to store all that in.
Besides, pirated video content was really the only thing that most normal people could fill up their hard drives with (okay, maybe GoPro people or people with huge Steam libraries, too), but the Netflix and YouTube and Amazon Prime have taken a huge chunk out of that.
But I think storage space still matters on mobile. I haven't come close to filling up my laptop after 4 years of use, but I have to be careful about putting stuff on my iPod. Still it can store dozens of hours of audio, so it's not the biggest issue.
I can't wait for a photo manager with some kind of AI that goes through the pictures I've taken myself and offers them based on "feel good vibes", "eerie", "cozy" and so on, and learning what I like as it goes.
Apple please?
Not so useful for me, since I don't care to upload my entire photo collection to Google. Something like this that either works offline or can progressively tag photos and insert those tags in a format Lightroom could read would be awfully useful.
Series:
House of Cards
Marco Polo
Breaking Bad
The Blacklist
Movies and Documentaries
Smurfs 2
Philadelphia
Jerry Macguire
Crouching Tiger, Hidden Dragon
Oceans (documentary)
Forests (documentary)
Flowers (documentary)
Edit:formattingEither way it's the content producers and publishers' job to create more higher-quality and make consumers aware of it and make it easier for them consume it.
In this Age of YouTube, some of us can play our part by recording (as some recent phones/tablets can) and uploading in 4K. Design wallpapers in 4K. Even just seeing the higher resolution in the list of available options will increase awareness and pique more people to try a 4K display out.
While we're at it, let's make 10-bit/channel content more common too, now that the latest Macs and iPads have 10-bit displays. Anime fansubbers have already been producing 10-bit videos for a while.
Ever experienced YouTube buffering on an average connection? And that's for the metropolitan US -- consider the developing world, or even rural US, the mobile internet caps, etc.
Not to mention that even if you have enough speed, still you don't get much benefit from 4K video for 99% of stuff out there. Diminishing returns. I should know, I have projected good 1080p video to movie theaters, and nobody would call it bad or inadequate. On an average 15-27" monitor? It's not even an issue...
Said the American. In fact, probably the US doesn't matter in this regard, as it has lower average speeds than most of Europe and SE Asia.
Are you talking about consumer storage? Because need for enterprise storage (which I sell), is only increasing. What is decreasing is the revenue/profit per TB (naturally).
Once content producers make more higher-quality content, and enough consumers start consuming it, then won't the content providers need the extra storage too to serve that content from?
Back then it was possible you could smash the plates and somebody could reassemble some of the data.
Then by 2005 or so the density of the data was high enough that the scanning probe microscope wasn't much better than the read heads, and at that point extreme methods of data extraction got much much harder.
I believe that is an urban legend. http://all.net/ForensicsPapers/2012-12-07-OverwrittenMagneti... describes attempts to track down such cases:
> To date I have found no example of any instance in which digital data recorded on a hard disk drive and subsequently overwritten was recovered from such a drive since 1985, when about 15% of the overwritten data was claimed to have been recovered from an modified frequency modulation (MFM) disk drive.
It cites "Overwriting Hard Drive Data: The Great Wiping Controversy" at http://www.vidarholen.net/~vidar/overwriting_hard_drive_data... which gives a best case example of a pristine hard drive, written once and then wiped once, and where you know the data is located before hand. Even then nearly all of the data had disappeared. If the drive was not pristine, it was not possible to recover the data. Quoting from it (emphasis mine):
> The purpose of this paper was a categorical settlement to the controversy surrounding the misconceptions involving the belief that data can be recovered following a wipe procedure. This study has demonstrated that correctly wiped data cannot reasonably be retrieved even if it is of a small size or found only over small parts of the hard drive. Not even with the use of a MFM or other known methods. The belief that a tool can be developed to retrieve gigabytes or terabytes of information from a wiped drive is in error.
> Although there is a good chance of recovery for any individual bit from a drive, the chances of recovery of any amount of data from a drive using an electron microscope are negligible. Even speculating on the possible recovery of an old drive, there is no likelihood that any data would be recoverable from the drive. The forensic recovery of data using electron microscopy is infeasible. This was true both on old drives and has become more difficult over time. Further, there is a need for the data to have been written and then wiped on a raw unused drive for there to be any hope of any level of recovery even at the bit level, which does not reflect real situations. It is unlikely that a recovered drive will have not been used for a period of time and the interaction of defragmentation, file copies and general use that overwrites data areas negates any chance of data recovery. The fallacy that data can be forensically recovered using an electron microscope or related means needs to be put to rest.
> if you looked close you might find the edges of bits that had been written before and weren't perfectly aligned.
GGP spoke about both recovering data that had been overwritten and re-assembling data from destroyed platters without the original drive mechanism.
I didn't look too into the question of how to recover the contents of a hard disk with microscopy because I figured it would be possible, but expensive. Looking now, I quickly found a MS thesis at http://escholarship.org/uc/item/26g4p84b which recovered data from a disk using MFM. While the performance was poor, the author attributes that to the experimental setup.
Ahh, and http://www.dataclinic.it/magnetic-force-microscopy.htm appears to provide a commercial service to extra data from a hard disk using magnetic force microscopy.
So in practice in production it's more useful to have smaller hard drives in more places to work on the data in parallel. And in the truly archival cases there are other concerns (like redundancy) that mean there isn't as much demand for a single massive drive.
In my case I was trying to get him to sign off on powering down some of the drives that could not be reached to save power. But even with the data staring him in the face he could not go there. Network bandwidth gets better, and that exposes more data to the pipeline, but if you want < 500mS request response you have to balance the system.
I don't think the economics works out, since at 1G/s it probably takes too long to load the data, and as this essay points out most people will stream what they want on demand. I also doubt there will be a standard content set which is around long enough to assure that my imagined on-the-fly compression model-building-by-corpus-reference will take root.
There was an estimate in 2011 that total storage of everything everywhere by everyone was more than 250 Exabytes, increasing by around 25% annually.
There's going to be a lot of duplication in that, and a lot of it won't be public. So as a ballpark guess a complete collection of public-only sources - including all available commercial content of all possible kinds ever recorded, academic papers, Wikis, news sites, forums, and such - is going to need 25-50 Exabytes, with maybe 25% compound of new content every year.
So you could get the entire Internet delivered by truck or two, but you probably wouldn't have anywhere to put it.
You could have indexes of hashes and store any chunks of anything in a big flat address space. You wouldn't even need to know what you have. Just a massive amount of archived chunks of storage. (OK, that is more than hand-waving, maybe arm-waving?).
Your premise is that anything you don't use on a regular basis belongs in archival storage, presumably in some kind of central archive.
Suppose there is 1PB of static data of which you access a different 50GB every day. You can call it archival if you like but it's still going to save you 50GB/day of network traffic to have a local copy.
So for example, Netflix could make a box that came loaded with all their content and new content is added using IP multicast or P2P during off-peak hours. The peak hours bandwidth savings would be immense and you would be completely immune to crappy or unreliable network connections.
Or, if you care capacity-constrained, you can use cross-server erasure coding. And even call it cross-server RAID, if you like.
Just text (loads instantly):
http://webcache.googleusercontent.com/search?q=cache:Jr34hZX...
Images etc. loaded from the site (seems to work, albeit slowly):
http://webcache.googleusercontent.com/search?q=cache:Jr34hZX...
I found the problem.
/s
Though, it does work really well.
I use Apache but falcolas' point of nginx (often in combination with Varnish and potentially hhvm; I've used this combo before with great success) is worth considering as well.
No replacement for good configuration of your database (try MySQLtuner.pl if you use MySQL/MariaDB) of course.
Now, when it comes down to it, do you really need to backup the OS? or your installed software? if the machine and the backups fail, you're going to be reinstalling anyway (probably) So, for the third copy i rely on 3rd parties. Different people have different needs, you might want to do something fancy in house.
It pretty much boils down to finding a service for your stuff. I have a couple of private github repos. Photos on iCloud and whatever Alphabet is calling Picasa these days. 20 gigs of music to Amazon or Alphabet (or both). Administrative stuff, like taxes, i just email to myself. It's probably smarter to keep that in dropbox or something along those lines.
The key point is, there are the things you make or capture that are irreplaceable, save those lots of places. There's a bunch of other crap on your computer to make it be useful. That stuff is trivial to reinstall. Well, ok, it might cost you a day or two to redownload and reconfigure emacs just so - but with a little planning you can put that config in git, so it's easy to restore or set up on a new machine.
It's almost better to think in terms of, if i had to upgrade tomorrow, what would i need to copy over? that's the stuff to be really fussy about.
I'd counter this with what I do with my laptop. The OS is considerably smaller than the data I actually care about (<20gb), it takes almost no time to backup and so it leaves me with a very quick ability to restore the system to a known good state in the event of some kind of failure. I don't do constant backups of it, but maybe once a month i'll update the backup I have of the OS.
It probably doesn't really matter as much as your intuition might suggest anyway because each drive can and will fail. And they fail at similar rates (as opposed to an exponential difference of a factor of 10x or more). So you need to take similar precautions for each type of drive.
For personal use - I use exclusively SSDs because they are much faster. The I put all the information I don't want to lose in dropbox.
For servers, all data that is important goes in a database cluster (Cassandra) with a replication factor of 3. Those drives are backed up daily offsite. For data that cannot be lost at all (even a days worth), I also copy each record to Amazon S3 every time it is changed. - I'm sure there are many other ways to tackle this problem.
https://www.dropbox.com/en/help/11
So, as long as you notice before the clock runs out, you should be able to recover from ransomware via file history.
[I work for Dropbox, but am not speaking on behalf of my employer.]
I made a page that shows hard drives and SSDs sorted by price per TB: https://edwardbetts.com/price_per_tb/
http://arstechnica.com/gadgets/2015/08/samsung-unveils-2-5-i...
> As the pace of magnetic disk development slackens, an alternative storage medium is coming on strong. Flash memory, a semiconductor technology, has recently surpassed magnetic disk in areal density; Micron Technologies reports a laboratory demonstration of 2.7 terabits per square inch. And Samsung has announced a flash-based solid-state drive (SSD) with 15 terabytes of capacity, larger than any mechanical disk drive now on the market. SSDs are still much more expensive than mechanical disks—by a factor of 5 or 10—but they offer higher speed and lower power consumption. They also offer the virtue of total silence, which I find truly golden.
This was done back in the 80s in http://www.cis.upenn.edu/~KeyKOS/ . A favorite demo reportedly was to pull the plug on a running computer then start up again. They took the need to redesign security as an opportunity to make it better.
I don't think this is just a security issue; it really breaks all of the assumptions that we like to make in modern programming languages.
I think he was rather saying that the OS could do it: persistent virtual memory as the primary abstraction. In Unix, files and processes are different kinds of things; in KeyKOS there were only processes; RAM was effectively a cache. As Unix directories have links to files, KeyKOS processes could be given capabilities to invoke other processes (passing capabilities and data as arguments). The different security model makes this analogy misleading, but you can see how you could emulate a filesystem.
What assumptions do you mean?
[1] Johnny Mnemonic - Official Trailer - https://www.youtube.com/watch?v=Uwl5MBzTCRQ
[2] Also, I guess Toshiba is in a lawsuit with Apple over their new VR headset - the EyePhone - https://www.youtube.com/watch?v=vXSqN7qXwpU
This is similar to how a cellphone camera of today can shoot video in a single minute that would completely fill hard drives from the 90s. (And even a single high-res image from today's digital cameras is larger than the install size of Windows 3.1.) Back then it would have been difficult to imagine these uses for storage.
http://techreport.com/review/27909/the-ssd-endurance-experim... "All of the drives surpassed their official endurance specifications by writing hundreds of terabytes without issue."
I expect to upgrade for increased capacity long before I reach that.
Well, it's coming either way, in the form of large scale NVRAM
that being said, for the slashdotting issue, I can't see how bad the graph is
50 gig per hour for a cinema screen quality setup (Most houses in the next 20 years...) would be 20,000 hours of entertainment, meh I might want access to that in a life.
Also remembering we might be heading towards an environment where we record everything at all times.
Certainly currently I'm buying a hard disk every year as quality goes up and it's easier than throwing stuff out.
Closer we get to Petabyte HD the better I say.
Probably not
As long as folks keep on reinventing the wheel, only bigger, hard drives are going to have to keep increasing in size.
What country are you from?