A 1970s Cray-1 hard drive has been imaged
blog.archive.org
blog.archive.org
We do use standard formats, and the standard formats are continually changed, and the formats are not always backwards compatible. It's a nice goal, but it actually doesn't work.
I have in fact electronic information that in fact goes back through many different computer systems. Some of it now I cannot access. In theory I could, or with enough effort, find people to decipher it, but it's not readily accessible. The more backwards you go, the more of a challenge it becomes.
And despite the goal of maintaining standards, or maintaining forward compatibility, or backwards compatibility, it doesn't really work out that way. Maybe we will improve that. Hard documents are actually the easiest to access. Fairly crude technologies like microfilm or microfiche which basically has documents are very easy to access.
So ironically, the most primitive formats are the ones that are easiest.
So something like acrobat documents, which are basically trying to preserve a flat document, is actually a pretty good format, and is likely to last a pretty long time. But I am not confident that these standards will remain. I think the philosophical implication is that we have to really care about knowledge. If we care about knowledge it will be preserved. And this is true knowledge in general, because knowledge is not just information. Because each generation is preserving the knowledge it cares about and of course a lot of that knowledge is preserved from earlier times, but we have to sort of re-synthesize it and re-understand it, and appreciate it anew.
Source:
http://blogs.computerworld.com/the_kurzweil_interview_contin...
The Disk surface shown here, meant to be a guide to the contents, is etched with a central image of the earth and a message written in eight major world languages: “Languages of the World: This is an archive of over 1,500 human languages assembled in the year 02008 C.E. Magnify 1,000 times to find over 13,000 pages of language documentation.” The text begins at eye-readable scale and spirals down to nano-scale. This tapered ring of languages is intended to maximize the number of people that will be able to read something immediately upon picking up the Disk, as well as implying the directions for using it—‘get a magnifier and there is more.’
On the reverse side of the disk from the globe graphic are over 13,000 microetched pages of language documentation. Since each page is a physical rather than digital image, there is no platform or format dependency. Reading the Disk requires only optical magnification. Each page is .019 inches, or half a millimeter, across. This is about equal in width to 5 human hairs, and can be read with a 650X microscope (individual pages are clearly visible with 100X magnification).
The disk still seems to be a work in progress, but the Rosetta project is concentrating on many Internet and audio-based initiatives (see http://www.nytimes.com/2011/07/29/us/29bcculture.html)
One could theoretically take this further by then explaining how we built our primitive computers, some simple math, and continue.
I would love an accompanying Rosetta project that was in just one language, but exposited our understandings of math, physics and computer science so that some civilization that discovered the twin discs could use the first one to learn English (as long as they knew or could decipher at least one of the languages) and the second one to reconstruct our understanding of math, physics and computer science and rebuild a 2000 AD era computer, and finally input to it a tar.gz dump of all of Wikipedia.
Just waiting for the comments asking "what does this have to do with startups" ...
Great article, it's incredible no more Cray-1 software has been preserved.
Looks like it was 140MB at highest density.
Now, excuse me a moment, while I try to get those pesky kids off of my lawn.
The unit has an average seek time of 30ms in 1973. Today, in common 7200 RPM drives, it seems to be around 8-9ms. There is only a factor of improvement of 3 to 4 times.
I find this pleasantly surprising.
This is going to have big consequences for the way that we design systems in the future. Transferring a large amount of data is going to be cheap and fast. Seeking, handshaking, back-and-forth and any other latency sensitive operations are going to be slow and expensive. This has already played a big role in algorithm design in HPC, and is going to start being felt to a much larger degree in the wider field over the next decade. As the latency/throughput ratio gets bigger, the tradeoffs behind optimal system design will change.
30ms for a Cray 1 was only 2400 clock cycles. 8ms for a modern CPU is around 30 million. That's a big change.
2.4 million clock cycles. (According to Wikipedia, the Cray-1 had a 80 MHz clock, which you also mention in your other post.)
Tape seek times (for half a tape) are closer to 40s, which is still 100x more. We'll have to wait a while yet before disk is the new tape in terms of cost of seeks, even when comparing the 70s to today.
That's only considering latency. The cost consideration is a completely different story.
It wasn't rare at all to find fair sized servos (as opposed to steppers in consumer grade stuff in the 80's) in those old disk pack units (the size of a washing machine).
You can't really compare 'common 7200 RPM drives' of today with a top-of-the-line medium from the 70's, physics didn't change at all in that time. That's why there are 'servo tracks' on the drive, they help with finding the right track (in a stepper scenario you don't actually need those, the stepper resolution defines where the tracks are).
Today enterprise level drives achieve < 4 ms average seek times at a price point which is a small fraction of what that drive cost in the 70's.
That's where the real improvement factor sits: performance (both in capacity, seek time and transfer rates) vs cost.
And that's many orders of magnitude.
Now of course the big question: what is on that disk?
Unfortunately, within 30 seconds of the heads being loaded a high-pitched whining noise began to be emitted from the drive, implying a potential head-to-disk contact was taking place. The drive was then powered down and the disk pack and heads were carefully examined. Thorough examination revealed that Head #4 on the drive (which reads the bottom surface of the lowest data platter) had 'crashed' into the disk surface and scraped away a concentric ring of oxide material, permanently damaging the platter. This is a good time to point out the advantages of not experimenting with your primary source material when performing digital archeology experiments!
Src: http://www.archive.org/details/2011-cdc-disk-archaeology-fen...
http://www.phys.huji.ac.il/~springer/DigitalNeedle/
It came surprisingly close. Perhaps with today's technology we could do better?
Q: Why do Cray supercomputers have a clear panel in the front (you can see it in the photos)
A: So you can Seymour Cray.
Whew!
(thank you, Cheshire Engineering!)