Stephen Hawking's Ph.D. Thesis Crashes Cambridge Site After It's Posted Online
npr.org
npr.org
A 72MB file being served 500,000 times over a 24hr period is 3-4Gb/sec.
The department probably has one or more web servers serving content off a network drive. Your average "rack of network storage" won't even blink at that.
There was probably a gigabit switch somewhere or that was being a bottleneck or the web server was simply misconfigured for this task.
Edit: It evaluate to `document.write('HIS_EMAIL at gmail')`
The vast majority of crawlers aren't that smart, headless browsers used to be painfully slow so no one used them.
In fact doing so probably violates the halting problem somehow.
If your server is configured for a low number of concurrent connections, it can easily seem swamped if a small number of people are concurrently downloading a large file slowly. All that's happening is that it's not accept()ing new connections until existing ones finish.
It's a java application with XLST in the page rendering pipeline, so not exactly optimised for speed.
What brought down the website?
The Greys.I'm guilt of downloading and hoarding things that seems interesting and never getting round to even opening them. "When I'm retired", I tell myself.
I wonder if they faxed it at some point too?
Note: Didn't go to the link, so just guessing.
This is Stephen's Hawking's PhD thesis. It is dated 1966. How could he have entered it digitally? UNIX research didn't even start until 1970s.
I believe this comment chain thinks the thesis was written recently. So you guys thought Stephen Hawking didn't have a PhD till now?
I quoted UNIX to show how primitive technology was at the time, rather than say he could've possibly used a computer. With ALS, using a computer would as hard as writing I would presume.
HN's guidelines say that one shouldn't ask if one has read the posted link, but I'm tempted to all the time.
But even so, while 1966 was indeed early for "regular" use of fax - the first "user friendly" Xerox fax machines hit the market around then -, the first transmission of facsimiles of images dates to the 1840's, and the first fax that used similar methods to "modern" fax machines of scanning line by line (the "scanning phototelegraph") dates to 1880. Commercial fax machines have been around since around 1900.
So it would indeed be possible.
One weird and wonderful product of early faxes (fax over radio predates "wired" fax machines): Finch Facsimile's [1] were used to transmit "newspapers" via AM radio in the '30's, that was then printed on thermal paper at the home of the subscriber.
From [1]: "Six hours overnight was enough time to print a six page two column news bulletin, delivered in time for breakfast."
[1]: http://www.theradiohistorian.org/Radiofax/newspaper_of_the_a...
"What do you mean everything?"
"TV shows. Movies. Even the japanese ones."
"How about older stuff"
"I've gone through early Chaplin work. I've seen Metropolis 17 times."
"I'm sure there's something"
(grabs his friend)
"You don't understand, Paul. I've been reading Shannon. A Mathematical Theory of Communication. I've run out. I've taken to begging strangers for a fix."
"That bad?"
(Guilty), "I just... " (resigned) "I just asked someone to put up a torrent of Stephen Hawking's Ph.D. thesis..."
Even the Eastern European ones. And all of Ingmar Bergman.
(shudders)
Magnet link:
magnet:?xt=urn:btih:e5878b9cdd55286310135419a69371a31195a32a&dn=PR-PHD-05437_CUDL2017-reduced.pdf&tr=udp%3A%2F%2Fexplodie.org%3A6969&tr=udp%3A%2F%2Ftracker.coppersurfer.tk%3A6969&tr=udp%3A%2F%2Ftracker.empire-js.us%3A1337&tr=udp%3A%2F%2Ftracker.leechers-paradise.org%3A6969&tr=udp%3A%2F%2Ftracker.opentrackr.org
If you don't have a torrent client: https://instant.io/#e5878b9cdd55286310135419a69371a31195a32a
Anybody know of mirror sites? A basic web search doesn't show any, and archive.org doesn't show it.
Cambridge's network probably isn't as hardened to spikes in traffic since they don't get much traffic. But still, it isn't 1995. They should have some form of load balancing or distributed/clustered web/data/file systems to handle temporary spikes in traffic and data requests. Serving simple static data isn't something that should "crash the site".
Even without issues, it often felt a bit sluggish when serving locally. The pages are quite large, and the whole pipeline from content -> webpage is rather tedious.(Java, XSLT -> html)
It shouldn't have happened - but I assumed it would.
disclaimer: I am a former contributor to the project [1]: https://github.com/DSpace/DSpace
Sample:
This implies that the universe is spatially homogeneous and isotropic since there is no direction defined in the 3- space orthogonal to Ua. In this universe we consider small perturbations of the motion of tl1e fluid and of the '.ifeyl tensore 1 Ne neglect products of small quantities and perform derivatives with respect to the undisturbed metric. Since all the quantities we are interested in with the exception of the scalars, µ, ~' e have unperturbed value zero, we avoid perturbations that merely represent coordinate transformation and have no physical significance. To the first order the equations (1) - (4) and (7) - (9) are
Stephen Hawking and his dissertation are high-profile as these things go. The NPR mentions other popular items generating 100s of requests per month. I've run across items with lifetime request counts in the double or triple digits frequently (and suspect I doubled the count on one particular item).
More often, though, the truth is that this material simply isn't available online. There are several thesis repositories (either Michigan State or University of Michigan are one, as I recall), and I can frequently turn up a shelf reference via WorldCat ... somewhere.
But there's work from surprisingly prominent names in numerous fields that simply isn't available in electronic format. The worst case is for materials from rougly 1924 - 1980: to late to be out of copyright, and too early to have been composed, or converted to, digital formats (and 1980 is an early cut-off date for that, though it's when material seems to start appearing in bulk).
This includes PhD dissertations, Masters theses, and numerous academic or other writings, often including government documents not under copyright. Thankfully with Sci-Hub, actual published academic journal articles can be found, freely, with a very high success rate. Particularly painful for me are popular magazine and newspaper items, for which even the indices are very frequently locked behind site-restricted or affiliate-only access.
The time-and-effort differential of being able to look something up online, vs. travelling many miles to a facility for access, is tremendous. And it absolutely stops a great many incidential queries dead.
See Rick Falkvinge's excellent rant about how the KRACK vulnerability was blocked behind corporate-only paywalls for over a decade:
https://www.privateinternetaccess.com/blog/2017/10/the-recen...
Note that the issues here are twofold. One element is the task of scanning and making available documents, and organising the results in a manner useful for search.
But much the harm is the direct consequence of the present regime of copyright and paid access to information, AS WELL AS the perverse incentives of advertising-backed media and media manipulation have created a media regime that is actively harmful to society.
I'd really like to see the elements of this addressed.