"Unknown or expired link." - Why?
google.com
google.com
(= dead-msg* "\nUnknown or expired link.")
(defop-raw x (str req)
(w/stdout str
(aif (fns* (sym (arg req "fnid")))
(it req)
(pr dead-msg*))))
If the fnid (function ID) isn't in the fns* list then you get the dead message. (def flink (f)
(string fnurl* "?fnid=" (fnid (fn (req) (prn) (f req)))))
In many places in the code closures are used to handle requests (see the flink code there). If the fns* list is cleared (say news.arc is restarted or harvest-fnids kills them) then you'll get the message.The use of closures in this manner means that the code needed to handle say a form submission is really compact and set up when the form itself is generated.
And if you don't care you might as well not add authentication because without signing it's just a fancy CRC - ie, totally replicable by an attacker. As cookies and links are sent over TCP there should be vanishingly few errors in transmission - you're far more likely to introduce false positives with buggy code, and ...
You need to sanity check your inputs anyways. Just do it. This is also how you avoid bugs normally.
Yes if you want to trust it you have to sign it and make sure you implement all the crypto correctly. But I don't see a need for that here.
Also TCP's checksum sucks.
How many corrupted web pages do you see because of CRC failure in TCP?
In a scary way
I would be interested to know if this is even possible to determine from a macro which variables are bound in the closure.
In particular, harvest-fnids has as maximum number of allowed fnids. If there are too many, it purges any fnids that are older than their expiration time, and the oldest 10%.
Thus, the more fnids created (i.e. the more users), the sooner fnids will get harvested and you'll get the expired error.
To see the potential, look at this code snippet from an academic paper on the topic. The web server presents a form asking for a number, then presents a form asking for another number, then displays their product. This technique makes event-driven web applications feel (to the programmer) like sequential imperative programs.
;; main body
‘(html (head (title ”Product”))
(body
(p ”The product is: ”
,(number→string (∗ (get-number ”first”) (get-number ”second”))))))
The paper: http://cs.brown.edu/~sk/Publications/Papers/Published/khmgpf...Using closures to store state on the server is a rapid prototyping technique, like using lists as data structures. It's elegant but inefficient. In the initial version of HN I used closures for practically all links. As traffic has increased over the years, I've gradually replaced them with hard-coded urls.
Lately traffic has grown rapidly (it usually does in the fall) and I've been working on other things (mostly banning crawlers that don't respect robots.txt), so the rate of expired links has become more conspicuous. I'll add a few more hard-coded urls and that will get it down again.
You should hard-code that one too.
Edit: I investigated further, and actually you're right, the problem was due to caching. It should be better now because we're not caching for as long. But I will work on making login links not use closures.
I ask because I'd love to be able to make a claim like "even Hacker News, which is written in a Lisp, managed to implement a modern password hash".
I see that newer versions of Arc run on Racket, but I have no idea if that's what HN is using or not.
I haven't seen a scheme powered PBKDF2 implementation so I'd guess that's out.
The only other expensive KDF I can think of is scrypt, but I would be pretty surprised if that's got a scheme implementation.
Of course, I guess pg could have decided to call out to the OS to run any of those functions too.
If not, what was the design goal?
If slowing down web login attempts isn't part of it, why not get a dedicated auth server and offload the crypt stuff onto it?
And if it is the goal, you could use CPU-friendly sleeps on the front-end to give increasing delays to the repeated guesser.
Probably: http://codahale.com/how-to-safely-store-a-password/
Hashing functions designed for speed are absolutely the wrong thing for passwords.
But I don't see the need to do the processing on the web servers.
When the HN server starts running out of memory, it drops entries from this table. When your browser asks for an entry that is no longer in this table, you get the "Unknown or expired link" error.
This is a crazy design, but unless someone would like to patch the source code and get PG to accept it, we're stuck with it.
In other words, there's a problem here, but it's not the programming model that pg chose.
Furthermore, the technique of holding important state authoritatively in memory like this is not a good web-development practice for various reasons. Doubly so if its state data which can be round-tripped. Links should not break when the web server or cache (I'm not sure which one it is) runs low on memory. So yes, there is a problem with the programming model that pg chose.
Your point seems to be that since Hacker News is a "hobby project," that we may forgive sacrificing a bit of robustness to make the programming exercise more pleasant. That point was not clear to me from your original posting. Rather, the point seemed to be that the technique was good because it was clever and fun, and I disagreed with that sentiment.
PG seems to be saying elsewhere that it was used as a rapid prototyping technique. That seems to be a fair justification of the technique, in my estimation.
Not necessarily. The way I would approach this is to keep the link cache in memory, but have the links contain the minimal necessary state to reconstruct the link from disk-based storage in the case where the cache is gone. That gives the same excellent median performance without any breakage.
It's not a cache, and you clearly don't understand the issue. These are callbacks to closures, embedding the state in the URL is what they attempt to avoid because doing that is tedious.
The fact that embedding state in the URL is tedious is neither here nor there, but in any case it's a pretty garbage excuse. You know what's more tedious than writing code to pass a few integers around in links? Thousands of people losing carefully written paragraphs of enlightened prose on a regular basis. In fact the more carefully considered, the more likely the text is to be lost. If it took 24 hours or even 12 hours for links to expire then maybe you could justify the approach, but it seems to be well under an hour on average before a given closure is purged. This site seems hardly so complex as to be gaining much from a pure continuation approach, and if you can't ease this problem in a lisp then are all of us building services for non-hackers doomed to life of bitter tedium?
Anyway, you're totally right that it's his prerogative to build a site however he sees fit, and it's my prerogative to leave, but it's also my prerogative to complain about it and call it half-assed. I do build websites as well, so I'm not just armchair commenting.
Poof the closures are all gone; the state of the articles and comments are of course rebuilt from disk state on an as needed basis.
> I do build websites as well, so I'm not just armchair commenting.
I appreciate that, I just use a similar framework and understand why one would choose to use callbacks and not bother ever replacing them. It has to matter enough to bother and to pg, it doesn't yet.
To you, but that's a value judgement, it obviously was the other way around for pg. Had he not taken those shortcuts, there would be no hacker news at all; be thankful he automating that boring stuff and bothered to build the site.
Building broken software is always much easier than building robust correct software so this is hardly a good argument.
I don't think you grasp the issue... the link is expired because the callback no longer exists to link to.
The real issue is that there's a whole bunch of saved state on the server for operations that could be (and should be) completely stateless.
I understand why it's designed this way -- he's got a tool meant for more complex tasks than what it's used for on this site. He used that tool because it's what he knows. But I really can't see it saving that much programmer effort in general.
Obviously replacing closures with directly linking is more difficult than just using direct linking in the first place.
Hell, you can get routing for free in PHP if you just name your files like: comment.php, topic.php, upvote.php, downvote.php, etc.
Look, I use both styles daily and I'm telling you it'll be over my dead body before I allow someone to take closure style ugly links away from me. You aren't going to convince me that manually routed URL are always preferrable.
Closures create a resource for every closure instance which isn't very scalable or efficient -- it is in fact the problem with this very site. It might be useful for building complex applications but it's just a liability and a waste for something as simple and busy as HN.
I'm trying to convince you that manually routed URLs are always preferable but they should be for a site like this.
You're right that pg did it this way in the beginning to save time but years have gone by and the site is now more central to his business - especially as a tech demo. This back and forth argument presupposes that there isn't a better fix than the naive 'use old-style code' solution.
Anyways, the discussion is worth having. Only by pointing out problems do you fix them.
Given that the HN code was written by an increasingly busy man in his spare time, his use of an unreliable but quick-to-write implementation technique may be an acceptable trade-off. But it doesn't make the design any less crazy.
No, the code works fine for an acceptable period of time, it doesn't just not work.
The software stores the current state in a closure. The closure gets cached. When the cache is full, the older closures get flushed, hence the error message.
Getting user experience right depends on the users. I wouldn't use this technique in an online store. Random online shoppers would be confused by expired links, and you'd lose sales. But HN users aren't confused by them. What HN users care about is the quality of the stuff on the site.
Since I can't work full time on HN, I focus on the things that matter most. What I spend my time thinking about is e.g. detecting voting rings. Those affect what you see on the frontpage, which is what users of this site care most about.
That said: I come for the community - and the community has obviously noticed that the site occasionally throws up an annoying 'error'. The fact that you've done something cool programatically has no bearing on what I get from HN.
So far I'd rate the user experience of the site around 3/5 and the content 5/5. You don't need to work any more on the content unless it starts dropping!
I could care less if this issue got fixed. I have never once felt, "man, I'd definitely jump to another site if it didn't have this expired linky thing happen".
(Now, politics stories on the front page, on the other hand... I've often wished for a site with as good a crowd as HN but without the politics...)
I'd rather have any kind of cool new feature, like messaging, than have this fixed.
TIL Windows has always been an amazing user experience. Just look at their numbers.
> But HN users aren't confused by them.
"About 2,200,000 results" -- says they are.
I have a question: What advice would you give to one of your YC startups if they were having this same issue?
"Ah, your users won't be confused. Ignore all evidence that says they are."
I can tell you exactly why people continue to use Hacker News. It's because it's YOUR SITE. They put up with the broken web design where links die after a few minutes.
Your site won't get beat out by a competitor because it has you, and you fund people. So people continuing to use the site is orthogonal to whether the user experience is any good. You have lots of feedback indicating that it isn't.
Aren't you the one who says "listen to your users"?
Obviously, it makes sense for someone who is versed in the news.arc internals to fix the problem; nonetheless this issue certainly bugs me.
1. Open Hacker News
2. Go to lunch
3. Come back from lunch and click next
Every. Time.
I have to ask, and I'll probably get down voted to hell because I'm naive or something, but what is so elegant about a coding technique that breaks under normal usage conditions? If I put out a customer facing piece of code, especially after 4 years, wouldn't it make sense to use an "uglier and less flexible but more efficient alternative" that doesn't break?
I understand your previous explanations of why this happens and of rapid prototyping etc. But at what point does the architecture actually get changed to eliminate this bug?
What's good about this technique, and about rapid prototyping in general, is that you can write an initial version quickly in very little code, then gradually make it more efficient as the demands on the app increase.
The rate of expired links says more about how busy I personally have been lately than about the desirability of storing state in closures.
My goal for a web site or web app is to have 0 expired links. Sometimes stuff you link to outside your site will go dead, and it must be fixed or removed or whatnot. But for your own internal stuff... I don't know... something doesn't feel right about an architecture that allows that systematically. How much time could you save if you didn't even have to worry about fixing any expired links? Any idea on what the ROI on your time would be?
Anyway, just thinking out loud. Thanks again for the site though. I do indeed enjoy it very much regardless.
Is this true? If so, where can I find the code?
This only happens on HN (at least to me).
Else couldn't one set up malicious scripts to up-or downvote many stories, or post comments under someone else's name?
EDIT: The only way I was able to login was to use the 'add comment' button on this post.
It also reliably happens clicking the "next page" link on the bottom of the front page; by the time I'm done reading the front page the next page link usually expires.
Please fix?
Perhaps one visitor is no great loss -- I'm hardly the personality around here that someone like patio11 is -- but I hope my contribution is constructive, and my comment scores have always suggested so.
However, subjectively, it seems like the quality of posting and voting has taken a sharp nosedive since the "Unknown or expired link" problems have become a several-times-per-session occurrence over the past few weeks. I can't help wondering whether long-standing regular contributors are being put off as a result. If positive contributors can't even log in to refute an objectively incorrect post with a verifiable link or downvote Redditesque diversions, a downward slide seems inevitable, and then the loss of high quality posting and voting becomes a self-sustaining decline.
That's not really the point, though, is it? The important thing is whether posters who want to offer a useful comment and/or mitigate a poor comment can do so. Once HN gets into unknown/expired mode at the moment, it seems common that even basic things like "More" links and logging in can fail as soon as you load/refresh a page, at which point the site is effectively unusable: you can't contribute even if you have something worthwhile to add saved away in your clipboard from the previous failed attempt.
Firefox: https://addons.mozilla.org/en-US/firefox/addon/lazarus-form-...
Chrome: https://chrome.google.com/webstore/detail/loljledaigphbcpfhf...
On the other hand, new users who might not be accustomed to The Way We Do Things Around Here would be more likely to get upset at the superficial inconveniences and leave.
It's probably even the case that improving the site or adding features to it would work against its best interests by making it more accessible. HN's implementation is such that the more traffic it sees, the more frequently those errors will occur -- and as an unexpected side-effect, the popularity and instability of the site will work against each other until equilibrium is reached.
A small barrier to entry like Reddit's spartan design or MeFi's $5 fee can go a long way toward delaying the onset of the entertainment-seeking masses.
It's a fine line to walk and it may not be wise to rely on programming bugs to point the way.
Besides the actual annoyance of the error, what's extra rankling is it is an example of privileging a neat trick over user-experience, which is one of my Least Favorite Things Ever that programmers tend to do.
(As an aside, I am extremely skeptical that increasing rates of this error occurring will help keep the original user community of the site -- it seems equally likely that longtime users will just get fed up and wander off.)
1. http://arclanguage.org/install
2. Extract arc3.tar
3. See news.arc
However, AFAIK pg forked from the latest public news.arc, so the current Hacker News platform is not open source.Wow, not sure why this is so deserving of downvotes. Trying to find a source, but I thought there was a previous discussion of many HN pages being statically generated and served quickly with links that expire after a certain time (or become invalid because of what may happen on the server side of things). But oh well.