Later on, the story reaches the front page on HN but it's a new HN link. So if I include the 'original' HN link, it's not useful to the reader, and may even mislead them into thinking there was insignificant HN discussion.
For example, I linked this story:
[Web Applications from the Future: A Database in the Browser](https://stopa.io/post/279)
in early May but it only recently garnered significant discussion a month later in early June (https://news.ycombinator.com/item?id=27424496).
So I'm not entirely sure how to resolve this.
In the case of your example https://hn.algolia.com/?query=Web%20Applications%20from%20th...
But not sure if that’s a good idea or not. Just suggesting it as one alternative to consider
I'm gonna think on how to solve this conundrum in a way that satisfies everyone. This is one data point I'll think about. Thanks for bringing it up.
1. It's interesting to see that a story has been submitted many times, for instance. This can be seen by clicking on the domain name in brackets after the story
E.g., a few days ago I posted a link to Heidi Howard & Ittai Abraham's 'Raft does not Guarantee Liveness in the face of Network Faults'
It was the fourth time this 6-month-old link had been posted, each time attracting upvotes but not enough to get on the front page. (I think it usually takes about five upvotes for a story still on the first page of 'new').
https://news.ycombinator.com/from?site=decentralizedthoughts...
2. Given that your newsletter mostly contains links that are less than a week old, the discussion link will still be live. You seem to think discussion is most likely to happen on a resubmitted link: I'm guessing your newsletter's readership is unlikely to be large enough to get the regular kind of front-page HN discussion going for many of the stories in your newsletter, but the opportunity to comment can still have value even if only a couple of comments are posted.
3. If a story is well-written, I generally like to glance at other things the author has written, and looking at stories that have been upvoted here on HN is often a good place to start.
It uses a giant Bloom filter with every article ever submitted to preserve user privacy.
If you normally enjoy being linked back to HN discussion, maybe you will like it!
Direct string comparison of the current URL to previously submitted ones doesn't work because there are many ways for two identical web pages to have different URLs. For example, the URL fragments can differ (the part after the "#" that may or may not be present). Also there can be tracking parameters (often—but not necessarily—prefixed with "utm_"), which don't change anything about the page. But the URL parameters can't be entirely disregarded because sometimes sites, forums in particular, rely on them – consider pages that use an "?id=..." parameter for different pages. Thus some parameters should be removed, but some shouldn't. The same website having different domains (or domains that change over time) further complicates the situation.
My solution was to "canonicalize" URLs by transforming them into a simplified form using some pretty rough heuristics for common sources of noise. The Python code to do that is here: https://github.com/jstrieb/hackernews-button/blob/master/can...
All of this to say that even though I've used my extension for months and have been quite happy, there will inevitably be false negatives.