Hacker News for Hackers
ontology2.com
ontology2.com
One aspect the author touched on is "I'd like to be suprised by relevant things that I don't know about." It's not immediately clear how to discover serendipitously items that are outside of the things you're already interested in, and therefore would already be in your feed. This is an area I'd like to know more about.
I'm also really keen on someone taking the time to automate curation of their own feed. One can still fall prey to the more negative aspects of Daily Me[1], but at least you're better aware of what's going into producing the feed you're reading and have the tools to update it if you find it's not serving your best interests.
I also don't particularly care for more Apple news, or some political rant, but the discussions that sometimes follow these can be extremely interesting, and enough so that I would regret their disappearance.
1) I found here, on Hacker News, a nice "kill sticky" javascript function to put in my shortcuts
2) In Firefox, I gave it a shortcut keyword
3) In my AutoHotKey script, I assigned a key to "type the shortcut keyword in the address bar" to run it quickly
4) when a pop-in occurs, I reflexively press my kill key
The script is
javascript:(function()%7B(function%20()%20%7Bvar%20i%2C%20elements%20%3D%20document.querySelectorAll('body%20*')%3Bfor%20(i%20%3D%200%3B%20i%20%3C%20elements.length%3B%20i%2B%2B)%20%7Bif%20(getComputedStyle(elements%5Bi%5D).position%20%3D%3D%3D%20'fixed')%20%7Belements%5Bi%5D.parentNode.removeChild(elements%5Bi%5D)%3B%7D%7D%7D)()%7D)()
Enjoy! var i, elements = document.querySelectorAll('body *');
for (i = 0; i < elements.length; i++) {
if (getComputedStyle(elements[i]).position === 'fixed') {
elements[i].parentNode.removeChild(elements[i]);
}
} .overlay {
display: none;
}It makes me sad when I see an interesting looking post, and it turns out that it is only a video or podcast, with no proper writeup. I just don't have time to watch an hour-long video, for content that I could read at my own pace in ten/fifteen minutes.
One of the thing I honestly dislike about comments on Hacker News is the advocation of the No True Scotsman-esque definition of "Hacker," where the only thing that matters is code and how it's used. In 2017, there's more to being a "Hacker" than just what low-level language is being used.
Always has been:
There's an historical pendulum with a cycle of many decades re monopoly/market power regulation which seems to be reversing itself right now. IMHO, this will likely be the most important shaping factor shaping the tech industry over the next decade, especially re small startups. So I don't turn up my nose at "Big Tech Companies Behaving Badly" articles, although I know there's a fair bit of repetition. I do want to know where that pendulum is, and whether it's really swinging back. Yes, it's law that will be the instrument of that change, but it's public perception that will necessitate changes in the law (say re patent misuse) and its implementation.
If you start from keywords you can use NLP (Word2vec) to measure how close articles are to your interests, so for instance "python, machine learning" gives you stuff on tensorflow, scikit learn, etc whereas "java, machine learning" gives you articles on spark, Deeplearning4j.
You can also measure how similar articles are to each other - getting the most dis-similar articles avoids the "piling on" problem mentions.
The ultimate goal was to save me time and effort in manually classifying them, as I and everyone else do when we scan through what we come across on a daily basis. Instead of manually doing that, the hope was the program could do it for me, and I could just focus on reading the interesting articles and not even have to deal with the uninteresting ones.
From that experiment I learned a few things:
- First, that I'd have to manually scan through all the articles anyway, just in case the classifier made a mistake and maybe dumped an article I found intensely interesting in the uninteresting pile.
- Second, that having to consciously think about which article was interesting or uninteresting in order to do the training, about whether the classifier was working or not and which articles needed to be reclassified, about having to re-train it when it messed up, and so on was a hell of a lot more work than just scanning through my RSS feed manually and deciding on which articles I found interesting or not myself.
- Third, my interests were not static things that the algorithm could learn and classify on correctly from then on out. My interests were constantly changing. Sure, maybe there were a handful of things I always found interesting or uninteresting -- but overall what I found interesting or not changed from day to day. It was also kind of unpredictable, even to myself.
The third point kind of argues towards the approach of the HN front page, which is un-classified and un-tagged. I've read all sorts of great articles on HN that if I'd been going by some pre-written list of interests that I had, I would have never have read. I do still often wish for tagging on HN anyway, just because there are certain types of articles that I really never ever want to read, and I'd love to be able to exclude them. But the vast majority of HN articles aren't of that kind (or I wouldn't be here).
That experiment with bayesian classification turned out to be rather short-lived, as I found the whole thing way too much of a bother to maintain and to retrain when it misclassified articles. I'm still reading RSS feeds the old fasioned way today, and am a little suspicious about any AI/machine-learning-like approaches to article classification.
There are a bunch of great curated email newsletters around specific interests (Javascript Weekly, etc) so I'm aiming to do something similar, but more granular. So far the ML thing has been promising, but it helps to start from a pre-vetted dataset.
But I'm sure there are interesting articles that didn't get classified at all. It doesn't seem like the end of the world if you lose a little wheat with the chaff.
> overall what I found interesting or not changed from day to day
This seems like the problem. If you rated an article highly yesterday, it doesn't mean you want to read the same article again today. "Interesting" largely means "novel" and it is hard to find that by looking at similarities to what was new in the past.
But is it losing a little or a lot? No way to tell without looking.
Even if it misses just a little, that little bit might have been crucial. If it misclassifies "NYC NUKED!!!" as uninteresting that's a single mistake that could make me oblivious to a hugely consequential event.
Of course, as a human classifier, I'll doubtlessly making my own mistakes, and maybe the AI classifier could help me out by pre-classifying articles for me, but it could also be misleading and will cost me in terms of spending time on training and maintenance. I'm not really sure what the right solution is here.
That's exactly what I found out a few years back when I tried to do something similar with modern machine learning techniques. The article we're commenting on even mentioned that, for this kind of classification to work, you need to maintain a stable viewpoint over time. I do know some people who maintain stable viewpoints over time, but they don't need fancy ML techniques to find things that are of interest of them. They tend to talk to (and be well known to) other people who have similar interests.
I can't help but feel that a better solution to this problem has already been envisioned and enacted about a dozen times already. Myspace, Twitter, Livejournal, and tumblr all did a great job of showing you things you find interesting by following people who post things you find interesting. The only reason I'm still on hacker news is there's a small group of people who always post things I find interesting. I have an RSS feed set up for each of them. I also positively filter posts and comments on a few topics I always find interesting (like RSS).
The downside to the method I use above is there are dry spells where there's nothing interesting to read. Is that a bug? Or a feature? A decade ago, I would have said bug, but not now. I've never had a fear of missing out, but I have been addicted to the novelty of getting a steady feed of interesting articles, essays and posts to read. I came to the conclusion I was addicted when I would spend hours reading and absorbing information that never gets used and is completely irrelevant to the tasks I need to get done and the goals I want to accomplish with my life. These days, I'm happy reading less and doing more.
https://gist.github.com/ivmirx/66a0015884d44297ea05a8c54d935...
That would be nice but unfortunately in practice all the big news make their way to the top of HN, despite the fact that we've already read and heard about them from countless other sources. For example some of the top posts at the moment are:
- Catalan parliament declares independence from Spain
- New Zealand to ban foreigners from buying existing houses
- How to Read the JFK Assassination Files
One issue I have with pure keyword filters when I apply them to e.g. RSS feeds is that they can't capture details well - e.g. I'd like seeing a detailed article about how V8 works and wouldn't to exclude an otherwise interesting thing just because it is in JS, then you spend quite a lot of finetuning the filters.
As you say, there is interesting deep stuff going on in the JS space, but there is too much average stuff for me to look at right now.
I'd like to be able to credit comments that got me to re-think or even change my position on a topic. Upvotes basically say, 'that's a good point', or 'I agree with you'. But I'd like to see a way to gauge 'influence', in terms of affecting what people think in a positive way.
Can that be worked in?
I would also argue that an upvote is not "I agree with you", but "this contributes to the conversation in a positive manner". I frequently upvote posts with which I disagree. Conversely, a downvote in my book says, "hey, you're kind of being a dick and dragging the conversation down" rather than "I disagree". I mean, I'll downvote something that is just demonstrably wrong, but more often than not it's "quit being a dick".
Sadly, I begin to conclude that right now most downvotes mean: "I really, really hope that isn't true, even though you have a bibliography and I just have an emotional reaction."
Maybe a Markov-chain based reputation system (a la Google for search) that makes downvotes unequal is in order.