Show HN: Blogosphere – Discover independent technical blogs
bilbof.com
bilbof.com
I made blogosphere last Summer in order to improve discovery. It’s seeded with personal blogs that have at some point appeared on the HN front-page.
The intent of the site is to recommend blogs and articles to you, with the objective that it’ll get better the more you use it.
The site uses a co-occurrence matrix to recommend articles and bloggers to you based on the articles you ‘star’ and topics you follow. If you’re not logged in, the site will show you popular articles. It doesn’t track what you read, and doesn’t have google analytics. I have some ideas for improving recommendations (i.e. people like you also like.., and improving the auto tagging of articles), but I thought I’d get feedback on whether people want this first.
The site is at an MVP stage, there will be bugs and missing features, but feedback on whether this is interesting to you would be useful.
The time consuming part was manually auditing blogs for quality (and to ensure it's a personal blog, not e.g. a BBC feed). I came up with some heuristics (i.e. blogs with github.io subdomains that got upvoted were often good so I blanket approved those) to speed things along, but it did take a while.
There are various recommendation approaches I'd like to try. If the site had more users (beyond the ~4 friends I showed it to) I'd like to try out a collaborative filtering approach (https://en.wikipedia.org/wiki/Collaborative_filtering).
Just wanted to say thanks for minimizing tracking of users.
> it’s seeded with personal blogs that have at some point appeared on the HN front-page.
also this https://hnblogs.substack.com/ newsletter (than I manage) send you everyday with non commercial blogpost of HN of the day, if that can help seeds your aglo
1. Repeated content.
I don't want to read what I have already read from a different place. I built a simple LSH based filter. I experimented with few ways to sort out text and process it. It worked.
2. Filter controversy.
I tried Bayesian fitler initially and moved to logistics regression using tf-idf. I settled on Bayesian because my dataset became very expansive. I used news-site corpus and manual entries from reddit/HN. I used sentiment analysis using a dictionary but it worked only in very specific cases. I do like some controversial and pessimistic content.
3. Filter clickbaits.
I couldn't filter the clickbait and gave up. There are ton of clickbaits on HN which I loved after I read them but ton of them are terrible and a huge waste of time. No reliable way to distinguish based on an article too. Length is not a good feature, negativity is not a good feature (I like to read strong opinions from say, founder of an open source analytics company criticising a big company for malpractices and how they fix those), sentence complexity is not good feature, and ton more.
4. Relying on user input is bad.
I read ton of nonsense everyday that I could go without knowing. I click on those links and that is a not me saying you should show me more of that. I don't want to do manual work of training something either. It's friction and I don't like it.
anyway, good luck on your site!
I found random posts from people's garages and amazing projects, one of them being a dad's nanosecond counting clocks to prove time dilation to his kids over the weekend.
I was pretty sure I bookmarked it, but I don't find it anywhere now.
Web rings are another thing that comes to mind. They were super popular in the 90s and often very niche. A site joined a ring and placed a widget on their page and you could click it to jump from site to random site in the same ring.
I think about these things and wonder why there seems to be fewer tools these days for that kind of random discovery. Then again, maybe I'm just nostalgic for a time in my life in which I could spend more time wasting time at a computer.
But there is so much noise, so very very much noise and utter crap out there that it is terribly hard to find what we are looking for.
I sympathise with you, I used to be proud in my Google-fu, saying if I came across something on the internet, I can find it again, but that's just not true anymore. It's not finding needle in a haystack anymore, now its finding needle in Mount Craperest.
Search engine : https://wiby.me/
Time dilation post : http://www.leapsecond.com/great2005/tour/
Search engine : https://wiby.me/
Time dilation post : http://www.leapsecond.com/great2005/tour/
Another recent one that came up is https://news.ycombinator.com/item?id=26618000
Is the one you look for listed in one of the threads?
It was wilby.me
This was the time dilation post : http://www.leapsecond.com/great2005/tour/
Thanks a lot for making this.
Something I'd love: a search engine for personal websites.
https://git.sr.ht/~sircmpwn/openring https://drewdevault.com/
One thing I thought would be cool would be an API for similar blogs or articles, pretty much opening up an internal part of blogosphere. I think one could use that to provide a webring.
But I can’t create an account on mobile (iPhone Pro Max). :(
My focus so far has been making sure recommendations are good enough, but it would be nice if one could use the site on a smartphone.
Thanks again.
One recommendation I have is Jacob Kaplan-Moss' blog: https://jacobian.org/index.xml