Show HN: The Sample – newsletters curated for you with machine learning
sample.findka.com
sample.findka.com
Some of the newsletters I found on my own, but most are submitted by users (there's a "what other newsletters do you subscribe to" question after you sign up). I set up an inbound-only mail server, and I generate a unique address for each newsletter, which I use to sign up manually. I approve each issue that comes in so that we don't forward welcome emails, promotions etc. (It only takes 5 - 10 minutes a day). Before forwarding I also scrub out any links with certain keywords like "unsubscribe", "manage your preferences" and so on. It's not a perfect process but it's good enough for now.
Long-term I want to build a genaral-purpose recommender system[2]. I'm starting with newsletters because I think it'll be the easiest way to grow fast initially (which seems to have been validated so far). The short explanation is that I've designed The Sample to be extremely effective at cross-promotion[3]. (If you have a newsletter, submit it at https://sample.findka.com/submit/ and I'll send you a referral link).
[1] https://news.ycombinator.com/item?id=24921127
[2] https://jacobobryant.com/p/why-newsletters/
[3] https://jacobobryant.com/p/an-algorithm-for-driving-newslett...
Interesting, and inevitable, project, but I have my concerns.
I think the concerns about ML causing echo chambers/other problems, while not completely unfounded, are overblown (perhaps due to the overall anti-big-tech sentiment). I think human behavior plus ease of sharing online is a much larger factor. I'm optimistic that ML filtering can actually help people get a much larger variety of information, which is one of my goals for this.
> I think the concerns about ML causing echo chambers/other problems, while not completely unfounded, are overblown
I disagree that it's overblown. This is a widely discussed topic in ML research, especially around recommender systems [0]. While I do agree that ML systems have enormous potential in augmenting human capability, we should be addressing possible flaws such as the one mentioned.
I'm personally interested in how to solve biases in machine learning systems, as it impacts so much of my professional work. But I also think bias, or echo-chamber, isn't unique to ML since we see it so much in the world and institutions that surround us, but we are in a unique position to address these problems directly on the systems we create.
[0] https://arxiv.org/abs/2010.03240 [1] https://link.springer.com/article/10.1007/s11390-020-0135-9
Fair enough. To be clear, I'm not saying ML bias isn't a significant problem; and certainly I'll continue to address it as the project grows. Based on this, it sounds like we might be in agreement mostly:
> But I also think bias, or echo-chamber, isn't unique to ML since we see it so much in the world and institutions that surround us, but we are in a unique position to address these problems directly on the systems we create.
I often run into people (usually not ML practitioners) who appear to think that ML inherently results in filter bubbles/bias (as opposed to specific ML implementations) and thus think the entire approach should be abandoned; whereas I think ML, with a proper focus on reducing bias, is one of our most promising options.
[1] https://news.ycombinator.com/item?id=27669863
[2] This experience has affected the way I perceive advice from people on the internet
I'm specifically interested in it because I'm following a need gap - 'Get me the news I need, not the news I want' on my problem validation platform - https://needgap.com/problems/156-get-me-the-news-i-need-not-... and your solution if applied for news articles directly could address that need gap.
There are a few reasons. The biggest one is distribution. I came up with the idea for The Sample because I was working on an essay recommender system[1] and I was trying to figure out how to incentivise people to share it. I've built a referral system[2] into The Sample, and whenever someone submits a newsletter, I send them a referral link. If they share the link, their newsletter gets forwarded more often. So far it's been promising; last I checked, about 15% of our subscribers came from referral links. As The Sample gets larger, the referral system will become more compelling (network effects). I've used a fair amount of paid advertising to get things going, but after some threshold, I think it will grow extremely fast with just the referral system.
The next reason is implementation-related. I prefer to use collaborative filtering as much as possible, falling back to content-based filtering only when there isn't enough rating data for collaborative filtering. But for collaborative filtering, you need long-lived items. If you're recommending news articles, by the time you collect much rating/interaction data, the article will be stale and you won't want to recommend it more anyway[3]. Recommending newsletters is IMO an excellent way to combine ML with human curation: humans figure out what news is good, and ML helps readers find curators who are good/relevant to them.
[2] https://jacobobryant.com/p/an-algorithm-for-driving-newslett...
[3] Unless perhaps you're a large news site and can collect lots of interaction data quickly. An interesting thing about my situation is that I'm trying to bootstrap a recommender system from the beginning, whereas most recommender systems today were built as an add-on for another product, after it already had lots of users.
The implementation related decision sounds interesting, I need to think more about this as I wonder whether news in newsletters are anyways going to have limited lifespan. Moreover investment of time on user side for newsletter delivering news increases with each newsletter vs One newsletter delivering what I want.
Congratulations for the launch, I feel recommendation systems need innovation and so I support your endeavor.
> (FYI looks like Needgap is down right now; Cloudflare is showing me a 521 response)
Sorry for that, It's been resolved now. You're welcomed to join the discussion on 'Get me the news I need, not the news I want' by explaining what Sample does to address that problem.
You might be interested in this article[1] and the discussion of subscribed vs. filtered sources. It'd be nice to have a convenient way to assign different priorities to newsletters/content sources. So for a few, you'd get every issue, and for the rest, they'd be combined together in some way.
> You're welcomed to join the discussion on 'Get me the news I need, not the news I want' by explaining what Sample does to address that problem.
I'm not sure I have much insight about needs vs. wants, but you might be interested in The Factual.[2]
[1] https://jacobobryant.com/p/the-pros-and-cons-of-algorithmic/
> Do you want to focus your attention on problems that are most likely to hurt you?
That's the difference between want vs need. I may want to read all the news posted on HN, But only a fraction of it might be what I need in terms of what I need for betterment of my well-being, career etc.
[1] http://www.bayesianinvestor.com/blog/index.php/2021/06/13/av...
NOTE: These are not criticisms so much as release 1.1 (or even 2.0) use cases....
The categories are too broad (and, for some use cases, the 21 days is rather long, but I get it...).
Re categories: While I am interested in science, quote unquote, I am orders of magnitude more interested in physics than biology, with the exception of biochemistry as it relates to evolutionary theory, and within physics I am at least an order of magnitude more interested in QM (two orders, if you can nail it down to relational QM) than I am in astrophysics, and my interest in astronomy is even lower, and my interest in exoplanets is mostly nil.
And that's my personal interest stuff. Professionally, I am interested in JS specifically, so that would be programming, but I really have zero interest in Fortran or COBOL or Haskell, unless somehow those relate directly to things I can or might want to do in JS (full stack React FE, node BE).
The timeframe side relates to that professional side as well: That JS focus will change at some point (startup life) and by the time 21 days have passed, well, I'll be about 2.5 weeks behind, eh?
To add to that, topic modeling/content-based filtering is only used as a starting point anyway; as you rate the newsletters you receive, your recommendations will start to be dominated more by collaborative filtering ("people who liked X also liked Y", without thought for what X and Y are about).
Also, the "21 days" thing is pure marketing/explanation :). It's just an attempt to help people (especially non-technical people) understand that it'll adapt to your preferences over time. It'll start adapting immediately, and it'll continue adapting after 21 days.
> a surprisingly difficult problem for people not familiar with recommender systems
Do you feel like open-sourcing and/or writing a post on how it works? I feel interested in recommender systems too :-)
Recently I've been factoring out the ML code into a separate recommendation service so it can different kinds of apps (I just barely made this essay recommender system[4] start using it for example).
I'm happy to chat about recommender systems also if you like, email's in my profile.
Each newsletter includes a link to rate it from 1 star to 5 stars.
I consistently rate the newsletters I receive at 1 star! About 5 good newsletters have been delivered over this time, these have been rated from 2 to 5 stars but I haven't received more of that type.
The recommendation system seems fairly ineffective, I just receive junk each day.
Or would this be reflective of the state of newsletters (98% junk, 2% good)?
FWIW the average rating is 2.5, which does seem kind of low, but still indicates that you're probably an outlier. That being said I've been relying mainly on regular newsletter engagement metrics to judge how effective the recommendations are, since rating can be a little unreliable (people are more likely to rate something when they've had a bad experience). Lately we've been at 46% open rate and 16% click-through rate.
The classification problem is interesting though. I ended up with a long list of hundreds of topics. Most articles fall in two or more. There's also a sub-problem of clustering news by subject.
That's actually one of The Sample's strengths--it leaves most of the curation to humans, and then it helps people find those curators.
> The classification problem is interesting though. I ended up with a long list of hundreds of topics. Most articles fall in two or more. There's also a sub-problem of clustering news by subject.
Yeah, certainly difficult. I'm doing it partially manually right now but also with fastText[1]. I'd like to switch completely to fastText soon though since more often than not the newsletters I add don't fit well with the checkbox categories on the landing page. Clustering the newsletters with fastText is pretty easy, but it might be hard to generate good tags for the landing page that match up with the clusters.