Claude code with sonnet was pretty good but needed a lot of back and forth to get to the right solution. Opus feels closer to a colleague, maybe not your absolute best colleague, but far from your worst.
2,703 karma · joined May 6, 2013
If I posted something you found interesting, I probably found it on Scour. Sign up in <1 minute to get your own personalized feed.
Writing at https://emschwartz.me
Claude code with sonnet was pretty good but needed a lot of back and forth to get to the right solution. Opus feels closer to a colleague, maybe not your absolute best colleague, but far from your worst.
When I did a little bit of contracting for TigerBeetle, I was working on these learning exercises and came up with some fun ingredients that could be used if you're trying to push what's possible https://github.com/tigerbeetle/tigerlings/pull/2
There isn't a way to block feeds (yet), but you can subscribe to specific feeds, which basically acts like creating your own list of allowed feeds.
Separately, I'm going to be working on letting you exclude content from certain domains (which was requested in https://feedback.scour.ing/33).
The same day Cloudflare had its unwrap fiasco, I found a bug in my code because of a slice that in certain cases went past the end of a vector. Switched it to use iterators and will definitely be more careful with slices and array indexes in the future.
It ranks articles by how closely related they are to your interests. You can import a set of RSS feeds or scour all 15,000+ sources.
I built it because I wanted to find the good articles among noisy feeds like HN Newest. I've also avoided RSS readers in the past because of that feeling of having thousands of unread emails.
You can also scour all 14,000+ sources for posts that match your interests.
It’s got just the features you need, is built by a solo dev, and it’s got a very fair split between free and paid features. I used it to put up my personal site and have been very happy with the experience.
Generalization:
> Maybe Chinese models generalise to unseen tasks less well. (For instance, when tested on fresh data, 01’s Yi model fell 8pp (25%) on GSM - the biggest drop amongst all models.)
> We can get a dirty estimate of this by the “shrinkage gap”: look at how a model performs on next year’s iteration of some task, compared to this year’s. If it finished training in 2024, then it can’t have trained on the version released in 2025, so we get to see what they’re like on at least somewhat novel tasks. We’ll use two versions of the same benchmark to keep the difficulty roughly on par. Let’s try AIME:
> Almost all models get worse on this new benchmark, despite 2025 being the same difficulty as 2024 (for humans). But as I expected, Western models drop less: they lost 10% of their performance on the new data, while Chinese models dropped 21%. p = 0.09.
> Averaging across crappy models for the sake of a cultural generalisation doesn’t make sense. Luckily, rerunning the analysis with just the top models gives roughly the same result (9% gap instead of 11%).
Cost-effectiveness:
> Distinguish intelligence (max performance), intelligence per token (efficiency), and intelligence per dollar (cost-effectiveness).
> The 5x discounts I quoted are per-token, not per-success. If you had to use 6x more tokens to get the same quality, then there would be no real discount. And indeed DeepSeek and Qwen (see also anecdote here about Kimi, uncontested) are very hungry.
It also works well for feeds that are too noisy to read through manually, like HN Newest.
https://scour.ing (I’m the developer)
This is the blog post and HN discussion where they announced the intention to go this direction:
I found Babbel to feel much more like an app designed by language instructors.
Hope this explanation helped explain why, at least a little bit.
It's not quite the same as capturing all of the queries used in development (or production), but it seems somewhat useful.
I'll also note that I had an LLM generate quite a useful script to identify unused indexes (it scanned the code base for SQL queries, ran `EXPLAIN QUERY PLAN` on each one to identify which indexes were being used, and cross-referenced that against the indexes in the database to find unused ones). It would probably be possible to do something similar (but definitely imperfect) where you find all of the queries, get the query plans, and use an LLM to make suggestions about what indexes would speed up those queries.
Let me know if you have any other feedback as you use it more!
I thought a lot of good stuff was probably getting buried in the fire hose, but I had no idea how well it would actually work at finding those hidden gems for me.
Feedback is extremely welcome! Feel free to email ideas to me (in any state of polish) or post them on https://feedback.scour.ing. Looking forward to hearing your suggestions!
Comments like these are very motivating, so thank you!
Are there specific applications you’re targeting where latency matters more than durability?
Any plans to support binary quantized vectors and hamming distances?