HNHacker News
TopNewBestAskShowJobs

RobinL

2,196 karma · joined February 9, 2013

Data scientist/engineer. Blog: robinlinacre.com Twitter: https://twitter.com/RobinLinacre Lead author of Splink, for record linkage at scale: https://github.com/moj-analytical-services/splink
submissionscomments
RobinL··on A sharded DuckDB on 63 nodes runs 1T row aggregation challenge in 5 sec
It's main purpose is to solve the problem of upserts to a data lake, because upsert operations to file based data storage are a real pain.
RobinL··on A sharded DuckDB on 63 nodes runs 1T row aggregation challenge in 5 sec
Yes - OLAP database are built with a completely different performance tradeoff. The way data is stored and the query planner are optimised for exactly these types of queries. If you're working in an oltp system, you're not necessarily doing it wrong, but you may wish to consider exporting the data to use in an OLAP tool if you're frequently doing big queries. And nowadays there's ways to 'do both ' e.g. you can run the duckdb query engine within a postgres instance
RobinL··on Solar energy is now the cheapest source of power, study
I'm generally pro nuclear and think it should be a significant part of the energy mix. But an interesting point is that the floor price of nuclear is the cost of the turbines, which are surprisingly expensive, and aren't really getting cheaper. Solar can go much cheaper than that, potentially at least. More on this point in the recent Casey Handmer interview of Dwarkesh.
RobinL··on AI-generated “workslop” is destroying productivity?
I largely agree. As a counterpoint, today I delivered a significant PR that was accepted easily by the lead dev with the following approach:

1. Create a branch and vibe code a solution until it works (I'm using codex cli)

2. Open new PR and slowly write the real PR myself using the vibe code as a reference, but cross referencing against existing code.

This involved a fair few concepts that were new to me, but had precedent in the existing code. Overall I think my solution was delivered faster and of at least the same quality as if I'd written it all by hand.

I think its disrespectful to PR a solution you don't understand yourself. But this process feels similar to my previous non-AI assisted approach where I would often code spaghetti until the feature worked, and then start again and do it 'properly' once I knew the rough shape of the solution

RobinL··on Are touchscreens in cars dangerous?
If that were the case they'd be no need for seatbelt laws
RobinL··on Cape Station, future home of an enhanced geothermal power plant, in Utah
This provides a lot of interesting info on geothermal: https://worksinprogress.co/issue/watt-lies-beneath/

And more: https://www.complexsystemspodcast.com/episodes/fracking-aust...

One interesting point made here is that the cost of turbines puts a floor price on any form.of generation which uses them, whether renewable or not, meaning in the long run solar has a big advantage: https://www.dwarkesh.com/p/casey-handmer. I don't know how accurate that is

RobinL··on Polars Cloud and Distributed Polars now available
100% agree. I've also worked as a data engineer and came to the same conclusion. I wrote up a blog which went into a bit more depth on the topic here: https://www.robinlinacre.com/recommend_sql/
RobinL··on Melvyn Bragg steps down from presenting In Our Time
May be of interest: https://www.braggoscope.com/directory (a categorised catalogue of episodes)
RobinL··on U.S. Emissions Rise 4.2%, China's Fall 2.7%
That stat is bonkers. China's GDP is only 5x that of UK. Total UK solar is about 19GW.

So even if you divide China's solar by 5, they added in a month what we have built in >10 years

RobinL··on Data engineering and software engineering are converging
We use aws glue for spark (but are increasingly moving towards duckdb because it's faster for our workloads and easier to test and deploy).

For Spark, glue works quite well. We use it as 'spark as a service', keeping our code as close to vanilla pyspark as possible. This leaves us free to write our code in normal python files, write our own (tested) libraries which are used in our jobs, use GitHub for version control and ci and so on

RobinL··on Data engineering and software engineering are converging
I think this may be a databricks thing? From what I've seen there's a gap between data engineers forced to use databricks and everyone else. From what I've seen, at least how it's used in practice, databricks seems to result in a mess of notebooks with poor dependency and version management.
RobinL··on Do I not like Ruby anymore? (2024)
There's a good interview with DHH, who is the creator of Ruby in Rails here: https://lexfridman.com/dhh-david-heinemeier-hansson-transcri... I have no skin in the game having never used Ruby, but I found his arguments interesting
RobinL··on The new geography of stolen goods
130,000 car thefts a year. That's over £1bn loss, probably closer to £4bn. In this context the total police budget of around £20bn seems remarkably low!

You'd have thought it'd be worth insurance companies paying people to track down the thieves!

RobinL··on Does OLAP Need an ORM
When I read the title my brain immediately jumped to a slightly different idea. With olap, I often find it annoying to figure out the joins from the fk/pk relationships, so I was imagining a tool that kind of automatically followed the links for you. A bit like how a orm gives you auto complete, but without the user having to manually enter the schema.

And I wanted it to emit the raw SQL because that's generally what I want for olap.

So I had to go at building it. If anyone's interested a very rough demo/prototype is here: https://www.robinlinacre.com/vite_live_pg_orm/

Load in the demo Northwind schema and click some tables/columns to see the generated joins

RobinL··on Hillary Clinton says she'd nominate Trump for Nobel Prize if he brokers peace
Misleading headline:

> she would nominate Trump for the award if he was successful in getting Putin to end his war and give back all the territory his forces took from Ukraine in the conflict

So the meaning of the comment is actually the opposite of headline. It's really a criticism of the basis of the negotiation

RobinL··on GPT-5
Hypothesis: to the average user this will feel like a much greater jump in capability then to the average HNer, because most users were not using the model selector. So it'll be more successful than the benchmarks suggest.
RobinL··on GPT-5 for Developers
Totally agree. At the moment I find that frontier LLMs are able to solve most of the problems I throw at them given enough context. Most of my time is spent working out what context they're missing when they fail. So the thing that would help me most is much a much more focussed ability to gather context.

For my use cases, this is mostly needing to be really home in on relevant code files, issues, discussions, PRs. I'm hopeful that GPT5 will be a step forward in this regard that isn't fully captured in the benchmark results. It's certainly promising that it can achieve similar results more cheaply than e.g. Opus.

RobinL··on “No tax on tips” is an industry plant
Because you're artificially favouring one specific industry, at a cost to all other industries.

And because the implication of your argument is we should never tax anything because that's a benefit to both the consumer and the business

RobinL··on “No tax on tips” is an industry plant
It's just supply and demand. It makes working in these industries relatively more attractive, increasing supply of labour and therefore reducing price of labour. So restaurant owners capture some of the benefits
RobinL··on Observable Notebooks 2.0 Technology Preview
This looks great. I love the idea behind notebooks and for a long time it was my favourite environment to program in. But slowly I stopped using them because it never quite felt like the code was entirely mine, and alternatives became easier due to llms. This looks like exactly the remedy I was hoping for. I'm excited to start using them again.
RobinL··on Ask HN: What are you working on? (July 2025)
A fuzzy matching community extension for duckdb: https://github.com/moj-analytical-services/splink_udfs

And I've been vibe coding some maths educational tools and games for (my) 6yo: https://rupertlinacre.com/

RobinL··on Claude Code Is a Slot Machine
I have had similar experiences to OP. All of these apps (mostly maths learning, but also the bus one) were coded using a mix of copilot and Gemini CLI: https://rupertlinacre.com/

I'm probably capable of building all of them by hand, but with a 6yo I'd have never had the time. He loves the games, his mental arithmetic has come on amazingly now he does it 'for fun'.

All code is here: https://github.com/rupertlinacre

Much of this built out of a frustration that most maths resources online are trying to sell you something, full of ads, or poor quality. Just a simple zoomable numberline is hard to find

RobinL··on ChatGPT agent: bridging research and action
This feels a bit underwhelming to me - Perplexity Comet feels more immediately compelling as new paradigm of a natural way of using LLMs within a browser. But perhaps I'm being short-sighted
RobinL··on AI slows down open source developers. Peter Naur can teach us why
The problem is that code often takes as long to review as to write, and AI potentially reduces the quality bar to pull requests. So maintainers have a problem of lots of low quality PRs that take time to reject
RobinL··on Grok 4 Launch [video]
This is what it says in the supposed system prompt see https://news.ycombinator.com/item?id=44517453
RobinL··on Libpostal: C library for parsing/normalizing street addresses around the world
There are many useful applications of libpostal, and it's an impressive library, but one I would caution against is for the purpose of address matching, at least as the 'primary' approach.

The problem is the hardest to parse addresses are also often the hardest to match, making the problem somewhat circular. I wrote about this more in a recent blog on address matching: https://www.robinlinacre.com/address_matching/

RobinL··on Spending Too Much Money on a Coding Agent
This is why unlimited plans are always revoked eventually - a small fraction of users can be responsible for huge costs (Amazon's unlimited file backup service is another good example). Also whilst in general I don't think there's much to worry about with AI energy use, burning $24k of tokens must surely be responsible for a pretty large amount of energy
RobinL··on Ask HN: What's the last non-obvious skill that made you better at your job?
Writing a blog. Many of my posts are things I've learnt at work, or arguments I've failed to make it meetings. By writing it down, I can pin down the argument better and share my thoughts in advance.
RobinL··on Building Accurate Address Matching Systems
Thanks - appreciate the kind words!
RobinL··on Ask HN: What Are You Working On? (June 2025)
I'm working on a free high performance address matching (geocoding) library. As it turns out I blogged about it just today: https://www.robinlinacre.com/address_matching/
← PreviousPage 4 of 18Next →