HNHacker News
TopNewBestAskShowJobs

RobinL

2,196 karma · joined February 9, 2013

Data scientist/engineer. Blog: robinlinacre.com Twitter: https://twitter.com/RobinLinacre Lead author of Splink, for record linkage at scale: https://github.com/moj-analytical-services/splink
submissionscomments
RobinL··on Why DuckDB is my first choice for data processing
Exactly - these huge machines are surely eating a lot into the need for distributed systems like Spark. So much less of a headache to run as well
RobinL··on Why DuckDB is my first choice for data processing
Author here. I wouldn't argue SQL or duckdb is _more_ testable than polars. But I think historically people have criticised SQL as being hard to test. Duckdb changes that.

I disagree that SQL has nothing to do with fast. One of the most amazing things to me about SQL is that, since it's declarative, the same code has got faster and faster to execute as we've gone through better and better SQL engines. I've seen this through the past five years of writing and maintaining a record linkage library. It generates SQL that can be executed against multiple backends. My library gets faster and faster year after year without me having to do anything, due to improvements in the SQL backends that handle things like vectorisation and parallelization for me. I imagine if I were to try and program the routines by hand, it would be significantly slower since so much work has gone into optimising SQL engines.

In terms of future proof - yes in the sense that the code will still be easy to run in 20 years time.

RobinL··on Why DuckDB is my first choice for data processing
Author here. Re: 'SQL should be the first option considered', there are certainly advantages to other dataframe APIs like pandas or polars, and arguably any one is better in the moment than SQL. At the moment Polars is ascendent and it's a high quality API.

But the problem is the ecosystem hasn't standardised on any of them, and it's annoying to have to rewrite pipelines from one dataframe API.

I also agree you're gonna hit OOM if your data is massive, but my guess is the vast majority of tabular data people process is <10GB, and that'll generally process fine on a single large machine. Certainly in my experience it's common to see Spark being used on datasets that are no where big enough to need it. DuckDB is gaining traction, but a lot of people still seem unaware how quickly you can process multiple GB of data on a laptop nowadays.

I guess my overall position is it's a good idea to think about using DuckDB first, because often it'll do the job quickly and easily. There are a whole host of scenarios where it's inappropriate, but it's a good place to start.

RobinL··on Why DuckDB is my first choice for data processing
It is true that for json and csv you need a full scan but there are several mitigations.

The first is simply that it's fast - for example, DuckDB has one of the best csv readers around, and it's parallelised.

Next, engines like DuckDB are optimised for aggregate analysis, where your single query processes a lot of rows (often a significant % of all rows). That means that a full scan is not necessarily as big a problem as it first appears. It's not like a transactional database where often you need to quickly locate and update a single row out of millions.

In addition, engines like DuckDB have predicate pushdown so if your data is stored in parquet format, then you do not need to scan every row because the parquet files themselves hold metadata about the values contained within the file.

Finally, when data is stored in formats like parquet, it's a columnar format, so it only needs to scan the data in that column, rather than needing to process the whole row even though you may be only interested in one or two columns

RobinL··on Ask HN: Share your personal website
https://robinlinacre.com
RobinL··on Find a pub that needs you
> In her November Budget, Chancellor Rachel Reeves scaled back business rate discounts that have been in force since the pandemic from 75% to 40% - and announced that there would be no discount at all from April.

That, combined with big upward adjustments to rateable values of pub premises, left landlords with the prospect of much higher rates bills.

RobinL··on Ozempic is changing the foods Americans buy
I was curious for a UK comparison so I looked it up.

At the start of 2025, about 3% of adults in UK had used GLP-1 drugs in past year in the UK. And "most GLP-1 for weight loss in the UK is from private, rather than NHS provision" [1].

[1] https://pmc.ncbi.nlm.nih.gov/articles/PMC12781702/

RobinL··on ChatGPT Health
Another interesting aspect is that the NHS app makes all your detailed health history (doctors notes, scan results etc.) available to you as the patient.

Which in turn means you have the option of feeding it into ChatGPT. This feels potentially very valuable and a nice way of working around issues with whether doctors themselves are allowed to do it.

I'm not sure this applies to every surgery, but certainly my dad had access to everything immediately when he had a scan.

RobinL··on Total monthly number of StackOverflow questions over time
I'm hoping increasing we'll see agents helping with this sort of issue. I would like an agent that would do things like pull the spark repo into the working area and consult the source code/cross reference against what you're trying to do.

Once technique I've used successfully is to do this 'manually' to ensure codex/Claude code can grep around the libraries I'm using

RobinL··on Tiled Art
The art creation tool is great: https://tiled.art/en/create/
RobinL··on Gemini 3 Flash: Frontier intelligence built for speed
Yeah - agree, Anthropic much better for coding. I'm more thinking about the 'average chat user' (the larger potential userbase), most of whom are on chatgpt.
RobinL··on Gemini 3 Flash: Frontier intelligence built for speed
Feels like Google is really pulling ahead of the pack here. A model that is cheap, fast and good, combined with Android and gsuite integration seems like such powerful combination.

Presumably a big motivation for them is to be first to get something good and cheap enough they can serve to every Android device, ahead of whatever the OpenAI/Jony Ive hardware project will be, and way ahead of Apple Intelligence. Speaking for myself, I would pay quite a lot for truly 'AI first' phone that actually worked.

RobinL··on Is it a bubble?
I've found Opus 4.5 in copilot to be very impressive. Better than codex CLI in my experience. I agree Copilot definitely used to be absolutely awful.
RobinL··on AI Is Breaking the Moral Foundation of Modern Society
Price and scarcity go hand in hand, not value and scarcity.

Diamonds are pretty worthless but expensive because they're scarce (putting aside industrial applications), water is extremely valuable but cheap.

No doubt there are some goods where the value is related to price, but these are probably mostly status related goods. e.g. to many buyers, the whole point in a Rolex is that it's expensive.

RobinL··on Geothermal Breakthrough in South Texas Signals New Era for Ercot
More coverage of the broader geothermal renaissance: Geothermal’s time has finally come https://www.economist.com/interactive/science-and-technology...

And i found this interesting: https://www.complexsystemspodcast.com/episodes/fracking-aust...

Seems fairly promising

RobinL··on Python is not a great language for data science
Out of the current options, I strongly agree - I even wrote a blog post! https://www.robinlinacre.com/recommend_sql/

But on the other hand, that's doesn't mean SQL is ideal - far from it. When using DuckDB with Python, to make things more succinct, reusable and maintainable, I often fall into the pattern of writing Python functions that generate SQL strings.

But that hints at the drawbacks of SQL: it's mostly not composable as a language (compared to general purpose languages with first-class abstractions). DuckDB syntax does improve on this a little, but I think it's mostly fundamental to SQL. All I'm saying is that it feels like something better is possible.

RobinL··on Python is not a great language for data science
I would argue that's about how the data is stored. What I'm trying to express is the idea of the programming language itself supporting high level tabular abstractions/transformations such as grouping, aggregation, joins and so on.
RobinL··on Python is not a great language for data science
I think a lot of this comes down to the question: Why aren't tables first class citizens in programming languages?

If you step back, it's kind of weird that there's no mainstream programming language that has tables as first class citizens. Instead, we're stuck learning multiple APIs (polars, pandas) which are effectively programming languages for tables.

R is perhaps the closest, because it has data.frame as a 'first class citizen', but most people don't seem to use it, and use e.g. tibbles from dplyr instead.

The root cause seems to be that we still haven't figured out the best language to use to manipulate tabular data yet (i.e. the way of expressing this). It feels like there's been some convergence on some common ideas. Polars is kindof similar to dplyr. But no standard, except perhaps SQL.

FWIW, I agree that Python is not great, but I think it's also true R is not great. I don't agree with the specific comparisons in the piece.

RobinL··on Nano Banana Pro
Thanks - this worked for me (some errors, some success).

Last week I was making a birthday card for my son with the old model. The new model is dramatically better - I'm asking for an image in comic book style, prompted with some images of him.

With the previous model, the boy was descriptively similar (e.g. hair colour and style) but looked nothing like him. With this model it's recognisably him.

RobinL··on Gemini 3
- Anyone have any idea why it says 'confidential'?

- Anyone actually able to use it? I get 'You've reached your rate limit. Please try again later'. (That said, I don't have a paid plan, but I've always had pretty much unlimited access to 2.5 pro)

[Edit: working for me now in ai studio]

RobinL··on UK gov to ban selling show tickets above face price
Here is an alternative idea: https://aeturrell.com/blog/posts/one-trick-to-stop-ticket-to...

The key idea is that only the original ticket buyer is eligible for a large rebate when attending the event. It prevents touting, but does not mean everyone who wants a ticket gets one.

Though in practice it is perhaps to techie, and in the end not dramatically different to what Glastonbury Festival does, which is that the ticket is only valid for entrance by the original purchaser, using photo id.

RobinL··on A new book about the origins of Effective Altruism
> many of EA's most notorious supporters.

The fact they're notorious makes them a biased sample.

My guess is for the majority of people interested in EA - the typical supporter who is not super wealthy or well known - the two central ideas are:

- For people living in wealthy countries, giving some % of your income makes little difference to your life, but can potentially make a big difference to someone else's

- We should carefully decide which charities to give to, because some are far more effective than others.

That's pretty much it - essentially the message in Peter Singer's book: https://www.thelifeyoucansave.org/.

I would describe myself as an EA, but all that means to me is really the two points above. It certainly isn't anything like an indulgence that morally offsets poor behaviour elsewhere

RobinL··on Britain's railway privatization was an abject failure
Salaries don't tend to be strongly correlated with bad working conditions or stress. In most industries (like software development) it's just supply and demand, and I imagine there are more people willing and able to work for £65k as a train driver than as a software developer. It's a bit different for train drivers because of the strong unions; my guess is that explains their high salaries more than lack of supply.

(Median total reward for TOC train drivers is £66,043) https://www.orr.gov.uk/sites/default/files/2022-10/review-of...

RobinL··on Britain's railway privatization was an abject failure
Is that true? My understanding is they command very high wages because their unions are strong and they have a lot of leverage: by striking they can impose extremely high costs on the wider economy (not to mention bad press for politicians).
RobinL··on GPT-5.1: A smarter, more conversational ChatGPT
Same:

Yes — the Romanian player is Costel Pantilimon. He won the Premier League with Manchester City in the 2011-12 and 2013-14 seasons.

If you meant another Romanian player (perhaps one who featured more prominently rather than as a backup), I can check.

RobinL··on Australia has so much solar that it's offering everyone free electricity
FWIW, here's a chart showing current prices across developed countries, showing UK is worst! https://fullfact.org/economy/uk-world-electricity-prices/
RobinL··on Australia has so much solar that it's offering everyone free electricity
I think it's best to view this from an economics point of view - in a nutshell price signals are usually the most powerful way to create behavioural change; in this case, we want people to shift demand away from peak times. Nobody is being forced to, they just have to pay more for the convenience of not bothering.

> IMO this is just setting us up for insane surge pricing for those people who don't do the good citizen thing of becoming nocturnal

It actually costs a lot more to produce marginal energy at peak times, the cost just reflects the cost of production. It doesn't seem unreasonable for me for the consumer to bear the cost, and also get the benfit if they choose to put their car to charge overnight rather than at peak time.

This also has a nice secondary benefit: anyone on agile tariffs who shifts demand away from peak time actually benefits those who don't want to bother, because the peak price/cost goes down, and so the overall average price of electricity goes down.

> I see this as just yet _another_ job the government/business is making us do instead of them

In most other market, people are expected to respond to price incentives. When local apples are cheap relative to imported cherries, people don't complain that government/business is making us do a job to push demand in the direction of apples.

> Is it too much to ask for my government to provide sensibly and simply priced energy so we can get on with our day, working, studying, raising kids etc?

The free market price _is_ the agile price. The government intervention is actually in the direction of fixing prices (e.g. by the energy price cap, which is sometimes below the free market price at peak times). In general, markets do not work very well when the government fixes the market

When you let the market clear and send out price signals, markets almost always become more efficient (which means that consumers benefit overall)

RobinL··on Australia has so much solar that it's offering everyone free electricity
Yeah - unfortunately the UK has some of the highest electricity prices in Europe.

Almost all households are on fixed tariffs, typically about 26p/kwh at the moment.

RobinL··on Australia has so much solar that it's offering everyone free electricity
In the UK, you can go on an agile tariff that does exactly this. I'm on one.

It's quite fun (and educational) with the kids to work out when to put the car on to charge, when to run the dryer etc, looking at the few days ahead forecasts.

Last month, we paid 11p per kWh on average, which is less than half what you'd pay on a standard tariff, and it's nice to be doing something good for the environment too. It's particularly satisfying to charge up the car when tariffs go negative.

Here's today's rates (actuals): https://agilebuddy.uk/latest/agile

Here's a forecast: https://prices.fly.dev/A/

RobinL··on Uv is the best thing to happen to the Python ecosystem in a decade
https://docs.astral.sh/uv/concepts/projects/init/#creating-a...

I use the 'bare' option for this

← PreviousPage 3 of 18Next →