HNHacker News
TopNewBestAskShowJobs

RobinL

2,196 karma · joined February 9, 2013

Data scientist/engineer. Blog: robinlinacre.com Twitter: https://twitter.com/RobinLinacre Lead author of Splink, for record linkage at scale: https://github.com/moj-analytical-services/splink
submissionscomments
RobinL··on Preview: Amazon S3 Tables and Lakehouse in DuckDB
Yes - agree! I actually wrote a blog about this just two days ago:

May be of interest to people who:

- What to know what DuckDB is and why it's interesting

- What's good about it

- Why for orgs without huge data, we will hopefully see a lot more of 's3 + duckdb' rather than more complex architectures and services, and hopefully (IMHO) less Spark!

https://www.robinlinacre.com/recommend_duckdb/

I think most people in data science or data engineering should at least try it to get a sense of what it can do

Really for me, the most important thing is it makes it so much easier to design and test complex ETL because you're not constantly having to run queries against Athena/Spark to check they work - you can do it all locally, in CI, set up tests, etc.

RobinL··on UK homes install subsidised heat pumps at record level
But the same is true of gas, and in the UK at least they cost about the same per kWh of output heat (electricity is more expensive, but this is balanced out by the higher efficiency of the heat pump)
RobinL··on UK homes install subsidised heat pumps at record level
I've heard this several but don't really understand the rationale. Can't you just put in a bigger heat pump? Or top up an air to water with an air to air?
RobinL··on Grok3 Launch [video]
Thank you!
RobinL··on Grok3 Launch [video]
This sounds great - would love to hear a little more about the prompts. Are you literally just asking 'write me a context.md that explains how feature x works' or something like that?
RobinL··on Grok3 Launch [video]
Apologies for possibly stupid question but where can you use it right now? Just on 'direct chat' on https://lmarena.ai/ or is there a better alternative? Or do you have early access?
RobinL··on Perplexity Deep Research
They were doing web search before open ai/anthropic, so they historically had a (pretty decent) unique selling point.

Once chat gpt added web browsing, I largely stopped using perplexity

RobinL··on Sam Altman: How I use AI in my own everyday life–it's great for 'boring' tasks
If I was a billionaire I'd have humans do all these types of things. I can imagine AI would be useful to give me an unbiased/non sycophantic view.
RobinL··on Why Blog If Nobody Reads It?
A big reason I blog is that I find writing something down is essential to clarifying my thoughts.

Once it's written down, I might as well put it online, and it has the added advantages that I can simply link to the material rather than having to explain it over and over again e.g. in emails.

RobinL··on GitHub Copilot: The Agent Awakens
Thanks - having tried Co-Pilot again today after a 6-month Cursor hiatus, I think this is a good summary
RobinL··on GitHub Copilot: The Agent Awakens
Update: I have used co-pilot agent mode for a couple of hours today.

It's definitely catching up with Cursor but not there yet. In particular: - Edits take quite a bit longer to apply, breaking flow - Autocomplete predictions (equiv. of Cursor Tab) not as good

But in the past 6 months or so it's gone from being pretty hopeless to very useful. If I was forced to use it instead of Cursor it wouldn't be a huge deal any more.

RobinL··on GitHub Copilot: The Agent Awakens
Can anyone speak to weather it's worth going back to co-pilot from cursor. On the face of it $10 a month for unlimited messages looks compelling. Is it really unlimited? From these videos it's starting to look pretty similar to cursor...
RobinL··on Introducing deep research
Interesting

You might find it amusing to compare it to: https://hn-wrapped.kadoa.com/timabdulla

(Ref:https://news.ycombinator.com/item?id=42857604)

RobinL··on OpenAI O3-Mini
Whilst I had tried R1 before, I hadn't paid attention to how fast it was. I just tried some similar prompts and was pretty impressed with speed and quality. I think o3-mini was still a bit quicker though.
RobinL··on OpenAI O3-Mini
Wow - this is seriously fast (o3-mini), and my initial impressions are very favourable. I was asking it to layout quite a complex html form from a schema and it did a very good job.

Looking at the comments on here and the benchmark results I was expecting it to be a bit meh, but initial impressions are quite the opposite

I was expecting it to perhaps be a marginal improvement for complex things that need a lot of 'reasoning', but it seems it's a bit improvement for simple things that you need doing fast

RobinL··on Arsenal FC AI Research Engineer job posting
Has anyone tried to link and dedupe the various datasets using a probabilistic linkage tool like Splink?

https://moj-analytical-services.github.io/splink/

(Disclaimer: I am the lead author, but the tool is FOSS)

RobinL··on Companies Need to Offer Current Documentation in a Single Document for LLMs
I also wish this was more widespread. I do it for my FOSS data linkage software here: https://moj-analytical-services.github.io/splink/topic_guide...
RobinL··on Ask HN: Is wind power financially viable without subsidies?
I'd also love to see some high quality, well researched information on this.

For instance I recently saw this exchange: https://x.com/i/bookmarks?post_id=1881090543956222293

I'm generally a proponent of renewables, but I think it's also important to take seriously quantitative arguments against them. But without good sources, it's impossible to work out who's right.

This is useful to suggest solar will win in the end: https://ourworldindata.org/grapher/levelized-cost-of-energy?...

But much less clear how well wind, especially offshore, can compete.

This is also decent but doesn't get much into cost: https://www.sustainabilitybynumbers.com/p/can-solar-and-wind...

RobinL··on No Calls
There's a theory an economics that says that the more different prices a provider can charge the more of the surplus they capture (ie they can tilt that percentage towards the seller and away from the buyer).

Of course, if they're a monopoly provider and the buyer really needs it, they have to cough up. But generally there are substitute products. So the buyer would do well to look for an alternative that doesn't do differential pricing to capture more surplus for themselves.

RobinL··on No Calls
Schedule a call is a huge red flag to me because:

- it implies differential pricing, meaning they will charge you as much as possible both now and in the future (when you may be locked in)

- it usually obscures what the product actually does

Differential pricing is really pernicious because if the product happens to be super valuable to you, they're likely to find out and charge you even more

RobinL··on What Will Congestion Pricing Do to Manhattan Dining?
There are two components to the price of driving:

1. The congestion charge, if it exists 2. The opportunity cost of the time taken

It's plausible to me that whilst total cars will go down (supply and demand) you may get more high income customers (because they have a high opportunity cost of time) and hence some businesses may do better.

It's also not obvious it should increase costs of e.g. deliveries to businesses because the the delivery driver will spend less time waiting in traffic.

I wonder if any research is being done on either of these effects

RobinL··on Apache DataFusion
> It's impossible to improve performance when you have to go out to the DB anyway;

That's not right. There are many queries that run far faster in duckdb/datafusion than (say) postgres, even with the overhead of pulling whole large tables prior to running the query. (Or use like pg_duckdb).

For certain types of queries these engines can be 100x faster.

More here: https://postgres.fm/episodes/pg_duckdb

RobinL··on Uv's killer feature is making ad-hoc environments easy
I've only recently started with uv, but this is one thing it seems to solve nicely. I've tried to get into the mindset of only using uv for python stuff - and hence I haven't installed python using homebrew, only uv.

You basically need to just remember to never call python directly. Instead use uv run and uv pip install. That ensures you're always using the uv installed python and/or a venv.

Python based tools where you may want a global install (say ruff) can be installed using uv tool

RobinL··on Advent of Code 2024 in pure SQL
You can use duckdb on a single machine. It's also indexless (or more accurately, you don't have to explicitly create indexes)
RobinL··on Joco almost died at launch. Now, it's a lifeline for e-bike delivery riders
Stepping back a bit, having these on the road rather than cars or petrol scooters must be a big win even if they don't last as long as we'd like
RobinL··on Tracking Down the Bulgarian Marketplace Scams
I got several people wanting to send a courier last time I listed something on Facebook. Checking their pages, they were all from eastern Europe with no obvious connection to my city. Good to know the mechanics of the scam, I wondered what they were up to. Don't understand why Facebook couldn't have auto detected the messages though - seemed like a pretty major failure of marketplace that the majority of the messages I got were scams.
RobinL··on Should you ditch Spark for DuckDB or Polars?
We're still using parquet. So we use the native duckdb format for intermediate processing (during pipeline execution) but the end results are saved out as parquet. This is partly because customers often read the data from other tools (e.g. AWS athena)

I'd be interested in hearing about experiences of using duckdb files though, i can see instances where it could be useful to us

RobinL··on Should you ditch Spark for DuckDB or Polars?
Interesting. So what does that look like on disk? Possibly slightly naively I'm imagining a single massive file?
RobinL··on Should you ditch Spark for DuckDB or Polars?
I submitted this because I thought it was a good, high effort post, but I must admit I was surprised by the conclusion. In my experience, admittedly on different workloads, duckdb is both faster and easier to use than spark, and requires significantly less tuning and less complex infrastructure. I've been trying to transition as much as possible over to duckdb.

There are also some interesting points in the following podcast about ease of use and transactional capabilities of duckdb which are easy to overlook (you can skip the first 10 mins): https://open.spotify.com/episode/7zBdJurLfWBilCi6DQ2eYb

Of course, if you have truly massive data, you probably still need spark

RobinL··on OpenTTD is an open source simulation game based upon Transport Tycoon Deluxe
My son loves openttd. Fun fact: the original transport tycoon was primarily written in assembly by Chris Sawyer
← PreviousPage 6 of 18Next →