HNHacker News
TopNewBestAskShowJobs

Lemaxoxo

265 karma · joined March 15, 2018

https://maxhalford.github.io/bio/
submissionscomments
Lemaxoxo··on Quack: The DuckDB Client-Server Protocol
Right, I get that usecase. You have to crunch numbers that sit somewhere, and store the outputs in the same place. DuckLake is great for that. But where does this DuckDB client-server setup fit in?
Lemaxoxo··on Quack: The DuckDB Client-Server Protocol
+1

I can't think of many use cases for this and Arrow Flight, other than moving data around.

Lemaxoxo··on Text classification with Python 3.14's ZSTD module
That's very cool, thanks for sharing. Our of curiosity, did you ever get to run on a Twitter/X stream of political tweets?
Lemaxoxo··on Text classification with Python 3.14's ZSTD module
You are correct. To be fair I wasn't focused on comparing the runtimes of both methods. I just wanted to give a baseline and show that the batch approach is more accurate.
Lemaxoxo··on Text classification with Python 3.14's ZSTD module
Author here. Thank you very much for the comment. I will take a look. This is a great case of Cunningham's law!
Lemaxoxo··on Text classification with Python 3.14's ZSTD module
Author here. Thanks for your comment!

Compression algorithms may have been supporting incremental compression for a while. But as some have pointed out, the point of the post is that it is practical and simple to have this available in Python's standard library. You could indeed do this in Bash, but then people don't do machine learning in Bash.

Lemaxoxo··on A 2.5x faster Postgres parser with Claude Code
Ok that makes sense! On my side I can get away with using it through WASM. But your performance needs won't allow that.
Lemaxoxo··on A 2.5x faster Postgres parser with Claude Code
I'm curious because I have a similar use case for a querying frontend. Did you consider using https://github.com/tobymao/sqlglot? If so, what was missing to justify writing your own parser?
Lemaxoxo··on Text classification with Python 3.14's ZSTD module
Hello HN. 5 years ago I posted an article about text classification via data compression. I got helpful and educative comments in response. Now that Python have shipped zstd in 3.14, I thought it would be time to revisit this approach. The throughput figures are much better. This means you can do baseline machine learning with Python's standard library!
Lemaxoxo··on Ask HN: Share your personal website
https://maxhalford.github.io/
Lemaxoxo··on Do LLMs identify fonts?
Op here. I tried what the font a bit but didn't mention it in the article. I didn't get good results with it. Although it's probably a good idea to ask it for a guess, and feed that to the LLM too.
Lemaxoxo··on Markov Keyboard: keyboard layout that changes by Markov frequency (2019)
Nice! I wrote about something similar for rectangular layouts: https://maxhalford.github.io/blog/dynamic-on-screen-keyboard...
Lemaxoxo··on Lea: Minimalist Alternative to Dbt
Cheers! Mainly a couple of things:

- I don't like to have to put {{ ref('source') }} everywhere. I think the tool should parse dependencies automatically. I wrote more about this here: https://maxhalford.github.io/blog/dbt-ref-rant/ - I don't like the idea each .sql file has to have an associated .yml file. It feels better to have everything in one place. For instance, with lea you can add a @UNIQUE tag as an SQL comment to unit test a column for uniqueness.

Moreover, although dbt brought a shift in the way we do data (which is great) it's very straightforward under the hood. It boils down to parsing queries, organizing them in a DAG, and processing said DAG. dbt feels bloated to me. Also, it seems to me some of the newer cool features are going to be put behind a paywall (e.g. metric layers)

Lemaxoxo··on Lea: Minimalist Alternative to Dbt
Hey there HN. lea is a tool we developed over the past year at Carbonfact. Carbonfact is a platform that helps fashion brands decarbonize. We believe in doing this in a data-driven way, and lea is a cornerstone for us.
Lemaxoxo··on Show HN: Want something better than k-means? Try BanditPAM
Hey, great work. Do you think this algorithm would be amenable to be done online? I'm the author of River (https://riverml.xyz) where we're looking for good online clustering algorithms.
Lemaxoxo··on Stochastic gradient descent written in SQL
Thank you so much for this! It's very generous of you to have taken the time.
Lemaxoxo··on Stochastic gradient descent written in SQL
Thanks, I wasn't aware.
Lemaxoxo··on Stochastic gradient descent written in SQL
Do you have some data/resources on this? I'm a total snowflake at this, but I'm willing to learn.
Lemaxoxo··on Stochastic gradient descent written in SQL
That's too bad, I would have expected it to work out of the box. Other than rewriting the query in a different way, I'm not sure I see an easy workaround. Are you still working on this?
Lemaxoxo··on Stochastic gradient descent written in SQL
Hehe I was wondering if someone would catch that. Rest assured, I know the difference between online and stochastic gradient descent. I admit I used stochastic on Hacker News because I thought it would generate more engagement.
Lemaxoxo··on Stochastic gradient descent written in SQL
Isn't that the point of common table expressions (CTEs)?
Lemaxoxo··on Stochastic gradient descent written in SQL
I agree. Databases are going to be here for a long time, and we're barely scratching the surface of making people productive with them. dbt is just the beginning.
Lemaxoxo··on Stochastic gradient descent written in SQL
I'm watching it, it's really good. Montana makes a great point: you can move data to the models, or move the models to the data. Data is typically larger than models, so it makes sense to go with the latter.
Lemaxoxo··on Stochastic gradient descent written in SQL
Postgres has excellent support for WITH RECURSIVE, so I see no reason why it wouldn't. However, as I answered elsewhere, you would need to set some stateful stuff up if you don't want the query to start from scratch when you re-run it.
Lemaxoxo··on Stochastic gradient descent written in SQL
DuckDB is what I used in the blog post. Re-running this query simply recomputes everything from the start. I didn't store intermediary that would allow starting off from where the query stopped. But it's possible!
Lemaxoxo··on Stochastic gradient descent written in SQL
Thanks for the links, I wasn't aware of them. The Russians often seem to have a step ahead in the ML world.
Lemaxoxo··on Stochastic gradient descent written in SQL
Yes I know some people look down on that. I hope it doesn't take away the merits of the article for you hehe.
Lemaxoxo··on Stochastic gradient descent written in SQL
Cheers, I learnt with the best :)
Lemaxoxo··on Stochastic gradient descent written in SQL
This is definitely a possibility. What I meant to say is that the implementation in the blog post doesn't support that.
Lemaxoxo··on Stochastic gradient descent written in SQL
Good question. I touched upon this in the conclusion. Basically, if you run this in a streaming SQL database, such as Materialize, then you would get a true online system which doesn't restart from scratch.
Page 1 of 2Next →