I’m surprised at the poor performance of python here. For reference there are several very brief R examples which are just 2-3 seconds. Eg http://blog.schochastics.net/posts/2024-01-08_one-billion-ro...
It seems they are not running against the full dataset:
> Moving on to the 100 million file to see if size makes a difference.
ggplot2::autoplot(reorderMicrobenchmarkResults(bench1e8))
One would also have to run both approaches on the same hardware for a meaningful comparison?