I’m kind of interested in the opposite problem, what is the simplest solution using a well known library/db that approaches the fastest hand optimized solution to this problem?
I’m kind of interested in the opposite problem, what is the simplest solution using a well known library/db that approaches the fastest hand optimized solution to this problem?
Java is also more mature, which means you are entering a massive package-bloat setup that has evolved over the years to work for everyones wild and varied needs. By the time you have your database, cache, http/other handlers, tests, fixtures, metrics, logging, tracing, etc... setup you're looking at a scary pile of dependencies spanning thousands of classes that would make even NPM jealous.
https://github.com/gunnarmorling/1brc?tab=readme-ov-file#run...
The Python version:
https://github.com/gunnarmorling/1brc/blob/main/src/main/pyt...
It seems they are not running against the full dataset:
> Moving on to the 100 million file to see if size makes a difference.
ggplot2::autoplot(reorderMicrobenchmarkResults(bench1e8))
One would also have to run both approaches on the same hardware for a meaningful comparison?