Nice! Does anybody know a project that combines this kind of techniques in a program that could be ran across billions of rows, most preferably something that scales to multiple nodes?
Probably some xargs sed awk hacks are also sufficient for results in milliseconds.
You start getting problems when dealing with >128 TB databases like the legacy crm databases at Bayer and Siemens e.g.
To answer the initial question - take a look at CitusDB that runs over PostgreSQL. They seem to be doing some cool work in real-time analytics/scalability: https://www.citusdata.com/