Over 2 orders of magnitude reduction in disk footprint and going from 20s to 20ms in query performance - those are incredible improvements. Would love to learn more about how you pulled that off.
Over 2 orders of magnitude reduction in disk footprint and going from 20s to 20ms in query performance - those are incredible improvements. Would love to learn more about how you pulled that off.
We just figured out exactly what we wanted, stored the data in chunks, compressed(we used 2 var-int encoding schemes, and snappy compression), indices(skip-lists) for each file and each chunk and for queries, we parellize access to as files required across multiple OS threads (scatter-gather). In fact, I am sure we could have gotten better performance if we wanted to spend more time on that problem. It wasn't novel or particularly interesting or hard anyway. Just something that needed to be done to help us solve a problem.