Sounds like an interesting approach, but just that I understand the scope or impact of the paper right: Surely data-aware indexing can't be the novel part, right? Or was it always so complicated to model the data distribution that no one managed to do it until now? It seems natural to try to adapt your index to the type of data you see more often than not.
Very cool idea though.