I’m curious too. We need to see 1b+ numbers or else it just seems like we’re leveraging compression here, which is great but it’s hard to understand why it’s novel and why it’s the best solution no contest.
I can archive my data just fine by using an access based storage tiering like S3 - have your cake and eat it too. Gets around the headache of actually querying all of the 1B+ records because realistically there’s very few use cases for that. If you ever really need it, spin up Kafka and run through the entire archive.
Parquet is also 13 years old and not the GOAT at compression either, there's been a lot of developments in research since that like Fastlanes and Btrblocks which Vortex implements to put it at the Pareto frontier