Super interesting! I’m curious how this differs from InfluxDB’s German strings implementation https://www.influxdata.com/blog/faster-queries-with-stringvi...
German-style strings/views are not a compression algorithm, they're just a way for storing string data and making it quick to compare them in-memory. You can in fact store views, while storing the corresponding full-length strings in compressed format with FSST. We don't currently do that but we're working on it.
[1] https://arrow.apache.org/docs/format/Columnar.html#variable-...
[2] https://db.in.tum.de/~freitag/papers/p29-neumann-cidr20.pdf