Reading data via index would lead to scattered memory reads on actual data, which is a dead-end from an optimisation standpoint unfortunately.
Using index would have been much simpler but also slower.
Using index would have been much simpler but also slower.
Once you flush your input buffer, everything becomes sequential so after your 10 second window you no longer have this problem. And even for the latest 10 seconds reading an index won’t be terrible as far as performance because your input buffer is relatively small (compared to your huge dataset).
As far as compaction, I suspect for the kinds of workloads you are using this kind of storage for you won’t get much in terms of deduplication. IoT devices sending sensor readings are unlikely to produce loads of stable readings and the time stamps will be ever increasing.