Another option if you have enough local storage would be to use something like JuiceFS that creates a virtual file system where the files are initially written to the local cache before JuiceFS writes the data to your S3 provider as larger chunks.
SeaweedFS can do something similar if you configure it the right way. But both options require that you have enough storage outside of your object storage.
https://github.com/seaweedfs/seaweedfs/wiki/Cloud-Drive-Bene...
https://github.com/seaweedfs/seaweedfs/wiki/Cloud-Tier
https://github.com/seaweedfs/seaweedfs/wiki/Benchmarks
https://github.com/seaweedfs/seaweedfs/wiki/Words-from-Seawe...
https://github.com/seaweedfs/seaweedfs/wiki/Amazon-S3-API
...your true issue is it seems like you're using the filesystem as the "only" storage layer in play, but you also need time and entity querying(!?!).
>> we need to be able to query on a per sensor basis and a timespan
...look at the "Cloud-Tier" wiki page. If you're truly in an "everything's hot all the time" situation, you really should be using a database. If you're pulling "usually recent stuff, occasionally old stuff" then fronting with something like SeaweedFS seems like it might "just" transparently reduce your overall costs.
Really, I'd nudge towards "write .txt ; compact ... ; SELECT ... && cat .txt".
Basically, keep your inbound writes cached to (eg) seaweed as unit files. "Compact them" every hour by appending rows to some appropriate database (I mean: migrate to using litefs, turso, postgres, something like that). When you read, you may need to supplement "tip" data from your incoming files, but the majority should be hitting a "real" remote database, there's plenty to choose from!
A nifty note, sqlite can connect to multiple DB's at once: https://www.sqlite.org/lang_attach.html ... https://stackoverflow.com/posts/10020/revisions
...something like `select * from raw union (select * from one_hour) union (select * from today) union (select * from historical) ...`
Of course you could also use Aurora for a clean scalable Postgres that can survive zone failures for a simpler solution
[1] https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-tab...
Hopefully this will come at some point. Product looks very cool otherwise.
Then spin up duckdb and do some performance tests. I’m not sure this will work, there is some overheard with reading parquet, which is why it is discouraged to have small files and row groups.
Do you ever go back and reaggregate older data into bigger, sorted files? That is, maybe you originally partitioned by hour, but stale data is so infrequently accessed, you could roll up into partitions per week/month/whatever. Depending on the specifics, you might save some space from less file overhead and better compression statistics.
No more files. You might be able to avoid per usage pricing just by hosting this on a regular vps.
> To use all fsspec features, either install via pip install ratarmount[fsspec] or pip install ratarmount[fsspec]. It should also suffice to simply pip install fsspec if ratarmountcore is already installed.
You can also do this with a landing table or even branches+WAP.