The scale of data for my work is modest (~50TB, ~1 million files total, about 50k files per "directory").
The scale of data for my work is modest (~50TB, ~1 million files total, about 50k files per "directory").
ExtFUSE (optimized FUSE with eBPF) [1] can offer you much higher performance. It caches metadata in the kernel to avoid lookups in user space. Disclaimer: I built it.
https://docs.aws.amazon.com/AmazonS3/latest/userguide/optimi...
This comes up because we store a lot of image data where time-to-first-byte affects the user experience but the access patterns preclude caching unless we are willing to spend $$$.
If the reason you were unable to use a CDN cache was because your access patterns require a lot of varying end serializations (due to things like image manipulation, resizing, cropping, watermarking, etc.), then this API could be a huge money saver for you. It was for me.
OTOH if the cost was because compute isn't free and the corresponding cloudflare worker compute cost is too much, then yeah, that's a tough one... I don't have a packaged answer for you, but I would investigate something like ThumbHash: https://evanw.github.io/thumbhash/ - my intuition is that you can probably serve some highly optimized/interlaced/"hashed" placeholder. The advantage of thumbhash here could be that you can optimize the access pattern to be less spendy by simply storing all of your hashes in an optimized way, since they will be extremely small, like small enough to be included in an index for index-only scans ("covering indexes"). (I have not actually tried this.)
A 4U server with 24 8tb 2.5” $400 consumer sata ssds is probably the best bang for the buck. Probably about $20k plus hosting fees. That’s 192TB of storage.
See https://www.nature.com/articles/s41592-021-01326-w figure 1 a and b for a demonstration of the distinct time-to-first-byte for uncached data (local POSIX, S3, and nginx/http). I'm trying to find a way to accelerate that TTFB for a dataset that lives in S3 and we don't know which parts could be cached before a user shows up and clicks on a dataset.
(I know it's a weird use case. It's not one I particularly want to support, as my own interests are more in the large-scale batch compute than the interactive user).
Looking at your figure, almost a second to download a 128KB chunk with HDF5 seems extremely high. How many requests are being made here? Is the HDF5 metadata being re-read from scratch for each iteration of the benchmark?
It might be interesting to trace log each http request being made and its timing during an iteration of the benchmark.
Using AWS CloudShell I see range requests for the last 128KB of a 512MB S3 file without authentication taking about 25ms when reusing a connection and 35ms when not.
> I'm trying to find a way to accelerate that TTFB for a dataset that lives in S3 and we don't know which parts could be cached before a user shows up and clicks on a dataset.
If latency from S3 turns out to be the problem I'm not sure there's any alternative but paying for the space your data uses on fast storage. The cheapest option on AWS will likely be EBS SSD volumes. That runs to $0.08/GB/month or $8000/month for 100TB (4x more than S3) before IOPS charges. You could try they EBS HDD volumes but they do not advertise latency figures and are still 2x S3 pricing.
We have already tried EFS and it worked for our needs but cost about the same as your EBS. We had to enable the "EFS go fast" button (at least it exists!) which greatly increases the cost.
Note, it’s not that the network is slow, it’s that object storages aren’t designed for low latency access (but with parallel reads can serve data at high bandwidth.)
Then my work must be downright embarassing.
I've used all the main FUSE cloud FS (gcsfuse, s3-fuse, rclone, etc) and they all end up falling over in prod.
I think a better approach would be to port all the important science codes to work with file formats like parquet and use user-space access libraries linked into the application, and both the access library and the user code handle errors robustly. This is how systems like mapreduce work, and in my experience they work far more reliably than FUSE-mounts when dealing with 10s to 100s of TBs.