Gcsfuse: A user-space file system for interacting with Google Cloud Storage
github.com
github.com
I didn’t know GCS supported appends efficiently. Correct me if I’m wrong, but I don’t think S3 has an equivalent way to append to a value, which makes it clunky to work with as a log sink.
Obviously, you need to make sure no write operation breaks a record midway while doing that. (unlike the posix write() API which can be interrupted midway).
If you need order then probably writing to separate files and having compaction jobs is still better.
0: https://cloud.google.com/storage/docs/xml-api/reference-head...
Even so, I avoid parallel appenders to a single file, it can be hard to reason about and debug in ways that having each process appending to its own file doesn't.
Doing so has huge benefits: Write performance is way higher, you can do rollbacks easily (just ignore the tail of the files), you can do snapshotting easily (just make a new file and include by reference a byte range of the parent), etc.
The downside is from time to time you need to make a new file and chuck out the dead data - but such an operation can be done 'online', and can be done during times of lower system load.
A WAL would usually disappear or truncate it's length after a while, and you'd only rerun things from it if you absolutely have two. Changes in business requirements shouldn't require you to do anything with a WAL.
In contrast, Event sourcing log would be kept indefinitely, so when business requirements change, you could (if you want to, not required) re-run N previous events so you can apply new changes to old data in your data storage.
But, if you really want to, it's basically the same, but in the end, applied differently :)
Bringing back the topic to what the parent was saying; since S3 is a pretty common system, and a distributed system at that, are you saying that S3 does support appending data? AFAIK, S3 never supported any append operations.
Not as elegant / fast as GCS’s and there may be other subtleties, but it’s possible to simulate.
Basically if your data is append only (such as a log), buffer whatever reasonable amount is needed, and then put a new version of the file with said data (recording the generated version ID AWS gives you). This gets added to the "stack" of versions of said S3 object. To read them all, you basically get each version from oldest to newest and concatenate them together on the application side.
Tracking versions would need to be done application side overall.
You could also do "random" byte ranges if you track the versioning and your object has the range embedded somewhere in it. You'd still need to read everything to find what is the most up to date as some byte ranges would overwrite others.
Definitely not the most efficient but it is doable.
You are correct. (There are multipart uploads, but that's kinda different.)
ELB logs are delivered as separate object every few minutes, FWIW.
The scale of data for my work is modest (~50TB, ~1 million files total, about 50k files per "directory").
I've used all the main FUSE cloud FS (gcsfuse, s3-fuse, rclone, etc) and they all end up falling over in prod.
I think a better approach would be to port all the important science codes to work with file formats like parquet and use user-space access libraries linked into the application, and both the access library and the user code handle errors robustly. This is how systems like mapreduce work, and in my experience they work far more reliably than FUSE-mounts when dealing with 10s to 100s of TBs.
Then my work must be downright embarassing.
ExtFUSE (optimized FUSE with eBPF) [1] can offer you much higher performance. It caches metadata in the kernel to avoid lookups in user space. Disclaimer: I built it.
https://docs.aws.amazon.com/AmazonS3/latest/userguide/optimi...
This comes up because we store a lot of image data where time-to-first-byte affects the user experience but the access patterns preclude caching unless we are willing to spend $$$.
If the reason you were unable to use a CDN cache was because your access patterns require a lot of varying end serializations (due to things like image manipulation, resizing, cropping, watermarking, etc.), then this API could be a huge money saver for you. It was for me.
OTOH if the cost was because compute isn't free and the corresponding cloudflare worker compute cost is too much, then yeah, that's a tough one... I don't have a packaged answer for you, but I would investigate something like ThumbHash: https://evanw.github.io/thumbhash/ - my intuition is that you can probably serve some highly optimized/interlaced/"hashed" placeholder. The advantage of thumbhash here could be that you can optimize the access pattern to be less spendy by simply storing all of your hashes in an optimized way, since they will be extremely small, like small enough to be included in an index for index-only scans ("covering indexes"). (I have not actually tried this.)
A 4U server with 24 8tb 2.5” $400 consumer sata ssds is probably the best bang for the buck. Probably about $20k plus hosting fees. That’s 192TB of storage.
See https://www.nature.com/articles/s41592-021-01326-w figure 1 a and b for a demonstration of the distinct time-to-first-byte for uncached data (local POSIX, S3, and nginx/http). I'm trying to find a way to accelerate that TTFB for a dataset that lives in S3 and we don't know which parts could be cached before a user shows up and clicks on a dataset.
(I know it's a weird use case. It's not one I particularly want to support, as my own interests are more in the large-scale batch compute than the interactive user).
Looking at your figure, almost a second to download a 128KB chunk with HDF5 seems extremely high. How many requests are being made here? Is the HDF5 metadata being re-read from scratch for each iteration of the benchmark?
It might be interesting to trace log each http request being made and its timing during an iteration of the benchmark.
Using AWS CloudShell I see range requests for the last 128KB of a 512MB S3 file without authentication taking about 25ms when reusing a connection and 35ms when not.
> I'm trying to find a way to accelerate that TTFB for a dataset that lives in S3 and we don't know which parts could be cached before a user shows up and clicks on a dataset.
If latency from S3 turns out to be the problem I'm not sure there's any alternative but paying for the space your data uses on fast storage. The cheapest option on AWS will likely be EBS SSD volumes. That runs to $0.08/GB/month or $8000/month for 100TB (4x more than S3) before IOPS charges. You could try they EBS HDD volumes but they do not advertise latency figures and are still 2x S3 pricing.
We have already tried EFS and it worked for our needs but cost about the same as your EBS. We had to enable the "EFS go fast" button (at least it exists!) which greatly increases the cost.
Note, it’s not that the network is slow, it’s that object storages aren’t designed for low latency access (but with parallel reads can serve data at high bandwidth.)
> From reading the docs, it looks very similar to `rclone mount` with `--vfs-cache-mode off` (the default). The limitations are almost identical.
> However rclone has `--vfs-cache-mode writes` which caches file writes to disk first to allow overwriting in the middle of a file and `--vfs-cache-mode full` to cache all objects on a LRU basis. They both make the file system a whole lot more POSIX compatible and most applications will run using `--vfs-cache-mode writes` unlike `--vfs-cache-mode off`.
https://news.ycombinator.com/item?id=35788919
Seems rclone would be an even better option than Google's own tool.
Also hi capableweb I think your name rings a bell from LLM's/Gen AI threads
Hello! That's probably a sign I need to take a break from writing too many HN comments per day, thanks :)
Instead, tar up the files in some random order, and put the tar file on a web server or bucket, then stream then in during the first epoch, while keeping track of their byte offsets in the tar file, which you cache locally, assuming ample local Flash storage. Then permute the list of offsets and use those when reading samples for the next epoch.
If you only have local HDD then you will need a more advanced data structure like the one provided by https://github.com/jacobgorm/mindcastle.io , which will allow you to write out permuted samples at close to disk sequential write bandwidth. See my talk at USENIX Vault 2019 for a full explanation, linked from https://vertigo.ai/mindcastle/
[0] Seafowl is an early stage open source database written in Rust. https://seafowl.io/
[1] https://github.com/splitgraph/seafowl-gcsfuse
Post https://www.splitgraph.com/blog/deploying-serverless-seafowl
[2] Previous post https://news.ycombinator.com/item?id=36215088You'd have to use your own cache otherwise. IME the OS-level page cache is actually quite effective at caching reads and seems to work out of the box with gcsfuse.
1. Cache of file attributes in the Kernel (this is controlled by "stat-cache-ttl" value - https://github.com/GoogleCloudPlatform/gcsfuse/blob/7dc5c7ff...) 2. Cache of directory listings 3. Cache of file contents
It should be possible to use (2) and (3) for a better performance but might need changes to the underlying fuse library they use to expose those options.
Let’s say I have a directory tree with 100MM files in a nested structure, where the average file is 4+ directories deep. When I `ls` the top few directories, is it fast? How long until I discover updates?
Reading the docs, it looks like it’s using this API for traversal [0]?
What about metadata like creation times, permission, owner, group?
Any consistency concerns?
[0] https://cloud.google.com/storage/docs/json_api/v1/objects/li...
PS, I'm founder of JuiceFS.
We've got some additional documentation on the differences and limitations between Gcsfuse and a proper POSIX filesystem: https://cloud.google.com/storage/docs/gcs-fuse#expandable-1
Gcsfuse is a great way to mount Cloud Storage buckets and view them like they're in a filesystem. It scales quite well for all sorts of uses. However, Cloud Storage itself is a flat namespace with no built-in directory support. Listing the few top level directories of a bucket with 100MM files more or less requires scanning over your entire list of objects, which means it's not going to be very fast. Listing objects in a leaf directory will be much faster, though.
Our theoretical usecase is 10+ PB and we need multiple TB/s of read throughout (maybe of fraction of that for writing). So I don’t think Filestore fits this scale, right?
As for the directory traversals, I guess caching might help here? Top level changes aren’t as frequent as leaf additions.
That being said, I don’t see any (caching) proxy support anywhere other than the Google CDN.
It has some nice features like streaming with block level caching for fast readonly access
Haven’t used it but it looks cool, if a bit immature.
It even supports GCS (as GCS has S3 compatible API)
https://github.com/s3fs-fuse/s3fs-fuse/wiki/Google-Cloud-Sto...