HNHacker News
TopNewBestAskShowJobs

daviesliu

215 karma · joined January 2, 2014

Happy hacking

Founder and CEO at Juicedata, Inc.

davies@juicedata.io

submissionscomments
daviesliu··on JuiceFS is a distributed POSIX file system built on top of Redis and S3
JuiceFS can be scaled to hundreds of PB by design, also is verified by thousands of users in production [1].

[1] https://juicefs.com/en/blog/company/2025-recap-artificial-in...

daviesliu··on Launch HN: Regatta Storage (YC F24) – Turn S3 into a local-like, POSIX cloud FS
Founder of JuiceFS here, congrats to the Launch! I'm super excited to see more people doing creative things in the using-S3-as-file-system space. When we started JuiceFS back in 2017, applied YC for 2 times but no luck.

We are still working hard on it, hoping that we can help people with different workloads with different tech!

daviesliu··on A Distributed File System in Go Cut Average Metadata Memory Usage to 100 Bytes
This post said another topic: how to backup the metadata of JuiceFS in readable format (JSON) and restore it into an empty database.
daviesliu··on A Distributed File System in Go Cut Average Metadata Memory Usage to 100 Bytes
The recover process is similar to what a database does after a crash, it loads the recent snapshot in disk and apply any newer transaction logs.
daviesliu··on A Distributed File System in Go Cut Average Metadata Memory Usage to 100 Bytes
The S3 API allow user to specify the hash of content as HTTP header, it will be verified by the JuiceFS gateway and persisted into JuiceFS as ETag.

With POSIX API or HDFS, there is no such API to do that, unfortunately.

daviesliu··on A Distributed File System in Go Cut Average Metadata Memory Usage to 100 Bytes
JuiceFS relies on the object store to provide integrity for data. Besides that, JuiceFS stores the checksum of each object as tags in S3, and verifies that when downloading the objects.

Inside the metadata service, it uses merkle tree (hash of hash) to verify the integrity of whole namespace (including id of data blocks) between RAFT replicas. Once we store the hash (4 bytes) of each objects into metadata, it should provide the integrity of the whole namespace.

daviesliu··on A Distributed File System in Go Cut Average Metadata Memory Usage to 100 Bytes
The company may want the product to be difficult to use, so they can sell more support, that's not win-win for both sides.
daviesliu··on A Distributed File System in Go Cut Average Metadata Memory Usage to 100 Bytes
Agreed, the all-in-one solution (Ceph) should be better, if you have to setup all the components.

If you already have the infra (databases and object stores), then JuiceFS is the easiest solution to have a distributed file system.

daviesliu··on A Distributed File System in Go Cut Average Metadata Memory Usage to 100 Bytes
JuiceFS is similar to HDFS/CephFS/Lustre, so it MUST has a component to manage metadata, similar to NameNode of HDFS or MDS of CephFS, this point of failure is the problem we have to address.

The underlying blob store systems is similar to DataNode or OSD in other distributed file system, could be slower than them a little bit because of the middle layers, the overall performance is determined by the disks.

So we can expect similar performance comparing to HDFS/CephFS, the benchmark results also confirm that.

daviesliu··on Gcsfuse: A user-space file system for interacting with Google Cloud Storage
If you really expect a file system experience over GCS, please try JuiceFS [1], which scales to 10 billions of files pretty well with TiKV or FoundationDB as meta engine.

PS, I'm founder of JuiceFS.

[1] https://github.com/juicedata/juicefs

daviesliu··on Open source cloud file system. Posix, HDFS and S3 compatible
Usually the meta engine or object storage can scale horizontally by itself, JuiceFS is middleware to talk to these two services.

To serve S3 request, you can setup multiple S3 gateway and put a load director in front of them.

daviesliu··on Open source cloud file system. Posix, HDFS and S3 compatible
Agreed, it's very hard, that's why GFS and HDFS had give up some parts of POSIX compatibility.

Per CAP, it's addressed by different meta engines (CP system, Redis, MySQL, TiKV) and also different object stores (AP system). When the meta engine is not available, the operation to JuiceFS will be blocked for a while and finally it returns EIO. When object store returns 404 (object not found), which means it's not consistent with the meta engine, it will be retried for a while, may return EIO if it's not recovered.

The file format is carefully designed to workaround the consistency issue from object store and local cache. Any part of data is written into object store and local cache with unique ID, so you will not go stale data once the metadata is correct [1].

Within a mount point, JuiceFS provides read-after-write consistency. Across clusters, JuiceFS provides open-after-close consistency, which should be enough for most of the applications, also provide good balance between consistency and performance.

[1] https://juicefs.com/docs/community/architecture/#how-juicefs...

daviesliu··on Open source cloud file system. Posix, HDFS and S3 compatible
Atomic file/directory renames/moves is the fundamental feature of JuiceFS, which makes it truely a file system rather than a proxy to S3, please check the docs for all the compatibility details [1].

https://github.com/juicedata/juicefs#posix-compatibility

daviesliu··on Open source cloud file system. Posix, HDFS and S3 compatible
This is an experimental feature to do this, still working on it.
daviesliu··on Open source cloud file system. Posix, HDFS and S3 compatible
It's doable to run a MinIO gateway on top of CephFS mount point, but that will has performance issue, especially for multipart-upload and copy. That's why we put MinIO and JuiceFS client together and use some internal API to do zero-copy uploads.
daviesliu··on Open source cloud file system. Posix, HDFS and S3 compatible
Yes, JuiceFS uses the Apache 2 fork [1] directly (master branch), but also provide a full featured S3 gateway (gateway branch) under AGPL for people' choice.

[1] https://github.com/juicedata/minio/tree/master

daviesliu··on Open source cloud file system. Posix, HDFS and S3 compatible
Yes, the data can be encrypted [1] by the client before sending to S3, but the metadata is not encrypted.

[1] https://juicefs.com/docs/community/security/encrypt

daviesliu··on Open source cloud file system. Posix, HDFS and S3 compatible
The button to switch language is at the bottom of right-top menu, we will fix that.
daviesliu··on Open source cloud file system. Posix, HDFS and S3 compatible
Apache Ozone is not POSIX compatible, even with the File System Optimized format [1].

https://ozone.apache.org/docs/current/feature/prefixfso.html

daviesliu··on Open source cloud file system. Posix, HDFS and S3 compatible
99.99999999% reliability means you will not loss more than one byte in every 10 GB in a year.

JuiceFS uses S3 as the underlying data storage, so S3 provides this durability SLA.

daviesliu··on Posix Compatibility Comparison: GCP Filestore, Amazon EFS, and Azure Files
Yes, we picked the default one in docs, which should be NFS v3, will redo the test against NFS v4 and update the article, thanks!
daviesliu··on Posix Compatibility Comparison: GCP Filestore, Amazon EFS, and Azure Files
Azure has Azure Files and Azrue NetApp Files, the later one is provided from NetApp. Azure Files was used in the article, maybe you are using NetApp Files?

We will update the article to make it clear, thanks!

daviesliu··on Open source cloud file system. Posix, HDFS and S3 compatible
Juicedata Inc is a US company, was registered in Delaware. The founding team are Chinese.

ps, I'm the founder of Juicedata.

daviesliu··on Posix Compatibility Comparison: GCP Filestore, Amazon EFS, and Azure Files
We had compared nfs of Azure in the article, I believe cifs/samba should be much worse.
daviesliu··on Open source cloud file system. Posix, HDFS and S3 compatible
Yes, JuiceFS is not a good choice for PG, unless if you don't care the performance.

One interesting use case is the backup of MySQL [1].

[1] https://juicefs.com/docs/cloud/backup_mysql_in_juicefs/

daviesliu··on Open source cloud file system. Posix, HDFS and S3 compatible
The JuiceFS Cloud supports ACL, but open source one does not support it yet.
daviesliu··on Open source cloud file system. Posix, HDFS and S3 compatible
JuiceFS supports create-if-not-existed by using the Java SDK (HDFS compatible), so I guess it should work well with Delta.io.
daviesliu··on Open source cloud file system. Posix, HDFS and S3 compatible
We have a fork of MinIO at https://github.com/juicedata/minio, which will be maintained by us.
daviesliu··on Show HN:JuiceFS 1.0, A POSIX, HDFS, and S3 compliant cloud file system
If you have good network, JuiceFS is the thing you are looking for. Multiple hosts can mount the same volume and it just works without conflicts.

If the network is slow and unreliable, the experience will be bad, all fs operations will be blocked if meta server is not available, since JuiceFS is a CP system.

Comparing JuiceFS to S3QL [1] https://juicefs.com/docs/community/comparison/juicefs_vs_s3q...

Comparing to ObjectiveFS, the big difference is that JuiceFS has faster metadata engine (comparing to S3), so it can provide low latency access AND strong consistency.

daviesliu··on Not a single car was sold in Shanghai last month
I live in Hangzhou, 150 miles away to Shanghai. There are a few positive COVID-19 cases in past weeks and we prepared to lockdown, but it did not come.

Right now, there is a 48-hour-test policy, and we should have a negative COVID test within 48 hour to enter a public building. Once you have that, almost all services in Hangzhou are open. We can work, shop, meet as usual. The COVID test is free, available in every community, could be collected (mixed with 10 people) in 10 minutes. There are still a few cases some day. I believe we can stop the wide spread of Omicron in Hangzhou in this way.

Also, this is how Shenzhen fight with Omicron two months ago, it contain it, but Shanghai got it spread in the same time. Right now, the cases in Shanghai is declining, will be open soon hopefully, finger crossed!

Page 1 of 4Next →