[1] https://juicefs.com/en/blog/company/2025-recap-artificial-in...
215 karma · joined January 2, 2014
Founder and CEO at Juicedata, Inc.
davies@juicedata.io
[1] https://juicefs.com/en/blog/company/2025-recap-artificial-in...
We are still working hard on it, hoping that we can help people with different workloads with different tech!
With POSIX API or HDFS, there is no such API to do that, unfortunately.
Inside the metadata service, it uses merkle tree (hash of hash) to verify the integrity of whole namespace (including id of data blocks) between RAFT replicas. Once we store the hash (4 bytes) of each objects into metadata, it should provide the integrity of the whole namespace.
If you already have the infra (databases and object stores), then JuiceFS is the easiest solution to have a distributed file system.
The underlying blob store systems is similar to DataNode or OSD in other distributed file system, could be slower than them a little bit because of the middle layers, the overall performance is determined by the disks.
So we can expect similar performance comparing to HDFS/CephFS, the benchmark results also confirm that.
PS, I'm founder of JuiceFS.
To serve S3 request, you can setup multiple S3 gateway and put a load director in front of them.
Per CAP, it's addressed by different meta engines (CP system, Redis, MySQL, TiKV) and also different object stores (AP system). When the meta engine is not available, the operation to JuiceFS will be blocked for a while and finally it returns EIO. When object store returns 404 (object not found), which means it's not consistent with the meta engine, it will be retried for a while, may return EIO if it's not recovered.
The file format is carefully designed to workaround the consistency issue from object store and local cache. Any part of data is written into object store and local cache with unique ID, so you will not go stale data once the metadata is correct [1].
Within a mount point, JuiceFS provides read-after-write consistency. Across clusters, JuiceFS provides open-after-close consistency, which should be enough for most of the applications, also provide good balance between consistency and performance.
[1] https://juicefs.com/docs/community/architecture/#how-juicefs...
https://ozone.apache.org/docs/current/feature/prefixfso.html
JuiceFS uses S3 as the underlying data storage, so S3 provides this durability SLA.
We will update the article to make it clear, thanks!
ps, I'm the founder of Juicedata.
One interesting use case is the backup of MySQL [1].
If the network is slow and unreliable, the experience will be bad, all fs operations will be blocked if meta server is not available, since JuiceFS is a CP system.
Comparing JuiceFS to S3QL [1] https://juicefs.com/docs/community/comparison/juicefs_vs_s3q...
Comparing to ObjectiveFS, the big difference is that JuiceFS has faster metadata engine (comparing to S3), so it can provide low latency access AND strong consistency.
Right now, there is a 48-hour-test policy, and we should have a negative COVID test within 48 hour to enter a public building. Once you have that, almost all services in Hangzhou are open. We can work, shop, meet as usual. The COVID test is free, available in every community, could be collected (mixed with 10 people) in 10 minutes. There are still a few cases some day. I believe we can stop the wide spread of Omicron in Hangzhou in this way.
Also, this is how Shenzhen fight with Omicron two months ago, it contain it, but Shanghai got it spread in the same time. Right now, the cases in Shanghai is declining, will be open soon hopefully, finger crossed!