LCFS: A New Container Filesystem for Modern Datacenters
portworx.com
portworx.com
ZFS's dedup uses grotesque amounts of RAM since it insists on holding the dedup table in it and only doing dedup in realtime.
btrfs however allows you to scan the filesystem and do dedup of existing data later (even while it is mounted), using no RAM at all. This is one of a couple of places where I think btrfs is actually becoming a better-designed filesystem than ZFS.
It's literally set it, and forget it. While you gain 50% of your storage back.
Docker is already doing de-duplication by keeping shared data in image layers and expect the storage driver (graphdriver) to manage shared data between many container application efficiently without instantiating multiple copies of shared data in memory and disk. That problem can be solved by well known file system snapshot and cloning technologies and that is what LCFS is doing, and a lot more efficiently than any existing storage drivers out there.
Also curious to know if anyone is using it in production and what their experiences are.
It is an experimental release at this point. We would like to receive feedback from the community and make it successful for everybody. Our tests so far look very promising.
The fundamental thing here is that (a) it's taking a specific approach [with layers], (b) we're in it for the long-haul, and (c) we want this to be in the open. We’ll update as the battle-testing, production progress continues. So 100% agree with the question/point.
Docker is already doing a good job of de-duplication by identifying shared data between applications and keeping a unique copy of shared data on disk. But many of the existing storage drivers today, would create multiple copies of such shared data in memory when many applications (containers) consume such shared data. Applications are running on top of file systems and each need a separate instance of the file system and Linux kernel does not know if all those file systems are sharing same data underneath. Each application also require a unique view of the file system and they may choose to modify their file system (thus their shadow copy of shared data) and such changes should be hidden from other applications (containers). That is where cloning technologies come into play. And there are many other things a docker storage driver need to solve. Rather than focusing on one aspect of the whole problem, it would be better to look at the whole picture and see if there are any solutions out there solving all those problems correctly and efficiently for a docker storage driver.
I think all the energy spent around designing image layer file system after image layer file system is good evidence that the Docker approach isn't as easy as it may seem.
We looked at the requirements of a docker storage driver and built a solution from scratch, just to address those efficiently and nothing else.
For example, what one wants with launching multiple instances of an image is to be able to clone the image. Today, to accomplish that, file system snapshots are used to simulate this feature. And that is where a lot of the problems start happening. LCFS does things purpose built for these use cases.
Currently, the install experience is not the greatest... We rely on the Docker v2 plugin interface and that is still in beta. We are working with the Docker team to help us smooth out some of these rough edges.
LCFS will also support other container formats and any help from the community is much appreciated.
Is anyone using LCFS in production? Why did you get started building it?
We started building this because there are a lot of issues running containers in production, specifically around image space and inode management and also with page cash usage.
This file system was specifically designed to address memory and storage management of container images
This project is currently in experimental mode. We would appreciate any feedback you may have.
#!/bin/ksh
typeset -A hashes
find . -type f -exec shasum -a 512 {} \; |
while read hash filename; do
if [[ -z ${hashes[$hash]} ]]; then
hashes[$hash]="$filename"
else
echo ln -f "'${hashes[$hash]}'" "'$filename'"
fi
done
Pipe to /bin/sh if you're happy with the results.see also jdupes fork of fdupes