Cryptsetup 2.0.0 introduces new on-disk LUKS2 format
kernel.org
kernel.org
> Write programs to handle text streams, because that is a universal interface.
A lot of the data formats I've written recently have used JSON. Partly because of that aforementioned principal, but mostly because it's just easier and there are very little downsides. Almost every language has support for it so it's fairly universal, it's self documenting, it's easy to debug, easy to manipulate by hand, and easy to maintain.
Put more simply: I've implemented several data formats using custom binary and several data formats using JSON. JSON was easier and faster every single time.
My recommendation for data formats: just use JSON; unless you have a really good reason not to.
And I do mean really good reason. For example, many might think concerns about data size would be a reason not to use JSON. But go ahead and run some JSON through a compression algorithm some time. You'll be amazed. Compressed JSON is very competitive versus a custom binary format.
You may not realize it, but you've presented a false dichotomy here.
The alternate to JSON isn't "custom binary", it's standardized binary encodings that allow the same flexibility (or additional flexibility) compared to JSON. Some examples include cbor & msgpack.
>> Write programs to handle text streams, because that is a universal interface.
cryptsetup manipulates block devices. I don't think that line can or should really be applied to something that is fairly similar to a filesystem's on-disk format.
Also, why use json for meta-data instead of simple C Structs? No human is supposed to read this kind of data anyways.
Unfortunately I don’t think that’s possible. If I scan your encrypted disk and see that 10% of blocks are zeroed out, I can assume that your disk is 90% full. So information has been leaked, which is arguably a vulnerability.
This would be on a level of a nice CS Bachelors/Masters thesis.
Personally, I would like to explore the idea of a secure enclave that keeps a map of which blocks are in use that gets referred to during write operations. This seems like a problem that is going to need to be solved with hardware.
I mean, I guess is depends on what burden of proof is needed in the case, certainly if it were balance of probability I'd be surprised if having the disk in a computer looking like it has been written to is sufficient to not make it plausible to deny.
LUKS2 format and features
~~~~~~~~~~~~~~~~~~~~~~~~~
The LUKS2 is an on-disk storage format designed to provide simple key
management, primarily intended for Full Disk Encryption based on dm-crypt.
The LUKS2 is inspired by LUKS1 format and in some specific situations (most
of the default configurations) can be converted in-place from LUKS1.
The LUKS2 format is designed to allow future updates of various
parts without the need to modify binary structures and internally
uses JSON text format for metadata. Compilation now requires the json-c library
that is used for JSON data processing.
On-disk format provides redundancy of metadata, detection
of metadata corruption and automatic repair from metadata copy.
NOTE: For security reasons, there is no redundancy in keyslots binary data
(encrypted keys) but the format allows adding such a feature in future.
NOTE: to operate correctly, LUKS2 requires locking of metadata.
Locking is performed by using flock() system call for images in file
and for block device by using a specific lock file in /run/lock/cryptsetup.
This directory must be created by distribution (do not rely on internal
fallback). For systemd-based distribution, you can simply install
scripts/cryptsetup.conf into tmpfiles.d directory.
For more details see LUKS2-format.txt and LUKS2-locking.txt in the docs
directory. (Please note this is just overview, there will be more formal
documentation later.)It seems like it could be better to use some easier to parse binary format that allows the same flexibility that json does, like cbor/msgpack/bson.
Using JSON here is a very strange choice.
I agree! Perhaps the userspace administrative tool (which links json-c) parses it and converts it to something more like an ioctl struct for the kernel.
(If you want to have a bad day, you can send a patch to Linus to put a JSON parser in the kernel.)
It's kind of shitty that it was called BSON, because it falls so far short of the generality that JSON has.
What the hell? Why would they do this? The only (tenuous) justification for using a parsed human-oriented text format in any protocol is so that humans can edit things by hand, but this will presumably not be the case for file metadata.
I don’t want to read too much into this since there might be a sane explanation, but this seriously makes me question the design of this security-critical system.
I don't know if that's the reason they chose JSON, but that's what I would think about IIWM.
This isn’t a use case for file system metadata. It only makes sense internally anyway. People interact with filesystem metadata through standardized APIs, not metadata dumps.
> It may also make debugging a lot easier since it is human readable.
It’s also basically guaranteed to introduce a ton of bugs; parsing and generating JSON is orders of magnitude more complicated than generating and reading an unambiguous tagged binary format.
https://github.com/libyal/libvslvm/blob/master/documentation...
https://gitlab.com/cryptsetup/cryptsetup/raw/v2.0.0/docs/LUK...