I want XAES-256-GCM/11
words.filippo.io
words.filippo.io
Inability to encrypt more than 64Gb with the same (key,nonce) pair.
Lack of commitment (whether key-commitment, or key+nonce+ad commitment).
If one is seriously considering breaking away from existing GCM standards to create yet-another-standard, such proposal would need to offer improvements in all areas (ex. a proposed standard for converting any AEAD into streaming chunk-based AEAD with practically unlimited message sizes under the same (key,nonce) and unlimited message counts.
GCM-256 is ubiquitous and is often the preferred choice for all the reasons mentioned by the author, but that very argument is what makes non-standard GCM with 11 AES-rounds silly.
In 2023 we should be working on new standards that "wrap" existing crypto-primitives (which are already implemented/available in countless hardware-accelerated libraries/APIs) to get additional features/benefits/capabilities - not musing about AES with 10+1 rounds or SHA-512-really-fast with 80-1 rounds..
I'm curious to hear more about what you've seen. My naive hope was that a proper streaming decrypt API would be enough of a pit of success that developers wouldn't be tempted to sabotage themselves.
Chunked streaming can make the difference smaller, but even that "small" difference is beyond what is relevant to say ... filling an L1 cache, or waiting a round-trip. Some of the cases of "read before auth" I've seen have been on very small messages, but in contexts where the incentives are even further driven up, like trading or bidding protocols. It just left me thinking that we should enforce AEAD mathematically. Many practitioners often assume it already is enforced!
I think a better way is to derive the content per file part and then use a ratcheting nonce to encrypt the subparts. That also gives you random access into the entire file in ~O(1) (i.e. no need to decrypt the entire file) and the ability to interrupt and resume decryption. Unfortunately, there's no standard that describes how the output should be serialized, so tool interop becomes a problem. Although, to be fair, there's no serialization standard for AES either (i.e. what do you do with the nonce?) so it's probably not a big deal.
So you can do 512Gb or 64GB under one nonce. Then simply increment the fixed field and run the next 64GB under the same key and new nonce (nonce+1) and so on. In essence, it's the same thing is just making the fixed field smaller and the counter bigger, but meets the letter and intent of the law. The "fixed field" can be anything the user wants, including being "constructed from two or more smaller fields". And it is not constrained to remain the same under multiple invocations. Still compatible with FIPS and common implementations. It doesn't have to be some fancy ratcheting scheme.
The initial "fixed field" or nonce could even always just be all zeros [1]. It doesn't matter, it's not secret.
If for some reason you want to encrypt that much under one key, which I think you really don't.
1: well, in most cases especially AES-256: https://crypto.stackexchange.com/questions/68774/can-a-zero-...
If all your keys are ephemeral this isn't a big worry, but if they aren't, you can end up talking about reliably keeping state between invocations of your whole program.
(Apologies if this is obvious!)
That's an exaggeration. Reusing the nonce in GCM allows decryption of messages with the same nonce. It does NOT compromise the key.
It's a cipher mode. You can use GCM with any block cipher. OK, I assume that you meant AES-GCM.
But GCM as a construction in itself is not vulnerable to chosen ciphertext attacks, as long as the underlying symmetric cipher is secure.
GCM will lose the authentication property, if you know the authentication key, which you _might_ be able to get if you can mount a chosen _plaintext_ attack under conditions of nonce reuse. Simply getting a couple of random messages with the same nonce is NOT enough.
AES-GCM as specified has a nonce that is large enough to not care about it in practical cases (e.g. TLS), and it can become a problem only in very unrealistic cases (attacker-controlled likely exabyte-sized plaintexts).
These cases are maybe _juuuust_ in the realm of possibility, if you have access to a supercomputer, and you want to specifically design an application that is vulnerable to an attack, and then allow your adversary to covertly connect to your supercomputer cluster. To be clear, we're talking here about repurposing the entire NSA computing and storage power to host this single application, and allowing the attacker (e.g. Russian troll farms) to completely control the plaintexts that it transmits.
Extending the nonce to 256 bits would move that from outside the realm of possibility even for a contrived scenario. It's not a bad idea, but it's also not at all an urgent one.
Yes it is. You simply XOR the two auth tags and then compute the roots of the resulting polynomial (with known coefficients). There typically aren’t that many candidate roots to test. This has been known since GCM was first specified, see eg Joux’s comments: https://csrc.nist.gov/csrc/media/projects/block-cipher-techn...
It’s clear from your comments here and elsewhere that you don’t know what you are talking about, so I’ll take tptacek’s advice and bow out here.
The authentication key is _derived_ from the AES key, but they're not the same.
And no, you can't recover the encryption key (i.e. the thing that allows you to decrypt messages) from any weakness in the nonce choice.
I think you should stop digging. Sean and Hanno gave a Black Hat talk whose slides were unwillingly hosted on a GCHQ website because of this problem.
just like any counter mode. it's vitally important, but not difficult to understand or implement.
the other point is, WHY NOT JUST ROLL THE KEY MORE OFTEN. nobody should be encrypting 64GB under the same key. and 96+256 is enough bits that can be chosen randomly to never worry about collisions.
If you create the opportunity to make a mistake remembering to freshen a nonce, even if that opportunity is remote, such that you'd never trip over it accidentally, you've given attackers a window to elaborately synthesize that accident for you. That's what a vulnerability is.
There is a whole subfield of cryptography right now dedicated to "nonce misuse resistance", motivated entirely by this one problem. This is what I love about cryptography. You could go your entire career in the rest of software security and not come up with a single new bug class (just instances of bug patterns that people have been finding for years). But cryptography has them growing on trees, and it is early days for figuring out how to weaponize them.
That's why people pay so much attention to stuff like nonce widths.
[0] "Linear time" means O(n), which is computer science speak for "f(n) = an, but we ignore the a because we're only interested in the fastest growing component of the function". In some cases the a is so large as to make the algorithm in question not feasible for ANY size. Let's just pretend the attacker doesn't have that problem.
...and rationally and scientifically assuming that the rate of the progress won't increase and that there will be no major breakthroughs in the future, we propose revised number of rounds for AES, BLAKE2, ChaCha, and SHA-3.
The crypto is already fast enough, thank you very much; many attacks work only precisely it's quite fast to brute force huge subspaces of key material.
Rogaway says he has released all his OCB patents into the public domain[1], so it sounds like OCB might be an option for greenfield projects? OTOH, I don’t see why you’d use any kind of AES in a greenfield project outside the surreal FIPS world, given how absolutely miserable constant-time software for it is. (ChaCha is kind of miserable when you don’t have at least a 32-bit ALU with a barrel shifter—that is on exactly the kind of platforms where table-based AES might actually be constant-time—but I’m not sure that’s a good enough reason.)
[1] https://mailarchive.ietf.org/arch/msg/cfrg/qLTveWOdTJcLn4HP3...
Encrypt-then-MAC remains the most conservative and theoretically secure option.
Leaving aside the (very serious) nonce reuse issue, the cracks on non-committing AEADs in general (such as AES-GCM) are already showing. Partitioning oracle attacks affect all of them: https://crypto.stackexchange.com/questions/88716/understandi...
There are also other minor GCM-specific issues (weak keys etc.). None of the issues are cypher-breaking, but I wouldn't say that AES-GCM is automatically the best choice for everything.
AEADs are obviously better than EtM, because EtM doesn't allow for authenticating the unencrypted context.
I wrote about turning CTR+HMAC into a committing AEAD and promptly screwing it up badly: https://soatok.blog/2021/07/30/canonicalization-attacks-agai...
The only thing you can do with an integrated AEAD that you can't do with a constructed one (with standard interface and security) is include authenticated and unencrypted context halfway through an encryption.
You can specify an EtM construction that accepts additional authenticated data. However, you can also do it insecurely (as the post I linked above describes) without realizing you did it insecurely. This is why most people prefer to use cryptographer-approved AEAD modes.
> In fact, you must ensure that the nonce, a piece of unencrypted context, is authenticated.
For CBC mode, sure. For CTR mode? Not really.
> Nothing stops you from throwing more stuff in there.
What prevents an attacker from shifting bits from the ciphertext field into the AAD field in the decrypt path and yield the same HMAC tag? Unless you have an answer to this question, vanilla "encrypt then MAC" is not sufficient. You need a better-engineered construction than that.
I'm pretty sure the linked post covered all of this nuance. Please let me know if something wasn't clear, or you feel it was missing.
Sure, I agree with this. But then the advantage of AEAD over a bespoke EtM is not that AEAD allows the authentication of unencrypted context.
>> In fact, you must ensure that the nonce, a piece of unencrypted context, is authenticated.
> For CBC mode, sure. For CTR mode? Not really.
If you don't, you do not get ciphertext integrity: decryption will succeed, but mostly yield gibberish, if the adversary changes the nonce in a decryption query. This may expose a padding oracle, with all the nice attacks those things allow, depending on details of the application.
>> Nothing stops you from throwing more stuff in there.
> What prevents an attacker from shifting bits from the ciphertext field into the AAD field in the decrypt path and yield the same HMAC tag? Unless you have an answer to this question, vanilla "encrypt then MAC" is not sufficient. You need a better-engineered construction than that.
Yes, you need a well-engineered construction.
> Please let me know if something wasn't clear, or you feel it was missing.
And yes, your linked post covers all this, but that is not the point: your summary of the linked post just claims superiority of AEAD (which I took as integrated AEAD modes) over EtM because of a functionality you claim is missing from the latter. But in the same way you need a well-engineered integrated AEAD to get any kind of security, you will need a well-engineered Encrypt-then-MAC-with-AD construction to get a secure construction. And here "well-engineered" means "ensure unambiguous parsing of decryption inputs," we're not talking about high-flying stuff that doesn't have standard solutions.
In short: I accept the point of your linked post, and I agree with it. But I reject the claim that a functionality mismatch is what makes integrated AEAD better than a constructed EtM.
Please describe the padding oracle attack against AES-CTR you're envisioning.
> In short: I accept the point of your linked post, and I agree with it. But I reject the claim that a functionality mismatch is what makes integrated AEAD better than a constructed EtM.
Okay, I don't think we disagree then. We're just debating semantics at this point. :)
But when people say "use AES-CBC + HMAC" and cite Signal as an example, and Signal's implementation does this: https://github.com/signalapp/Signal-Android/blob/main/app/sr...
Well, when that happens, I feel the need to pipe in :)
If you're careful enough to not implement a naive protocol that stitches AES+CBC and HMAC-SHA2 together (or, as tptacek put it in a podcast episode, throw some crypto potions into a cauldron and see what happens), you're probably the minority of crypto-savvy people.
That's very vague and therefore not very helpful. Could you say what exactly is wrong with the code you linked?
They provide AE, not AEAD.
They feed an IV and ciphertext into HMAC. They don't feed additional authenticated data.
If someone followed Signal's example, they either wouldn't have AEAD, or they're likely to make the exact mistake described in the post I linked above.
I don't know how to be more helpful here. I've been only repeating myself.
AEAD modes let you bind a ciphertext to a context without increasing bandwidth. This is super important for database cryptography. Read more: https://soatok.blog/2023/03/01/database-cryptography-fur-the...
Whether "it's not AEAD" matters for an application depends on many factors. Signal doesn't need it.
The IV can be any length up to 2^64-1. The reason for picking 96-bit IVs is that other values require an extra invocation of GHASH (page 15). The document recommends 96-bit for interoperability but that's by no means a requirement.
The X part is thus in theory already allowed.
That said, I find the wording of dedicated counter space a bit bizarre. There are only 128 bits per block in AES - in fact I'm not aware of a widely deployed cipher which a larger block size. Whatever goes through ghash becomes the initial plaintext passed to aes under the key (this is the counter) part and then you just increment. This is just a limitation of counter mode in general: the whole IV is technically counter space and is treated as whole when evaluating the birthday bound issue.
The critical part mentioned in the post is that XChaCha "hashes the key+nonce into a fresh key". A very similar technique is used un: AES-GCM-SIV (https://datatracker.ietf.org/doc/html/rfc8452#section-9). The improved security here comes not from any "counter space" or nonce extension per se but the fact that changing the IV changes the effective key used in the block primitive (i.e. the raw aes-encrypt function) as well as the first block of plaintext fed to AES by deriving these from nonce+key.
So I guess this is asking for a somewhat different construction. Personally aside from the fact it is already widely deployed I'm not sure I'd keep GHASH.
> The AEADs defined in this document calculate fresh AES keys for each nonce.
I guess write an RFC for extended nonce only?
So yes, using GCM with a 128-bit random nonce is already good enough for most of these cases.
However IMO all of this is a distraction anyway. One of the most devastating real-world attacks involving nonce reuse was the KRACK attacks, and that involved a protocol error allowing the attacker to force nonce reuse. No amount of extra large random nonces would have saved from that. (And using random nonces in such a protocol significantly bloats the wire format).
What we really need to do is move away from hugely fragile polynomial MACs. For 90%+ of usecases a more robust PRF is perfectly performant enough - eg note that the impact of KRACK was less severe against CCM than GCM. Heck, even CTR/CBC+HMAC is perfectly fast enough for many use-cases. Stop with the premature optimisation already.
Sure, no more than 128 bits, but indeed better than 96.
The only thing to potentially be aware of is that the randomized block counter may end up overflowing if it happens to end up with a large initial value (or you encrypt large messages). That should be fine, but it's quite likely that some GCM implementations are not expecting that and either blow up when the counter resets to 0 or do something else unexpected. So although I think this is theoretically a fine thing to do, I absolutely wouldn’t trust my sensitive data to it.
EDIT: my question is probably more clearly asked as: what's wrong with XChachaPoly that is solved by this new construct?
I feel like I just “actually it’s GNU/Linux”ed you there though... I feel bad, I’m sorry.
EDIT: I'm wrong, ChachaPoly has an rfc, but not the X variant
https://soatok.blog/2022/12/21/extending-the-aes-gcm-nonce-w...
I never considered only using 11 rounds, though. That'd have a significant performance impact if we could.
The design I sketched out extended the 96-bit GCM nonce to 224 bits, which is longer than the 192 bits of XSalsa and XChaCha. That's also the maximum that's supported by the algorithms as used.
If we supported arbitrarily longer inputs to AES-CBC-MAC, it's going to get mixed down into an AES block (128 bits long) anyway, so the benefit of arbitrary-length extensions over a 128-bit extension is unclear to me.
I mean, sure, if you really want to, you can already do that with the GCM part. I would hesitate to do that to the AES-CBC-MAC part.
Your proposal would then be to dedicate the first 16 bytes (128 bits) to the extension, and the rest to GCM.
(That’s gigabytes per second, not gigabits)
https://datatracker.ietf.org/doc/draft-irtf-cfrg-aegis-aead/
https://github.com/jedisct1/aegis-128X
Many implementations are available, including for Go: https://github.com/jedisct1/draft-aegis-aead#known-implement...
The conversation about nonce sizes is 100% on point however. Why isn’t it at least the size of the key?
Specifically, "nah with OpenSSL on AMD or Intel boxes from the last decade":
AES-128 (ctr mode) 9.7 GB/s
AES-256 (ctr mode) 8.6 GB/s (90 % as fast,
not ~70 % as 10-vs-14 rounds would naively suggest)
AES-128 (gcm mode) 12.4 GB/s
AES-256 (gcm mode) 10.8 GB/s
(plain CTR is slower than GCM because the GCM implementation in OpenSSL has received more attention than the counter mode implementation, simply because standalone counter mode is used a lot less)And this is just one core anyway. How many 100 GbE ports are you running per core?
I keep reiterating this because crypto=slow is still a very common idea from the days when AES-128 would peg a core running at 40 MB/s, rcp actually was 3x faster than scp, http:// instead of https:// would actually meaningfully reduce load, SMB/NFS encryption would actually slow things down and it was a bog to actually use a system running with LUKS-encrypted drives. But that's not the case any more. Encryption is really really fast and has been for about ten years, depending on what you're looking at exactly. Nowadays even many microcontrollers have AES acceleration because - internet of things mainly.
Asymmetric encryption remains “slow” but symmetric is within an order of magnitude of a memcpy.
It's measureable, but in most applications i wouldn't optimize this part.
This is satire right? The computational and storage requirements to preform such an attack to just get a small probability of decrypting one message seem ludicrous.
In cryptography, you don't want "ludicrously" infeasible, as in the NSA can just about afford the hardware and do it, you want astronomically infeasible.
Its not enough to just make it astronomically infeasible to attack everyone.