* Intrinsically simpler than AES
* Easier to implement
* As an ARX design, doesn't need S-boxes, and so doesn't leave a cache footprint
* Has free key setup
AES is:
* A global standard
* Available in hardware on most platforms (extremely important)
* A conventional block cipher for which a bunch of modes (in particular: wide-block and AEAD) are already defined
But unlike Salsa, AES:
* Has relatively complicated key schedule (you have to expand its key input to a series of per-round keys, which imposes a cost when you switch keys)
* Relies on S-boxes for security and so must carefully avoid microarchitectural side channels
* Is much harder to implement
* Is not a native stream cipher, so requires an adapter (usually: GCM mode) to use safely.
AES is usually faster on modern systems because it's implemented directly in silicon. Salsa is usually the fastest pure-software option. Both are so fast that the speed difference is not particularly important, but most systems will prefer AES when hardware support is present.
Salsa is almost certainly the better choice for new designs just because of its simplicity. It's harder to screw up Salsa20 or its derivatives than it is to screw up AES (it is very easy to screw up AES), and its performance is more than satisfactory.
ChaCha is standard enough to make it into TLS and IPSec
So now there's three variants of ChaCha20
* ChaCha20 (256-bit key, 64-bit nonce, 64-bit counter)
* IETF ChaCha20 (256-bit key, 96-bit nonce, 32-bit counter)
* XChaCha20 (256-bit key, 192-bit nonce, 64-bit counter)It has? With test vectors and all? I want that, do you have a link?
Edit: Shit, considering it further, what I was remembering was the recent paper on BLAKE2X, not XChaCha20.
Without hardware support timing attack resistant AES is not so fast.
(and then there is the adventure of many motherboards shipping with hardware AES disabled in the bios...)
The safe way to use AES is by using a hardware implementation, like modern x86 and some ARM CPUs.
The best software implementations use bitslicing and SSE, but are still slow. The best I saw is an Emilia Kasper and Peter Schwabe paper[1] from 2009 on bitsliced AES-GCM has 21.99 cycles/byte performance for constant-time implementation authenticated AES-GCM.
For comparison, Intel shows[2] 0.77 cycles/byte for same with a hardware implementation, albeit on a newer CPU.
Chacha is fast on modern general purpose CPUs without the need for a hardware implementation of chacha. One reason it's fast is that it was designed so that a normal compiler can generate machine code from regular-looking C code in such a way that it uses vector (wide) registers and uses independent operations to use as many operations in the CPU in parallel in the same clock cycle, without requiring an assembly wizard to do that. Intel can afford assembly wizards (i.e. Shay Gueron), other people can't.
Modern TLS stacks prefer AES when running on a CPU that has AES hardware and fallback to chacha otherwise. They of course fallback to either a slow or an insecure implementation of AES if the other side doesn't support chacha.
1 - https://eprint.iacr.org/2009/129
2 - https://software.intel.com/en-us/articles/improving-openssl-...
https://sourceforge.net/p/pbkdf2/code/ci/default/tree/pbkdf2.c
If "ultra-efficient code" means what could be produced by a programmer highly skilled in some amd64 implementation (intel core2, amd bulldozer, ...) for that implementation then yes I doubt GNU C produces it. But the odds are GNU C's output runs faster than that's guru's code on other amd64 implementations.Being a stream cipher you can also precompute the keystream. This reduces encrypt/decrypt to a simple XOR when handling the message - depending on message length of course. And yes, AES-CTR can also be used like this.
all things being equal, in practice AES is often implemented in hardware which gives it an advantage.
This makes them very useful for, say, file encryption with random access.
It is most assuredly NOT a block cipher under the hood.
Indeed, one reason for using AES in counter mode is this random access, which among other things enables parallel encryption. The same strategy works with Chacha20.
I'm no expert of course so I don't even know if there's an AEAD that can bring you integrity on parts of the input; at least I know that minilock (https://github.com/kaepora/miniLock/blob/master/README.md#-m...) builds some kind of counter mode where each chunk is properly encrypted and has everything needed to check its integrity.
You can also just combine Salsa and HMAC.
It's true that you need to authenticate your data, but this is true for any cipher that you use.
It's a bad idea to implement your own cipher code, no matter what you're doing. If you're looking to include Salsa/ChaCha in an application, use Nacl, which refuses to give you unauthenticated ciphertext.