Google open sourced PSP (hardware cryptographic offload)
cloud.google.com
cloud.google.com
That should be around mid 2013 after Ed Snowden revelations. I remember some frustrated Google engineers on G+ commenting on the leaked slides from NSA where a packet capture of Google cross data center traffic was shown.
https://www.washingtonpost.com/world/national-security/nsa-i...
...so...no nic vendors mentioned?...what are we supposed to do with PSP but wait for a private company to build a PSP nic?
> They have published their architecture specification, a reference software implementation, and a suite of test cases.
Presumably at least one network card manufacturer has already made hardware (or firmware) for google and would be close to making a product for general sale. Now that the spec is out, others can follow.
This figure is lower than I thought. Granted, I deal a lot with memcache style workloads, where each RPC basically just does some hash table lookup/insert and a memcpy of hundreds of kilobytes. I personally see up to 5% of CPU time spent on cryptographic operations (mainly AES-NI instructions). If you aren't Google and can't afford custom silicon to do hardware offload, but you can't drop the encryption-in-transit requirement either, what would you do to improve performance?
> each NIC has two 256-bit AES keys, called master keys, not shared with any hosts including its own, or with any other NICs. The master keys are "critical security parameters",which are kept ephemerally in on-NIC RAM, and must not be stored on any persistent medium.
I take that to mean you do the asymmetric key stuff outside of PSP, then the symmetric key stuff is offloaded to PSP. Assuming you send a lot of data per connection, the symmetric key part will be much larger, so the expensive part is offloaded.
The number of organizations that this is relevant for is very, very small.
For most organizations, buying slightly more hardware is likely going to make far more sense than the amount of engineering effort required to implement something like this, if only because it's a much simpler, safer bet.
It's a little unfortunate that they chose a name that was already proposed as a L2 encryption method, Packet Security Protocol (one of the first hits I got when I tried searching for any analysis or prior discussion):
https://scholarworks.rit.edu/cgi/viewcontent.cgi?article=938...
> it took ~0.7% of Google's processing power to encrypt and decrypt RPCs, along with a corresponding amount of memory.
> As of 2022, PSP cryptographic offload saves 0.5% of Google's processing power.
eh?
TLS is a record based protocol and you can use it over bits of paper transported by pigeon if you like, can't you? I've looked into it extensively (though not for a few years I'll admit), and I've seen TLS implementations that don't include any transport mechanism - that bit it left to the application. I know that it must receive things in-order, and perhaps that does rule out UDP... but there's DTLS :)
That doesn't mean this isn't interesting, and I have no idea how 'offloadable' TLS or DTLS might be. This scheme seems to use derivable keys, so that packet sequence isn't important, but it looks also like we have a mac-then-encrypt scheme. I presume it's been analysed properly...
Definitely worth digging into.
> I know that it must receive things in-order
TCP? No, it can handle dropped, duplicated and even reordered packets, but doing so is problematic for many applications (either due to real-time requirements or adverse network conditions that a given TCP implementation doesn‘t handle well).
> TCP does not support non-TCP protocols efficiently
I'm going to assume you meant TLS for the first of those :)
> TCP? No, it can handle dropped, duplicated and even reordered packets
And TLS is what I meant here, it generally doesn't like things being out of sequence, and TCP usually at least hides that from the application layer where TLS lives. Which I guess is why it is mostly dependent on TCP, but can be implemented (albeit less efficiently) over other reliable, in-order comms mechanisms.
That's presumably what Google is referring to with their statement about TLS not working well for non-TCP (more precisely: non-stream-semantics) traffic: Encapsulation of datagrams in streams is possible, but very inefficient. That's why VPNs using TCP generally suck for performance sensitive operations. (They usually also do TCP-over-TCP, which is even worse than stream-over-TCP.)
TLS does require stream semantics to work efficiently; its security guarantees don't break in the presence of reordering, duplication or datagram loss, but all of these would be interpreted as active attacks or severe uncorrected lower-layer failures and probably either cause abysmal performance or even application-level errors due to unexpected/unhandled alerts.
And why go through the effort if there is already DTLS, which is to UDP (and other datagram layers) what TLS is to TCP?
wrt IV reuse the protocol doc says the NICs use a picosecond timestamp counter -- do NICs really have picosecond resolution clocks, or is it nanoseconds + monotonically increasing counter within the nanosecond?
edit: an even better RFC for this question is 4106 which is about aes-gcm in ESP, it calls the tag an ICV also.
The document acknowledges this but basically leaves it to other aspects of the network stack to defend against this (maybe there is some extra protection provided by the ICV check). Google's stack seems carefully designed to be secure in this way but it feels brittle.
Wouldn't it have been better to require a checksum of some of the exterior headers (source IP??) inside the encrypted section to block attempts to repackage the same encrypted content inside another packet. Or is that somewhere in there and I'm missing it?
Sharing a config file doesn't make it the same connection — in fact, the whole security model is that it's an easily grouped, but cryptographically ensured separate connection — no?
So it looks like the inner data can be either TCP or UDP. So you have the option of making it reliable or not.