Trickles - Stateless High Performance Networking
cs.cornell.edu
cs.cornell.edu
So, I'll try to answer some of the questions here and provide some of the insights and background that do not appear in academic papers.
The intuition behind trickles is that the packets act like continuations, the same continuations you may be familiar with from programming languages like Scheme. Instead of a server holding state, it pushes that state to the client. To get service, the client presents the continuation to any server host, which can reinstate its state, and provide the service.
Since the servers are stateless, you can direct a client request to any host. This kind of handoff is, let's just say, not easy to do when you have stateful TCP connections to maintain.
Not every state machine can be converted into a format where it can run over Trickles. But I think the community was surprised to see that something as complicated as TCP could be trickle-ized.
"... shortcut information for finding a particular object, such as a file system inode."
How do you guarantee that the client isn't handing you something malicious?
* They're encrypting and MAC'ing the client-held state, the same way IIS/ASP.NET does with ViewState.
It's not "more" secure than just holding the state serverside, but if it's implemented correctly, it's asymptotically close. (FWIW: the code is pretty messy [the relevant stuff is grafted on to the Linux ipv4 stack], I'm finding it challenging to reason about, and it's using memcmp to check the MAC, and the HMAC key is a charstar --- but this code isn't really the point).
Their replay attack prevention mechanism (http://www.cs.cornell.edu/~ashieh/trickles/security.php) seems to be a little bit weaker at first glance. They're preventing replays by enforcing freshness and keeping short-term state about packets that have been seen recently. That seems strong enough, but also seems like exploiting the limited size of the short-term state store may be DoS attack vector, or a vector that allows you to do replay attacks during high-traffic periods. I haven't read their proposal in enough detail yet, though.
Regarding this particular implementation: it's not documented well enough to tell, is it? It does use an HMAC. It does use AES. In the few minutes I gave myself to go through the code, I can't even figure out what mode they're running AES in, though, or what order the operations are applied in. How are keys generated? What keys are shared between client and server, and what keys are held serverside?
The latter questions are irrelevant to the academic project of proving the Trickles concept, so I'm not criticizing the team. I'm just saying, there's a lot of detail you'd want to have before saying they've got things covered.
The prototype did use AES to compute and check nonces -- AES is faster than HMAC for such small input sizes.
Since the server is effectively sending encrypted data and MACs to itself, there is no need for key distribution. Only the server holds keys; the client treats the encrypted data and MACs as opaque data.
Regardless, in truly stateless protocol, a replay attack shouldn't be a threat - who cares that someone can regenerate the same response from the same request? If it's really stateless, that replay can't _do_ anything. It turns into a threat when you put something stateful on top of this stateless machinery, which simply means that applications using this would need their own replay prevention. Isn't the right place to put the responsibility for preventing potentially dangerous erroneous state-transitions where you administer state in the first place - and not here?
The protocol is stateless, not necessarily the application and/or session. The state is serialized to/from the client as needed.
> If it's really stateless, that replay can't _do_ anything.
A stateless pipeline may front a stateful backend. An application commit point may be reached which should invalidate all possible previously valid requests. It's definitely a necessity.
The replay protection is intended to prevent clients from tricking the server into sending more than permitted by TCP congestion control. I.e., it protects the network and other clients from DoS.
Protecting TCP congestion control prevents a couple of attacks: * a selfish client may consume more than its fair share of the bandwidth * a malicious client may use the server to amplify a DoS attack. E.g., in the absence of replay protection, an attacker limited to only 1 Mb/s of bandwidth could cause the server to consume 10-20 Mb/s (other examples of amplification include smurf attacks; some TCP implementations from the 1990s could also be tricked). With replay protection, the attacker would actually have to have the full 10-20Mb/s of bandwidth to cause that level of damage.
Note that the HMAC protects against related attacks that are based on spoofing the client IP.
The amount of state stored at the server is proportional to the bandwidth. E.g., a server with a fatter pipe will need larger bloom filters to handle the higher packet rate. In Trickles, the amount of bandwidth-proportional state is mathematically clean to compute -- typical Bloom filter collision equations. By comparison, TCP will hold some fixed overhead per connection, consisting of TCB (TCP control block, for congestion control state) and socket/fd structs. Every TCP connection also buffers a variable amount of sent but unacknowledged data (proportional to window size).
The fixed per-connection overhead of TCP alone requires asymptotically more server-side state than Trickles. Since the size of TCP send buffers varies according to window size, which is determined by protocol dynamics, it’s tricky to construct a model of how much server state will be consumed by socket buffers. In our experiments, the socket buffers dominated server-side memory consumption.
- Alan Shieh