Smashing the state machine: the true potential of web race conditions
portswigger.net
portswigger.net
I hadn't thought of using these race conditions to trigger security vulnerabilities. That makes all sorts of sense, and I'm going to use this work to justify my persnicketyness in future projects.
It doesn't help that popular ORMs either don't expose transactions at all or they consider transaction handling to be an advanced feature. It looks like mongodb only (finally) added transactions in mongo 4.2.
The general rule for how to avoid this problems isn't locking. Its using a database transaction per HTTP request, or "operation" - however you define that. Transactions are perfect because if any part of the request fails you can rollback the entire transaction. And if you really need to, you can emulate something similar yourself by having a "last modified" nonce on your record (a random number will do). If you ever do a read-write cycle - say from a web frontend - then the read should include reading the nonce, and the write should fail on the server if the nonce has changed since whatever value you previously read. (And you can quietly fix this by simply retrying the entire operation).
From a security point of view, this is a fantastic find. I anticipate a massive amount of code out there is at best buggy, and at worst vulnerable to these sort of attacks.
I mean, you're right, and it's a bugbear of mine, too. But let's not pretend that SELECT FOR UPDATE and out-of-DB manipulation of said data fixes everything. (Nor that a direct UPDATE statement would.)
Ultimately, either your updates are linearisable or they are not. Changing someone's first name with or without this methodology does not alter the fact that "last one to change wins". Which change is the right one is in this case in the eye of the beholder and the domain you work in.
But, yes, updates that build on the existing values to derive new ones are common data races because people, indeed, do not understand databases, transactions or isolation levels. There are easy fixes for this type of problem.
But it's never quite as simple as making everything SERIALIZABLE and FOR UPDATE and expecting that to resolve version conflicts because two different things have different expectations of what a row value should be.
Can you give some examples where making everything serializable is a poor choice?
I think most of the vulnerabilities and bugs happen because people don't think about isolation / atomicity at all. And to me, SERIALIZABLE is probably the right default for 99% of regular websites. The example that comes to mind for me is realtime collaborative editing, though you can build that on top of serializable database transactions just fine. In what other situations do you need a different strategy?
Of course the callers may be surprised at being out-raced, so you still need to deal with that. Often it's enough to return the updated value.
What I mean by "transaction sharing" is for separate connections to be simultaneously using the same transaction.
Consider something whose logic is like this:
begin_transaction()
select_something()
update_something()
compute_something()
update_some_more()
commit_transaction()
where compute_something() includes calling a service running in another process or on another server to return some data, and that service uses the same database, and we want it to see the data that update_something() updated.With transaction sharing it might instead work something like this:
tid = begin_shared_transaction()
select_something()
update_something()
compute_something(tid)
update_some_more()
commit_transaction()
and in the service that compute_something(tid) calls it could take tid as a parameter and it would do something like open_shared_transaction(tid)
...
close_transaction() version = txn.get_read_version().wait()
compute_something(version)
Then on the remote computer, you can create a new transaction and call: def compute_something(version):
txn2 = db.create_transaction()
txn2.set_read_version(version)
# ... Then use the transaction as normal.
This will ensure both transactions are looking at the same snapshot of the data in the database.Docs: https://apple.github.io/foundationdb/api-python.html#version...
Maybe you could do it with a database proxy. The whole approach feels unclean, though.
But generally I agree. This pattern should generally be built into everything.
Whenever I’m setting up an application server I usually whip up some middleware which takes care of this for me. The code is database specific - the semantics can differ slightly depending on what database you’re using. And it only matters for requests which issue multiple commands to the database before returning. (But this happens all the time in many applications. Especially if you’re using an ORM.)
The article mentions this example:
Some data structures aggressively tackle concurrency issues by using locking to only allow a single worker to access them at a time. One example of this is PHP's native session handler - if you send PHP two requests in the same session at the same time, they get processed sequentially! This approach is secure against session-based race conditions but it's terrible for performance, and quite rare as a result.Modern databases are incredible feats of engineering.
Years ago, Stripe had a capture the flag related to distributed systems (I believe it was the second one they ran). One of the final levels involved attacking a system that would make calls to 3 other systems in turn that each held part of the key before calling back to your webhook. You could tell how many systems it had spoken to by watching for incremented port numbers between the requests.
I got killed by the jitter, but I seem to recall someone saying there was a way to stream your requests at a lower lower level on the same socket. Caveat, 10 years ago, details hazy.
Found the HN thread from the time though https://news.ycombinator.com/item?id=4453857
>> Every pentester knows that multi-step sequences are a hotbed for vulnerabilities, but with race conditions, everything is multi-step.
This is something that we spent a lot of time thinking about and why we decided to upgrade SocketCluster (an open source WebSocket RPC + pub/sub system https://socketcluster.io/) to support async iterables (with for-await-of loops) as first-class citizens to consume messages/requests instead of callback listeners.
Listener callbacks are inherently concurrent and, therefore, prone to vulnerabilities as adroitly described in this article. It's very difficult to enforce that certain actions are processed in-order using callbacks and the resulting code is typically anything but succinct...
Some users have still not upgraded to the latest SC version because it's just so different from what they're used to but articles like this help to confirm our own observations and reinforce that it may be a good decision in the long term.
For all of its benefits, though, one of the gotchas of a queue-based system to be aware of is the potential for backpressure to build up. In our case, we had to expose an API for backpressure monitoring/management.
I'm building an app on top of django where I have to worry about this, if you're using django check out there support for select-for-update, and if you're database supports it nowait=True can be a great thing that will fail a read if a select for update is already run:
https://docs.djangoproject.com/en/4.2/ref/models/querysets/#...
Also worth mentioning optimistic locking if you're looking to solve the issue in a different way, there is more involved from the application side but it has some advantages as well. I tend to prefer select for update with nowait=True since it's simpler on the application side, but I have used optimistic locking in the past with great success and some systems support it OOTB. Here is a description from AWS for those curious:
https://docs.aws.amazon.com/amazondynamodb/latest/developerg...
It seems a valuable way to spend the rapidly smaller training budgets
thank you
Why do you think that?
Finally, use the single-packet attack (or last-byte sync if HTTP/2 isn't supported) to issue all the requests at once. You can do this in Turbo Intruder using the single-packet-attack template, or in Repeater using the 'Send group in parallel' option.