How not to sign a JSON object
latacora.micro.blog
latacora.micro.blog
I don't care what you throw in the bytes; I can figure that out easily enough. I want to know who generated those bytes and that they haven't changed.
For this reason also, I think the "reach for symmetric first" is bad advice. The primary benefit I can think of for HMAC is if you want to store client-side state for more stateless browser-targeted services. It makes it easy to throw an obscure cookie to your client and then validate it on the way back. It seriously complicates things if you want to allow other people to validate provenance of those bytes. With symmetric crypto, as soon as you can validate you can also generate. That's not the system we want.
If you do need third parties to validate, a much simpler API is to just ask the thing that holds the HMAC credential over TLS, and have it return true/false. That makes it harder to demonstrate non-repudiation, but I'm not convinced that's a property you generally care about. Even in OIDC (which mandates a JWT, but for reasons that defy understanding doesn't mandate cryptographic domain separation between (IdP, RP) pairs), leading OIDC providers have long recommended you just talk to them over TLS to get userinfo. (GSuite has recently muddled this in their docs, which irks me.)
(Disclaimer: I'm the author.)
But my alternative suggestion also addresses that: in a model where you don't trust the host with the key but you do need to help it make a determination if a particular thing is valid or not, you could just go ask the server you got it from instead. Basically: the key doesn't have to live on the machine doing the validation.
One critical point is that the two servers do not have to be the same. You might distribute your files via a CDN like CloudFront (or if you're in the 90s and a Linux distribution, a ragtag team of servers that don't generally implement https). The server responsible for delivery can lie all it wants; the server responding if something is valid is what actually matters.
https://theupdateframework.github.io/
Also see how we took it further at Datadog:
https://www.datadoghq.com/blog/engineering/secure-publicatio...
We don't do token introspection (basically what you're referring to) and instead use JWT/JWE(soon) to reduce round trips for the RP.
Could you expand on the OIDC issue with cryptographic domain separation? Not sure I fully understand what benefits you're looking for there.
(They seem to have flipped that in the last, I dunno, 3 months or so?)
And if you're connecting to the original party anyway might as well not even use HMAC or any crypto at all. When "signed" data needs to be sent to the client don't send it at all, instead store the data locally under a GUID and send that GUID to the client. When client takes their GUID to another party, that other party connects to the original party and retrieves perfectly authentic data. Free bonus - faster client performance.
Best crypto is absence of crypto.
Such as? Any of them different than that of the suggestion I was replying to?
Obviously you need to build and hash the pairs and not just key values. But the order is obviously important too.
Is there some standard for this I don’t know of?
OTOH: base64(jsonString(envelope)) '.' base64(jsonString(claims)) '.' base64(sig) works fine for JWT, because JSON has a very much limited choice of encoding (one option, zero alternatives!) and a JWT has no requirement to still be JSON on both sides of the transformation. It's just an opaque 7-bit-safe string to any party that doesn't know what the contents are.
I do think whoever is coming up with the scheme for assembling those bytes should care a lot. Specifically, whenever you're signing something, you have to pay attention to whether that something still has exploitable ambiguities. Most of those come from how you delineate the fields in the thing you're signing, which is a classic "you find out what you missed when someone sends you a CVE" type of problem. So you're right back to the problem of "if not JSON or XML then you're still rolling your own."
I agree with "reach for symmetric" in this sense: what are you trying to accomplish? Prove the data came from a specific source or that they are who they claim to be? Can you just use TLS with mutual auth? Prove the data came from yourself? Symmetric is probably okay.
When you need verifiable attestations of data from parties that don't talk to each other... welcome to the jungle.
The bitcoin malleability problem was the same issue, just with ASN.1 instead of JSON.
{
"previous": "%XphMUkWQtomKjXQvFGfsGYpt69sgEY7Y4Vou9cEuJho=.sha256",
"author": "@FCX/tsDLpubCPKKfIrw4gc+SQkHcaD17s7GI6i/ziWY=.ed25519",
"sequence": 2,
"timestamp": 1514517078157,
"hash": "sha256",
"content": {
"type": "post",
"text": "Second post!"
}
}
gets signed like {
"previous": "%XphMUkWQtomKjXQvFGfsGYpt69sgEY7Y4Vou9cEuJho=.sha256",
"author": "@FCX/tsDLpubCPKKfIrw4gc+SQkHcaD17s7GI6i/ziWY=.ed25519",
"sequence": 2,
"timestamp": 1514517078157,
"hash": "sha256",
"content": {
"type": "post",
"text": "Second post!"
},
"signature": "z7W1ERg9UYZjNfE72ZwEuJF79khG+eOHWFp6iF+KLuSrw8Lqa6
IousK4cCn9T5qFa8E14GVek4cAMmMbjqDnAg==.sig.ed25519"
}
which means that you have to be careful about how you remove the signature in order to verify the original. The "main" node implementation does this by parsing the json, removing the field, and then re serializing, forcing any alternate implementation to exactly match the node serialization in order to be compatible {
"original":"json",
"goes":"here"
}
Signed: {
"original_json_contents_base64":"ewogICJvcmlnaW5hbCI6Impzb24iLAogICJnb2VzIjoiaGVyZSIKfQo=",
"hmac_sha256_of_base64":"bf1f4cb95ce8633aff46888e1717873e32bb2a770b3d4b5b74a59e5e9adefeda"
}
This way you have full control over the raw bytes you want to sign (by forcing them into Base64 where other systems can't get their dirty paws on them).I guess the problem here is if intermediate systems want to do stuff based on the payload (but without validating it), they won't like this.
But if the problem is just intermediate systems barfing on non-json, this might work!
p.s. enjoyable blog post - as they always are! ;-)
You also correctly identified why that is different from the other schemes: they don't change the structure of the outer object.
{
"data": {...}
"signature": "z7W1ER..."
}Yep, JSON has no round-trip guarantees
Everyone is on the same page that at some point bytes go into a hash function. That's not the problem.
You mention JWT, but JWT only does external signing, which doesn't trigger the problematic case several people are describing to you. Perhaps an example would be more useful. If you start with a JSON like:
{"a": 1}
how do you build a JSON like: {"a": 1, "tag": "deadbeefdeadbeefdeadbeef"}
with a signing and verification algorithm that works?> Also, you don't validate a signature by re-creating the bytes in question... that's a flawed approach, and not the approach for example JWT takes.
Can you describe an HMAC validation process that doesn't involve recreating the bytes in the HMAC tag?
{"body":"base64-json","tag":"hash"}
It seems to me that trying to do what you're saying is a flawed approach.That said, I certainly hope that it is, generally speaking.
The stringify function returns a String in UTF-16 encoded JSON format representing an ECMAScript value, or undefined.
Of course, that makes the header effectively useless in practice.
You are definitely right that if for some reason you must do JWT, the way to do that is to strip as much of it away as possible. In particular, if you wanted to do safe HS256-only, you'd ignore the header, decode the body and tag, and validate the tag.
Also, it was literally only validating the tag and ignoring the header... it was using an asymmetric key for signing
1. You don't know what you don't know, and there is a lot to know about cryptography beyond the minimum needed to interoperate with other systems.
2. If you're engineering seriously, you know people are going to inherit your code and your design down the road, and if you're relying solely on a minimal feature set without a coherent, informed design, those people are building on sand.
Rolling your own JWT is a bad plan.
It's not terrible to "roll your own JWT" if you don't actually care about interoperability. And he's right, it does sidestep a lot of issues because JWT and corresponding libraries are designed to handle far more use-cases than what he may need it for and therefore if you don't fully understand it all, you may be shipping with unsecure configuration.
… with a well-typed library. Otherwise a lot of people end up comparing it in a timing-unsafe way. :(
> his post is mostly about authenticating consumers to an API. ... you’re trying to differentiate between a legitimate user and an attacker, usually by getting the legitimate user to prove that they know a credential that the attacker doesn’t.
The recipient of the API key doesn't need to verify their object. There's no attack from being able to give someone a fake API key - any attacker in a position to modify the API key in transit, which would just be a DoS, is also in a position to drop the connection, which is also a DoS. Such an attacker is probably also in a position to steal the API key silently, which is a bigger problem. (If a client is really curious whether they have a valid API key, they can just make an API call with it and see if it works, they still don't need to actually check the signature.)
Can anyone enlighten me what point is the author trying to make? JWT is pretty damn standard so it's my go-to for signing objects.
The short version is that there are flaws in the JWT specification that make certain bugs likely. A classic example of the "you have to parse a header to use the JWT" problem is the HS256 vs RS256 confusion bug, where your JWT library would interpret an allegedly-HS256 (HMAC) JWT using RS256 (RSA) key material. The JWT would get validated using the public key of the RSA pair, interpreted as an HMAC key. But the public key is, you know, public! So the impact of the bug is that everyone can forge JWTs. That is not a problem that can happen in well-designed schemes.
We do have a blog post from last year that tells you what we think you should do if you want to be safe and you know what kind of abstract thing (e.g. signing, MACing, etc) you need: https://latacora.micro.blog/2018/04/03/cryptographic-right-a...
(Disclaimer: I'm the author.)
Thankfully, JWT being quite standard and having momentum behind it I think there's a lower risk in recommending someone to pick a popular JWT library than telling them to roll his own simple scheme on top of HMAC (which is the article's recommendation and what JWT will end up doing regardless), specially when scale is considered.
I know, I know I'm arguing for "worse is better", but I honestly can't imagine a clean solution to this that would be so good it justifies dropping a seemingly decent standard, leaving mountains of legacy and requiring the entire developer world to learn about yet another crypto scheme. But then again I'm no expert in that area, and I would love to read a post about it from someone who is.
Currently using a terrible method with hashed query strings based on date, path, and a secret, which are then validated and have an expiration. Also, HTTP, so yeah it’s broken, but it at least prevents drive by scraping.
At this point we have no need for assymetric (don’t need to identify the requester, just need to prevent spoofing). What method of securing would you recommend?
2. Step two is a bit more complex. I assume your clients hold the secret, know what path they want to hit, and compute the key that way?
Tell me a bit more about the clients: what are they implemented in? What do they run on? Can you reliably ship updates?
(I'd prefer to have this discussion here but if you can't discuss publicly I'd be happy to take a look. My contact info is in my HN profile.)
How did you know! ;)
I’ll probably hit you up on the side, but there are multiple clients each with their own technical debt, and it’s an old solution, but I’m putting things in place now to make changes more possible, maybe per client app using separate hostnames, for example, so that we can transition to the new without breaking the old.
I think most of the client apps actually hit a manifest API first that gives out signed urls. This already happens over https in most cases. Some may generate their own, but I’m not sure all the usage scenarios.
I can’t give more details on the clients here, but let’s just assume they are diverse and difficult but not impossible to upgrade. It’s the kind of thing we want to get right the first time and has to work for a decade.
Probably the most useful thing for me to know is: what's the algorithm for signing a URL? If everything uses a manifest API, can you just make that a random token instead of a signature, and store that token in a database somewhere with an expiry?
Now somebody needs to manage the resulting PKI.
Even what looks like the no effort case, where you punt to the Web PKI and have all the clients use Web PKI certs (ie client1.example.com needs a cert with SAN dnsName client1.example.com like a web server would have but making sure the EKU for client auth is asserted) means now you have to either keep pace with us, or risk an impedance mismatch if our policies change in a way you're not OK with.
If you use one or more private CAs there's a tension between on the one hand the convenience of it not being your problem, and on the other hand the risk that it turns out you were just engaged in theatre and there's, for example, an unsecured SCEP server that will happily give any adversary with network access an authorised certificate with whatever identity they ask for...
TLS terminating LBs/WAFs/<things> that cant authenticate the client cert or pass the public key through to something that can, dealing with key/cert expiry, nobody to run the PKI infra with any interest managing identities of things that aren't AD computers, you name it.
Encrypt the files using AES-GCM. If you trust that your own client software is distributed securely and won't run in hostile environments and be reverse engineered, just ship that AES key with your software. Otherwise it will get complicated fast.
A little bit off-topic, but what is your opinion on NSS versus OpenSSL? Or LibreSSL and BoringSSL?
As a general rule: OpenSSL, BoringSSL if you can get away with it, don't use NSS or LibreSSL. But that presupposes that using a library like OpenSSL was the right answer to begin with :)
One of the things that makes "The JWT Problem" such a hard post to write (and not one we've already just written) is that the cryptographic flaws are but one problem; architecture implications are another. So, to say "yes go use PASETO" instead implies that something "like" PASETO or JWT is even the right answer, and in the vast majority of cases where we've seen JWTs applied at Latacora that has not been the case.
The issues folks have around JWT's header field always seem to be interface issues: libraries need to understand what keys they can use to validate, and what types those keys are, and validate them against the header.
(Though, I do not know why there is a "none" algo. That does seem like a folly. Could we recon that in the RFC?)
So, if my application supports v2 and you "forge" a v1 token, my application will not validate it.
(or whatever algorithm you wish) of course only "internal" clients would understand the jwt that isn't according to the spec.
"JWT, but you must roll your own implementation to avoid security risks" is strictly more dangerous than "Our custom signing system," because at least with the custom signing system, people are going to read your spec and not JWT's.
In other words if you stop utilizing JWT, you won’t have JWT specific problems.
So knowing JWT exists, this is still of interest to me.
The funniest line from this article and my word of the day.
1000 times this! If you think you need canonicalization, always remember: "no you do not!" It is not a hill you want to die on.
Unless you 100% need to use signing for the use case like a client side ACL, there is genuinely no need to overengineer your web app with your own authentication scheme.
But if your scale already mandates doing operations without server side ACL lookup for performance reasons, doing it over HTTP and web stack might already be inefficient by itself for the task.
One thing that comes to mind is that either you have to check revocation on the server side, or you have to re-provision client certs frequently, and in most client libraries that's actually difficult / annoying. If the server just sends me an updated token, I can just put that in a local variable and call it done.
There's a reason most APIs have moved to "you can just put the token in the request payload". It's hard enough to ask people to set HTTP headers, asking them to set client certs will be super complicated.
You can add a client cert with just 2 clicks on Windows. To my experience, that's way easier than most API authentication schemes
Two, I'm talking about client libraries, not web browsers. Every client library (including those used by web browsers) is perfectly capable of passing a query parameter. Most can pass cookies or custom headers. Not all of them can pass client certs.
If you're going to insist on symmetric key signatures, yeah. Otherwise a symmetric signature would be the same as using symmetric encryption to store user passwords, wouldn't it? You have to have the secret key to verify the user's signature.
The server _could_ just store the keys plaintext in a database but I'm assuming we can agree that's a horrible idea. The best it can do is encrypt them symmetrically before storing them, using an extra-special secret key that needs to be protected very carefully.
With username/password authentication, we don't store symmetrically encrypted passwords, do we? We store a one-way hash of the passwords instead, because symmetric is deemed not-good-enough.
If the server generates the token, and all it's effectively doing is verifying that the client has the correct token, then what is gained by making that token be a signature of any sort?
Add an expiry time to your signed payload. Tokens can have metadata.
> If the server generates the token, and all it's effectively doing is verifying that the client has the correct token
I think you might be understanding "token" as the API tokens that are just a random bunch of bytes that are later matched to authenticate like they were a password (let's call them "passtokens", I don't know if they have a name).
Tokens can be far broader than that. For example, JWTs contain arbitrary data. Verifying that the token is valid is just verifying its signature. You wouldn't check that the passtoken matches (in fact, you wouldn't have a passtoken at all). The payload is where the sauce is at.
A payload can include anything, from an user id (which is similar to the passtoken use case, authentication) to a list of grants (so you wouldn't even have to hit the "users" table to check for permissions... or even have access to it! As long as you can verify the signature.)
> to be revoked
That's true though. This is usually handled with short-lived tokens that must be renewed periodically.
Alternatively you could have a Token Revocation List of some sort (which isn't O(n) storage since 1. not all tokens will be revoked and 2. you can purge expired tokens). But then you get the problem of synchronizing the TRL across services (or centralize the token verification in a service which IMHO kinda defeats the purpose).
You still have a central server to track account status, but now it can be more like "a text file with a list of usernames on a single box running Apache, if it crashes we reboot it" and less like "a distributed, high performance, highly-available in-memory K/V store that's in the critical path for every request," which is going to make you a lot happier operationally.
Or you can push the list of usernames to revoke to each server, or something.
Correct. Usually if this was in a app and they use a symmetric key and they wanted to sign requests from a first-party app, in the real world this key is heavily obfuscated to ward off reverse-engineers from lifting the secret. The server will then decrypt this request with the same symmetric key to determine if it is indeed from the client and not a third-party.
As the author has outlined in the article, some of these API services use standard algorithms to do this such as HMAC while other services go to the extreme to use whitebox crypto + obfuscation. This is just security by obscurity, but it is for the purpose to slow down the attacker.
For example, when a app developer release a new app that uses a new API version, they can rotate the keys to slow the attacker down and can keep compatibility with the v1/v2/v3 versions with different keys and can choose to deprecate a endpoint without breaking the app.
That worked out just fine, but I can see the argument that it's much harder to get to a canonical JSON representation than it is to get to a canonical "tree of files" representation. Indeed, it was easy enough in my case that the repo contained a shell script one-liner that would compute it, and that was the reference against which the "real" python implementation was validated.
To be clear: I think that's a niche use case and while I think ObjectHash does a great job of exploring it, I don't expect the median startup to need an ObjectHash implementation.
(Disclaimer: I'm the author.)
Or is the argument that even though this is worse it's so useful we might really want to do it anyway?
If there's anything I said that made you think otherwise, let me know: I would like to amend that so no-one else thinks I could possibly mean that. The initially recommended (unless you can do otherwise) approach in the blog post is clearly "tag at the end" and every other approach also validates first. If you're referring to ObjectHash: like I said, it's a very niche application, I don't expect people to use it, and yeah, it enables new use cases.
(I expect you'd still really be authenticating the ObjectHash somehow -- e.g. by sending it over TLS -- but that's out of scope for ObjectHash itself.)
"I don't expect people to use it" just seems like the sort of awful excuse you'd usually be jumping on people for. It's like someone built a github project with a bunch of crypto red flags to check whether their new "Search github for projects with crypto red flags" idea works.
Don't get me wrong, it's clever, and I like clever. But I have learned in cryptography to only accept clever when it is clearly in the service of a specific pre-identified goal, and not just for its own sake. Isn't that normally a philosophy you'd subscribe to? What's the _pre-identified goal_ for this thing?
While we're here, another red flag. Mentioning Certificate Transparency as a model for some other X Transparency. Certificate Transparency isn't a model for anything. People have been saying to themselves almost from the dawn of CT "Oooh, this is clever, I should do the same for X" and it's always a bad idea. Someone might need a Merkle Tree. I'd argue they shouldn't use ObjectHash anyway. But the chance they need all the other paraphernalia from CT? Basically non-existent.
I think it was pretty clear, by calling it a specific niche that ObjectHash does a good job of exploring, that I am not making a recommendation.
Perhaps a more familiar case where this happens is SAML assertions with inline signatures?
Luckily the number of times I've had to invent signing schemes or even integrate SAML is limited. :)
There have been some poor implementations, but the method is pretty sound.
Create a json file
{"msg":"I am a God. My name is Bob.",
"sha256":"78A873E..."}
where the hash is the checksum of the file including the checksum. It would be equal to finding fixed point in cryptographic hash function that happens to be checksum to your message.> Canonicalization is a quagnet, which is a term of art in vulnerability research meaning quagmire and vulnerability magnet. You can tell it’s bad just by how hard it is to type ‘canonicalization’.
The flaw?
There's nothing here not already well known, so this isn't an insightful piece for those already well versed. Therefore, this is a piece written for those that aren't so practiced, for the sake of discussion: junior dev or devops, or non-security devs.
The content itself is a good discussion, and digestible (lol).
But, as a piece that junior folks are expected to get a takeaway from, the introduction is a disaster:
> This post is mostly about authenticating consumers to an API.
ie, not service-to-service auth.
> Unless you have a good reason why you need an (asymmetric) signature, you want a MAC.
A MAC/HMAC requires the signer and all verifiers to have the key. As stated just prior, this is about "frontend" signing. A novice reader might not realize they have to guard the key very well, and might even send it to the client browser. "Unless you have a good reason" is not a sufficiently strong warning for a post that is written as an instructable, more or less.
More architectural introduction (bonus with diagrams) is required. As is, this post is a footgun.
The only slightly JSON related content in there is a constructed scenario for 'in-band' signatures (the regex thing) which can just as well be achieved by a bit of string processing. Any JSON object will start with {, end with } and have some more information between it. Replacing the initial curly brace with '{"hmac": "foo",' gives you a valid JSON document. You can remove that easily before parsing and place no restrictions on the object's keys. You can handle edge cases like JSON literals or arrays by wrapping the whole thing in {"hmac": "foo", "payload": yourstring} if you feel like it.
The proposed solution is not in the numbered list. The numbered list describes how to sign a JSON blob from the outside. The rest of the doucment describes what you do if that's not an option, and you need to sign the blob in-line. The very next paragraph after said numbered list describes how to do that.
(I'm the author.)
Whee, I can stop rage-typing 'how not to self-attribute your comment about a thing you wrote' now!
This is a bit like an even more persnickety version 'nation state' in that the trivial fix is just dropping the pointless ceremony.