Show HN: Prefixed, dual-token, base58 encoded API Keys
github.com
github.com
Edit: Actually, the short token is more than just a token ID - it's also required, right? Perhaps replace:
When we receive an incoming request, we search our database for hash(long_token)
with: When we receive an incoming request, we search our database for hash(long_token) and short_token. A token can be blocklisted by its short_token.
So I think I prefer "short token" to "token id" but perhaps there is a better name for it. If it did get renamed it would probably make sense to rename "long token" as well. I'll defer to experts on this.edit based on your edit: Added your recommendation to the README, thank you!
> A token can be blocklisted by its short token.
With:
> An API Key can be blocklisted by its short token.
- You can promisify randomBytes once and reuse it rather than twice for every invocation
- There shouldn’t be a default value for the company name or people will end up using it.
- The company name isn’t validated so it could contain underscores which would cause issues with the short token parsing as it assumes it’s the second “chunk”
- The equals comparison of the hashes for the secrets is not timing safe. It’s not as bad as if they were plain text but it does short circuit due to how string equals works. Use the actual built in timing safe equals on the Buffer hash (not the stringified hex).
Thoughts from an ergonomic perspective:
- Base58 is nice and makes easy-to-handle values which is nice
- Naming of shortToken, longToken, and token are confusing to me, because it makes it sound like all of these values are tokens/secrets of some sort, rather than components of the token. Not to bikeshed, but what jumps to my mind is more like "prefix_id_secret", which helps make it clear what role each one plays.
- I don't think keyPrefix should have a default, I think your function should raise an error if it's not provided. Otherwise more than one user of your library is going to end up with "mycompany" keys in the wild.
- You probably ought to validate that keyPrefix does not contain an underscore and raise an exception if it does, otherwise that way lies pain for people trying to handle these keys.
Thoughts from a security perspective:
- You should probably use crypto.timingSafeEquals(buffer, buffer) [1] instead of comparing strings
- If you're trying to store the secret's hash, isn't something like bcrypt/scrypt preferred over raw sha256 these days?
[1] https://nodejs.org/api/crypto.html#cryptotimingsafeequala-b
Not for this type of use case. If a secret is long and generated from a cryptographically secure source (eg /dev/urandom), then any cryptographic hash function is fine as brute forcing the secret itself is not feasible. A single pass of sha256 is fine and presumably you’d be doing many of these operations each second in a high throughput application.
“Slow” hash functions like bcrypt or scrypt are for user provided secrets that might not have a large amount of entropy. It’s fine for something like user authentication but would be way too slow and pointless for API keys.
It seems this one, at 24 base-58 characters, is only roughly twice as long as it needs to be, right? How short is too short?
I think your thought on bcrypt/scrypt vs SHA256 is super interesting here. The long token is treated a lot like a password, so we should treat it similarly and use slow hashing. However, unlike a password an API key is repeatedly used for authentication instead of being exchanged for a session token. I don't think this meaningfully changes anything- so I think you're right that bcrypt/scrypt would be a better choice!
Edit: Also see koomla's answer!
I'd consider using a weaker bcrypt/scrypt/argon2 for API keys than would be used for login. Perhaps one that takes a hundredth or a thousandth of the time.
It could be unnecessary though.
Here's the main scenario: someone snagged a hashed long token from a database backup and wants to get the unhashed long token so they can use it to access something behind the API. They can do all the brute forcing they want and the server owner will never know about it. There are 1 with 42 zeroes worth of potential long tokens to try (58 * '51FwqftsmMDHHbJAMEXXHCgG'.length). Seems unlikely even though it's very very cheap to hash a potential long token. The tokens this is using are pretty short, but still not short enough to make cracking it feasible.
If you used argon2 maybe you could cut mycompany_BRTRKFsL_51FwqftsmMDHHbJAMEXXHCgG down to mycompany_BRTRKFsL_51FwqftsmMDH. Shorter token!
I would perhaps make the long token 58 digits long just because it would have more than a googol (10^100) possible values but still be shorter than 80 characters with the prefix and short token, and maybe swap the SHA hashing for a low cost (memory and CPU) argon2.
If you wanted to have really short API keys you could get creative with argon2. This is relevant for magic links.
https://startdebugging.net/2013/10/counting-up-to-one-trilli...
[1]: https://docs.github.com/en/developers/overview/secret-scanni...
edit: See also the discussion from last time: https://news.ycombinator.com/item?id=28296864
Other disadvantage I forgot to mention earlier: base58 is variable length, which is a foot gun that will bite you eventually.
GitHub has secret scanning features built in now, with a prefix of your product name or company name you could easily create a regex of sorts to find when someone has uploaded an API key for your application to their GitHub repo and revoke the token or email them.
As an implementation observation, I am also on a life-long campaign to rid the world of `split(...)[-1]` type manipulations, because they lack the context of a more rigorous "parsing" style. In this specific case, it allows attackers to smuggle almost arbitrary characters between the shortToken and longToken:
checkAPIKey("alpha_BRTRKFsL_and this one time at band camp\r\nContent-Length: 0\r\n_51FwqftsmMDHHbJAMEXXHCgG",
"d70d981d87b449c107327c2a2afbf00d4b58070d6ba571aac35d7ea3e7c79f37")
Also for your consideration, the code currently has duplication: export const extractLongTokenHash = (token: string) =>
hashLongToken(extractLongToken(token))
// ...snip...
longTokenHash: hashLongToken(extractLongToken(token)),
// ...snip...
) => hashLongToken(extractLongToken(token)) === expectedLongTokenHash
and my experience is that it's so easy to remember to update one and forget to update the othersMaking an issue :)
With prefixed-api-key, the hash of each token is stored in the database, so the timestamp can easily be added there.
Most of the time I don't see the utility in the client having the timestamp, outside of a scenario where you have a third party validate on their own (e. g. JWT w/ RSA keys). The best way you see if a token has expired is by trying to use it.
With UUIDs I prefer UUID4 to UUID1 most of the time.
I prefer to only include the relevant information and for API keys the client doesn't usually need the time the key was created.
(I work for MongoDB).
Sounds like some salt should be used here. Search the DB for short_token, then use the discovered salt to check the "password".
Both are valid approaches of course, I'm just interested to hear your thoughts on the relative tradeoffs.
// Store the key.longTokenHash and key.shortToken in your database and give
// api.token to your customer.
I'm pretty sure that should be "key.token" instead.What parts of tokens does the server store in the Database (everything but the raw long token I think?)
What is sent to the consumer, and what do they need to send with the request?
[1] https://datatracker.ietf.org/doc/html/draft-msporny-base58#s...
I assume the performance is acceptable when you're dealing with very small API tokens, but it would be totally unsuitable as a replacement for base64 in the context of, say, email attachments. Even using it for something like a PKI certificate would probably be asking for trouble.
It's not really obvious (IMO) from the wording of the specification, but step 2 of the decoding algorithm is actually a nested loop.
> Base58 is designed with a number of usability characteristics in mind that Base64 does not consider. First, similar looking letters are omitted such as 0 (zero), O (capital o), I (capital i) and l (lower case L). Doing so eliminates the possibility of a human being mistaking similar characters for the wrong character. Second, the non-alphanumeric characters + (plus), = (equals), and / (slash) are omitted to make it possible to use Base58 values in all modern file systems and URL schemes without the need for further system-specific encoding schemes. Third, by using only alphanumeric characters, easy double-click or double tap selection is possible in modern computer interfaces. Fourth, social messaging systems do not line break on alphanumeric strings making it easier to e-mail or message Base58 values when debugging systems.
I use something similar but a fixed length for the prefix and uppercase
Basically a copy of stripe tokens