User IDs probably shouldn't be passed around as ints (2018)
rachelbythebay.com
rachelbythebay.com
Of course we could require a developer to "know what they are doing", many work environments do. However, you won't see many posts about that. First of all because it doesn't scale, and secondly because it doesn't make for interesting reading.
I also suspect that many devs today are learning platforms first and software skills last. This was the reverse for many devs that came from building it the hard way and then using platform tooling to simplify. Newers devs are looking for the tooling to provide the core skills guardrails.
2. use typed schemas for APIs, i.e. GraphQL with a custom scalar type for UserID. Other typed schemas for APIs: OpenAPI, AsyncAPI, protobufs/gRPC, etc.
// add GraphQL parsing and serialization code
scalar UserID
3. Make the implicit - explicit. Don't use naked primitive types in your code (similarly as you will not use naked literal constants, i.e. use PI instead of 3.1415...). Use a PL with the string static typing (preferably Algebraic Type System). Define type for UserID which cannot be mixed with the integers, i.e. type UserID = UserID int
4. in dynamic PLs you can use tagged tuples or tagged maps/structs, i.e. {:user_id, 1234}
or
{"_type": "UserID", "id": 1234}
5. validate all external inputs (even those comming from the DB or Message Broker): validateUserID: int -> UserID
6. The examples above assume integer representation of the user id, here's if we switch to UUIDv4: type UUIDv4 = UUIDv4 string
type UserID = UserID UUIDv4
validateUserID: string -> UserIDs/string static typing/strict static typing/
as opposed to the weak static typing in languages like C/C++/etc.
struct user_id { uint64_t val; };
void ban_user(struct user_id id); #include <stdio.h>
#include <stdint.h>
typedef struct user_id {
uint64_t val;
} user_id_t;
void ban_user(user_id_t id) {
printf("user with id %lu was banned", id.val);
}
int main() {
user_id_t someUser = {123};
ban_user(someUser);
return 0;
}
output: > clang-7 -pthread -lm -o main main.c
> ./main
user with id 123 was banned
And here's the code with the bug: #include <stdio.h>
#include <stdint.h>
typedef struct user_id {
uint64_t val;
} user_id_t;
typedef struct group_id {
uint64_t val;
} group_id_t;
void ban_user(user_id_t id) {
printf("user with id %lu was banned", id.val);
}
int main() {
user_id_t someUser = {123};
ban_user(someUser);
group_id_t someGroup = {456};
ban_user(someGroup);
return 0;
}
Compiler output: > clang-7 -pthread -lm -o main main.c
main.c:22:12: error: passing 'group_id_t' (aka 'struct group_id')
to parameter of incompatible type 'user_id_t' (aka
'struct user_id')
ban_user(someGroup);
^~~~~~~~~
main.c:13:25: note: passing argument to parameter 'id' here
void ban_user(user_id_t id) {
^
1 error generated.
exit status 1The only real defense here is language-level enforcement. Allow declaration of a subtype that is not assignment-compatible with the parent even though it's identical.
I don't do any web-facing stuff but the only bits of code that know about things like IDs are the database stuff. All the logic works with classes that contain the ID and relevant data--you always pass the class, not the ID.
In binary, UUID can be encoded as Big Endian or Little Endian or even Mixed Endian, and while the string encoding is usually lowercase with a certain hyphenization rules, I've seen variations of that.
The (hexadecimal) text encoding is also quite inefficient compared to more modern standards like ULID or kSUID.
RFC 4122 section 4.1.2:
> The fields are presented with the most significant one first.
> while the string encoding is usually lowercase with a certain hyphenization rules, I've seen variations of that.
RFC 4122 section 3:
> The hexadecimal values "a" through "f" are output as lower case characters and are case insensitive on input.
The hyphenation is normative, per the same section
UUID = time-low "-" time-mid "-"
time-high-and-version "-"
clock-seq-and-reserved
clock-seq-low "-" node type UserID = {
userID: number;
};
let validateUserID = function (x: number): UserID? {
...
};
Note, that my original example was missing the fact that not every input value is a valid user ID, i.e. the function should return Maybe<UserID>, Option<UserID>, or Result<UserID,SomeErrorType>.In TypeScript one can use optional types xxx? instead.
Your example: type UserID = { id: number; } Would result in more allocations. But I think it's the only way at the moment.
```ts
const validateUserId = (x: number): x is UserId => {};
```for the `UserId` type declaration, I'm actually a little unsure of the best way to go about it. But I feel like using enums, typed string literals, or classes might be better approach. Would be interested if anyone has any thoughts on this
It's a simple hack that takes a type and adds an additional field to it. You can then use a User-Defined Type-Guard to ensure a value is valid: https://levelup.gitconnected.com/user-defined-type-guards-in...
type Brand<K, T> = K & { __brand: T }
Looks like a great solution, also the constant syntax is very readable: 10 as USDYou need a test here because types won't tell you that a user was banned. And that test would also have caught this error.
- informal methods:
- unstructured time to think about the problem (Hammock Driven Development)
- write things down = writing is thinking
- Rubber Duck Driven Development
- design reviews
- lightweight formal methods:
- Decision Tables
- FSMs
- FP / Algebraic Type System with rich scalar datatypes
- DbC - Design-by-Contract
- PBT - Property-based Testing
- TDD - example-based testsUse 64-bit ints. Prefer not exposing them to users, but it’s not a problem for most software.
If you ever get so big you need to shard, use Snowflake or a 64bit scheme that encodes the shard.
Don’t overcomplicate things. You’re already using either SQLite or PostgreSQL and it already gives you auto-incrementing integer keys by default, without the need to encode/decode UUIDs in whatever software you’re writing to interface with it.
I don't imagine there are many valid use cases where you want a list of users sorted by their database record ID though, and if you're suggesting an auto-incremented int then the creation timestamp will give you the same order anyway.
When near 1 million users (yagni), reset sequence and do the same with 10 million (or one billion).
Doesn't solve the upside of 128-bit random numbers (ala uuid): the ability to generate remotely and expect no collision.
I don’t see how remote generation is an upside. If you’re using UUID as a database key, you’re hitting the DB anyway, so it doesn’t save a trip to the DB.
There are plenty of uses for UUID like peer-to-peer apps where that makes sense, but the article is talking about database keys.
By the time I compared UUIDs to sorted UUIDs (in the special pattern that was supposed to work for SQL Server) and ints, however, I didn’t find any difference in performance.
Perhaps by the time I tested it had been fixed, or perhaps it was only an issue in certain circumstances, but never an issue with a few million rows for me.
In theory there are all sorts of negative consequences to using integers that you’ll run into once in a while and will require a google for a solution, however, everyday all of your queries and all of your inserts will be faster than using anything else.
Also, it will be widely supported by any 3rd party code you use to build your project which is probably more important than anything else.
The problem is that 'ban_account' changes data in a user record. Hence it should be a method of a 'User' class. And the 'Message' class should have a way to fetch an instance of 'User' for the sender. Here's the right way:
ban_senders_of_messages(messages) {
for (i = 0; i < messages.size; ++i)
messages[i].sender.ban();
}Avoiding manual indexing doesn't prevent you from sticking other stray integers in place of user IDs (like specifying message IDs instead of the user ID that sent the message), but most of them are not biased toward the start like indexes are.
Still, I think it's prudent to avoid these kinds of errors at all if you can. Perhaps a good reason to switch to UUIDs for all primary keys, even if the normal concerns about enumerability don't apply.
It can be overwhelming because the correct way seems obvious, listen to the experts and do as somebody think you should do - but it never is as easy as that. You need to balance your structure. Everything you do has a performance hit in some way or another and you need to consider the impact said practice will make in your environment.
Strongly typed ID’s are a thing advocated for in general [0], yes.
And they aren’t even inconvenient in some languages. Modern C# for example makes it very easy to use them [1].
[0]: https://andrewlock.net/using-strongly-typed-entity-ids-to-av...
Also curious because I don't actually know: If a format spec is GPL, does that encumber implementations of said spec?
Most ULID implementations are still actively maintained and people are still using it in production. Most implementations also have liberal license like MIT so no, you should not worry about the GPL spec (unless you plan to distribute the spec with your app).
I'm using ULIDs for the cases when the entity is publicly orderd by time: any timestamped event, e.g. a chat message, a log entry, or a sensor measurement/metric.
But need to be careful that the timestamp resolution is detailed enough.
Also, while ULID might be good for optimizing RDBMS indexes, it might create hotspots in NoSQL K/V stores (i.e. all entities will be created on the same node in the cluster).
How often is this really a bad thing? Are you worried about someone enumerating the entire space of possible ULIDs for every millisecond without ever rate-limiting them? Not many people are building anonymous, privacy-first websites and there's plenty of other ways to determine when a user first started using the site regardless.
It's not about guessing user IDs, but about deducting some useful information about them.
For example if an attacker may deduct wether an employee in a company a senior or a new one, and will know when exactly they joined the company.
I'd probably start with LinkedIn first :)
User IDs probably shouldn't be passed around as ints - https://news.ycombinator.com/item?id=16946557 - April 2018 (84 comments)
On the other; Legal Names change: for lots of reasons. It makes a LOT of sense to store a UID the way *NIX systems often do. An explicit indirect lookup table entry.
It may happen that the data type in the database is an int... because the database does math on it (or its source - seq.nextval type thing).
However, when you get it, you don't need to have it be an int that you can do math with. Even if you keep the same underlying implementation of it, pass it around in a way that doesn't let you add two IDs together or mutate it.
The easiest way to do that is to store it as a String.
One of the advantages of taking it to a String in Java is that it becomes immutable and so can cleanly be used as a key to a Map and prevents the easy math operations on it without taking it from a String to a numeric type (and back). If you see someone doing `Integer.parseInt(someId)` it becomes clear that they're doing something wrong with that.
A Display Name (what humans interface with) and the 'internal identity in the system' (be that a UUID, a bare number, an exact string, whatever) should BOTH be available.
As an example, a backup / archive of a project is created. Years later it is restored, but now several employees have new names. Maybe some automation accounts got renamed too. Should the restoration fail when no user ID is present? Or should it instead restore the internal / native IDs? What if a user by the same name does exist, but they got re-added or the users migrated to a different authentication method and now have different internal / native IDs?
So you store the lookup table. Maybe even use a third tiny integer column if there aren't that many users. Store the platform ID and also the name. Though also have an option to not store or flatten that information.