User IDs probably shouldn't be passed around as ints
rachelbythebay.com
rachelbythebay.com
You can avoid these problems by wrapping the integers in objects that you use solely for referencing entities of the relevant type. For example:
class UserRef {
id: int;
}
Then you define functions like this: function ban_account(user: UserRef) {
// ...
}
And a static type checker will pick up incorrect uses. For dynamically typed languages, you could instead use a unique field name such as user_id, to achieve the same thing (though getting a runtime error instead of a compile time error).Obviously you'll still be using ints or strings as an external representation, but as long as you do the conversion at the point where the identifier enters the program, the type system will take care of the rest for you.
I don't think it's silly, you just need to protect your parameters. There's a way to do this as part of the basic programming framework. See my other post: https://news.ycombinator.com/item?id=16947546
Yes, this reminds me of "stringly typed programming", i.e. where the language may offer strong types, but the program just uses `String` everywhere. String injection attacks are examples of this: SQL injection can only occur if it's possible to concatenate "SQL statement" with "user input"; if these are both represented as `String` then it's easy to run into such problems; if they're represented as different types, then the only way to combine them would be with a designated conversion function, which is exactly where we can put the neccessary escaping.
See also: http://blog.moertel.com/posts/2006-10-18-a-type-based-soluti...
If getting such dangerously awful code deployed to production is likely then sequential IDs are just one of your many problems!
Sequential IDs for key data can be good to avoid for a few good reasons but awful code isn't one of them. Testing, code reviews and not having bad programmers should be in place to fix that.
Much of our code review is not about "is this code correct" - that's obviously important, but rather "will this be easy to review changes to in the future" or "if I came back here in a year, would I get it wrong". I think that's just as valuable.
There's always a trade-off with complexity, and I'm not sure whether this one pays off, but designing _for_ programming errors is important for any team/company/product that is growing, and will introduce new developers who weren't around when the decisions were made.
Instead, you can write this as (js):
banSendersOfMessages(messages) {
messages
.forEach(({senders}) => senders.forEach(banAccount));
}
In the ruby community, there is the principle of least surprise. Your design you should not introduce astonishing things.int64 are numbers, so the author cannot be saying "don't use numbers".
Instead, it's specifically numbers generated by something like an auto-incrementing primary key in the database.
For me the article's negative was that it focused on one edge case reason to not use IDs, when there are others.
While everyone makes mistakes I'd hope that using an array's size instead of its value is nigh on impossible to end up committed, let alone deployed to production.
As everyone makes mistakes for me the main reason for not using numeric user ids is because you're much more likely to accidentally expose an API or query string param that can be used to look up other user ids. When that happens being able to enumerate ids makes for a massive data breach, whereas a GUID stops that.
(There may be some reason why you'd be forced to use numeric IDs in a data store for performance at scale reasons but I imagine that's relatively rare.)
https://cloud.google.com/spanner/docs/schema-design#choosing...
If your url looks like https://awesomeunicorn.com/userProfile/123, it's pretty obvious that I could start poking around and trying to get user 122 or 124. If that url is https://awesomeunicorn.com/userProfile/123456-dead-beef-abba..., I don't really have any idea what the ID of the next user might be.
Disclaimer: not a web dev
Unlisted, private, public snippets all get links such as https://gitlab.com/snippets/1712835
If you have an instance with few users, you can just create a new snippet and then auto-decrement.
e.g. I just found https://gitlab.com/snippets/17128 this way.
Not only will you never run out but you don't need a round trip to the keymaster, horizontal scaling, clustering and vertical sharding even are all way easier due to that. This removes a single point of failure entirely when the keymaster is no longer needed.
Using ints for keys for profiles/users and many other things were needed way back when processing/db/disk/memory and performance from that were a problem, no longer.
Do your part, join the UUID revolution. Also, if you were a piece of data, would you not want to be unique across all the databases, storage and services? You've heard of Roko's Basilisk right? Do not disappoint.
It seems unthinkable now but who knows...
Another method, "Version 4", is simply a 122-bit random number.
- client can generate the UUID, allowing eventual consistency and retrying using distributed DBs (other DB is on phone for example)
- merging of DBs is easy
- allows capability based access control (you can’t guess the UUID, you don’t have access)
- can use a bit or teo to encode production vs development IDs, so cross contamination of systems is less likely
https://begriffs.com/posts/2018-01-01-sql-keys-in-depth.html
HN discussion: https://news.ycombinator.com/item?id=16050047
So, not a great example, and not a very convincing argument to me to stop using integers.
You might be surprised how much of a pain in the arse it is to even realise these bugs exist, because during development and initial tests these tables have a nasty habit in many situations of containing IDs that are the same as the index. And when everything is an int, or similar, fixing them can be quite painful too. It just takes one bug in one function for a set of subtly broken workarounds and/or misunderstandings to spread throughout the code. Lots of places where functions take an "id" and then pass it into a function that takes an "index", or vice versa... just what was the intention here? :(
This shit is the worst kind of bug.
My usual solution:
struct ThingID {uint64_t id;};
typedef struct ThingID ThingID;
struct OtherThingID {uint64_t id;};
typedef struct OtherThingID OtherThingID;
And that's it. When at all syntactically inconvenient, it's a sign you're possibly doing the wrong thing.Then we have had completely different experiences in software development and are unlikely to agree on the importance of the content in this post.
Thinking about it, maybe I shouldn't have opened with "That's interesting", which here means exactly what it says, but is a phrase sometimes deployed with malicious intent.
(That is, thing_do(foo.other_thing_id) still compiles, but at least it looks wrong.)
1/ stuffing random ints into your user functions : that problem can be solved with typing and tests. I'm hoping that extensive testing of any piece of code which would have a drastic effect on a user would get extensive testing before going into production.
2/ ID canary : seems a rather good idea, like stack canaries commonly used when you don't have much stack space and you might get a collision with your heap. It's only a problem for languages for which point 1/ couldn't be a solution.
3/ Using UUIDs to avoid disclosing information about your user count : I think that's a separate problem. You should avoid disclosing unneeded information in general. If you use int IDs, always have an opaque public_id field that you use publicly and for interoperability with third-parties. But it does not mean you have to use UUIDs internally. However they do have a number of advantages, mainly that you don't need a central authority to distribute new sequence numbers, you can just generate UUIDs where you need them which will save you DB round-trips and make sharding of your DB easier. Also will help avoid issues such as this one : https://blog.travis-ci.com/2018-04-03-incident-post-mortem
A User has an ID of type UserId.
A Message has an ID of type MessageId and a sender of type UserId.
ban_senders_of_messages would have a parameter of type [Message]
ban_account would have a parameter of type UserId
When I write code, I don't use "i" or other single-letter variable names. I write long names, like "currentItem", "currentMessage", "currentUser", etc. If I reference an object, I usually name it "thisItem", "thisMessage", "thisUser".
Compilers shrink executables, so it's not a size issue. I don't know why people want code to be shorter; I prefer it to be easy to read and debug.
Real problem is that it is possible to accidentally do this type of mistakes. If you can avoid doing it by leveraging type system, you should. Relying on humans to never make mistake is futile.
public enum UserId { Unset = 0; }
public class User
{
public UserId Id {get;set;}
}
This works with various ORM and other mapping tools since enums are ints underneath, but you get static checking.Then there's the issue of protection when passing around ids in web apps as parameters, in cookies, etc. I devised Clavis [1] as an experiment for protecting URL parameters via an HMAC. The idea works pretty well in practice, but it's current incarnation is a little too cumbersome to use.
[1] A url http://foo.com?userId=1234 becomes http://foo.com?-userId=1234&clavis=asdbwef67t34rfbs, where the 'clavis' parameter is an HMAC of the URL's protected parameters, and changing any of them causes the request to fail. Unprotected parameters are also supported, so GET form submissions are still possible. See: http://higherlogics.blogspot.ca/2014/01/clavis-rebooted-secu...
To reduce risk of mixup of different kinds of IDs in the system I used different increment values in Postgresql sequences (e.g. 13 for users, 7 for categories), so the IDs quickly went out of sync and had little overlap.
Also a stupid number of APIs I deal with like to use one or more leading zeros in their identifiers. The meaning isn’t different, but it gets annoying when trying to do search, because of course the end user typed those in and wants to be able to look up whatever as 021.
This argument is, in essence, an argument for strongly typed languages. There are many argument against this historical argument, and it is by no means a settled issue – on the contrary, it is very slowly, as the years go by, looking more and more like the strongly typed languages are on the way out.
What? I'd argue the complete opposite. Above a certain level of complexity, lack of type checking becomes so onerous and bug-inducing that dynamic languages start introducing stronger typing. Typescript is paradise compared to Javascript, and even Python has added an optional type-checker.
E.g., https://en.wikipedia.org/wiki/Politically_exposed_person
That and some sanity in your account-handling ops.
Granted, one thing languages could do is provide easy type containers so it's hard to misconstrue an I'd as referring to a wrong type. I once tried to do this with generics in C# but it wasn't worth the effort.
ban_senders_of_messages(messages) {
for (i = 0; i < messages.size(); ++i) {
ban_account(message[i].sendersManager);
}
}
The only way to catch that is to test, and because the manager and sender will have similar data types, it will compile/execute just fine. The point is, this is probably a more likely error than the one mentioned in the article, and needs manual inspection and testing to correct. If you're carrying out that process anyway, the added inefficiency of a more elaborate data type just for user IDs seems redundant.In languages that support value types, you'll typically make them integer size, so the cost is likely to be the same as passing an integer. In languages that support reference types only, you're just passing around a pointer anyway.
If manager IDs and user IDs truly are the same type of thing, then there's a limited amount the type system can do for you in this respect. Maybe you'll have to stop at this point and just accept that you'll have to exercise a certain degree of care.
But there's a big gap between stopping there, in my view, and what you appear to be advocating: deciding that since manager IDs and user IDs are the same thing then you might as well give up entirely and just decide that they may as well be the same thing as ints while you're at it.
As for the initialisation problem, it's true that it can't be structs all the way down, and at some point you will have to create one of these objects, probably from a primitive with a non-meaningful type such as int, or string. But my experience is that IDs and the like tend to be created in a small number of places, and then reused, copied and passed around. Far easier to find and check all the places where one is created than all the places where one is used!
For example, if you simply want to have functions that works for both of them (so that you don't duplicate code), you can either create function from Manager to User (so that you can reuse functions for users), or use whatever polymorphism stuff your language support (polymorphic function, OOP, ...).
If you want to mix User/Manager in same collection (or have function that returns any of those), OOP can help too (Manager is "child" of User). If your language have sum types, you can use those (have additional type "User or Manager", and accompanying matching/extraction function).
In some languages you can do this with no runtime overhead (i.e. the additional type will be erased during runtime, as it's already type checked).
Solution: use a modern language like Go or use pointers to structs like everyone else.
Obfuscation of how data is queried and stored is pretty low hanging fruit, security-wise.
This code is so easy to write a unit test for. I hope it didnt even get committed, let alone deployed to prod.
My understanding of numerical types is that they exist to perform math. User IDs are not used for math, they're a completely arbitrary vanity system to assist with identification, so they should be strings, equally arbitrary.
Personally, I think E-mail addresses are the best user identifiers these days. Back in the day when there were like 5 websites everyone used, having your username was a cool thing. These days there's a billion websites and nobody uses the same ones and there's zero inter-user interaction on most sites. From the perspective of user friendliness, E-mail addresses are the easiest because you kill two birds with one stone (contact method + username + password recovery).
If you want a numerical ID, what about using a hash of the E-mail address? Or perhaps a combination of things, email, full name, sign-up date.
Oh god no. You don't want all your IDs changing when a user changes their email address.
You probably want your ID (e.g. a UUID) and your user-friendly lookup method (e.g. an email) to be separate.
That's a pretty passionate response, can you explain your logic? What are you doing with your usernames that you can't afford to let users change them?
Or am I misunderstanding your perspective?
I know that some services have a public-facing "username" and a behind the scenes unique identifier (which is a great UX model), I'm just focusing on the unique identifier. Which I would think should always be it's own column, whether it's also used for the public "username" or not.
> Because if you change that, all your relations between tables will break.
Okay, that is not a response to my question, which is why would you ever use the row ID for anything in your program. If you never use it, then it cannot ever be changed. Also SQL allows relationships based on more than one field, so it seems such a disaster could be easily avoided.