All you have accomplished really is to move the descriptive name from the variable's name to the variable's type, and at what cost? You now have to create a whole new datatype. The programmer must look up that datatype and see oh, it's just an int. And now you need conversion routines or worse yet casting to convert that type to a simple int for interop reasons (database, UI).
Experience has shown that it is very productive to have a small number of generally useful datatypes which are augmented by custom datatypes.
Creating a new datatype for something as simple as an int or String defeats that and makes it more cumbersome to work with. (It's especially unwarranted for a variable that is an immutable variable serving as a simple id.)
If you are using a strongly typed language, you can just right click on the accountId type and find everywhere in your codebase where the “accountId” is being used.
The “accountId” is not “just an int”. An accountId has semantics that would be different than an int. An accountId is not the same as a “customerId” that may also be represented as an int. The accountId also has different uses than an int. You’re not going to take the sum, average, etc of an accountId.
An accountId that happens to have a value of 1 is semantically not equivalent of a customerId of 1. A method that expects a list of accountIds that are just ints will just as happily take a list of customerIds.
But a method that takes as a parameter List<AccountId> will cause a compile time error if you pass in a List<CustomerId>.
Why use a strongly typed language and then ruin one of the benefits of it if you don’t use domain specific types?
If you can explicitly refute those, then your position would be stronger.
Why should the developer care what the underlying type is?
It should be an opaque type. When you serialize it, your serializer should call your ToString() method. When you deserialize it, your deserializer would call the constructor that takes an int.
But why are you saving an id to a database - which would never be used as int - as an int?
But even in most CRUD apps, your domain model with rich types would be different than your view model which would probably also be different than your DB model.
You are going to be mapping back and forth regardless - hopefully using a tool like Automapper.
And as for the "type shyness":
> Creating a new datatype for something as simple as an int or String defeats that and makes it more cumbersome to work with.
It is never "just an int". Soon you will ask questions about the int, and in a sufficiently large application you will forget the answers to the quesions easily, and scatter the same question (i.e., the same boolean logic) all over the place. Consider an application where you need to filter something by year. A year, you know? As in 1974, 2018, etc. An int is bad, because an int could be -5, which does not make sense. Ok, a uint then? Still bad, because then there is a year 0 and - if your application deals with birthdays of real people, 1800 would be an invalid birthday. The semantics (or pragmatics, linguistically speaking) emerge from how you interpret the integer, and this interpretation can happily live in that one file with that one class that encapsulates "just" that integer.
The AccountId probably shouldn't be < 0, and a specific type could be the one place where you make sure that this is always the case.
What if you start with a 32bit integer and suddenly you realize that your weird Internet-Of-Things-Sensor reading database has grown and the 32bit integer is getting too small? You could just use 64bits. With "just an int", you now get to change every single function that expects an int for the purpose of a SensorReadingId (or an account id or whatever) and change it to size_t, uint, int64 or whatever. OR, you just tell the object that its internal representation uses 64bits now, because the AccountId class is the one single place that deals with account ids. The type provides a consistent interface that allows you to infer the semantics of its use. An int doesn't do that.
This idea predates java by some ~20 years [1]. An int is not an account id. An int is not a year. An email is not a string. The result of working like that is considered a code smell [2].
Still, coming back to the example I made: It was primarily about commenting. I was taught to always comment, and StyleCop enforced comments on each and every field of a class, leading to noisy, superfluous comments.
/// Gets or Sets the Name
string Name { get; set; }
does not add anything.As for returning an AccountID: I'd like to point out that the Account is a strong and independent object, who can make his own decisions and does not need to return anything. Returning its id like that is actually breaking encapsulation. Still if you do it, it probably shouldn't be an integer.
Why are people scared of creating types and objects?
[1] https://link.springer.com/chapter/10.1007%2F978-1-4612-6315-...
[2] https://sourcemaking.com/refactoring/smells/primitive-obsess...
Most languages don't have the plethora of integer types that C has. Java, for example, has just has 2 in general use: int (signed 32 bit) and long (signed 64 bit). (Nobody considers byte a general integer type.)
In C/C++, you have to define so many things about the types you use because so much is "implementation dependent." So for C/C++, you may be correct. Most other languages define their types more stringently and include a smaller number of general purpose types.
https://docs.microsoft.com/en-us/dotnet/csharp/language-refe...
But even then, a year is an int, but it would have certain semantics you would want to enforce outside of integers.
You deal with a lot of data that is of only a few simple types, and as I said in my post, interoperability means you can't get too crazy with your class definitions. They need to stay simple or the SYSTEM will become more complicated.
As for a variable never staying a simple type, again, see my example in my previous post. A string id is going to stay a string id and its data and behavior is dictated by being an id to remain supersimple.
I am not against defining new types; I am against defining a new type for every single variable "just because."
Are you honestly suggesting making a Year class to represent years in a C++-like language? And you want to forbid Years from being -5 or 1800, because they "do not make sense"?
What if I want to interpret the integer in a way you didn't think of when you wrote your Year class? Maybe I want to know whether my friend John was older when his sister Sally was born than Jesus was when St. Paul was born?
> What if you start with a 32bit integer and suddenly you realize that your weird Internet-Of-Things-Sensor reading database has grown and the 32bit integer is getting too small? You could just use 64bits. With "just an int", you now get to change every single function that expects an int for the purpose of a SensorReadingId (or an account id or whatever) and change it to size_t, uint, int64 or whatever.
The work needed to push the buttons on your keyboard to make the changes is the least of your worries when you want to change the underlying data types used in a complex program. If you have a special SensorReadingId type and you change what it is, all your code breaks at once and you get a hell of a mess to untangle.
It may not be practical to enforce "ints that represent sensor reading ids are SensorReadingId type" across all of your code. You'll need to unpack your type to a primitive type whenever you talk to an external library or serialise SensorReadingId's to storage or the network. Even if you did enforce SensorReadingId everywhere:
- Arrays of SensorReadingId now double in size. Perhaps your code crashes at startup because some stack frame grew too large.
- If you serialise SensorReadingId's anywhere, changing the type broke binary compatibility with your old code. Or the network protocol you're supposed to be speaking. Or the library you're using.
- You probably had to make a toInt() or unpack() or similar function anyway so that people who want to get at information inside the value of the SensorReadingId can do so. You still have to look at all the call sites of these functions.
In reality, you're going to do this sort of change incrementally. You push the 64-bit integers through more and more code, taking care that the high bits of the integers make it through everywhere you've made the change.
> As for returning an AccountID: I'd like to point out that the Account is a strong and independent object, who can make his own decisions and does not need to return anything. Returning its id like that is actually breaking encapsulation. Still if you do it, it probably shouldn't be an integer.
Why is an Account making its own decisions? This sounds like the "circle that knows how to draw itself" antipattern. My friend and I want to talk about a certain Account; how do I arrange things so that the Account makes the decision to put an identifier for itself into a certain field in a certain protobuf message before I send the message to my friend?
That’s why languages don’t represent dates as just as struct with ints - there is a Date class that knows how to semantically add days to dates taking leap years into account, time zones, etc. That supports our point....
Then will your class take into account that the calendar changed in the mid 1600s(?). If you need to calculate dates and compare dates.
If you have a special SensorReadingId type and you change what it is, all your code breaks at once and you get a hell of a mess to untangle.
How so? All of your other code should be treating it like an opaque type. If I was comparing two sensor ids for equality outside of the class and I changed the underlying type, I should modify the equality comparison operator to work with the new semantics.
It may not be practical to enforce "ints that represent sensor reading ids are SensorReadingId type" across all of your code. You'll need to unpack your type to a primitive type whenever you talk to an external library or serialise SensorReadingId's to storage or the network. Even if you did enforce SensorReadingId everywhere:
You shouldn’t have to do it in but one place. The serializer would call your .ToString() method anyway. You just override it.
If you serialise SensorReadingId's anywhere, changing the type broke binary compatibility with your old code. Or the network protocol you're supposed to be speaking. Or the library you're using.
Versioning and deprecating APIs and interfaces is a solved problem. But you shouldn’t be using the same model for internal use as your external model anyway.
Arrays of SensorReadingId now double in size. Perhaps your code crashes at startup because some stack frame grew too large.
I thought we were specifically not talking about low level languages and you mentioned Java. When would that be an issue in Java?
Why is an Account making its own decisions? This sounds like the "circle that knows how to draw itself" antipattern. My friend and I want to talk about a certain Account; how do I arrange things so that the Account makes the decision to put an identifier for itself into a certain field in a certain protobuf message before I send the message to my friend?
I wouldn’t be sending my Account class that probably has business logic in it across the wire. I would be mapping it to a dumb AccountModel POCO/POJO with something like Automapper.
In C, `struct tm` is literally a struct whose fields are all `int`s. But people typically represent time in C programs as an integer of some width measuring time in some fixed unit since some fixed point in time.
> - there is a Date class that knows how to semantically add days to dates taking leap years into account, time zones, etc. That supports our point.... > > Then will your class take into account that the calendar changed in the mid 1600s(?). If you need to calculate dates and compare dates.
I can't anticipate those requirements---maybe our date class has to bomb out if you pass in, say, 1800 or -5 as a year because someone on my team insists they're invalid.
People work around these classes when they have to do date arithmetic because they're too slow, or they depend on the machine's locale and the locale isn't under control, or they take timezones into account in fundamentally broken ways, or any number of other reasons.
>> If you have a special SensorReadingId type and you change what it is, all your code breaks at once and you get a hell of a mess to untangle. > > How so? All of your other code should be treating it like an opaque type.
That's a fantasy. When you wrote your reply, you deleted all the parts of my post explaining "how so"---why not all of your other code is going to treat it like an opaque type.
> You shouldn’t have to do it in but one place. The serializer would call your .ToString() method anyway. You just override it.
I don't want to convert your boxed int to a string and then parse it as an integer just to do math on it. That creates a ton of syntactic noise, not to mention being slow.
> Versioning and deprecating APIs and interfaces is a solved problem.
Yes, and the solution is a pain in the ass! So don't create new interfaces to version and deprecate for no good reason.
> I thought we were specifically not talking about low level languages and you mentioned Java. When would that be an issue in Java?
I didn't mention Java.
> I wouldn’t be sending my Account class that probably has business logic in it across the wire. I would be mapping it to a dumb AccountModel POCO/POJO with something like Automapper.
And I'd just send an integral id over the wire, probably documenting that it's the id of an account using something like the field's name in the message. No extra tooling needed, and I don't have to make a fuss about how the id() function (that I had to add to the Account object anyway) is for the big boys and should be avoided whenever possible.
People rapidly get themselves into deep shit when they find out their overcooked codebase is too slow to be useful. The class names remain but their structure and function and any semblance of the clean abstraction they used to provide are the first things to go.
Even C has a bunch of functions to work with the date struct as an opaque type....
I can't anticipate those requirements---maybe our date class has to bomb out if you pass in, say, 1800 or -5 as a year because someone on my team insists they're invalid. People work around these classes when they have to do date arithmetic because they're too slow, or they depend on the machine's locale and the locale isn't under control, or they take timezones into account in fundamentally broken ways, or any number of other reasons.
And then you use a better class. In Java you use something like JodaTime in .Net you use NodaTime - a .Net library written and maintained by Jon Skeet (the person with the highest ranking on Stack OverFlow). If that’s not good enough you write a custom extension or mixin to handle your special snowflake handling and still let all of the consumers treat it as an opaque type.
I don't want to convert your boxed int to a string and then parse it as an integer just to do math on it. That creates a ton of syntactic noise, not to mention being slow.
We were talking about an accountId that just happens to be represented as an int. You’re not going to be doing math on an account id. Also neither .Net with generics or even C++ templates have the boxing and unboxing issue.
Yes, and the solution is a pain in the ass! So don't create new interfaces to version and deprecate for no good reason.
If there was a business case to change the underlying structure - by definition there was a “good reason”.
And I'd just send an integral id over the wire, probably documenting that it's the id of an account using something like the field's name in the message. No extra tooling needed, and I don't have to make a fuss about how the id() function (that I had to add to the Account object anyway) is for the big boys and should be avoided whenever possible.
Even so, you should still not send your internal representation that may change across the wire and do some type of mapping. Are you going to send the same model that came from your ORM that’s tied directly to your database across as the representation to the outside world?
Yes, for conversion to/from time_t, formatting, and parsing. That's my point; you use the integral type whenever you don't care whether it was a Thursday because the integral type isn't going to waste your time.
> And then you use a better class. In Java you use something like JodaTime in .Net you use NodaTime - a .Net library written and maintained by Jon Skeet (the person with the highest ranking on Stack OverFlow). If that’s not good enough you write a custom extension or mixin to handle your special snowflake handling and still let all of the consumers treat it as an opaque type.
Stop with the hero worship. You should do things that make sense in context. Recommending that everyone use a library in all contexts because it's written by Jon Skeet is a ridiculous thing to do. (BTW, it seems like the Joda-Time people recommend you use the standard java.time package instead.)
> We were talking about an accountId that just happens to be represented as an int. You’re not going to be doing math on an account id. Also neither .Net with generics or even C++ templates have the boxing and unboxing issue.
Who are you to say I won't do math on account ids?
> If there was a business case to change the underlying structure - by definition there was a “good reason”.
In the situations the two of us have discussed, there is no business case to create these interfaces in the first place. Where is the good reason?
> Even so, you should still not send your internal representation that may change across the wire and do some type of mapping. Are you going to send the same model that came from your ORM that’s tied directly to your database across as the representation to the outside world?
What ORM? What database? If I'm using an ORM and a database, high-performance time operations are likely out of the question.
What else are you going to do in C? But even most real world implementations of C have all sorts of structs and typedefs that you shouldn’t make any assumptions about and you should treat as an opaque type.
Who are you to say I won't do math on account ids?
Is that really what you are going to argue? That in a real world use case you’re going to be doing math on an accountId?
What ORM? What database? If I'm using an ORM and a database, high-performance time operations are likely out of the question.
So you’re going to pass a raw key/value record set around your code and program like it was 2006? So now you’re going back to mapping your sql resultset to an object that makes sense anyway. Going right back to my point that you’re going to have a mapping between your raw results -> domain model -> view model (where the view model is serialized for external use) and back again anyway.
Subtract two times? (BTW, struct tm isn't an opaque type.)
> > Who are you to say I won't do math on account ids?
> Is that really what you are going to argue? That in a real world use case you’re going to be doing math on an accountId?
You seem to think that's absurd? Maybe I want to store a set of accountids. So I sort them (bit-extraction or comparison!), take deltas (subtraction!), and encode the deltas sensibly. Maybe I have two lists of accountids and I want to check that none of the things in the first list show up in the second list. Some sort of hashing scheme (bit-fiddling!) is going to be my friend here.
> > What ORM? What database? If I'm using an ORM and a database, high-performance time operations are likely out of the question.
> So you’re going to pass a raw key/value record set around your code and program like it was 2006? So now you’re going back to mapping your sql resultset to an object that makes sense anyway. Going right back to my point that you’re going to have a mapping between your raw results -> domain model -> view model (where the view model is serialized for external use) and back again anyway.
What SQL? What resultset? What object? I have an integer. Why do you assume it doesn't make sense?
You’re going to do “bit extraction” to sort? If you treat accountId as an opaque type, you are going to have a comparison operator as part of the type and the rest of your code is going to just use a < or > symbol.
take deltas (subtraction!), and encode the deltas sensibly. Maybe I have two lists of accountids and I want to check that none of the things in the first list show up in the second list. Some sort of hashing scheme (bit-fiddling!) is going to be my friend here.
Or in a modern language since you have already defined equality between the two types and overridden the GetHashCode() function....
var deltas = accountList1.Except(accountlist2)
Why are you doing that all through your code instead of encapsulating the concept of equality, less than and greater than in one place?
What SQL? What resultset? What object? I have an integer. Why do you assume it doesn't make sense?
You were referring to CRUD apps and having to serialize the Account object. Either you are using an ORM, you are getting results from a database as a record set - which is usually represented by a dictionary of key value pairs and mapping it to your Account class or you are sending back a raw record set.
More than likely, you are mapping from your domain model to your externally exposed view model anyway.
You (i.e., all commenters in this thread somehow disputing the AccountId object version of the story) are completely right, it is difficult to decide up front what you need. Of course, depending on your domain, it may be perfeclty valid to have -5 years or 1800, but I am not asking you to never implement years like this, I am merely suggesting that it is useful to have one place where such decisions, invariants, roles - your understanding of a year, perhaps specific to your domain - are located.
How your understanding is structuring is a different concern, because of course, a class with 10000 LOC is horrible to work with. But you can always use means of abstraction, and compose independent modules. An int does everything, maybe in one context a year shoudln't be 1800, maybe formatted as 4 digits. A year class / type / object is the place to store these decisions. As I outlined, such decisions are both technical (how many bits?) but may also be domain specific (year must be > 1800). Maybe you have an int internally, but abstract it with a non-zero int.
It boils down to interpretation. An int allows for many operations, but they may be invalid (in the sense of "implausible", or "undefined"). For example, an account ID of 1 can be multiplied by 2 and then by 2 again, and so on. Integers form a relation this way, but not all extensions of this relation might be "meaningful.
Integers as ids are a great example for this. Technically, they are great, because they are simple numbers, can be typed on a keyboard, and be readily interpreted, but not like integers.
6 people two times as many than 3 people. In Germany, grades are marked 1 - 6, (1 = very good, 6 = very poor). But the relationship is merely ordinal. A grade of 4 is not "twice as bad" as a 2, although the numbers are relatable that way. An Account ID of 6 is not twice as "good" as an ID of 3. For ids, you want them to identify, in that they are exclusive and exhaustive - but integer IDs are not supposed to be ordinal (in that their sizes are comparable in relation to each other, 6 > 5 is an invalid statement). Of course, you can retain this interpretation, e.g. an auto increment id would indicate that id 6451 was created WAY later than ID 6, but this is difficult to interpret; because it doesn't really tell you how much later, and also deletion of integers in between a range may be reissued (id 1, 2, 3, 4, delete 3, 3 is a missing rank, 3 could be reissued). So ids aren't ints, because they are exhaustive, but they aren't ordinal, or at least interpreting them ordinally is dangerous, and some operations aren't allowed even if you interpret them to be ordnial; for example, it would not make sense to calculate an arithmetic mean from integer ids, although mathematically this operation is allowed. In statistics, this concern is discussed as the Skalenniveau (German, meaning Level Of Scale), in English it's called level of measurement [1].
In sum, it does not matter what your usecase is, an int is an int, but depending on your use case, it might be worthwhile not to pass an uninterpreted integer around, but actually wrap it in some kind of object, where you localize all your decisions how to interpret the integer.
None of should be derogated as some form modern hipster javascript, where none of us highlevel kids don't know how to bit-bang a set difference from some account ids; but rather, these ideas are really old. Even in C, data abstraction is useful, in that you don't fiddle with integers but define a set of methods, possibly in a module that interact with a hidden internal representation. Deciding on where to cut these modules appears to be a difficult task, but we all have known this for a while now [2].
As I said, nothing wrong with bitbanging, but encapsulating interpretation in types, classes, objects, modules or functions is a useful strategy to reduce the complexity, and you lose all of it when you just pass integers around and cross your fingers and hope the next developer won't calculate a mean from your integer ids.
And finally: The performance argument. Tell me how many requests/ops per second you need and let's find out why your program can't do them. Make it work, make it fast, not the other way around.
I have had too many arguments where people talked about "performance" without stating numbers. I had a colleague who argued that joins were bad, because they were slow (which conceptually, they are), but then your database is a highly optimized processor to do exactly these operations. Their fear of table-joining yielded a database with few tables, each of which had very long columns, each of which contained values separated by two semicolons, which they would then manipulate using string manipulations. Also the table grew, because the lack of normalization caused a lot of redundancy. I have seen many sins in the name of performance, and I will kindly ask about some numbers. Performance without numbers is not a good argument against data abstraction.
Also it can usually be handeled. If your high level Python program becomes too slow for some reason, feel free to implement the slow parts in a really fast language and add a clever algorithm in assembler; see for example np and scipy.
These approaches aren't mutually exclusive. "Please don't do it" and "considered harmful" doesn't get us anywhere.
Anyway, coming back to the topic of the thread: How does code become unmaintainable? By following rules of thumb without thinking on their contextual requirements, and most often by passing around integers in the name of performance, while ignoring locality of decision, locality of code, and ease of comprehension. An int is memory, and account id is an intentional interpretation of this memory (an int also being an interpretation of memory already, I get it, this is about abstraction - and finally, true comprehension comes from world-reference. The int doesn't care whether it is 5. You do, see above.)
RAM and CPU power are less expensive than two developers wasting their time trying to understand some low-level code and identifying which functions expect an int64 and which need a size_t, and whether they are equivalent, and so on, while they could just be passing around some thing with a stable interface, and a localized world-reference (you know, a name).
I wouldn't argue that all of this object/type stuff it is THE way to go, but these treatments were all invented specifically to solve the issues of raw-integers. And certainly I am quite ok with the fact that not everything uses Java; however the ideas we're talking about here are not specific to Java, and Java in particular often provides a very poor version of the story.
I would, however argue, that it is important to be congurent to the unit of expression of your programming language. Treating C as if it were object-oriented will give you a bad time, and not using objects in C# and Java will also give you a bad time. If your language natively provides an optimized iterator pattern, as does Python, coding in the C-for-loop-index idiom will give you a bad time. Most unmaintainable Java comes from not understanding Java, because you mistake it for C without & and *, but with Objects.
Unmaintainable code is about people. Code doesn't maintain itself.
</rant rel="sorry">