How Developers Choose Names
arxiv.org
arxiv.org
No one is going to confuse e7693160-b5cf-4761-9202-de019cfd0fc9 with c3d8b9ac-d0da-4bbc-912b-025ce4e47f62
arr1 and arr2 might use subscripts for 1 and 2 with no side effect on plan9.
Why use "delta" when writing a delta symbol is so easy?
How did you know?! Actually it's probably not far from the truth... most of the code I write is more like scripts in R & Python and small utilities than fully-fledged applications. And lots of SQL. I glue a bunch of stuff together, and after that is where my real job begins. I used to have a bit more fun writing C code to work around the limitations of an ancient ERP: that system's foundations literally extended back to before the moon landing and long before I was born.
The output of nm for the function originally named "write" from musl libc might look like
000000000006ee20 T 000000000006ee20That was not fun.
I've seen way too many times how much confusion is created and time wasted when a developer not too familiar with the project comes up with a brand new verb for a common action, and couples it with a new synonym for an existing business domain entity – or worse still, uses the same name previously meant for a completely different domain entity.
The unfortunate burden with this is that a new developer might be right that their name is better for their use case, which then leads to laboursome debates about which is more important: consistency or slightly crisper name for the one use case.
That's why good API design is so hard. Every name you assign to something becomes a convention by default.
It is not beneath you to hit thesaurus.com while picking a name.
It's one way to achieve job security for life though: https://github.com/Droogans/unmaintainable-code
Hey, think yourself lucky. It's only a pretty recent occurrence that my team has more than just the two "Prod" and "Non-Prod" environments...
I remember the era when sys admins would name machines or environments after nerdy shit like Tatooine, Dagobah. And you had to memorise which was the test environment and which was the mail server.
Thankfully I have never seen that in code.
Edit: eschew -> assign
There are very few situations where you can definitely say that there isn't meaning, order and correctness as opposed to not knowing if you are just too dumb to find them.
Eg it’s more helpful for readability to do “onClick => updateDate()” rather than “onClick => handleClick()”
(Not the best example and I’ve seen far more egregious, but those examples are escaping me now)
EDIT: sorry meant to reply to the parent comment.
But this whole thread brings up something about coding that has always irked me, it's very opinion based. When two people strongly hold the opposite opinions on the same team, it can be a massive hassle for absolutely zero benefit.
It's good to know the difference and it gave them the space to challenge me on some of those methods, which in turn helped me write better code.
On teams, or in life in general, IMO when there is no consensus, or even an anti-consensus, the only rational global decision is not to make a global decision.
Solving an issue that is intractable via discussion is when you escalate. Probably to a tech lead, or even to a manager.
This sounds like exactly the kind of pre-mature optimization that leads to overly-abstract function names throughout the codebase. Your point is a good one, once the function does something more than `updateDate()`. But until then, just call it `updateDate()`. As a bonus, this way any time you do choose to use `handleClick()`, the name is conspicuous enough that you pause to see what side effects it might have.
Besides, there’ll be less work if the requirements do in fact change, and there’s nothing wrong with accounting for the possibility of change when you’re writing code because what you write now constrains what you can write in the future.
Microservices and JSON APIs.
Common library code (eg. common session handling) incorporated into dozens of microservices.
Monorepos (which actually make this problem somewhat easier).
I have a C# .net MVC web project. All of the JS, .cshtml templates and C# controllers are in the same visual studio solution. So it has complete vision into the code.
So the boundary is from cshtml template to javascript function and then another one from the js function (ajax call) to C# controller.
Visual studio has complete vision, but loses the connection at each boundary, so intellisense says that the C# controller function is never called.
I am not criticizing Visual Studio, it is a great tool, but this is a hard problem.
The downside is lots of little functions, the plus side is smaller behaviors. (Ie avoiding 5 levels of nesting in a single function.)
There’s nothing wrong with 5 levels of nesting. It’s much easier to follow than 5 separate functions.
When someone (or you) is reading the code later, you see a method call and know exactly what it does. This prevents you from having to click in or find it just to figure out what it does.
Frequently, if you click into a vaguely named method, it doesn't do the thing you were searching for and you wasted a little time and need to backtrack, only to repeat the cycle with other vaguely named methods.
Also when you are reading the reverse, it’s not clear on what on what updates date. The worst scenario is when two listeners share a handler.
updateDateCallback()
updateDateFromClick()
These show how it's used and what it's doing.
(actually, "up-date date" kind of irks me,)
handleDateChangeClick()
dateChangeModalClickEvent()
And what does this thing actually do? Is it updating data, or making the update UI visible? It's kind of unclear.
If you're showing/hiding UI rather than performing the model update,
showChangeDate[Field/Modal/...]()
beginChangeDateFlow()
...
Code should read kind of like a book.
In that case what I find preferable is onClick => handleClick()
handleClick = () => { stuff in handle click updateDate(); }
on edit: I see lots of others made same point one node lower.
In your example, assuming we're looking at a React codebase (since this seems to square with React's style of events, etc), the resulting data being set into the updateDate function would be an event.
This, unfortunately, doesn't make any sense for a function by the name of updateDate to receive. I would expect to receive at least a new date to update with in the parameters, or ideally in a functional world, both the state to update and the new date we want applied to it. Anyone thinking they could simply reuse the updateDate function somewhere else is going to be woefully disappointed they largely cannot, since it would have been constructed around receiving an Event object.
In that case, I find the "handle" nomenclature to be very useful, as it appears to be largely shorthand for functions design to handle an _event_ (and we tend to see this pattern being used in various React libraries and documentation). React does have a few of these shorthands it tends to use (such as useEffect largely being short for useSideEffect).
Ultimately, I recommend using both functions. One, a handleClickUpdateDate function (notably not a hyper-generic handleClick function which conveys nothing) that receives the event and pulls out the necessary information out of that event to determine what the date is that the user has selected. It then will call the updateDate function with only that date, which creates a nice, reusable updateDate function if we need it anywhere else.
This roughly squares with the idea of having functions that handle the data transport layer (such as receiving data via HTTP or gRPC, etc) whose responsibility it is to receive those events, pull out the necessary information out of the package, route the request to the correct controller, and ultimately return the result in a shape that satisfies the original transport mechanism. In this case, our handle* function is responsible for receiving the original transport of data, then routes the request through to the appropriate controller which is entirely unaware of the means of data transport.
It also means we have a nice, easily testable unit of updateDate to verify our state modification is working to our liking without needing to assemble an entire Event object.
Anyhow, that's how I think of these things ;p
> getInstances(); // ok, returns an array of instances
> getInstanceId(); // ok, returns an id
> getInstancesId(); // ??, returns an array of id
> get InstancesIds(); // ?? returns an array of id
> getIdsOfInstances(); // ?? returns an array of id
> getAllInstancesId(); // ?? returns an array of id
> getAllInstancesIds(); // ?? returns an array of id
How do I conjugate that ? And what about the possessive `s` ?
Only ones with an "s" at the end of the id should return an array if you have the existing "getInstanceId" either that or invert that function to be getIdOfInstance and subsequently getAllIdsOfInstance for the array version.
Generally speaking don't use the possessive "s" it is implied by being next to it. Also I usually use "instance" in this case more like an adjective
Edit: usually I think of it as (current subject) verb adjective object.
So server get instance IDs becomes getInstanceIds() or zoo clean active dogs becomes cleanActiveDogs()
> getInstances(); // ok, returns an array of instances
Correct.
> getInstanceId(); // ok, returns an id
Correct.
> getInstancesId(); // ??, returns an array of id
Unusual, but would return a single ID referring to a collection of instances.
> get InstancesIds(); // ?? returns an array of id
Correct.
Though more common would be getInstanceIds() since you don't need to pluralize twice. E.g. we don't say "blues shoes", we say "blue shoes", so these aren't "instances ID's" but "instance ID's".
However, in naming functions or variables you'll sometimes see plurals repeated for clarity. E.g. getInstanceIds() is ambiguous because it could refer to the ID's of multiple instances, or multiple ID's of the same instance. While getInstancesIds() makes it clearer that it's the ID's of multiple instances, even though it's not actually grammatically correct.
> getIdsOfInstances(); // ?? returns an array of id
Same as getInstanceIds(). Both are correct, though getInstanceIds() would probably be more common -- most people wouldn't put in the extra word "of" simply because it's an extra word, but sometimes it might be clearer for consistency with other function names.
> getAllInstancesId(); // ?? returns an array of id
It would mean the ID for the collection of all instances, though that would seem very unusual. Perhaps if there were a global account ID for a service or API you used.
> getAllInstancesIds(); // ?? returns an array of id
Sams as getInstanceIds(), but "All" presumably means it doesn't include a parameter to filter -- so I'd assume this implied the existence of another function such as getCollectionInstanceIds() or similar that it was contrasted against.
> How do I conjugate that ? And what about the possessive `s`?
Obviously you can't or shouldn't use apostrophes in variable names, so people generally try to avoid the possessive s. That's why we generally change nouns to noun adjectives -- instead of "computer's ID" (computer = noun) we use "computer ID" (computer = noun adjective).
Completely agree with all your points (as a native English speaker who also minored in English)
However, for the sake of clarity (since we're trying to disambiguate between places the 's' could indicated possession vs. plurality) "instance IDs" should not have an apostrophe.
The only time an apostrophe goes in a non-possessive plural is in a contraction. E.g. "the '90s"
Then we also have fun words that probably should never be used which get an apostrophe on both ends (contraction at beginning + possessive plural)
e.g. " ... the '60s' countercultural attitudes ..."
But for clarity it might be better to go with
"... the countercultural attitudes of the '60s ..."
Some style guides insist plural all-capital abbreviations should never have an apostrophe ("IDs"), other insist they should ("ID's").
Your personal usage is absolutely valid, but it should be recognized as opinion and not fact. "ID's" is equally valid in modern usage.
Edit: You can read an extensive discussion of it here:
https://en.wikipedia.org/wiki/Acronym#Representing_plurals_a...
I've seen it for one-letter words where you don't want to switch case: e.g. "dot your i's and cross your t's"
But I haven't seen it in upper-cased abbreviations followed by a lower case plural 's'. Or if I have, I've assumed it's wrong. Would love it if you can refer me to a modern style guide which doesn't recommend dropping the `'` in this usage, as my own background has drilled into me that "IDs" is preferred for clarity.
> ...whereas The New York Times Manual of Style and Usage requires an apostrophe when pluralizing all abbreviations regardless of periods (preferring "PC's, TV's and VCR's").
URL was carefully chosen, since you can't make a consistent ruling about initialism vs. acronym, I've worked with people who pronounce it "you are ell" and people who say "earl", and sometimes both! Absolutely no one is going to start spelling it U.R.L., either.
Especially given the abundance of non-native speakers, let's not add a peculiar and ambiguous edge case to one of the more frustrating rules of English grammar.
First, I don't know what it means to "shadow" the possessive. Do you mean we shouldn't hide it? Although it always has an apostrophe so we're never hiding it.
Second, I don't understand your examples. Your second example clearly necessarily uses the possessive, while the first is a plural, but the first would clearly be plural even with an apostrophe:
> "this script retrieves JSON's from all API URL's"
And since a possessive necessarily precedes another noun in this type of construction, it would still be clear:
> "this script retrieves JSON's from all of the API's URL's"
It's clear in speech and clear as written, even though the "'s" serves two grammatical purposes here.
“Adopt a plural form identical in form to the possessive so as to create avoidable ambiguity between the two.”
(Incidentally, that’s why the Times shouldn’t do it either, even if context will usually disambiguate it.)
> And since a possessive necessarily precedes another noun
Not necessarily a simple noun, though, it could be a noun phrase. In practice, this will usually be disambiguated by context still, but merging the possessive and plural means that there will be more situations where you have to process more context to disambiguate meaning, which distracts from whatever your purpose was in reading the material in the first place.
Reducing the inherent redundancy in natural language reduces quality and clarity of communication.
I don't know why an instance would have more than one id, probably more plausible with different nouns, but a similarly named function might return an array of arrays, or a flattened equivalent, possibly uniqued.
Probably a bad idea though. I have been guilty (though pleasurably, pridefully guilty) of abusing plurals -- day, days, dayses etc. If you can let the types do the talking that's obviously better, and more descriptive (and less eccentric) names can help a lot with communication and clarity, but there is also sometimes a place for brevity, and for levity.
So FetchInstances() is expensive (each time it's called) but the property Instances is cheaper or only expensive on the first access.
Not necessary tied to it's inner workings but more tied to how you should use it. There are good reasons that a function might not cache it's value -- maybe because you want that fresh data each time. The name is the signal to the user how they should use it. I expect an "Instances" property to give me the same values each time. However with "FetchInstances()" I would store that result in my own variable if I want to keep referring to those instances.
> What if you change the implementation so that it returns a cached value?
That's a different function then. The semantics have changed significantly.
(1) How would I explain face-to-face what this code does to another developer?
(2) Use the key words from that explanation to make a name as short as possible without losing clarity.
Example:
Step 1: This class handles all the navigation work for the sign-up flow in our mobile app.
Step 2: SignUpFlowNavigator
It´s a funny story, but naming means ALOT when you are developing, or building systems or processes that requires a name or an identificator.
Also don´t be afraid of changing the name if it no longer reflects on its purpose.
For example in graph algorithms, I used to write descriptive names like:
weight = adjacenyMatrix[node][neighbor]
Nowadays I write it as: w = G[u][v]
Because in graph theory, an edge is usually (u, v), weight is usually w, graph is usually G, etc.It's basically the same reason why people don't name their indices "index" and instead use "i", "j", "k". And instead of point.first and point.second, you use point.x and point.y.
Naming conventions are awesome when people can agree on them.
This zealotry a lot of devs have over descriptive names needs boundaries.
Likewise, when programming, if I reuse some variable very often in a short space, I will temporarily rename it to something single-lettered, e.g. "s = volumeScalingFactor; x = s; y = s; z *= s".
I think maths style vs. programming style is only one factor here, with the other simply being this application of DRY: describe one-off things well, but shorten their names if they're oft-repeated within a given context.
Maths tends to reuse the same variable for a lot of different operations in a single context whereas programming doesn't as much, which naturally leads to this single-letter name vs. long name difference in styles.
https://hsm.stackexchange.com/questions/5563/where-does-the-...
https://www.joelonsoftware.com/2005/05/11/making-wrong-code-...
This seems flawed, when I look to name something, one of the most key things is what other things are named in that context, and patterns and consistency between them.
It appears the study described a problem, instead of putting them in preexisting code?
http://pdinda.org/Papers/ipdps18.pdf
surveyed ~150 and concluded therefrom. the more hilarious thing being that they drew conclusion completely unscientifically i.e. by just interpreting vague plots (i.e. without performing t-tests or anything like that).
It takes up far too much tought-space to the point I would gladly follow some ugly set of conventions merely to have to just never think about it.
Not really, because naming is 98% convention when it's done right.
The challenge is establishing all the conventions for a project, and then sticking to them.
If there are really good conventions in place, well known, then naming is considerably easier.
In fact, even if the conventions are not super good - they can be powerful when they are very well adhered to.
Example: recently broke a rule and decided to use a known naming anti-pattern by suffixing variable names with the system type. This is normally not good. But within this module, the meta typing was ambiguous - by adding the suffix, the code was magically more clear. That little convention, very easy to apply, solved a clarity problem far more so than any issues around what functions should be called. So we used it for the module and that module only.
Apple has some pretty hardcore naming conventions that I don't really like, but what's more important then whether they are good or not, is that they are very consistently applied - in other words - a lot less to think about.
would not be surprised if some dev somewhere considered naming his/hers first born "/tmp/first_born" until coming up with a better name
we can also add mathematicians and physicists to the group of people that are exceptionaly bad at naming stuff
But don't know what anything does.
How do you name things so they're obvious to the eye?
Descriptively,
descriptively,
descriptive... L-Y!
Like they say, there are 10 kinds of people, those who understand ternary, those who don't, and those who mistake it for binary.
Your unit test of a function comparing dates relies on server time and usually fails if run around 12pm. Whatever. Nobody uses this app around lunch time anyway. Just tell your teammates in the to ignore that. Wait, why is the test failing every time today?! Did we just switch to DST? Uh oh...
There is an initial request to create a session that has to pass the desired expiration time of the session. Unfortunately, the vendor requires the time to be in Eastern Time. The poor, naive soul that originally implemented this just got the current date, added 15 minutes, and converted it to "EST". As soon as daylight savings hit a few weeks ago, the expiration time was automatically being set to 45 minutes in the past, the vendor was responding with "Invalid Session" and the company was unable to take payments from customers.
Within a day, I got a bug report from a coworker that it rendered incorrectly for them. I was pretty surprised, since I had been using it for a long enough that I would have expected to encounter all the edge cases.
I dug into the code, and realized that this bug only happened if you went to sleep before midnight, which I never did...
Perhaps there is an answer lurking out there in the universe, but it feels more likely it'll only be invalid pointer Aborted (core dumped)