CamelCase vs underscores: Scientific showdown (2011)
whatheco.de
whatheco.de
Case insensitive, of course
In Python the logging library is written in camel case (as it was heavily based on log4j) unlike most of the rest of the standard library, but the developers have let it be because fixing it would break backwards compatibility, and in any case PEP 8 [1] says consistency within a module is more important than global consistency.
In my opinion they should have fixed it in the move from Python 2 to 3, but I guess they had enough headaches from that.
Looking at you, C++ STL. Almost nobody names structs/classes all lowercase. This leads to silly style guidelines eg. in the Google c++ style guide:
- Name classes PascalCase (even if it's templated)
- Except if your class happens to be a templated container type, in which case name it all lowercase.
Postgres in particular has case-sensitivity quirks if you try to force the use of camelCase table names.
My general preferences beyond that are:
- Opening braces on the same line (I see the article has examples where it's on a new line);
- Two space indent, no tabs
- No trailing white space. This should be automatically removed so it doesn't generate extraneous changes on commit;
- A reasonable line length between 80 and 120 characters, depending on the language. You need to be able to look at 3 files side by side without wrapping.
- Don't put the return type on a separate line. This is a really old school (K&R) C style
- Don't align function parameters with opening parentheses. Change the function name and you generate a bunch of changed lines for the parameters.
- No space before semi-colons eg for (i=1; i<100; i++) not for (i = i ; i < 100 ; i++ )
While you could apply the same kind of automated logic to naming, the risk of collision is non zero and moreover would likely break runtime mechanisms like reflection, etc.
Generally yes, but stuff like
ReturnType<SomeOther<S, TypeConstructor<Nested<S, T, U>, T>>, U, HmmLetsAdd<V, W>>
can be on a line of its own.1. You're writing C++. You've really lost half the battle laready :) and
2. Writing correct templated code is difficult and should generally be reserved for when you're writing a library; and
3. For complicatred types like this, one should strive to increase readability by using type aliases, assuming your language supports it (eg C++ and Hack do). It's not always possible.
ReturnType<
SomeOther<
S,
TypeConstructor<
Nested<S, T, U>,
T,
>,
>,
U,
HmmLetsAdd<V, W>,
>Also
(i = i ; i < 100 ; i++ )
Whoever does that do not change it, they are probably a psychopath. Dont risk your life correcting them.
French generally adds a space before punctuation.
And here I was agreeing... But that was because of the "i = i".
Every developer can control how those tabs are rendered to their preference.
My suggestion:
> There seems to be a basic assumption of ASCII identifiers. Hey, this ain't the eighties!
> Let's have us an XᴍʟHᴛᴛᴘRequest. Absolutely clear with no scope for misunderstanding in either direction: Gᴄ; Rᴄ; Aʀᴄ; SimpleHᴛᴛᴘServer.
> Monospace font support is a little poor, but I'm sure they'll fix that up once the desire is demonstrated.
> Q and X don't have small-caps variants in Unicode, so acronyms will be banned from having a Q or an X in the middle.
(My email wasn’t just trolling; I also added meaningful arguments in both directions to the discussion. But small caps was just too fun a concept to not mention. I also notice that Unicode 11 in 2018 added U+A7AF "ꞯ" LATIN LETTER SMALL CAPITAL Q for some reason (subhead “Letter for Japanese phonemic transcription” and I haven’t looked any deeper), but there’s still no small caps X.)
I’m glad to say that Rust stuck with the .NET rather than Java style—it’s easier to reason about, apart from anything else, because of having fairly unambiguous rules, and supports tooling better in a similar way to snake_case, because of unambiguous word separation. I’m also glad that no one ever challenged Rust’s use of snake_case for variables and fields and such.
some_var.some_fun(some_param, another_param)
Instead with CamelCase, they are immediately visible: someVar.someFun(someParam, anotherParam)
But my preferred syntax is Lisp: (some-fun some-var some-param another-param)
(when this-looks-appealing
(setf you-like-lisp-syntax true)
(vote-poll 'camel-case-formatting))
IMO, "this-looks-appealing" is more readable than "thisLooskAppealing", but "-" is more space friendly respect "_". some-var.some-fun(some-param, another-param)
the "-" hides the ".", but it should be the contrary, because "." is a stronger separator.And yeah in Japanese it's fine, there's a clear visual situation between the Kanji and Kana.
whenwordseparationbecamethestandardsystemitwasseenasasimplificationofromanculturebecauseitunderminedthemetricandrhythmicfluencygeneratedthroughscriptiocontinua
I also got unnaturally irritated by PowerShell's insistence on hewing so closely to CamelCaseOrthodoxy that common abbreviations and acronyms like ID, IP, and DB became Id, Ip, Db.
Left to my own devices, I'll use OTBS + lower_snake and it drives some of The Youths on my team insane.
Configure TAB width = 1 character, and there you go.
This, and the fact that variable names are allowed to be implicitly defined, lead to the famous bug:
DO 10 I = 1.100
declared the variable `DO10I` with a value of 1.1, instead of the loop from 1 to 100 and declaring the "statement label" 10: DO 10 I = 1,100
SUM = SUM + I
10 CONTINUE> It is good to have names containing multiple words, but there is little agreement on how to do that since spaces are not allowed inside of names. There is wun [sic] school that insists on the use of camel case, where the first letter of words are capitalized to indicate the word boundaries. There is another school that insists that _ underbar should be used in place of space to show the word boundaries. There is a third school that just runs all the words together, losing the word boundaries. The schools are unable to agree on the best practice. This argument has been going on for years and years and does not appear to be approaching any kind of consensus. That is because all of the schools are wrong.
> The correct answer is to use spaces to separate the words. Programming languages currently do not allow this because compilers in the 1950s had to run in a very small number of kilowords, and spaces in names were considered an unaffordable luxury. FORTRAN actually pulled it off, allowing names to contain spaces, but later languages did not follow that good example ... I am hoping that the next language does the right thing and allows names to contain spaces to improve readability.
Yes, ML style function application is a problem and treating newlines as "normal" whitespace without having line separators (aka semicolons). And keywords used as infix operators, like another post reminded me of.
Since the space of valid syntax becomes so much larger, typos are more likely to result in valid but incorrect programs. Especially in dynamic interpreted languages like python.
It makes sense, you would never want your variable to be line wrapped anyway, right?
The challenge is that _some_ languages define the space to be unicode whitespace, not the space character.
CamelCase vs. underscores: Scientific showdown - https://news.ycombinator.com/item?id=9138156 - March 2015 (57 comments)
CamelCase vs underscores: Scientific showdown - https://news.ycombinator.com/item?id=5224531 - Feb 2013 (6 comments)
Also:
CamelCase vs. underscores revisited (2013) - https://news.ycombinator.com/item?id=34525139 - Jan 2023 (257 comments)
As someone who has battled RSI, I stopped using snake case and underscore-prefixed member variables because of the added stress all those underscores place on the weakest fingers.
But you do have a point because not everybody has access to programmable keyboards almost all the time. Maybe snake_case isn't that ergonomic now that I think about it, although I find it easier to read than camelCase.
Ultimately, the winner is still kebab-case, which is both aesthetically pleasing and is not a double-pinky keystroke!
Madness. The only thing worse that developers willingly tolerate is prettier's sacrilegious linebreaks [0].
0: https://prettier.io/playground/#N4Igxg9gdgLgprEAuc0DOMAEBXNc...
foo
modified_foo
foo
modifiedFoo
With underscores, "foo" looks the same whether it stands alone or not. foo
fooModifiedI still don't understand how comes that in times of programming fonts with fancy ligatures there is not a single one that would attempt to make camelCase more legible by slightly separating lowerUPPER sequences (presumably by constructing a "anti-ligature" for each unique pair that would have still width of two glyphs).
Or try to do anything similar to ease reading theseAnachronisticCharacterSequences.
In Swift (where I spend most of my time), it's CamelCase, or, quite often, dromedaryCase.
Probably not a good idea to propagate it.
;)
I think that there is an official name for lowercase-prefixed CamelCase, but I don't remember it.
Thanks!
But then again, camels have heads at about the same height as the hump, so...
Also allowing hyphens generally leads to issues when interoperating with other languages that don't support hyphens. Probably the best example of this is CSS which does allow hyphens, and Javascript which doesn't. `background-color` in CSS gets translated to `backgroundColor` in JS.
It's an annoying paper cut that trips up beginners and makes code less greppable. I generally avoid hyphens wherever possible for that reason.
You shouldn't be able to start a name with a hyphen, so `a -b` and `a - b` can be both valid.
But I mostly just came to acknowledge your fitting username ;)
Like what the sibling comment said: `-b` should be illegal. And I don’t think this would be a big deal in practice for experienced users (used to these rules).
For beginners you could build in an error check: give a dedicated error message if you write `a-b` but you happen to have both variables `a` and `b`. Then the compiler can tell that you probably meant `a - b`.
> Also allowing hyphens generally leads to issues when interoperating with other languages that don't support hyphens.
Make underscore illegal in the language. It’s not like you need them anymore. In turn you have a bidirectional translator.
Sure... but I didn't use `-b`. Maybe I've misunderstood.
> give a dedicated error message if you write `a-b` but you happen to have both variables `a` and `b`
I guess, though the idea of having mutually exclusive sets of identifiers sounds like a nightmare. You can have `a` and `a-b`, or `a` and `b`, but not `a`, `b` and `a-b`... Maybe it wouldn't come up much in practice but it's still a pretty big WTF.
> In turn you have a bidirectional translator.
But the problem is that you have this translation in the first place. It makes the identifier ungreppable. Ideally any place you have an identifier it looks exactly the same through your whole codebase and if you have to translate it then that is no longer true. For example if you want to update all background colours in your project you might search for `background-color`... but miss `backgroundColor`.
Note: this is not a good idea
The first restriction might make this a problem. I am not saying it is a good idea, but it is not obvious to me.
foo and bar == true
means foo && bar == true
or foo_and_bar == true
?You could fix it with ugliness like making the keyword `@and` or some such, or the variable `foo @and bar` but that's not an improvement.
Like I said, I can’t picture the requirement, but am not sure it is a good idea.
DO 10 I = 1.100
as DO10I = 1.1
and DO 10 I = 1,100
as the beginning of a loop ;)I've never written a compiler, but I don't see how the last lines are harder to parse:
if ( thisLooksAppealing )
{
youLikeCamelCase = true;
votePoll( camelCaseFormatting );
}
else if ( this_looks_appealing )
{
you_like_underscores = true;
vote_poll( underscore_formatting );
}
else if ( this looks appealing )
{
you like spaces = true;
vote poll ( space formatting );
}You have to be careful with infix operators, but there is always Haskell's solution of adding backticks for infix names.
So add(a, b) is the same as a `add` b
But without joking, a _real_ solution to infix operator names would be a backslash as prefix, as pseudo-Latex-style, so a \add b
That's not a problem at all, as long a a newline isn't treated as space and you don't use ML style function application, where a space is used to separate the argument(s) from the function name - `f x` instead of `f(x)`.
var initial factory = abstract configuration factory factory. configure new factory() var search result = users. find all by name (name);
if (search result. is present()) {
return ok(search result. get());
} else return not found();
It’s weird and syntax highlighting is absolutely necessary to read this. search result := users @ find all by name (name)
if search result @ is present() {
return ok(search result @ get())
} else {
return not found()
} search result := users, find all by name (name)
if search result, is present() {
return ok(search result, get())
} else {
return not found()
} // comma is optional, used for disambiguation
search result := users, find all by name. // omitted parameters match words
if search result is present, // empty brackets can be omitted
return ok(search result value); // API change for better readability
else
return: not found. // colon is another optional delimiter
I actually like it so much, I would try to write a transpiler to Java… PS> ${Valid characters for an identifier? Eh, ¯\_(ツ)_/¯ whatever (`}) you want.} = $True
PS> ${Valid characters for an identifier? Eh, ¯\_(ツ)_/¯ whatever (`}) you want.}
True
Jesting aside, I have actually used that syntax before to prefix variable names with '&' and '@' to differentiate between virtual and physical addresses in some code for patching a binary.ThisClass thisVariable = null
The above is "obvious", as is the below:
this_class this_variable = null
However what is
this class this variable = null
Is it a variable named "variable" of type "this class this"? Is it a variable named "this variable" of type "this class"? Is it a variable named "class this variable" of type "this"?
These are false time savers. We’re writing code, just recognize that and do what’s natural, snake_case. Any conventions to make code “more readable” to make it like written language seem like fool’s errands to me.
Anyways, I think you should follow the specific guidelines of the language you are currently using.
C, Odin :- my_function("hi")
D, Java :- myFunction("hi")
C#, Pascal :- MyFunction("hi")
Lisp, Scheme := (my-function "hi")
Even these are not 100% accurate. For Odin, I would create a struct like MyStruct. In C, it would likely be my_struct_t.
etc.
In future, I might start following Anti Pascal Case for all languages
mYnEWfUNCTION("hi")
val somethingRepository = SomethingRepository()
Is much more visually satisfying and balanced to me than:
something_repository = SomethingRepository()
But honestly it’s most likely because I was socialized on Python. Feeling lucky personally that Rust follows the same convention.
That got me to learning about all the ways that markup can (and should) be used to convey _both_ the content _and_ the structure of the information on the screen, for the benefit of vision-impaired readers.
That in turn led to the epiphany that the tab character _is_ markup for indentation, and in the world of programming, where indentation is so significant for understanding (especially in whitespace-sensitive languages like Python), I wondered why we were making things harder for vision-impaired users by focusing so much on _visual_ consistency (which can still be achieved by syncing editor "how do I render tabs" settings)