Exploring the impact of code style in identifying good programmers (2022)
arxiv.org
arxiv.org
I grant that conformity does signal education of the norms. Which can be seen as signaling effort. But I expect it to be so noisy that it will have little predictive power.
As such, it is seen at both ends and provides little predictive power.
Agree. To take this a step further, I would say good programmers follow the style that exists in the codebase, if applicable, for consistency and therefore maintainability's sake.
It' doesn't necessarily mean they are a good programmer, but it means they are fairly meticulous, which can be a trait of a good programmer.
The "good"/experienced programmers are the ones that will break established styles to improve readability in corner cases, instead of following the rules blindly. On top of that, they've probably been doing it long enough they have their own preferences and idiosyncrasies, some of which may go against established styles.
On the flipside are the "bad" and new programmers, who don't really have their own style nor have learned an established one, so they just do whatever.
It's the average ones in the middle who only follow established styles that I would expect to be easy to pick out.
I think the value of standardization is much higher than the value of "break[ing] established styles to improve readability in corner cases". I just find the marginal value of these little stylistic improvements to be extremely low, and the value of code looking quite similar everywhere to be quite high.
That said, using a programming contest as a judge is also likely to bring out the more unconventional styles, like minimised whitespace or indentation.
https://gcdnb.pbrd.co/images/hgHbrzHh2mWT.png?o=1
Over the years I came to use more and more custom vertical alignment of short code lines (whatever the language is). This makes semantic similarity/symmetry stand out, elicit functional decomposition, and enables one to have a complete one-screen bird's eye view over complex problems like this MQTT pulser that trigger interrupts in unison (because musicality of what appears in my MQTT console matters too):
If I was tasked with changing anything in there, or analyzing the code, the first thing I would do is run autoformat on it. even if I have to preserve the code style, I would just translate it back once I'm done
One thing I notice myself doing a lot is using lots of empty lines in my code. Sometimes I go really overboard with it, to the point that every other line is empty. I feel like it helps me read the code more easily. This seems sort of like the antithesis of that.
One trick for that is making your height line higher in focused mode in the editor.
An ex co-worker had this on VSCode and looked very useful.
If you are working on your own you do you. If you are working in a team pick a style that is okay and has great tooling and support in different editors.
As this exchange proves everyone has opinions about this stuff. Standardisation is more important than optimal choice especially as everyone thinks theirv choice is optimaland yours is terrible. It doesn't have a huge impact, solve it once and make it go away so you can focus on everything else.
Just set git to format everyone's code on checkin and never talk about it in a code review again.
Arguing over indentation is functionally the same as arguing over unix principles, or frameworks or languages. The fact that we're always looking not just for the solution, but the right solution, on every level of the process, is part of what makes programming so beautiful and enjoyable.
Law is actually a perfect example of the opposite; without that obsessive concern for correctness, the systems created are hopelessly confounding and difficult to navigate.
I mean that seems reasonable to me. If someone submits a giant bill with no indentation/paragraphs/punctuation etc. then it should be rejected. Being able to grok the bill itself would be an important part of then determining whether or not to pass it.
This could be done with local reindentation/deindentation. You'd need some way to map back to the original style, maybe treating the whole process like a git edit, putting a newline before each new addition, highlighting it, with a checkmark button to "commit" them once the indentation is adapted to the original style, and by "commit" I mean make the highlighting go away.
However the criticism about the right part of the code being visual noise I see in your reply and others is somehow unfair, the second pic bordering on the limits of what I find acceptable myself. It's still linguistically readable as in you'd still be able to replace the parens and braces by reading that one line of code aloud. It's not code that's supposed to be readable in a typographical sense, with harmonious spacing meant to catch the wandering eye like an ad, it's code that was written to be thought through. Sit and think. For 30 minutes before hitting 'build'. After minor compilation errors and the first successful build, I had just a single one-line fix to make before having the code run correctly.
Whenever I need that "high-level view" you mentioned, I enjoy seeing the whole module without having to scroll/search. In codebases where I gotta jump around too much, I feel disoriented and slowed-down. Like for example on Rails classes where single-use variables are replaced by single-use private methods, that are often far away from it's original usage point. I understand the rationale for doing it but it still tax me mentally when trying to get a holistic understanding of the code.
But then again, some people get used to sparse code and for then I assume such terseness can be overwhelming. Oh well. I just wish they could see my point of view too.
My preferred form really depends on why I'm looking at particular piece of code at a particular moment. I like my code sparse and modular if it eliminates details I don't care about. I love my code dense, tight and inlined, if it lets me fit all details I need on a single screen, or otherwise minimize the amount of scrolling and jumping between files. Depending on my particular goal, the same code can be perfectly readable, or a horrible mess.
In practice, I find most of the code I deal with too sparse, but that's again the function of the work I typically do.
(Which the example source file I found was. Btw HTTPS error)
I've been programming >40 years now.
However, if you are the only person maintaining this code, chacun à son goût!
For an amazing comment by someone discussing and justifying very similar goals to yours, check out https://news.ycombinator.com/item?id=13571159
Purely my opinion, not a critique of your style at all, but I don’t love vertically aligning the open braces of functions, unless they’re one-line and similar (like readTDS & readEC), nor do I love starting code or data way off to the right just to align it with the first paren or brace. I want my code/data block to have a single indent, not a random content-dependent amount. It gets unbelievably stupid when function names are longer, it wastes half the page/screen, and it drives me nuts how clang-format does this by default (yes I know there are different profiles, but I’m stuck with the one my work peers chose to use). For me, the opening braces aren’t a semantic similarity that I want to see or emphasize, because the opening code for each function is expected to be different, and because they’re rarely on subsequent lines.
All that said, the most important thing when working with other people is to have an automatic formatter, and eliminate arguments over style! It’s nice to have a datapoint in the article that style has no correlation with quality of GCJ entries.
Aside from the impossible to read without printing it out on ruled paper, change the length of one variable and your auto format pollutes git history with a five line change.
Most statements aligned like this aren't even related, they just happen to be colocated but then the (forced by the autoformatted standard) vertical alignment makes it look related.
Coding style (more precisely whitespace) is able to say a lot about the code without you even reading the code. This code style paints a picture that isn't there.
I’m a fan of only aligning things that are actually similar. The important part IMO is to make it easy to see differences in the middle of something that’s a pattern, not to extract similarity out of semantically different things (which is the reason I don’t love aligning the opening function braces). I’m also a fan of using the -w flag with git diff, and of avoiding hyperbole, but there is a valid and very good point in there about causing merge conflicts! ;) I don’t get to use my grid aligner all that often, because using clang-format is more important, so I don’t use it around other people very much. But all that said, I never hesitate to re-grid something if a variable changes, as long as I know it won’t cause a source control conflict.
* edit BTW I just realized the better argument which is that clang-format already has the very problem you mentioned: renaming anything can and will often cause reflow of text. Avoiding grid-alignment of text has very little bearing on whether this is going to happen to you. The only way to avoid formatting changes in your source control is to never reformat code, which isn’t something my team wants to prioritize, so we make do with selective application of clang-format. It’s sometimes useful to separate code changes from whitespace/format updates.
Specifically, the workflow here that allows intra-line whitespace without worry is to merge with whitespace disabled, and run the auto-formatter after merging. This means that vertically aligning things is almost innocuous, but having certain tools re-wrap text can still be problematic, at least as far as I know. If there are good ways to handle that second problem, I want that.
Learn your tools: git provides "-w" to ignore whitespace changes. SVN provided a similar option. This isn't new.
I don’t think vertical alignment is going away, I think it will become far more common, and that our tools will catch up. My prediction is that git and other source control systems become more tolerant of formatting changes, and that developers in general will learn more workflow techniques for allowing whitespace churn.
For me, having the alignment and making the code more readable to me is more valuable than the minor inconveniences with source control that it might cause, and on top of that I’ve learned that there are ways to mitigate most of the issues. I’m not arguing that it doesn’t cause any problems, I’m only saying that it’s worth prioritizing for me, that there are things you can learn to do about it, that the tools give you options and will probably improve, and that vertical alignment in particular does far less damage that line-breaking reflows like clang-format impose… and use of clang-format is pretty widespread and standard in my experience.
Code does change over time. Refactoring happens. Variables are renamed. Comments change over time. We don’t ask people to avoid writing code or adding comments because the changes appear in git diff or grep, or even because they might cause merge conflicts… those things are considered an acceptable cost of doing business in software. I don’t see any reason to treat good formatting differently, within reason.
(Edit: I see that yes git grep can search history if you pipe in a git rev-list command. Is that what you do? Note the trivial workaround with git grep is to put variable whitespace in the regex you’re searching for. If you aren’t looking for a substring with spaces in them, maybe use the much simpler git log -S command?)
Good news: the pickaxe command ignores whitespace changes by default, it only shows you when the substring itself was changed. And git blame takes a -w flag so you can ignore these whitespace changes. -M can help too. First Google result for “git blame ignore whitespace”: https://coderwall.com/p/x8xbnq/git-don-t-blame-people-for-ch...
Git gui has a menu option to ignore whitespace changes too, if you’re spelunking through changes on a given line an don’t have an exact substring change you’re looking for.
So hypothetically, if you had version control tools that truly and completely took care of not showing you whitespace changes, and you were free to reformat your code any way you like without causing any trouble at all, if you had complete freedom, what would you want out of your text formatting tools?
> what would you want out of your text formatting tools?
Consistent white space rules when reading from left to right.
I’m not sure I follow, what does consistent mean to you, can you elaborate? Is that consistency something you have a hard time getting right now because of source control issues? I guess I was wondering about whether you had a wish list for formatting, things that you don’t have but would take if you knew it wouldn’t cause any issues with source control. Do you use clang-format, or any other auto-formatters on your team?
I was somewhat afraid of the enforced indentation and newline policy, but I must say I love it, at least so far. Consistent and unified code style reduces my cognitive workload and lets me concentrate on the business logic of the classes I am working on.
When everyone has their own style, as was the case in all my previous projects (C++, C, Java, PHP), the extra friction costs some non-trivial amount of energy and time of everyone involved.
Of course there is more than just formatting to a "programming style", but still.
Perhaps the paper should be re-titled "Exploring the impact of code style in identifying programmers who can configure their IDE"
Modern IDEs have hundreds of knobs and dials you can tweak so that your code can confirm to any style you fancy.
Therefore if you are identifying programmers by coding style, what you are most likely doing is identifying people who can correctly configure an IDE.
Or, in other words, identifying people who can RTFM. ;-)
I am I suspect being grouchy, but I don't think that saving me some typing strokes is worth all that effort.
I love linting and really like real-time linting and syntax catching. But they usually just underline my typos. Anything else (like trying to prevent me from making the typos as I type, or worse adding in brackets just annoys me.
I don't want to invest time in learning an IDE with a million knobs - i can write my own functions if I see the need - just mostly I need to work on actual work.
Thanks machine learning.
I’d be more interested in the commonality of design patterns and standard data structures and algorithms.
Or does it mean style-of-semantics (if that's even a thing), for example in C#, preferring to define explicit constructors instead of a parameterless ctor and trusting the user to set properties in the post-ctor initializer correctly?
Or use of higher-order/FP techniques (like map/filter/reduce/fold) over explicit for-loops?
The summary says this (note the last remark):
> The results demonstrate that good programmers may be identified using supervised machine learning models, despite that no particular style groups could be attributed as a good style.
Obviously they're not saying, for example, that one particular bracing style is better than another (1TBS tho!), nor that co-incidentally "good" programmers prefer one style over another; but despite knowing what they're not saying, I don't know what they are saying with that...
Anyway, they looked at C++ code and measured things like the number of tabs, number of spaces, number of white lines, avg line length aso (table 2)
Most of which get standardized automatically by the IDE these days.
Randomly incoherent layouts, spacing and so forth are just signs of (a good programmer exploring and or a bad one who thinks it's finished)
But then there are other layers - too much going on, more than one thing happening, having to go back and work out what that call did. Just failing to follow a thread.
Then business layers and oh I give in. Determining "good programming" is like determining "good literature" or "truth". You know it when you see it. And you vomit on the bad stuff. But you have to read a lot of it.
This enters in the category of taste, I’m surprise that someone has investigate the assumption that could be correlated is something meaningful.
Although when comes to programming there is still different styles, compositions over inherence, condition verbalization, etc…
I wonder if that makes the difference. I believe makes better code, but some times I struggle explaining the value.
Literal full screen worth of newlines between thoughts. Indentation the reasoning behind only he knew - if there even were any.
I've never seen anything like it before or since.
I'd say otherwise he was a pretty average programmer.
Interesting job, that. We all had our own projects fully ours unless we went on vacation, it made weird formating choices like that largely not a problem.
Was insane trying to read the code.
This relates to a long-term, diverging trend in the practice of programming.. Vast numbers of highly skilled coders are paid money to meet requirements of a much larger system.. old-school is corporate but the cloud in general could be thought of this way..
Many strong programmers are on the opposite end -- dreaming of patterns and applying them diligently; much more like weaving or industrial arts.. not "paint by numbers" like the requirements work.
There are many, many gray areas in that analysis. Like that physics class demo of a growing balloon where all points expand and diverge at once.. the practice of programming will never be contained into one or more categories.. But to call out "good programmers" without recognizing the breadth and depth, is shallow and at worst counter-productive to the craft. This labeling gives fodder to the "Henry Ford"s out there looking to make other humans work as efficient machines themselves.. personally I can do without that.
- fanatical about manual placement of whitespace
- copy&paste&mutate large blocks of code all over the place
- checks in blocks of commented out code in case they're useful later
Identifying good programmers is much harder, there's a lot of variables in that.
That way, everyone would get to see the code the way they want, and no one would have to live with other people’s mostly arbitrary choices.
Basically, prettier, except you‘d have one (personal) setting for how things are displayed on your IDE, and one (shared) setting for how things actually are written to disk.
The main issue I see, is that it’d be introducing one more error source.
It's not clear whether this paper was accepted at the conference, but if it was, it is a testament to the low bar of the examining board. I like to follow the "if you don't have anything nice to say, don't say anything" mantra, but this paper deserves an exception.
Persuasive! I wonder if we could apply that insight to any other field of endeavor…
This is literally the goal of the study. Judge programmers based on the superficial details of competitive programming results.
I didn’t know! I’m part of everyone
Last time we did that, we got the Google/Leetcode interview that too many people are still cargo-culting. (Not realizing that it was originally what a professor thinks is important in industry software development, and then morphed even worse, into stereotypical university fraternity class filtering and hazing. Now you have other people, who should be scrappy street kid fighters, deciding who can join their gangs based on how good their posh accent, because they want to make money, and posh accents is what rich people seem to do.)
Academics should stay in their lane. And students should be very skeptical when academics swerve the school bus into oncoming traffic.
I think it's just a runaway process where some metric that's easy to measure was chosen, and it's become a worse metric that's more burdensome on candidates every year as it has become more and more gamed.
So, led by students, you were stuck with a student's idea of what software development is about. Which was heavily influenced by their classes thus far (and unaware that, traditionally, a new college grad is usually not much use until mentored).
Then, pretty soon, "tech" jobs became seen as a big-money career, which meant gatekeeping and smugness.