Linus Torvalds: 'I Do No Coding Any More'
linux.slashdot.org
linux.slashdot.org
This is high on my list of code craftsmanship points. It's very difficult to explain to young programmers who have never worked on an old code base how valuable this is when done well. In fact, often you hear complaints about how a code base "is crap", but more often than not I'd wager this is just a result of the context at the time not being known or appreciated. Frankly, all code we write is heavily governed by context we take for granted at the time, but is in precious short supply 1, 5, 10 years later. If you come back with a different use case later, the original code may very well be unsuited for that purpose. We can argue all day about good judgement and YAGNI, but at the end of the day no one can see all ends, therefore the best we can do is document clearly why we did what we did.
Taking the time to rebase, cleanup, and explain your changes in detail will pay huge dividends for any long-lived project. I've literally had successors write me a thank you note for a commit message from over a decade ago because I take this so seriously.
In my previous job, I received some feedback from my manager that some people complained about code quality of my work, I was dumbfounded, I have no problem with being criticised, but I at least wanted to be able to learn from it and adapt. I went back through every single PR I did at the company and found little to no comments and it made me feel personally attacked with very little evidence. I then went on my peers PRs and see what I could learn from them, most PRs included highly non-descriptive commit messages and descriptions often only included a link to a JIRA ticket which most of the time was only visible to their direct team members and not me. From that I took the position of being more demanding on not just code quality during reviews but also descriptions, I didn't manage to influence a single developer to improve documenting efforts. That was a major reason why I eventually left the company.
I continue to be a believer in taking your time to write good commit messages, I have yet to collect dividends from it however.
Writing good comments is far more important. Don't explain what the code does - that's what the code is for, explain the why and the background information in the code. Additional, explain what the code does at the function level or at the module level - at a much higher abstraction level basically than the line of the code.
No one ever looks at the commit messages when trying to figure out what the code does as it convolves history and state of the repository, having to contextually wrap your head around when and how this commit was added.
Inline documentation? Perfect.
Bisect leads to the commit, the commit message explains why the change was made. I take the explanation from 10+ years ago and determine if the reason still applies today.
Or, I bisect and there's a low quality commit from 10 years ago from a person long since departed that simply says "fixup" and I'm dead in the water.
(In case someone points out a comment is more suitable: comments are less likely to survive refactors, and people who write terrible commit messages also tend to be terrible commenters. Informative commit messages supplement good comments.)
Commit messages are far easier to parse and understand compared to a moderate size diff. By reading the commit message before looking at the diff, it really helps in terms of understanding the context of the diff and what to expect.
Commit messages provide intent, and should reference the original change request to see what the commit was intended to fix.
I would say the whole chain matters.
A good change request (Jira) with defects, and how to recreate the fault, as well as relevant requirements (by reference, not cut-n-paste)
Good comments on the code that is changing. Don't document the defect and how you fixed it, just comment what the code does (if the code is not perfectly clear),
Then a good commit message, referencing the change request, to tie it all together.
It all matters.
Commit messages aren't mutually exclusive to the inline documentation.
I am making the case that inline documentation is far more important than commit messages.
Git was made with the intention that commit messages look like whole emails, describing not only why the change was made but also the thought process behind it, why this particular solution was chosen instead of some other etc.
What you describe is a decent header, although most people would probably prefer "Add ability to" (not "Added ability to", this is a custom that goes way back before git).
Inline documentation, or comments, is something else entirely. That is a moving target that can describe intended usage, remote APIs, and hard to understand passages. Commit messages describes a particular changeset, at a fixed point in time.
Google it, and see that most people's ideas of a good commit message goes way beyond the header ("shortlog") line.
If you use markdown Syntax in your commit body, github and other tools will gladly render it correctly. You can grep for commit messages with "git log --grep <pattern>" and it will also search commit bodies as well. It's beautiful!
When additional detail is helpful about the motivation for or gotchas related to a commit, starting a new paragraph and elaborating there is totally fine.
But I agree with you, they aren't mutually exclusive to inline documentation. The inline documentation is more important when the explanation is relevant to understanding the code which results from the commit, whereas a clear commit message is more important when the information is to illuminate the reasons for (or history surrounding) the change itself.
This also follows what I've seen at a few companies since then. Would you say that's a general rule?
There was only used by one person who gave no context on why he used that format, so it's probably no coincidence in such an example that to me the extra structure seemed to add only opacity and no real value.
I can absolutely imagine that being different in a company that used this convention widely and built tooling around it. Doubly so when a lot of the staff is junior enough that it's helpful guidance for structuring thoughts around what the commit does.
Though, I'm not sure it's a helpful restriction for commits that clarify, refactor, or clean up code. One can certainly reference an issue tracker number as the {feature/bug} element that gives the necessary context, but the strictness spreads the information around more widely than is natural.
If someone was adopting that convention as a workaround for helping team members learn how to communicate, they're probably going to find that it's a very incomplete workaround.
When merging you can pretty easily see what ticket each commit was for. Seems superfluous to put the same information in every commit messages.
By contrast, referencing the ticket number (but not necessarily ticket title) is common practice in the PR message, and seems much more helpful.
Commit messages are about the change, whereas comments are about the state of the code.
You are assuming here that all programmers even "do" git. A lot don't. A whole lot of them. I worked with some companies that would have some horrible horrible version control system only and I just wanted to cry. Sure, git has its warts, but once you get it, it's a tool that can be _incredibly_ productive. I hate having good tools taken away.
I read them quite often. Whenever there is a regression detected between 2 software revisions that have been deployed somewhere the first thing I do is checking commits and commit messages for packages that have changed to find the more obvious reasons for a behavior difference. Checking code diffs at that point would be a lot more effort.
You're projecting. YOU never read them. I used to be the same way. Now I see other people reading them and I've started to myself. I follow other people's efforts and progress. It's a great way to learn not only the code but to communicate better.
Is there any data to support your or his claim as more representative of the population?
Do people not use git blame to check what they're changing before they change it? That's how I check the history around the changes I'm going to make before I make them.
And am frequently disappointed.
The comments I put in commits are things you'd be horrified to see littering your codebase (though I did it that way once upon a time too). They are the "whys" behind a change. Sometimes they have relatively little to do with individual lines of code, and don't need to be maintained like comment blocks should be. The temporal binding is absolutely intended.
Comments are often wrong, for the wrong reasons. People that change code, but don't update the comments to reflect the code change are evil. But then, people that put to much superfluous information in comments, that should just be left to the code, are also evil.
BTW, if your code implements a part of a standard, try putting the standard, and what subsection, your code addresses. Just the reference not the text. 802.11G for example (old fart) or TA55, or RFC 1536. Whatever. Then list the date of the document and the subsection.
I had to modify a Reed-Solomon implementation many years ago, and the original two authors based the implementation on a particular textbook. They gave the ISBN and then chapter/paragraph/table/illustration references on each block. I went and got a copy of the book, and it all lined up perfectly. Best documented code I have ever encountered.
Almost all "code quality" discussion in PRs is lightweight value judgements with no rigor or consistency. Just what random bikeshedding objection that person happened to think of, maybe they had a bad night's sleep? Who knows...
Remnants of an era where most software was low stakes and had low number of users. We now have many more "-ilities" we are judged on, that are measurable and that actually matter, so ignore the code quality trolls stuck in 2005.
Good example from a couple days ago: someone made a helper method for compressing something, with the signature
string Compress(string value)
It was actually doing gzip followed by base64. My comment was to rename it to something like Base64Compress as well as change the documentation comment to state it was doing both. I wouldn't want someone else to inadvertently call that in the future if they stumbled across it, not realizing it was returning base64 and either double-base64 encode it, unnecessarily have base64 when not needed, or even end up with a bigger string than input because of the 33% increase. In a future PR, misuse of that Compress() method would be impossible to spot unless you happened to remember what it was actually doing.Other times it'll be something like a constant 30s timeout hard-coded, and I like that to have comments like:
// This usually takes <1s, so 30s timeout for this is generous
or // 30s is max or (calling code) times out anyway
Why? Because if we need to change that, the comment helps let us know it's arbitrary and safe to modify or not. Without a comment (or good non-squashed commit message) someone in the future (maybr you) is doomed to waste time rediscovering a bug you already know about or researching if it's safe to change.If you're interviewing for a job that involves any sort of maintenance over a long-lived project, organizational attitudes towards code quality are a useful smell test to how miserable it might get. I bear the scars of a couple of wholesale refactors (that cowboy coders wanted no part in, to no ones surprise) to improve readability/maintainability. Merging those gigantic PRs and having everything work was satisfying; it almost made up for the hair pulling. Also satisfying was catching a surprising number of hidden bugs caused by sloppy programming (some logical errors are obvious when you give functions, arguments and variables good names)
I agree, and I've found the same adding tests to existing code. It often uncovers bugs that have been there for years, and there's a certain schadenfreude from doing that.
Sometimes it turns out that the bug doesn't really occur only because that function never happens to be called in a certain way. I've also run into cases where when I'm trying to figure out how to reproduce the bug for real that I uncover a another bug that is blocking it. Both of these are examples of brittle code: the right change or separate bug fix is going to expose this, and hope your QA is good enough to find it.
I've also been able to identify and fix a couple long-standing bugs this way: the kind where it's been observed in production a handful of times and in more than one deployed environment, but no one had ever figured out reliable reproduction steps. It's a great moment when suddenly you realize all this crazy behavior can be fully explained.
During a JS project refactor, I ran into a couple of those "This should be broken, why is it working?" headscratchers. In my case, 4 out of 5 times was that there was an earlier "quick-fix" further up the call-chain that papered over the bug for certain use-cases, but did not generalize well; I feel that that sort of bug is harder to catch using (unit) tests.
Sure it'll be inconsistent, and it'll lead to some changes that are idiosyncratic to the reviewer, while leaving other issues in place. But that's true of copy-editing as well, and I've never heard someone argue that it's not worth getting editorial feedback because that feedback might depend on the person.
Perhaps you had some experience with someone who cared about "code quality" in a way that rubbed you the wrong way, and they also designed a shitty system. Or perhaps your idea of what code quality means just differs from mine. I've spent enough time reading and fixing up left-behind cowboy code to spare a thought for the next engineer who'll have to come along and touch that system.
Style is in my opinion very arbitrary. It's good to be consistent, and for that reason I guess it's good to have rules that are enforced, but with the number of times I have to revisit my changes because apparently I'm violating some obscure style rule, and therefore the build server rejects my change (or sometimes it even rejects it because someone else's change violated something, which makes me wonder how that ever got through), it feels like it's a waste of time to be too strict about it. A bit of personal style isn't going to kill anyone.
Just because others don't use a tool for documenting code doesn't mean the tool shouldn't be used. I make sure the team I work with documents the what was done and why in their commit messsages. I also make sure that, when applicable, they describe the bug that was fixed and how it was fixed in the commit, or how performance was tested by applying a certain fix in the commit message.
Several times over the years, people have gone back to those commit messages by finding them via git blame and figuring out what was done months or years ago and preventing regressions by reading the commit messages.
> Never once did I have any indication that someone took their time to read descriptions or commit messages.
I do admit that I have been tempted to do things like:
git commit --allow-empty-message
for changes and see if anyone would notice.
Repo commit history is an artifact the team produces, as much as the code it contains.
If your team insists on 1 PR == 1 commit (or if you use an authoritarian tool that allows enforcing it like gh or gerrit) then be abundantly pedantic about making sure that only changes related to a single unit of work are being introduced in new PRs. I find that calling people out when they make sloppy changes starts to get them thinking that it would be really nice to just keep this typo fix in an isolated commit and not require a new PR for it. When I have seen that model be successful, it’s usually the case that the team agrees that it is more important to ask someone to break out an unrelated change than to merge a sloppy PR.
Yeah that's the scenario I'm thinking about. The trouble I have with this is two things:
- People inherently don't make the PRs as granular as a commit would be, because it puts load on other processes (like CI, review overhead, etc.) and slows things down that way. So you end up with PRs that are hundreds of lines instead of self-contained commits that are dozens of lines long (or less). They still fit under the same goal (and hence PR), but aren't as granular as they could/should be, which makes them harder to revert.
- It's far easier for people not to care about their commit messages and commit granularity than to care, and honestly, with CLI tooling, I can't claim it's easy. I use TortoiseGit and rebase all the time to keep my commits organized, and it's far easier for me to shuffle around commits that way, though it's still non-negligible overhead.
On the other hand, CI is precisely one argument I get for squashing: because rebasing/merging would imply some commits are untested, others specifically don't want them in the tree (they find it misleading when they bisect). I don't really find this compelling—I find it harmful to lose so much granularity information just because you didn't run a test against it (especially when it's already not the case that every commit passes all the tests) because it makes it impossible to look back and revert a tiny commit later, but I can't really convince others the trade-off is worth it. And to make this work well and keep a good history, you need to rebase to specifically avoid spurious commits like "fixed lint" or "removed whitespace", which generates additional overhead (and makes for harder reviewing when you rebase).
My position is it should be up to the author to decide whether a PR should be merged/squashed/rebased, because the answer can be different in different situations and the author would know best, but I have a hard time convincing anybody... people would rather have a uniform workflow instead.
git rebase --exec "test_command"
where test_command is one or more commands that you run to test the code for that commit. I typically do this for unit tests and integration tests.If the tests fail for some reason, you can amend the commit to fix the test and then run
git rebase --continue
to test the next commit.Rebase and push -- along with some squashing of redundant work-- is great. There's some degree of annoyance if intermediate states don't build, so it requires care to know that all the individual between states are probably safe (because CI, etc, generally don't help you here).
Merge commits are fine, though they do require a higher degree of acumen for people exploring the history.
And there's plenty of times when a developer submits a moderate sized 3 commit PR that whoever looks at the pull can decide that it's better to squash and make into one commit.
Of course circumstances vary and it depends on verbosity and idioms of your preferred language, framework and codebase. Maybe 150LOC is the minimum it takes to express a logical change unit.
My point is not about specific “maximum line count in a commit” but about a commit expressing a single logical change concept, and not a whole chapter of a book.
> making better individual commits is not going to help with understanding those … "jumbo" PRs of 1K-2K LOC.
This is where we disagree, it seems. You appear very categorical here, could you clarify why you think this is the case?
Also, I agree with aeontech that artificially mandating PRs to be squashed is very counter productive and often hinders bisect. Large features should be implemented in atomic units as a series of commits.
Also it gives me a chance to figure out what I actually worked on.
In the end before merging, we'll run
git fetch orgiin
git rebase -i --autosquash --keep-empty origin/master
git log -p --reverse origin/master..
git diff @{u}..
The first command will update the remote tracking branches. The second one will start the interactive rebase and order the commits based on the fixup! and squash! tags we used during the review. The third one will show the commit messages and associated diffs for the branch so that we can check that things make sense. The last one is to verify that we didn't inadvertently introduce a change during the rebase (though we will see a diff if someone else merged a PR into master before this rebase).
will verify that no diff was introduced during the rebase compared to the upstream branch as referenced by the PR.
Then, when doing the interactive rebase with the --autosquash and --keep-empty flags, you'll get to a point where your editor opens up with text saying that "this is a combination of X commits". At that point, you can delete the text down to the text of the last squash commit, exit the editor, and it will replace the commit message for the corresponding commit with the text that was in the squash commit message.
Using a collaboration tool to figure out what actually you worked on is placing this burden on other people instead of yourself.
When you spend a few days deep in a problem, it's very easy to explain the finer details of it while forgetting the big picture of what does this module even do, what real world problem were we fixing, etc, because that has become completely self evident to you.
It's a bit like giving a street address without saying what city it's in.
I have trained myself to look at both code and comments in a "what if I was a new hire 6 months from now" mode, and it's an indispensable tool for me.
Too many people think that commit message is to explain HOW you did something. Unless you did something extremely clever (and in any commercial project, 99% of clever solutions are wrong - simple is the king), it should need more than few words.
WHY you did it this way is the most important part. Given known and understood set of constrains, a lot of engineers will come up with similar solution, or at least understand it. And you can’t work effectively on something you don’t understand.
Not having this context is reason for so many rewrites. Developer thinks a solution is supper crappy, goes to rewrite it, with grand vision. Only to learn along the way about all the edge cases and reasons for hacks.
>WHY you did it this way is the most important part.
What was done is better suited for the commit message than how and why it was done, in my opinion. The latter two are better explained by comments in the code. I agree that why something was done is important for future reference, but when I'm tracking down code changes I'm looking for hints of the change itself. Understanding why comes next once you have the commented code in front of you to reference.
A "why" comment in a commit message, on the other hand, only tells you what the "why" was at that time, something that cannot become outdated.
No, but you seem to be implying that commit messages serve that purpose. If your code changes why wouldn't you add new comments and remove outdated/misleading comments?
>A "why" comment in a commit message, on the other hand, only tells you what the "why" was at that time, something that cannot become outdated.
What is version control for if not to retain state of the code as it was being written? Any popular IDE has simple interfaces to compare commits and see what was changed.
The point is that the code might not change. When you wrote it, maybe it was to cover use cases A and B. Sometime later, a new use case ”C” is added to the requirements, and this particular part of the code happens to already support C without making any changes.
Now, further down the line, some developer comes along to work on this code. If there’s a comment saying ”this code was written to support use cases A and B”, the developer will likely not take C into consideration when making their changes, and risk subtly breaking things.
>Sometime later, a new use case ”C” is added to the requirements, and this particular part of the code happens to already support C without making any changes.
Can you give an example of a new use case being introduced without the code changing? Why would this be documented in the code in the first place if nothing has changed?
> Can you give an example of a new use case being introduced without the code changing?
I once wrote a framework for processing data in different proprietary formats from different vendors into a unified format. At first, each vendor had its own code module, and any idiosyncrasies between different data files for a given vendor were solved within that module.
However, later on, we had to split one vendor into two modules. Now, some key parts of the code had to be changed to account for this, as it could no longer be assumed that one module would handle one vendor. But most parts of the code was already abstracted over the modules, and didn’t care.
> Why would this be documented in the code in the first place if nothing has changed?
That’s what I’m asking you! :)
As I've been saying comments in the code should describe how that specific section of code works, and why it's done that way if the "how" isn't enough. The new parts of the codebase would have comments describing how that code works. The unchanged code still functions as it always did. There's no reason to make comments on how another part of the codebase works. Code comments shouldn't be architectural descriptions of an entire codebase.
> The unchanged code still functions as it always did.
...the why of that unchanged code, i.e. the reason it exists at all, may change. And for that reason, the "why" (the business reasons) should not be in the code.
Totally agree. Commit metadata is nice to have, but ultimately I think good code mostly explains "how" by the code itself, and "why" should be comments. Especially for local context, where the why has to do with some other bit of code. Individual commit messages are also secondary to the pull request merge commit, which should have a better explanation of the goal, or tie it to a task, etc.
If found that when we starting using Github flow, with pull-requests and etc, commit messages stopped really matter, in favour of PR descriptions. Plus when you use "squash" strategy, and use PR description as commit message, history looks great.
When you look at commit history through Github tooling, all PR ids turned into links, and it is very easy watch for the history. Additionally you have a good integration with Github features, like discussions (which sometimes as important as description itself).
Yes it is a vendor lock-in, but lets be frank, Github most likely will survive your project.
The thing is this isn't obvious why it's useful until you don't have it. I don't do it often, but the things that make me go back looking through git blame and history are trying to identify where hardcoded constants came from, or when (time or version) a particular bug was first introduced and why. This is only necessary when there aren't decent comments in the code already, and in my experience, the types of developers that don't put in these comments are also the ones that don't make highly detailed PR descriptions. Good commit messages is also a hurdle, but an easier one to get over than highly-detailed PR descriptions.
On top of that you get all the other significant downsides of a squash merge strategy (eg: not being able to tell the difference between a local branch being merged or not pushed; confusing merge conflicts if you ever merge branch->branch before merging one to master).
As an aside, IMHO, "clean history" is not at all a useful feature. Only developers are looking at git log to begin with, and they are quite capable of using `git log --merges` (or equivalent GUI operation). I guess the most useful thing is to make a CHANGES file (release notes) -- but honestly, the wording is different anyway (eg, "Fixed several minor UI issues" is adequate for release notes, but that might be across several PRs that contain more detail like "Fixed button alignment in modal dialogs", "Upgraded bootstrap to version x.y.z" etc). Go back and look at your PR descriptions from a year ago and see if you can make sense of them -- I bet the context of many will be lost.
I put docs into the source tree (sometimes in the form of comments; sometimes in dedicated doc files) for things that are suited to live next to the code, but often my commit messages contain more discussion of what used to be and why I chose a certain implementation approach.
I generally think that documentation in code should describe what the code does and why, and commit messages should describe why critical choices were made and how the new approach differs from previous behavior.
This assumes a good version control system that follows history across renames and moves etc., of course.
(Sometimes absurdly so, as in the case of gcc, where flamewars have been fought to preserve commits that were broken to begin with.)
That is important, and that much care for the code probably helped them survive for so long. Botched VCS migrations is not something that should be accepted anywhere.
I'm a big fan of literate programming, and VCS is not an integral part of that.
In my niche of embedded work, the "host" machine stays the main interface to work with code almost always (and in the exceptions the "embedded" system is typically PC-grade hardware running Linux)
With care, it's possible even to merge different git repos into a single repo while preserving history. But it takes care; it's not a pit of success.
A lot of features take more than a single logical commit to implement, so it makes sense to have multiple logical commits associated with a feature update. Squashing them down to a single commit makes it harder to review in my opinion.
Also, this is all predicated on the business case for the software being well understood that it could be written out beforehand by a entry level engineer. Which is not a situation I've seen outside of a processor design house. Otherwise, your thinking for this or that choice will be built on shifting sand anyway.
A lot of companies run code in production that was written years if not decades ago. Being able to look at the commit history and blame output helps in terms of getting some context about what was done and why it was done they way it was when the original authors have long since left the company. In my company, we still have some code written back in 2005 still running in production.
Predetermining code longevity is challenging.
Translation: "Not written in the particular arbitrary style that I'm used to", of course if it were written in that persons preferred style another arbitrarily selected novice would call the result crap.
I've found this to be almost universally true especially when the claim is given forcefully but not immediately backed up with a litany of actual user-impacting serious problems.
This is not to say that code can't be crap, but generally people with the wisdom needed to make that determination know better than to waste their time yelling about it and know enough to suggest how to make it better instead. They recognize that good code can be written in a multitude of styles and that there is very little objective evidence for most style choices vs most others. Or, at least, they know enough to limit their forceful remarks to the most objective considerations-- such as code which is actually a source of frequent problems.
> Taking the time to rebase, cleanup, and explain your changes in detail will pay huge dividends for any long-lived project. I've literally had successors write me a thank you note for a commit message from over a decade ago because I take this so seriously.
Likewise, though more frequently I've been frustrated to see people make grave errors "fixing" correct code that they would have understood if they bothered to look at the message on the commit which introduced the "error" (or read the adjacent comment).
The culture of writing good, informative commit messages and comments needs to go hand in hand with a culture of reading them. Neither works in isolation.
There's certain code smells that most experienced developers are going to identify that would be almost universally considered bad. Things like poor variable/method/class names, use of magic values, duplicated logic, high coupling, and lack of unit tests.
None of these have immediate user-impacting problems, and arguing that's the only thing that matters is very short-sighted.
All of these issues will slowly degrade a codebase over the long term -- as the requirements change, new developers start working on the code, and bugs compound other bugs -- and the user impact will be slower delivery and lower quality. Eventually the code will get to a point where the only save is a major rewrite which causes even slower delivery and higher risk of regression in both features and bugs.
Writing good variable names at the time the code is created is significantly easier than fixing it months later. Writing the code in a way that can be unit tested is easier up front (and sometimes impossible later, without a rewrite and all the dangers that go with that).
There's also a big difference between a code base that's maintained by the same focused person (or small group) for a long time, vs being one of several code bases maintained by a team of people (who may change over time). In the former you can get away with things you can't in the latter case, and moderately or highly successful small projects tend to shift to the latter, and are the ones people will call crap code bases despite their commercial or user-facing success.
Like "arbitrary styles" such as "a function shall not be 500 lines long".
http://number-none.com/blow/john_carmack_on_inlined_code.htm...
Previous HN discussions about that:
https://news.ycombinator.com/item?id=12120752
At what age am I allowed to become a senior, oh wise one?
As I gain experience, I find that I understand in more details existing codebases. This increased my chance of a successful refactoring. Especially when replacing a core abstraction without a full rewrite.
FWIW, academia have never been good at teaching these skills, so the answer to your question is "You'll stop being junior with regards to refactoring complex code bases whenever you have enough experience of refactoring in big old code bases". Mentors pushing you in the right direction can also help. But this is very much a skill acquired by actually doing.
This has happened with SVN and also is happening with Git.
Unfortunately you can't argue with those ones since are in lead position and there's no one with technical skills above them in the C-level who can say anything (we don't have a CTO or a CIO, just CFO and CEO who can't even use their own iPad)
I'm not sure I understand what they're doing -- they overwrite some of the changes in their local working tree, and push to the repo, without explaining what they did/what they changed or why?
Thanks for explaining
When you do have those, then it's the ticket that gives the why's and how's. The commit messages do not have to be essays. And, at least with JIRA, if you name your branch such that it starts with the ticket id, it automatically picks up any pushed commits for that ticket. It's very useful.
I've been working at my current place of employment long enough to see them transition between several different issue trackers. The links to old issue tracker tickets no longer resolve, so it's definitely better to include a description of the issue directly in the commit message in addition to the link to the issue tracker.
Just look at the "imperative" that almost everyone recommends, try for enough times to cherry-pick commits that claim to do something and you'll realize how misleading they are
(Now if I could just get them to squash, rebase and rewrite commit messages before publishing...)
This also goes for certain programmers that have been in the industry for decades. And if you deal with one of those, you'll find that they are even less likely to change their practices, so I'll take an impressionable young programmer to convince any day of the week.
Of course, this topic extends to many other practices, such as using descriptive variable names, writing good comments and documentation (instead of the Captain Obvious approved ones), designing good APIs,…
On the discoverability: Do you use `git blame` much? Because in projects that live a long time, that will most likely be the way you'll discover commit messages the most. And it is very handy in one of those WTF moments to see why and how a particular piece of code was altered, as opposed to "update" or "improve logic".
Also: On your notion of "abstract what": That's supposed to be your commit's subject line. The rest of the message then will be the "why". I like to think of it as an email to some future maintainer (e.g. me in a year) on why it was necessary to change a piece of code.
Edit: It's also important to see the differences workflow between your project and Linux, which Git was originally built for. In Linux, your change most likely is an email to some maintainer. If it sucks, your code won't be merged.
And there's something to be said about being able to vocalize your reason for changing code. I still get the impression that many people who riff on the "return on investment" part of commit messages don't factor in that you're very likely not understanding a problem if you cannot describe it readably in text.
There's likely a threshold (a 10000 hours type thing) where you've researched enough code history to value great commit messages and descriptions and docs/readmes, issue descriptions and comments etc.
But how do we impart the importance of this to the other devs around us that will make it stick? Past "war" stories?
I seem to be the only one in my team doing this, though.
(Obviously I also include a short but meaningful description of the change in the commit message itself. Again, not something everybody always does.)
The only real argument for squash-merges is that history remains linear and visually "clean", but if you want to see a "clean" history of only the merge commits, you just can just use the `--merges` option on the command line.
Squash locally if you really want to or need to, but I'm a firm believer of disallowing it as the merge method for a changeset.
for bob's old ERP software company, really won't make a difference
I'm closing in on forty years of coding and still get a kick out of it. I'd still be doing it if I weren't getting paid for it. But my skills are at best, uh, modest compared to Linus. I'd think that doing a thing so much better would only be that much more fun. But this is a data point against that correlation.
Metaphorically, he used to generate bricks himself, then add them to his castle. Now people bring him bricks and he choses the ones he likes and which to discard. So at the end of the day, perhaps this is still a castle of his design. meh?
I have recently adopted a large, sprawling side-project of which I am the sole developer. Having a big space to play in on the side has really helped me keep my interest pure and exciting. Not answering to anyone but myself is easily the best aspect. The codebase is private and no one but me may ever see it.
I am also a lot more pragmatic about the codebase at work, as I am not trying to squeeze in fun/innovative ideas at the expense of productivity or the experience of other teammates. Now, most of my "blue sky" ideas wind up being things for my side project. Feels like I have a pressure relief valve that I can rely on now. There have already been a few cases where I implement a crazy idea on my side project and determine the abstraction is also an excellent fit for something at work. Being able to iterate on these ideas in a completely separate domain from work concerns allows for deep (and enjoyable) refinement without taking the rest of the enterprise for a trip every time.
People get so hung up on how cool the thing they just did is and not at all on was useful. Sometimes you just gotta let it rest for a month to get over that honeymoon phase.
Tool flame wars, constant bike shedding, “that’s too abstract” premature optimization wars (apparently a list comprehension is too abstract for some eng, if-else all the way down!). Every Medium post by a tech somebody becomes a religious soapbox.
I quit it all but kept my job. People complain I skip meetings and take Friday mostly off.
But my stories are done by Weds, so really I take Thur off and just do work log updates Fri to make it look like I was still working.
And the reason they’re done by Weds is cause I do the programming and skip everything else.
Job life feels like being in high school anymore. It’s about social appeasement, debating well trodden concepts at length, and exploring little new ground.
Import toolkits and go home. Like most creative endeavors, writing code to express my own mind is fun. Writing code to satisfy the pretense of our job worshipping culture is mind numbing.
Edit: Oh almost forgot my “favorite”: environments where open source is leveraged heavily. Millions of lines of imported code, at best 5% was written internally. But somehow the business environment is that they’re engineering thought leaders for building a “tech biz” around a traditional business model. Smh
This is definitely a serious issue. Imo, unless the project is tiny, your own codebase should be at least the majority of whatever binary is shipped, otherwise you're adding very little value.
Maybe somebody is exploring an idea and doesn't really care about maintainablity or something else.
So I don't think he said anything of the sorts that he is not having fun programming. And in fact when the interviewer asks a question in the similar vein, Linus is quick to point out that he would be really bored if he wasn't the maintainer.
Edit: All I am trying to say is I wouldn't speculate on what Linus feels and doesn't feel when he can speak for himself (and which he does in the video). At the same time the general discussion in this subthread is very interesting.
You can, however, use those same muscles and skills via other mediums and activities. It’s different which is different.
He still deals with plenty of code; he just doesn't write original code. Calling it a "managerial or administrative type role" as the previous comment did is somewhat lacking in nuance, IMO.
By now it's very hard to sustain anything remotely resembling motivation, especially having been around long enough to have been around the block with a few technology waves. This article, circulated on HN a while back, resonated with me more strongly than anything I've seen here in years:
https://frankchimero.com/blog/2018/everything-easy/
I'm completely burned out on it all, and I don't even do more than a few hours of coding a week many weeks. Even that takes some serious psychological effort to get into, and some weeks I don't manage it. Honestly, I'd much rather sit in meetings, answer e-mail, and advise others on architecture. I like talking and teaching, and people issues interest me fairly indefatigably. It very well might pay better, too, depending. Alas, I picked a business model that economically rewards me for coding--that other stuff is just non-revenue overhead.
I really hope a day comes soon when I can say, "I don't really get involved in code much anymore," before RSI and disinterest foreclose upon the possibility of a life worth living.
Don't get me wrong, I come from the same background you do and have similar wishes, but isn't this just "I want to do the things I like and not the things I don't" (which is true for anyone anywhere in any job)? Any job is going to have unliked things (that's why it's w.o.r.k and not f.u.n).
I tried to write a little bit every day, cause it is something that gives me joy, but I never have a project taking off like those guys, so I don't know how much free time they have to start new projects or to program just for fun.
My guess is Linus is still L&M (if people didn't enjoy his lean, they wouldn't appreciate his mean), but perhaps he's not in his element right now (and just taking a breather).
Either way, I'm a fan (not a professional) of a lean (looking) codebase with perhaps a new "overlay" commenting system being developed to accommodate a wide range of proficiency (one size fits all is either too terse or too bloated).
I am not sure of the difficulty of linking comments by line number between revisions, but outside of having an official repository of comments, being able to toggle through multiple perspectives from the same line/set of code would offer valuable insight and chronicle mindset/POV in a way that a global scratchpad could never deliver.
Writing terse or bloated comments should not be top of mind, but rather a frictionless byproduct (expect terse, but be pleasantly surprised on occasion by the detail and multiple perspectives).
The overlay thought was similar to the map concept, where you can load and toggle data layers (comments of various verbosity or from a dynamic source) easily.
This is funny to me, because I feel like Linux got ahead over BSD specifically because it had a more accepting policy in the beginning.
I think this says something important about project maturity and lifecycle. BSD was already a "well established" system by the time the Linux kernel appeared, and as such had more gatekeeping to meet their standards.
https://en.m.wikipedia.org/wiki/UNIX_System_Laboratories,_In....
And Wow, Slashdots back button disabling is the strongest I’ve ever seen, felt like I was in the Deathstar’s tractor beam trying to get back to HN.
Why commit messages rather than code comments?
That's simply not true, not in a git world where rebasing to rewrite history is standard practice.
Suppose you're making a change that has a few paragraphs of commit message explaining what you're doing and why. As well as a large well-commented change in a primary file, you've made some small incidental changes to other files. You're not going to include comments with all the context and explanation on each of these small changes.
If I'm looking at one of the files you incidentally touched later and don't understand why your change is there, I can check the git blame and have a good understanding of why the code looks like that now.
Now, why a particular piece of code came to be, why the code should have this expected behavior, why something is being added or removed or changed, may be unrelated (or unnecessary) to what this particular code is doing.
So there's two kinds of documentation here: 1) "this is what this code does", 2) "here are the series of events that led to this code". You need the former when you are editing code. You need the latter when you are reviewing code for merge, as well as in the future to figure out how the f*!$ this code came to be (which is useful to avoid re-running into a bug, troubleshooting/understanding a legacy system, or removing important code). The latter may also immediately become obsolete the second time the same block of code is edited, so we keep this historic information in e-mail or code commit messages.
1. If one needs to understand the history of the code to understand the code in its current state, the code is bad.
2. But if one is one is already familiar with the old code, it is easier to leverage that to understand the new code than start from scratch.
So we write commit messages for 2, but per 1 we can't have those commit messages be the only way to understand it, nor can we have the code comments refer to the old versions of the code that are removed[1].
[1]: except for the vaguest "other xxx you might think of were tried and didn't work".
it's seems to be more of an excercise into pleasing the priesthood to their, current, very important way/shape/form "canon" of what IS right at the moment than actually get any 'coding' done. And given the time it takes, i wonder if ANY of the top level maintainers actually program anymore, and are just "patch gateways".
And yes, there is a massive caste of priesthood of the patch series, discussing and making decisions on stuff they often borderline understand, and on criteria that are in fact quite far from the "it make sense" line of thinking.
I know, I've spend months re-writing drivers which were perfectly fine, neat, written with the upmost care of "upstreaming" as our team could figure out, to have to rewrite them completely and entirelly in the whim of someone halfway across the world, who didn't couldn't possibly have known what we were working on (as we were the designers, mostly) but who decided that version of the "canon" was not matching ours.
It's a stupid waste of time, money, brain cells, and the result is actually lesser than the original code.
For no good reason as the canon will change next year anyway.
For many years, I thought that the Linux kernel ecosystem was pretty much self-sufficient, but it's grown way too big, to the level of MASSIVE bloatware level nowadays. Massive corps have dozens of guys working on it, and many of them are also 'maintainers' -- full time; threading patches, some of them with conflicts of interest, or lack of understanding, or both, or just having been left behind a while back while still being innundated with patches.
I think it's doomed, in that form anyway.
Odd that there doesn't seem to be enough time to read beyond 140 characters, yet there's plenty of time to type up lengthy opinions...
...don't leave us hanging.
Yes he does so write code....
https://news.ycombinator.com/newsguidelines.html
(added recently: https://news.ycombinator.com/item?id=23648867)