Reflections on 10k Hours of Programming
matt-rickard.com
matt-rickard.com
Here's some more color on the reflection on comments, straight from the Linux Kernel documentation that capture my intention much better:
"Comments are good, but there is also a danger of over-commenting. NEVER try to explain HOW your code works in a comment: it’s much better to write the code so that the working is obvious, and it’s a waste of time to explain badly written code.
Generally, you want your comments to tell WHAT your code does, not HOW. Also, try to avoid putting comments inside a function body: if the function is so complex that you need to separately comment parts of it, you should probably go back to chapter 6 for a while. You can make small comments to note or warn about something particularly clever (or ugly), but try to avoid excess. Instead, put the comments at the head of the function, telling people what it does, and possibly WHY it does it."[0]
[0] https://www.kernel.org/doc/html/v4.10/process/coding-style.h...
Never is a strong word, sometimes the algorithm is just inherently complex and it's worth it to explain what you are doing.
Why is much more important, but it's also partially solved automatically by git blame and adding the task number to each commit message (and writing good commit messages which IMHO is more important than writing good comments). Each line of your code has commit message associated with it whether you care about it or not - make it useful.
> Know when to break the rules.
I think that we fundamentally start to lose the mental distinction between the language semantics and the idiomatic code. It’s all part of the scenery. So for a seasoned programmer, the what and the how blur together. There is no how to getting a value from an object, a struct, a tuple, or an array. You just do. Unless your language doesn’t have one of those things.
Non-idiomatic code is full of how. So what we seek is code that just tells us “what” and comments are a bargain we make with each other to guarantee a lower bound on that. Meanwhile the self documenting code people want a different arrangement, where you have to stoop to documentation less often by putting that energy into sustaining a constant output of idiomatic code.
It's actually funny how all the focus in CS seems to be on modeling the business part in more natural ways, while the "solving problems" part is still done like it's 1970s.
You can make the algorithms clearer with good variable names and extracting some parts to well named methods (and you should, usually) but it's often not enough.
Example taken straight from Wikipedia:
function BellmanFord(list vertices, list edges, vertex source) is
// This implementation takes in a graph, represented as
// lists of vertices (represented as integers [0..n-1]) and edges,
// and fills two arrays (distance and predecessor) holding
// the shortest path from the source to each vertex
distance := list of size n
predecessor := list of size n
// Step 1: initialize graph
for each vertex v in vertices do
distance[v] := inf // Initialize the distance to all vertices to infinity
predecessor[v] := null // And having a null predecessor
distance[source] := 0 // The distance from the source to itself is, of course, zero
// Step 2: relax edges repeatedly
repeat |V|−1 times:
for each edge (u, v) with weight w in edges do
if distance[u] + w < distance[v] then
distance[v] := distance[u] + w
predecessor[v] := u
// Step 3: check for negative-weight cycles
for each edge (u, v) with weight w in edges do
if distance[u] + w < distance[v] then
error "Graph contains a negative-weight cycle"
return distance, predecessor
This is pseudocode with some stuff abstracted away, yet they still added "what" comments. You can remove some of them and split it into several methods but for Step 2 I still kinda feel it's not enough information to make it obvious what is happening. What does it mean to "relax an edge"? Why do we need to do it V-1 times? Does order matter?I don't think fancy programming language features help here. You can change iteration into tail recursion or point-free functional code or whatever is fashionable but the underlying complexity remains.
And there are algorithms out there that are much trickier.
Source: https://en.wikipedia.org/wiki/Bellman%E2%80%93Ford_algorithm...
I would argue that the only comment really useful here would be one pointing to the paper/wikipedia article/whatever that explains the algorithm.
If no such article/reference exists and/or your implementation has nuances you want to explain, that's where technical documentation should come in. For instance, check the linux kernel's explanation of circular buffers and how to use them [1].
[1] https://github.com/torvalds/linux/blob/master/Documentation/...
> What does it mean to "relax an edge"? Why do we need to do it V-1 times? Does order matter?
These are questions related to understanding the algorithm, not the implementation. There are much better ways to convey that information (including examples, diagrams, lengthy explanations, etc.). Trying to convey that information through comments in an implementation is an exercise in __not__ using the best tool for the job ;)
I would chalk the complexity of the example you provided up to the inability of half of developers to handle four nested conditionals, and for most of the rest to handle five.
I sort of handwaved over the Death By A Thousand Cuts scenario, which I definitely believe is important. It doesn't matter how idiomatic your code is if you have enough of it. I don't maintain that idiomatic code is free by any means. It just doesn't have the factorial surface area of spaghetti code. It scales better, but scale absolutely matters.
If you count conditional branches, you'll see that parts of this code are pushing up against those limits. You're definitely into 'how' territory, but you've also shunted it off to a function that just 'does stuff'. So my code that relies on this for step 1 of a bigger answer doesn't give a shit (as long as the answer is correct, including boundary conditions).
> NEVER try to explain HOW your code works in a COMMENT
Comments are for explaining why, not how. If some code is algorithmically complex enough to need explanation, that should go in documentation, not an inline comment.
I've never seen or heard of a team that keeps documentation for their complex algorithms. It's just not done.
And if it was, I'd bet that they get out of date very fast
People love to hate on Bezos (less here than elsewhere, though, I guess) but this is exactly the tack he used to convert Amazon's IT setup to services, and I think both were shrewd moves.
Explain *why* you’re doing something a certain way. This also includes a little bit of how and what. Other than that, just make sure your business terms are reflected in the code, or at least have a proper definition of things, which you can easily reason about.
Once you start writing tests you find that the size of test functions will reflect the size of the function you're testing. The test functions start to have huge set-ups in order to test a small bit of logic, and you end up spending a lot of time fixing tests every time you make a change to that function.
This is why it's often better to find ways to make functions smaller.
There is a difference however of people refactoring code into smaller classes / functions vs what I'd call "hiding" code in classes / functions. If all someone has done is broken a large function into a chain of functions then of course that is not good.
When people refer to small functions / classes they refer to breaking up the concepts that a class / function represents into smaller concepts that build upon one another. This also increases the reusability of code.
So in short - I agree, longer pieces of code can be better for reasoning about, but I disagree that that necessarily makes them better.
Edit: He also refused to ever delete any code he'd written "because it might be useful later." I once spent three hours tracing the source of some values, only to discover that the output of the entire call tree I was working my way through was being quietly thrown away at the end.
(There's also a video where I explain the reasoning behind the mock-up, but I guess the motivation is already blatantly obvious to most readers of the above discussion, so feel free to skip the first chapters: https://youtu.be/wdf6S6oumH8 )
Then I learned that the Smalltalk editor Pharo has a similar feature! Seems many people have the same idea. Today's major IDE makers should have copied this feature already!
However, when writing the code if a section of code is repeated more than 2 times, it should be considered to break it out into a separate function with a useful name.
My point is just that you shouldn’t normally have to keep both sides of a call site in mind at once.
The WHY often will be elaborate enough that it deserves a design document that includes things like: who is the owner of those requisites, what is the rationale behind them, who are the stakehilders who signed off, what interactions it has other components, what alternative designs were considered, etc.
Comments should explain the decision process (if one was necessary) that lead to this variant. Because if there's no comment explaining it, the next developer looking at it will say "oh, but it's much simpler if I just do it that way", and will run into trouble because he didn't realize the pitfalls.
I only ask because of bits like the TCP rate estimator sometimes have step by step comments that might be interpreted as explaining how the code is doing something.
For example take the app limited detection function: https://github.com/torvalds/linux/blob/master/net/ipv4/tcp_r...
Otherwise, clearer code, smaller methods and better variable/method names are the way to go imo.
This makes me sad. Think about others who will have to reverse engineer your code. Not every functionality should be split into a function, unless you want to add 10 layers of abstraction and misdirection.
The only harm is, when you miss the important bits over too much blabla, or when you forget to update comments. There are not many worse things, than wrong documentation/comments.
"you want your comments to tell WHAT your code does, not HOW" but I agree very much to this.
Otherwise, there are many techniques you can use to make the code self-explanatory. I've taken a comment before and restructured the code to literally read like the comment.
That also extends to how. If the algorithm is interesting or complex enough you should move it as well, as the point of the host function is not to show how clever you are, but to accomplish a task.
Empirically I’ve found that I get much less pushback on algorithmic tweaks when I isolate them to a separate function. The reader doesn’t have to include them in their local reasoning, and I suspect but cannot prove that they tell themselves that they can always revert that code cleanly, so let’s just leave it be for now and see what happens.
I enjoyed reading this post. Cheers!
I mostly managed to dodge bans on my most used accounts at the time, but interestingly when I checked back a year or two ago the majority had been permanently suspended. I wonder if they have some method of retroactively determining if someone was botting?
But "why the fuck" is a daily occurrence.
Self documenting code explains how and what. If you need a comment to explain how and what its a code smell. Sometimes it's ok but ideally it should be rare.
I try not to overcomment but when I write code that looks confusing or clever I try to explain why I've done it that way.
If I'm optimising, I'll explain why.
If I'm making an assumption about something related but elsehwere I'll explain that.
If I'm leveraging some context that isn't explicit, I'll explain that.
...
Also, type of code matters, glue code rarely needs exhaustive comments.
Library code often gets line by line commentary so when someone comes in to change something for their use case they have a clear understanding of what they're really changing, and who and what might break because of their change.
I notice library code tests often get more comments explaining what and why I'm adding an assertion.
Glue code test rarely needs comments in tests.
Like many arguments, disagreements below are mostly not disagreements about facts but disagreements of definitions at heart, only masquerading as disagreements about facts.
Beyond a trivializing definition of "whatever follows a #", each time the word "comments" appears in the discussion below, it seems to actually refer to something entirely different from one person to the next.
IMHO The only way to tell if you are able to comment well is to do this: Work on some project - then in 2 years come back to that project and add a feature to it. You will learn real quick how good your comments are. Spoiler alert: You wont care about "how" - you will want to know "why".
And also Why.
I personally like an executive summary of the code, ie, I just got to this file, what am-I looking at.
Sometimes you just need a comment. For example, I was writing some code to an API and the company was offering a private experimental feature. The API was JSON, but this one feature was enabled by adding a small stringified JSON object to a field of the larger JSON post body. If a developer looks at it, they'll think "this must be wrong" but with a comment that briefly mentions why, everyone will be saved some time.
https://github.com/dlang/dmd/blob/master/src/dmd/cparse.d#L5...
Ever since the Intel CPU spec went online, I started doing this with the code generator - providing links to the man page for the instruction being used:
https://github.com/dlang/dmd/blob/master/src/dmd/backend/cgx...
And for bug fixes, reference the issue which often gives a detailed explanation for why the code is a certain way:
https://github.com/dlang/dmd/blob/master/src/dmd/expressions...
Ever since I enhanced the editor I use to open the browser on links, this sort of thing has proven to be very, very handy.
Even so, the utility of the links far outweighs the possibility of needing to adjust them.
P.S. I've found using links to Microsoft documentation to be extremely frustrating, because not only does Microsoft move them around all the time, but will wholesale permanently delete vast swaths of reference material.
Notice that the github links are to line numbers. Line numbers are ephemeral, I don't know a way around that.
If you go to your link, under the button with 3 dots there's a "copy permalink option" that creates a link to the specific commit you're looking at, ensuring the line number doesn't get out of date. Example:
https://github.com/dlang/dmd/blob/9003e2ae7e11faed433638e9cb...
Also, how do you notify a later viewer that there's a "why" they should check? A comment that says "check commit messages for this line"? :)
Not to mention: you're losing the notes as soon as you move repositories, or worse, version control systems. Yes, I've hit the "SVN git migrate" wall way too many times. For that reason I even started leaving issue numbers in comments for when it's particularly important, in case we lose the commit<->ticket link down the line.
I don't notify anyone, nor do I need to. When I encounter a WTF moment, I look at the commit messages. When there's nothing there, I curse the last developer.
I guess it can be fun spelunking through commit messages, though, to see when the comment was added and if it still applies.
What is important is to clearly communicate the intent of code to the next developer (incl., yourself) who comes across it. Comments are one tool provided to do this. I write comments. I approve pull requests containing comments. But comments are not the only tool, and they're not the best one.
Now most of the time there's no particular difficulty establishing the intent of code. Sometimes, though, there's something weird, something that looks "wrong" but isn't. In those cases, I isolate the weird change using abstractions (at least a variable, but probably a function, a method, a class); I name it and any components well -- long, descriptive, using appropriate conventions; I put automated tests around it, which will fail if anyone changes it without understanding it, providing helpful and explicit error messages (where tools allow); and I write a developed, full explanation in a commit message (which IME will not be too hard to track down, if you've isolated the "weird" change using an abstraction).
That will almost always communicate the the intent. If it doesn't, I (reluctantly) comment it up.
For many other things though, it's so easy to "cover up" commit history. Changing how you indent things (or splay one line into multiple), correcting a spelling error, etc all make it dramatically harder to follow than an in-code comment, even with good tool support (which is usually mediocre at best). Of course, you can keep those changes separate, and add them to an "ignore these commits" list for git... but few teams are capable of maintaining that in the long run.
(I assume/hope you just got downvoted in a burst of emotional bikeshedding. certainly doesn't seem negative-reputation-worthy to me)
Though I don't think your example is an exception. I would refactor that part out to a seaprate function, and give it a docstring like "json in json, because the api is weird, [link to api docs]." Basically, weird things like that should be refactored into their own spot, instead of called out with a comment in the middle of something else.
/* DO NOT DELETE the above test
If you're on localhost and opening/closing sockets
at high speed, you might end up being connected
to yourself by accident. That's what you get
when everything is a file and all files are
represented with integers
*/
For example.I find especially in learning projects, where I might write some quite dense code, it helps to describe what each line does and answer all the "idiot questions", the ones that you can answer a few minutes later or perhaps hours later with a face palm, when everything is clear again, that arise, when you have to cold boot your brain to pick up the project again or use the code in a real project later on.
- setting up a type (with a comment linking to that API's documentation)
- setting up a shape validator & a test that asserts that that that field is a string with the correct shape
- Using a variable name that makes it clear that something intentional is happening: `const weirdlyStringifiedJSONObject = JSON.stringify(...)`
A comment might still be the simplest thing to add, but it doesn't mean it's the only way to document that weirdness. Comments can be fragile compared to the same documentation expressed in code
[edit: list formatting]
Glad to see DRY called out here. I've seen so much crazy code simply to avoid breaking, The Rule.
Is the magic number 1024 the same in all instances, for example. Or is it merely a coincidence that it's the same in a few places. Is the apparently repetitious code:
context = create_context();
context.action(params);
In several places really the same? Or is it a coincidence that right now none of them pass any parameters to create_context and use the same params in their call to action?Avoiding the repetition makes it hard to even ask that question. And disentangling the different cases later is riskier than keeping a few pieces of repeated code around for a while to understand the situation better.
> The DRY principle is stated as "Every piece of knowledge must have a single, unambiguous, authoritative representation within a system".
I don't avoid copy-pasting as a rule, but would deduplicate a block of code if they represent different occurrences of the same "knowledge" of a shared procedure, but create separate functions or duplicate code if each one encodes "knowledge" of a different process to be followed. Of course it's hard to define and codify "knowledge" in an objective form that different people can use the same way, but in my experience, I still feel this rule (and the rest of my gut feeling for refactoring) hasn't provided bad guidance in the situations I've encountered so far (though it's not always applicable, takes domain understanding and experience to identify "knowledge", and it's sometimes difficult nonetheless).
GP's example of the constant "1024" present in multiple places doesn't necessarily mean they represent the same knowledge; one could be a bit mask, another could be the default number of items to show on a page, and another could be a buffer size. DRY-ing these different pieces of knowledge could be disastrous if the requirements for one of the uses changes. New requirement: only display 10 items on a page. Good luck with your bit mask reading the second and fourth bits now and filling your 10-byte buffer 100 times more often. New requirement: We added more flags, the bit mask is now 0x80000000. Good luck serving pages with 2 billion items and your memory usage randomly spiking by 2GB.
Our industry is full of dogma and false prophets. It’s unbelievably frustrating.
Interestingly, I just looked at the wikipedia page for DRY (https://en.wikipedia.org/wiki/Don%27t_repeat_yourself) and see they also mention another acronym - AHA (avoid hasty abstractions); I might use that in the future.
DRY originated from The Pragmatic Programmer, and is not about avoiding duplicated code. In fact, quite a bit of duplicated code doesn't break the DRY rule as originally defined.
If I make my own list, I'll be sure to put "Understand what DRY is before complaining about it." :-)
I disagree with this one. It can be bad, but if it helps syntax get out of the way to reveal the intent of your code, I think it's pretty good
Syntactic sugar complicates code for me, it may help to reveal the intent what the code is doing but I *have to* know how the code is doing it. Of course you can just learn what the compiler will do in every of those instances but that's a lot to learn and it increases exponentially. On the other hand in C you have assignments, conditionals, loops and function calls - nothing else.
In the end the code will do the same thing but I'm willing to spent few more keystrokes and maybe an added comment just for the peace of mind of knowing what is actually happening. (One may say that the assembler output is still a blackbox but the general structure of the computation will remain intact)
It may be outdated, I agree, but it's just something I can't shake off
With Rust I think it can be a bit tricky, because a lot of it is context dependent, and the compiler will often make a lot of things easier for you until it can't, at which point the error which got you into a situation may be pretty far from the code you just changed. That's why I think syntactic sugar is best when it's quite dumb and local. I.e. X is just another way of writing Y.
I think you nailed it here what I couldn't put in words. I don't mind context-dependency when it's in the code I'm writing (calltree, variables, flags etc.) but for some reason if it's context-dependent in the language itself, hidden from me, it feels uncomfortable because I cannot change it, have to learn it and then have to debug it.
First example (C):
for(i=0; i<N ;i++) { t = data[i]; ... }
Second example (python):
for t in data: ...
The difference is that in the second example I'm telling the compiler what I want to *have*. It will give me exactly that ('t' will contain consecutive values from 'data') and it will do it's best automagically. In the first example however, I'm telling the compiler what to *do* and it will just and only follow orders (iterate 'i', copy the value from data+i to 't').
Insignificant difference most of the time but an important one to me. And sugar makes it more complicated to distinguish what I told the compiler to do vs. what I wanted to have. Helps me with debugging and getting cycle count under control
but then who is the doer and haver c vs asm complier vs assembler
It makes no sense to do it most of the time in most circumstances, and I think downvotes I'm getting represent that (though I'm not preaching it or anything, just commenting about my point of view), but if you have a hard deadline of X cycles you think about code differently.
Would you be able to shed some more light on this in the opposite direction? I would appreciate it because I'm moving to a higher role and I would like to develop tools my team can use to streamline development but I can't implement them myself and I'm very anxious to take over the tools team because my brain is wired in a very different way (I know what they need but not how it should be written) and what I think is important is probably last year's snow there so I'll be slowing down things unnecessarily. Any advice?
Values is just a container, and we'd like to grab one of what it contains one at a time to do something with it (I'm assuming native English speaking here, other language's pluralization might not be as natural to this convention). It might even make sense to start with some kind of notation like, for value in valueList:, to help internalize, but this can be a crutch. Because of the ducktyping nature of python, values might be a list today, but the same code can easily handle any iterable (I've done enough maintenance programming where it was clear some variable had never been renamed, it was once a list, and was now something else. Type annotations on newer codebases help here). The other thing to realize, is that heterogeneous collections is one of the things that Python excels at, and something that is likely to make your head hurt in C (I'm not even sure how I'd implement mixed type arrays in C without doing something gross with void*). I'm not saying you necessarily want to mix types since it can often lead to a bad time, but it still helps to think about just how amorphous the data can be.
Another thing that will feel really gross, is that Python is going to be orders of magnitude slower than your used to in C (with similar memory bloat). Just get over that, and know that if you absolutely need to speed up parts of the code, you can write those parts in C fairly easily.
Anyways, good luck.
Syntactic sugar is often better in high-level languages because it provides a level of indirection and allows for more flexibility in API design.
In languages like C# and Java, I've found syntactic sugar to be very useful (e.g. lambda notion)
> async/await is syntactic sugar on top of the promises and provides a way to handle the asynchronous tasks in a synchronous manner.
as is stated on this website I just found with one second of googling to support my thesis: https://dmitripavlutin.com/javascript-async-await/
If that is true, I believe it's incredibly useful syntactic sugar.
I'd avoid cases where extending the language is done in a non-obvious and ambiguous way.
…mumbles something about prying those from my cold dead hands.
Like if you have super long expressions in your ternaries, or you have a bunch of nested/chained ternaries it can be hard to parse, but in a lot of cases it's so nice to get something onto one line instead of needing an if/else
const foo = bar ? 1 : 2;
Without ternary you wouldn't be able to use const specifier on foo.
Note: These lines should be split, but many of them are not. How do I do that? I read formatdoc, but that did not help.
; # see where text starts; mut = ((substr($0, lf1-length(lf1str), length(lf1str)) != lf1str) \ ? 1 \ : ((substr($0, lf1, 1) == " " ) \ ? (lf1+1) \ : lf1) \ ); # start of string;
---------------------------
# return file and line number with optional text; function filineg(ffilename, ffilerec, fltext) { ; # first version always gives a line number, second does not for line 1; return (ffilename \ (ffilerec == "" ? "" : (":" ffilerec)) \ ((fltext == "") \ ? "" \ : (" in \"" fltext "\"") \ )); }
------------------------------------
outstr[2] = (head \
( (lnum > 1) \
? (substr(bigsep, 2, lnum-1) " ") \
: (substr(spaces, 1, lnum)) \
) \
(header2==""?" ":header2) null("output space if null - 5-20-98") \
substr(bigsep, 1, \
ltocout-lhead-lpn-1 \
-(lheader-lheadchop)-lnum \
-((null_loc > 0) \
? lnullstring \
: 0 \
) \
) " " pn); # parens make it line up;
----------------------function telldepth(deptxt, dep1, dep2) { ; # change depth from dep1 to dep2; ; # good luck figuring this out in a month!; ; # tell how depth changes; recnumprint("{" deptxt " " \ ((dep1=="") \ ? ((dep2=="") \ ? ("leaves depth as is") \ : ("sets depth " dep2)\ ) \ : ("changes depth " dep1 " to " dep2) \ ) " in " filine() "}"); }
In Java it's kind of a syntactic sugar (lambda is a shortcut for an anonymous class with 1 method) because of historic reasons, but I'd prefer if it wasn't the case TBH.
It was always a dirty hack.
We can go on and on. Syntax sugar is a compiler-level of abstraction.
It’s good
A lot of stuff on this list has been tribal knowledge for decades (well, except for the part about build pipelines, language choices, and Stack Overflow, etc.).
My first code mentors in the 80's & 90's said some of the same things in this list, and I passed them down to my mentees (?) as well.
I think in the strictest sense deliberate practice is picking something that you’re weaker at, and spending time deliberately working on that.
I imagine that some part of the author’s 10k hours was working on projects outside their comfort zone, while others were within it. The argument would be doing stuff that you’re comfortable with or don’t have to think too hard about would not be deliberate practice.
- Working on a large open-source gave me access to high-quality reviewers from around the world.
- Likewise, at Google, I was surrounded by some exceptionally smart people.
- Finally, like you said, working on projects outside the scope of my work and expertise.
I'm not saying you're wrong - just implementing code at work by itself isn't deliberate practice - but I think it can be if you work just outside your comfort zone. Personally, I find work incredibly dull when I'm not learning something new.
My summary, and it won't be useful at all, is that a lot of programming decisions come down to judgement. Comment or not? Config file or DSL? Is it ugly? Is is a rare feature of the language? All of these things are the kind of thing that you could argue if you wanted, but an experienced programmer will likely have better arguments.
One thing I still don't agree with is the first one. There's a lot of things, mostly trivial, that are easier to find on SO than in the source code. In fact something like "what's the idiomatic way to concat a string in $lang?" is best found on SO rather than the source code, because the source code will allow more than one way to do it.
For number 2 problems ("In many cases, what you're working on doesn't have an answer on the internet."), the answer is again judgement. For this you want to have a network of programmers you've built up over the years that you can ask. I have a couple of good friends on chat that I can just pop a question to, and it saves a heck of a lot of time. Hard to find though, they have to be someone who is basically gonna work for free for you, and you have to provide a similar level of service when they have a question for you.
>That usually means the problem is hard or important, or both.
Really ? If a problem is important someone likely already tackled it. I mean 15 years ago this was less likely, but these days there's so much work in the open and search is very good, when I find there aren't any references for my problem it usually means I misinterpreted the problem or I'm doing something very niche.
Surely there is copius documentation on a significant number of hard problems.
Now that I think of it, the right answer would have been several 10k hour professions. I can't say that the products I designed in 2020 were more complex or amazing than those in 1980. It would have been a chance to achieve high levels of proficiency in multiple fields.
I think the programming field is large enough to learn and improve your whole life. You probably have to change domains, programming languages, frameworks, architectures, companies, and switch from web to mobile, from frontend to backend, from serverless to embedded -- but there are always new ways of experiencing the beginner's mind.
Granted, this may be because the areas I work in are less technically complex than, say, writing low-level code or IoT instructions.
Now, if you developed using a single language targeted on a specific platform for 10,000 hours while challenging yourself at a high level, you would have a very strong level of expertise in that area.
Furthermore, Gladwell made the distinction that the hours spent should be deliberate and tailored to improve skills. Working on tasks handed down to you by your superiors at Google is not deliberate practice.
I think it boils down to each 10k hours of programming on (or mostly on) a certain language will give your different reflections that we think are general.
It's DNS
And, when you're sure it's not DNS, it's DNS> Most recently, I worked as a professional software engineer at Google on Kubernetes
If this is not a world-class expert, who is?
The "I'm not an expert" mindset probably comes from working with many other people who, collectively, are certainly 'more expert' than the single author.
I can break rules like DRY, if the repeating myself makes it easy to change.
Do that for a couple of years in a single codebase and the beautiful abstraction design has changed into a state machine with a transition between almost every pair of states
I don't quite understand this one? Does it refer to domain name server? So if you're having network issues it's always DNS problems?
> Corollary: Most code out there is terrible. Sometimes it's easier to write a better version yourself.
I never cut-and-paste directly from stack overflow. I cut from several sources, combine the best practices and adapt everything to make something novel, and hopefully "better". But i think it's important that you understand what it does before you use it. This might seem self-evident, but my experience tells me not everyone agrees with this methodology. If i find inspiration to a solution, i usually leave a link to the original source or stack overflow page.
I'm not sure reading the standard library for language X is necessarily a good way to learn the conventions and practices of a language. Typically a standard library contains a lot of complex corner cases that don't need to be worried about when writing ordinary code.
(Expanding). The standard library authors don't know the users of the code. While when you're writing a piece of code with a limited number of developers (could even be thousands) you might, for example, say "all our accessors for class X return const results" while the standard library authors have to handle the non-const case"
For me, it’s usually meant I’m doing something so wrongheaded it’s never come up before.
I still don't know much.
I agree now, after switching from Python to Go.
I agree with this, looking back at my first year of professional coding I just realise how important code review is. Even if you are working for a small software agency / consultant, you have to force them to do code review before merging. Otherwise, you will spend a long time trying to make your code look better which mostly you won't be able too because lack of experience
> If you have to write a comment that isn't a docstring, it should probably be refactored. Every new line of comments increases this probability. (For a more nuanced take, the Linux Kernel Documentation)
https://course.ccs.neu.edu/csg107/design-recipe.html is a random link that describes all the steps, with something like 2. being the one mentioned (I think, I'm not OP.)
For example, working with GUI applications and C/+ can make your code a big pile of garbage really quickly because representing your data in something like an ORM is not standard, you can't do the many tricks like getters/setters in C/+, but you can in Python, or C#, etc. The benefits of VM languages are, in my opinion, not appreciated enough.
In my opinion knowing the right tools for the job is far more important than how you comment your code, how you name your variables, or anything else. I wish there were more posts where people impliment an application with multiple tool sets and compare them, give insight into what some things are good for and not. "Tool Benchmarking" might be a good term for it.
I recently started working on a project like this. It's hard work.
Building a complete tool in one language takes a long time. Not to mention 2, or 5.
Moreover, many (most?) programmers are only "fluent" in one or two languages. Even if you "know" a language, you might not know the full build toolchain, or you might not know common idioms, etc.
Most people probably want to work on building new/interesting/fun stuff, not reimplementing the same toy program in several different languages.
Not always, but sometimes, an interesting knowledge arbitrage opportunity when fields start to collide.
Is DNS here just referring to Domain Name System, as in to reference problems with systems out of our control?
When I read this, ES6 classes came to my mind :)
Great post btw!
- "Browsing the source is almost always faster than finding an answer on [the web, Stack Overflow didn't exist yet]." This would have taken a very long time as a newbie, not being familiar with C-style languages (I was working in PL/SQL and XForms, mostly) and the architecture of big software. Nowadays, depends how many levels of dynamic dispatch (the devil) the code uses. If it's more than one, the code is so unreadable as to be worthless to try to read.
- "Know the internals […].", same as above.
- "Syntactic sugar is usually bad." Depends how much more intuitive the sugar is than the salty version. Still the same opinion nowadays.
- "While rare, sometimes it's a problem with the compiler. Otherwise, it's always DNS." Disagree both as a newbie and now, but then I'm probably not working on anything similar to OP.
- "Some programmers are 10x more efficient than others." Certainly I was a <1x programmer as a newbie, but I don't think I've ever seen a 10x programmer. I've seen programmers which get features "done" by committing so many programming horrors that we were still dealing with the tech debt years later while they were at a FAANG, and I've also seen programmers which can whip up excellent code quickly but are unable to treat colleagues as adults.
- "There's no correlation between being a 10x programmer and a 10x employee (maybe a negative one)." I wouldn't have thought so as a newbie, but this rings true now. Visibility, agreeing with the boss on whatever they think is cool, being able to serve up banter on request, and joining all the "social" events are important because nobody is able to gauge programmer productivity yet.
- The "Heptagon of Configuration" is an interesting observation which I don't think I ever agreed with, but for different reasons. As a newbie because for most systems whatever we were using was usually decent enough, and now because I don't think this trend is cyclic but instead chaotic. We go from environment variables to Bash to INI to flags and so on. Usually this change is because we adopt some language or framework which staunchly refuses to treat anything but the Chosen Language as a valid configuration format, and so the existing configuration has to be adapted to work with N+1 opinionated (for the wrong reason) systems with as little pain as possible.
Sometimes this can mean writing a very small amount of glue code calling external libraries. Sometimes it can mean avoiding a library/framework and rolling your own solution which solves a specific subset of the problem, enabling a smaller footprint and less dependencies. No silver bullet, really.
Yikes. Can't read past that.
The test will not forget, and has a docstring explaining what and why.
But crucially, it is automated, and does not rely on mistake-prone humans to check (as is the case of a comment).
I think most people should take the advice of experts like him - people that didn’t start from the bottom - with several grains of salt.
Yes.
Right off the bat:
> Well, I'm certainly not a world-class expert, but I have put my 10,000 hours of deliberate practice into programming.
Okay, I don't know the authors exact history here, but I seriously doubt that they've had anything that's even close to 10k hours of deliberate practice of programming. Why? Because programmers basically /never/ do any deliberate practice. We don't have anything that's even close to what a piano player does when they practice fingersetting, chord transition, scales, or the like. I'm not even sure what this would look like, but I suspect that it is fundamentally different. Most musical instruments have a physical or mechanical side to them that is completely disjoint from the "musical" part of it. For instance knowing how to play a cool solo a guitar (meaning knowing which notes to pick and how long to hold them for), and being able to have your left hand fingers in the right positions in the right times and your right hand picking the right strings has almost nothing to do with each other. The way it usually works (I'm an amateur player who hasn't played in years, so maybe some professionals can take over this analogy) is that you practice very slowly and ramp up quicker and quicker untill "magically" your muscle memory takes over and it all kind of just happens. The result is that you don't really pick each indivitual note anymore, you kinda just play that part (when practicing you've decomposed the whole piece into small parts), and each part is atomic to you. Starting in the middle of a part doesn't quite work.
The only programming analogy I can think of (from the top of my head) is how effortless it feels to write `for (int i = 0; i < n; i++)`, or something like it. But we only spend a negligible time on actually writing out the lines of code when programming, so while being able to effortlessly write out a bunch of standard `for` loops doesn't really push the needle in any way. For a performing musicial, of course, this is very different, because the /have/ to play it right! Imagine having to write all your `for` loops in sync with the beat of music.
> 4. Syntactic sugar is usually bad.
It's hard to really get exactly what this means, because if taken literally it is obviously not true. Pretty much any loop or control flow structure is just syntax over goto, but I assume the author doesn't really think we should manually write out the gotos. Having recognizable patterns in a codebase is good. Cramming in some esoteric syntax quirk of your language into your code because it technically fits is not good.
> 5. Simple is hard.
Again, difficult to get something concrete out of this since it reads as a zen mantra. If they mean that finding a simple solution to a difficult problem is hard then I agree. If they mean that a system being simple means it is hard to use then I disagree.
> 7. Know the internals of the most used ones like git and bash (I can get out of the most gnarly git rebase or merge).
Knowing the internals of e.g. `bash` is too far for me. I'm happy if I even knowing the surface of it! I guess they mean that cargo culting these tools usually doesn't lead you anywhere good and that invesing the time into them will pay off, as opposed to "cookbooking" it, like with git in that xkcd.
> 9. Only learn from the best. So when I was learning Go, I read the standard library.
I think this is important, but it is very difficult for programmers to do this,
> 10. If it looks ugly, it is most likely a terrible mistake.
I don't think this is good advice because it suggests that all code should "look nice" and if it doesn't it's just wrong. In my experience the number one metric (if you can even call it that) one should be concerned about is flexibility. If your code is "rigid" and difficult to change it really doesn't matter a whole lot how nice it looks, because, presumably, you're not really sure exactly what you're doing, and the code you write will probably not last very long in the codebase. It's hard to put a number on this, but I'm pretty sure most of the code I write does not last a month. All the time I'd spend worrying about whether this code was "ugly" is effectively wasted, because in a month it will not exist any more. Make code easy to replace.
> 11. If you have to write a comment that isn't a docstring, it should probably be refactored.
I also don't agree with this. I think it's completely reasonable to have comments in function bodies expalining what this part of the function does. See Jon Blow's related blog post[0]. Docstrings are for users of a function, but you might also want comments for the people reading the function body. Sometimes this can be avoided by writing Good Code, sometimes not.
> 12. If you don't understand how your program runs in production, you don't understand the program itself.
Very true; I'd phrase it as "understanding the problem", in Mike Acton lingo [1]. If you don't understand The Problem, then you don't understand what your program, which is (a part of) the solution to The Problem, does either.
> 14, 15, 16
This is good advice, and I imagine it's difficult for beginners (or novices) to make sense of what to do here; some circles advocate for always pulling in dependencies, others for always inlining 3rd party trees in your own codebase. I actually think this is one of the parts of programming as a field that has the potential for some kind of a paradigm shift. I'm not convinced that in 25 years we'll be still trying to write generic "libraries" for others to use, and other people will download semvers of libraries to ensure compatibility. There has to be a better way.
> 17. Know when to break the rules. For rules like "don't repeat yourself,"
Thank you! DRY is probably the worst of the n-letter programming mantras. DO repeat yourself! Do whatever you need to quickly get a better understanding of your problem!
> 18. Organizing your code into modules, packages, and functions is important. Knowing where API boundaries will materialize is an art.
This is also very true. Given an API, filling in the function bodies is often close to trivial, and coming up with the API in the first place is indeed an art.
[0]: http://number-none.com/blow/blog/programming/2014/09/26/carm...
Well, there is "learning new stuff(practices, tools, approaches) by applying it on a toy project" or "read the source code of $popular project", which I would classify as deliberate practice. But I agree, getting to 10k hours with only that is hard.
The feedback loop for noticing and weeding out behavioral errors is simply too long and even determining what led to a success/failure is too hard to be able to practice in a classic way.
Yes.
10,000 hours of code kata might be "10,000 hours of deliberate practice".
And in-any-case "Malcolm Gladwell got us wrong: Our research was key to the 10,000-hour rule, but here's what got oversimplified"
https://www.salon.com/2016/04/10/malcolm_gladwell_got_us_wro...
Code kata are a good example of people trying to stretch metaphors from other domains to fit software engineering. It usually doesn't really work. For athletes and musicians, the "performance" part is a small part (in terms of time) of their job. They spend way more time training. For software engineers, that's not the case. There is a spectrum of things more or less important, for sure, but nothing as clear cut as with music or sports.
The difference between deliberate practice and "mindless practice" is usually using a moment to reflect on what you did, and how things went. In scrum, this is often the sprint review. So by that definition, if you do your sprint review correctly, you're doing at least some form of deliberate practice. Same thing with code reviews, you have the opportunity to have other people look at your code and evaluate it. You can of course add to that your own for of review. But in general, as a industry, I'd say we're very focused on practicing on the job.
OPs — "Reflections on 10k Hours of Programming" — inappropriately stretches popular metaphors from elsewhere.
I would just phrase it as "understand the tools' fundamental concepts and philosophy really well".