I’m a productive programmer with a memory of a fruit fly
hynek.me
hynek.me
Most programmers believe a number of blatant falsehoods about documentation, with the most prevalent being "comments go out of date quickly, so there's no point in investing in them". Maybe I'm just hyper-aware of it because my short-term memory sucks, but code comments have saved me on so many occasions that they're simply not optional.
You can document your code. You can keep it up to date. It isn't that hard. You just don't want to.
All this is doubly important when designing an API, where the comment (also) serves as the externally visible documentation for that API. It helps me put myself into the shoes of the caller reading that documentation without knowing the implementation details, which I believe helps me design a better API.
I hear that “comments go out of date” a lot, but in my case, if a comment doesn’t match the code, it’s usually the code that needs fixing.
"The literate programming paradigm, as conceived by Donald Knuth, represents a move away from writing computer programs in the manner and order imposed by the computer, and instead gives programmers macros to develop programs in the order demanded by the logic and flow of their thoughts.[4] Literate programs are written as an exposition of logic in more natural language in which macros are used to hide abstractions and traditional source code, more like the text of an essay."
The key feature of literate programming is separating the order that code is written in from the order that the compiler requires, not comment to code ratio.
Also, difficulty writing a clear, concise comment is a strong code smell indicating that you don’t sufficiently understand the following block of code. Much better to realise this at the comment stage than at the debugging stage.
Yes! Rewrite that comment until it's sound! Being able to write in spoken language what the code is doing is so often related to the writer's ability to write that code succinctly.
But comments that describe a paragraph/block of code have a separate value. Being able to skim code faster because of a series of well chosen sentences is an aide beyond clear code. As the OP described it is a good way to design before the code is written.
Comments about why something is happening are in my experience much more useful.
An example:
# Using an if / elif rather than a select here
Rather than:
# Use a large if / elif statement here due to a regression in TCL 8.4 which made select statements significantly slower. See GitHub.com/tcl-bugs.
def write_report():
db = connect_to_database()
data = get_data(db)
summary = summarise_data(data)
save_report(summary)
After this, do the same for each of these functions.Verbs in Javaland are responsible for all the work, but as they are held in contempt by all, no Verb is ever permitted to wander about freely. If a Verb is to be seen in public at all, it must be escorted at all times by a Noun.
http://steve-yegge.blogspot.com/2006/03/execution-in-kingdom...
Especially in Haskell where you can stub those functions as undefined and get type checking along the way.
<cycle through params and start a service for each one>
…
<set up an ocr subprocess and handle errors>
…
if (<is a numeric string>) {
<trim whitespace>
<extract integer part>
<send it to /path/to/route>
}
expands into a corresponding full-page code and more details when you click on these parts. Can’t simulate fonts and colors on HN, sorry. This isn’t exactly literate programming, but more like a programming markdown. I wish it was every language’s way to go. I know you can simulate that with functions and IDE, but having a continuous context has its own benefits (you don’t have to pass it around, which is tedious and again technical), and also naming things properly and in a unique non-polluting way is hard.readingCode likeThisKind reallySeemsTo absolutelyDdosMyBrain especiallyAsTheNamesGetLonger.
This is the stuff that a software engineering professor will tell you to do, you ignore it for years, and then later on, after you’ve matured, you find yourself picking up again because it’s just one of those incredibly simple and responsible ways to break down complex problems.
Then I might start sketching the implementation with function signatures docs, types, loops and calls, and gradually start filling it in beginning with the parts that are less clear to me.
I program the same way and wish more people would consider the audience when writing code.
So things like a quick "let's improve this part later, but this is fine for now" become items that are highlighted as not "DONE" and then are seen as eyesores or unprioritized tickets, which leaves a bad feeling for devs.
While I think these features of automatically tracking TODOs can have their uses, like most things if they become metrics they start to lose their value. Sometimes it's better to just have a keyword you can easily grep for when you do find some time, having these results show up as cases against your abilities to finish the work is unhelpful, but that's what the manager and maybe coworkers see, unfortunately.
I believe (but I'm not sure) I do get what you re trying to say, but not every TODO reflects some eternal trade-off. Some are for real clear improvements (like the "there isn't any code here" into "the code was written" on my example) whose history won't add any value after they got done.
On larger teams the likelihood than comments get out of date is probably higher - particularly if the original comment was very long and detailed and the original author has moved on.
Funnily enough, most of the outdated comments I can remember were dead docs links despite that being a very common 'solution' to the problem of outdated comments.
Eg, were other options considered and this one chosen? Why? Is this piece of code non-standard or particularly clever? Why? Is this piece of code more complex than it seems like it needs to be? Why? Does this piece of have non-local dependency that may not be obvious at first glance? Are there security concerns with this piece of code which are non-obvious to a novice or an outsourced worker? Etc.
There's the old saying, "since debugging code is more difficult than writing it, then if you write the most clever code you are capable of, then you are by definition incapable of debugging it". Pretend you're explaining your code to future you who has forgotten why you did what you did and isn't clever enough to debug it.
1. Comments that are effectively headers:
# If the user is foobar, do thing
user = User.find(current_user_id)
user.preload_bizbaz!(cache: false)
is_foobar = user.bizbaz.foobar.status == "yes"
if (is_foobar) { do_thing(user) }
Yes, you can read the 4 lines of code to see the "what' - the comment is duplicative. But when I'm skimming through a large codebase trying to find or understand something, these headers can be invaluable.2. Comments on dense code
# Matches any email like foo@bar.com (only .com, .net, .org email addresses)
return email.match(/^.*@.*\.(com|net|org)$/)
Yes I can that regex, but I can read that comment a lot faster.Here is the thing: Header comments are not really about explaining their following code. They are about reducing the lines you need to read to navigate through the code by a factor of 5 to 10. Huge time-saver, and helps a ton newcomers to become productive quickly.
// The following block should have the same
// result as this (simpler) code:
// <code>
// but for <reason> this version is better
This is great for things like opaque optimizations, things that mimic a library function except for one critical difference, blocks that become really convoluted to deal with special cases, and similar.I agree. However, "what" comments need a little more skill and judgement to know what to write and what to leave out.
Having a terrible memory helps develop that skill, because you kind of get a sense of what you'll want explained to you again in six months and what's hard to decode.
Regex is the worst language that everyone needs to know.
That always gets trotted out, but it's sort of like "don't write code unless you can guarantee the code won't have any bugs, so never write any code."
The trick is write code and comments, but be a little skeptical of both. If you're disciplined (especially with commit messages), it's usually pretty easy to figure out a mistake or omission after the fact.
"...only from %s domains..."
There are definitely places where that's hard or impossible and your regex code is a good example of that. I write a fair amount of C at the moment and you can have some pointer arithmetic that's not even that complicated, but a comment on it makes scanning through the function much simpler.
I'm mostly in agreement, so I'm going to focus on my disagreement too! ;)
This view is pretty common in the ruby community. So much so that I think they take it too far at times. I sometimes find myself reading a class that could have been a single 24-line method but is instead one method with method calls, each of which calls 0-2 more methods.
This style is good for when I'm trying to quickly understand what that class is supposed to do - I just read the top-level method that looks like:
thing = getThing()
munge { |i| thing[i] }
is_foobar?(thing)
But it's terrible when I need to understand what that function is actually doing - eg to debug it or to find some underlying function call I'm looking for. I have to bounce all over the file to mentally reconstruct a linear sequence of code which could have just been one function with some headers.Of course, this style is a reaction to impenetrable 500-line functions which are also terrible. I'd definitely prefer many small functions to that! I think it's a matter of judgement and experience to know whether some code is better as small stanzas with comments or small functions.
- a nice explanatory name to it
- if the functionality changes then it is more probable that the name would change
- the code is hidden away if you don't care about the why (maybe you are looking for something else)
I feel like there are a bunch of coding habits that work well for solo developers but not devs on a team.
Second one is not a good example, because I think it's a very common pattern and email.match with com|net|org provides enough context by itself. It's just bloat imo.
I generally prefer the word "context".
The trickiest aspects of writing and editing software tend to boil down to two subjects:
1. What's the underlying system we are trying to manipulate? How is its data structured, and what kind of loops and logic branches can be hijacked for our new functionality?
2. Where and when is our new code interacting with the underlying structure?
3. To tie the first two together, what data/functionality are we exposing to the user? How and when? What is the user?
Most contemporary programming over the last two decades has been focused on nouns. The last 5-10 years has seen a shift in popularity to verb-based (functional programming) paradigms.
We can find insight from noun-based programming to answer "what?", and verb-based programming is good at answering, "how?".
Neither really do a good job of answering "where?" or "why?". Those have to be peiced together like evidence in a crime scene, and by the time you figure it out, you've probably learned everything there is to know about the entire codebase.
(I don't write code for a living these days btw, only for myself. It's much more fun, and I get to not care anymore whether other people think I'm wrong about things like these).
> I never actually trust that documentation to be accurate or up-to-date. I expect it to be misleading and time-wasting, and so I only trust the code itself.
Two things here:
1) I don't implicitly trust comments. Obviously, the code is the definitive source of reference, and where the comment differs from the code, it only means that something interesting happened. But...
2) I believe in commenting "why" or "who", and less often, "what" (and rarely "how"). My experience is that "why" comments age well -- even if the code drifts, the intent of a method or class changes infrequently.
For example:
# I (timr) wrote this method because I needed a way to
# invert the index for {situation}, and {method a} and
# {method b} didn't work because {reason}.
is a better, more evergreen comment than: # this method inverts the index for {situation}.
which is far better than: ##
# inverts index.
# @args foo, bar, baz
#
Unfortunately, nearly all doc generation software encourages the latter, and so many comments are pretty darned useless. The first example is great because even if {method a} and {method b} and {reason} fail to be true in the future, some other programmer can come along and read it and say "ok, I understand why this was written this way, and the preconditions motivating it are no longer valid. maybe I can refactor."This is hugely valuable.
I appreciate having this kind of information available, but it often gets too verbose for my taste to keep as inline comment. For this, I typically push this kind of documentation to the commit messages. This however requires very disciplined use of git: you need to "massage" your commits so each of them is a self-contained change, and OFC avoid squashing during PR merges. Then, a git blame + git show will bring up the relevant information.
I actually completely agree with your ordering of least-useful to most-useful comments. I still tend to view these most-useful comments as just kind of interesting historical tidbits, rather than something that really helps me do my job, but at least we agree on the order :)
Nah...I'm joking about you mocking me. We definitely had this conversation a few times, though.
There are comments that are categorically always bad. These are the "this code does X" kind of comments. Those go out of date really quickly in a shared code base. They don't even make sense for myself on a project I'm doing myself. The code should read like the comment would. Almost like prose. If it doesn't, I haven't made the code good enough yet. We no longer live in a world where the code needs to only be machine readable and unreadable to most mortals. We have the luxury of being able to use good abstractions and extract functions without worrying about code running slow because of too many indirections, stacks that are too big etc. We can optimize for direct code readability!
Then there are the "why" comments. Those can be invaluable. If I assume I have the kind of good code I can just read as to the "what is this doing", I can sparingly add information to things that might seem unusual or weird or inexplainable.
Same with tests mentioned in some sibling posts. Tests should be written in a self documenting way. I like to name my tests after what they're testing. Like "should behave in X manner when doing Y to Z" from a user's point of view, not on a technical level. User not necessarily meaning end user, say if you're testing an API or library function. Different language make this easier or harder but I do it in all of them. Armed with that documentation of what my tests pre-requisites are and what the expected outcome is, I can write the actual test. I should be able to deduce the expectations of the test from the test's "name", thus double checking that I am testing the correct thing. Many tests I find in shared code bases are utterly unreadable, have way too many expectations and side effects and test too many things at once. With the above technique there's usually only one or very few expectations. If I parsed out the test names from all my tests and just gave them to you as a document, it should almost read like a documentation of all of the expected behaviors of my piece of software.
but I never actually trust that documentation to be accurate or up-to-date
This is like saying food is bad, because it can spoil. Sure, but we have ways of preventing that and figuring out when it happens.A simple git blame or history (or the equivalent) should quickly answer a number of questions. Is the code significantly newer than the comment? Who can I ask for verification? etc.
It's not perfect, but significantly better than the alternative.
Similarly, presumably code changes are approved by reviewers, who should be preventing the merging of code that invalidates its own inline documentation without an update to the comments.
1. Variable and function names. These should be descriptive and never deceptive. For example, I've seen metric tons of code like this `const json = makeSomeApiCall(params)` where the contents of `json` is a the decoded response data and not a JSON string. This code is deceptive and obfuscatory. But if you write `const decodedFooResponse = makeSomeApiCall(params);` then you are accurately describing what the value is. This is particularly import in untyped/loosely typed languages where the value of the variable could be anything.
2. Code structure and layout. When writing prose, we achieve clarity through structure. Code is no different. Keep related things together. Avoid run-ons. In other words, write code that has smallish functions, with smallish interfaces, it's easier to reason about how the code will behave. Avoid unnecessary mutation of values--repeated mutation of the same value is particularly pernicious. Use white space to create visual separations of distinct ideas. Use abstraction and encapsulation to keep code focused. Avoid deeply nested conditionals, flattening the tree whenever possible. And a million other strategies that can all help.
3. Tests. Tests can be documentation, but they certainly aren't always. They also aren't convenient documentation--what they offer isn't going to show up in your editor when you inspect a method. Unit tests are a nice way of recording verifiable expectations for behavior. But if the test code is complicated or poorly organized, all it does is compound the misery of working on a messy codebase.
4. Comments are distinct from documentation here. Comments are little side notes like "this implementation is a bit of a hack. It is brittle because .... but we decided to keep it because X. See someURL for contemporaneous discussion." They tell us why a thing is the way it is or about some sort of risk or unexpected detail.
4. Documentation is absolutely a part of making code understandable. These docs are primarily going to be used in an IDE/editor when viewing function signature data. Ideally you can write JavaDoc/TSDoc/POD or similar inline docs. Examples are worth their weight in gold.
We should be using all our tools to write communicative code. We don't need every technique for every line or block, but over the scope of a package or project, we should judiciously employ them all. Every line added comes with a concomitant maintenance cost--we must ensure that every line we add is worth that cost.
Just want to add my extra take on comments:
Comments are absolutely necessary when explaining complex logic that is hard to read for the casual reader (or yourself 6 months later).
Comments are not necessary when the code itself explains what it does by being clearly named, isolated and organized.
Comments are not documentation and shouldn't be used as such. They are tools to improve code clarity for the person who is changing the code. They are not tools for someone who is using the code. If the intent is to teach someone what the code does so they can use it somewhere else, write documentation in a wiki with examples and link the wiki URL to that point in the code.
This doesn't scale, because as the size of the project and the number of docs increases, we start refining generic knowledge into more refined slices of the problem domain. At first the docs are so far apart that you rarely miss, but as time goes you miss more and more. So while your outline might suggest that finding docs grows logarithmicly, it's linear at best.
The first line of any doc should answer "where am I?" and "why do I care" because the odds that they don't care go up over time, and people put a time limit on self-service. After 2-3 wrong pages in a row they start getting impatient, unless they got through those wrong pages in 5 seconds apiece.
Initially I would just link to a wiki page in a comment, but occasionally these links break, so in my experience it's better to include the notes directly.
Even if things go out of date, it does not matter, because it still is somewhat close to what used to be there and can still help the future reader (me or someone else) figure out what might have happened since then. It's a million times better than no documentation.
Had a series of strokes. Mostly recovered but now short term memory is one thing at time.
I have leaned to document like crazy as I’m working.
Todo lists are my life now.
It's meant to warn against letting comments become false, not against writing them. To me, at least.
More broadly, a problem with our cognition is that when we remember something right now, we often think it is unlikely or even not possible that we wouldn't remember it a day, a week, a month from now.
// Test that uploading data with XYZ function works
TEST(UploadingDataWithXYZFuncWorks) {
...
}
Like WTF! The function name says all that needs to be said. There are is no more info needed but I had to appease some over zealous programmer.Meanwhile, in an existing test in the same file I see code like this
// Set the dimensions
SetDimensions(0, 0, width, height)
Again, WTF!Comments are important for explaining why something exists, what its assumptions are, edge cases, etc..., but you can also go overboard with comments and you can make your code more readable by using readable names for functions and variables.
The "code changes way too often" excuse is basically a restatement of the myth that I wrote above. Yes, code changes over time. It's pretty easy to change the comment with the code, when necessary. The team members who don't do this aren't doing their jobs.
But at the end of the day, even if code deviates from comments that's still fine. Having an outdated, well-written comment is a historical record of what the code was supposed to be doing. That's useful.
It’s similar to writing your own flash cards, but by rephrasing the code into a comment you’re visually and linguistically using a form of repetition to stick the code inside your head with a human readable explanation. This can make you a better technical communicator when discussing any code with others.
There is a lot more than “I will read this comment in 6 months and it will save my ass” going on when you practice writing good comments.
I really like this because in my opinion, half of my job is good communication, another quarter is planning/design, and the final quarter is actually writing code.
Documentation that is close to the things it is talking about isn't hard to keep up to date. What is hard is knowing, when I make a change here, that someone (maybe me) talked about this thing over there without any pointer from here to there. Can I do a careful search of everywhere there might be documentation every time I make a change? Yeah, I guess; it'll definitely slow me down in the short term and I'll still miss things.
At worst, yes. I've increasingly seen docs in Notion, but Confluence was common before and probably still is. I agree that a big part of the solution is "get it all in the repo" but I don't think that's enough.
Even within a repo, writing can be pretty distant from the code it discusses. In my experience, programmers and reviewers are great about comments on or adjacent to changed lines (not coincidentally, what shows up in git diff and code review software). Both are still pretty good about comments that they can structurally expect to be present, like function documentation at the top of a function.
It falls off a lot for programmers when it's a comment talking about something more than maybe 20 lines away, whether that's a quick sketch at the top of a loop or a block comment somewhere in the file about invariants or caveats or gotchas. At that point, reviewers simply will not catch that the programmer didn't update the documentation, unless they happen to be unusually familiar with the file in question.
Comments that reference code in other files are thankfully rare if you have reasonable modularity, but will also be missed by both programmer and reviewer unless the task in question touches both files in a relevant way.
Documentation in the repository that is not a part of code is hit and miss. Developer setup guides are often brittle to particular system configurations, but usually get some effort every time the new hire trips over something. Runbooks similarly, and less limited to the new hire, and hopefully you look them over proactively on occasion. Large scope architecture documents are doomed; I am more optimistic about something like ADRs which should be relatively narrow and aren't intended to be updated beyond deprecation.
Much of this can be helped with tooling, but where it exists its poorly standardized.
And that's patently absurd on the face of it.
I always force the onboarding process to go through our docs, and I spend a little time with each new person observing their progress looking for regressions in the docs. You can't get that with old-timers because of the echo chamber/curse of knowledge effect.
This breaks down when you have a place that never hires new people. And rather than thinking that's a flaw in my process, I'm starting to think that's a flaw in the business itself. Without fresh ideas and feedback a project stagnates.
I wonder why not more people do this? Not just for code, but for everything. I remember where I learned everything I know, I don't trust anything unless I remember the source. How do other people think, do they think that whatever pops up in their head is the truth without any source? Then how do they know it is a fact and not just a hunch or a guess?
If I am in a rush to implement something and then am sidetracked because of a stupid small bug, then I just fix that small bug in the code. And then another one. In these states I only pay attention to code, and not text. If I would also have to read the text, then I would forget the original task.
So sadly no, comments do not automatically stay in sync with code for me, unless I put in the extra effort of cleaning up afterwards.
What if the small bug that side-tracked you eould have been prevented by a comment?
But seriously, in the cases I remember - not really. Most bugs are a cause of lack of higher level understanding of a certain module. Some clear comments can help with that, but better is proper higher level documentation and the time to read and maintain them.
Slow down, it'll get done faster.
And you are free to be as slow as you want, but my flow state works a bit different and context switches are expensive. Which is why I want the minimum of comments and rather have self explaining code and proper higher level documentation.
That's only true if everyone can check in code willy-nilly without any approval.
If you require code reviews then that's just part of the code review. No documentation update = no approval, just like missing tests or bad code.
Yeah I can or an individual can, that misses the point of the advice which is intended for groups on large projects with deadlines. Like once a manager says get this done now and we'll find time later for the rest... That never happens
This is maybe a close #2 on the list of documentation falsehoods that programmers believe.
Test can be written to tell a story but most people aren't taught this way (or simply don't buy it or don't write tests).
Imagining a specification out of this is like solving those problems of "what is the next number on this sequence: 1 18 31". It's simply can not be done, you can guess something, but you will never know if it's the real answer.
I hear this a lot, but I haven't found the same with myself. I can read and understand old code I wrote.
I attribute it to my IQ being mostly held up by reading and writing comprehension skills. The way I process information makes it easier for me to remember and understand old code (I believe), compared to a lot of programmers whose inherit skills align more closely to things like math and logic.
Documentation can be a rabbit hole for me, and I often don't feel I benefit from it within the code.
I appreciate this is as good tactic for many people, but implying it's for everyone is too dogmatic.
However, comments are typically better at answering the why.
If you are diligent about documenting AND updating
I don't understand why this is viewed as challenging. Writing a sentence or two here and there is orders of magnitude easier than writing the code itself. And any code review process should help to prevent situations where the code and comments are out of sync.Lastly, I feel like there's a larger human issue. I write comments to explain certain why's in the code because I care about my teammates and I care about the project.
If others don't do the same, I think it speaks to a lack of care for their fellow engineers and the work itself. I think, "I just spent five hours figuring out something that you could have explained with a single 30-second comment?"
I'm baffled that some engineers think that this is okay.
Tests do have a slight overlap with documentation, but it's that, only slight.
If a piece of code has some weird non-obvious behavior, the presence of a test for that particular behavior is a signal that it's actually intentional, not a random bug.
But, it doesn't tell me anything why that design choice happened. That's what the documentation is for. So facing such code, I sure hope it's well documented.
With Ada, it's not only easy, but encouraged, to encode so much information in about how things are modeled into the program itself. Not only does it function somewhat like documentation, it also lets the compiler helpfully yell at me when I still manage to forget how things actually work. It's saved me so much stress and debugging time.
Now if only any of these 'safer' languages would add even just strong typedefs. Even if they don't particularly encourage their use, it'd be something.
The key difference is where the described thing lives.
A comment describing the next few lines of code or some loop that follows, or similar, won't go out of date since it's easily updated together with what it describes.
A comment living in another file, describing loosely some far-reaching but still-evolving aspect will likely be outdated soon, if the thing it describes changes and that comment is out of sight and needs extra effort to be remembered and updated.
The author would code while wasted and couldn't remember anything if it wasn't on the page in front of them, hence the comments.
Having that handy you can seem like a god at times to your colleagues when they are trying to solve something. "Oh, did you know you can do this? I'll message it to you."
I think my working memory isn't very good, but there is something in there that is very good surely, because how else would I have been able to computers for this long?
// this does x
do_x()
is a bad comment, however... // this does x
OMG_WTF_function_call(rocket science maths & bitmask);
is a good comment :)Joking aside, documenting as a habit alongside coding is like a superpower. I find that writing can act as validation against my understanding of a problem; if I struggle to write about it, then it's likely that I don't understand the problem as well as I thought.
YES
Yep, exactly. It only gets "complicated" when you start having these heavy-handed doc generators that are parsing your code and breaking the build.
I feel pretty strongly that sphinx, rubydoc, python docstrings et al. are fine tools, but you have to have a light touch with how they're applied. Autogenerated external docs are a separate problem, and shouldn't discourage developers from commenting their code.
All your documentation can do is make it more ambiguous. Usually the documentation is wrong as well, but that might be because the programmer didn't know how to specify clearly what he wanted.
I take back my other comment where I said that "tests are documentation" is the #2 falsehood about documentation that programmers believe. This is the #2 falsehood.
Yes, "formally" the code is the spec. That doesn't help you very much though, squishy human, because you're not a computer.
// set i to 1
int i=1;
and this is a useful comment: // there is a method overload which accepts an array,
// but it's buggy and crashes. workaround: cast to tuple first.
...
then you seem to be assuming people only want to write and read the first kind. Because no matter how well you 'know all the fine details of the programming language', comments can tell you things the code can't.Even aside from the pointless gatekeeping of "competent programmers" - ok what about people who aren't employed as programmers but still need to read and write code? What about the people who just have to deal with it being an unfamiliar language because nobody who knows the language is available?
For example, how to get the length of an iterable in a given language. I may not remember the function name (or if it's a top-level function vs a method) but I know I can search "string length in $LANGUAGE" and find it. This scales better than memorizing every language feature I'll ever use.
---
PS: Dash is a great mac app, highly recommend. Your job will likely cover it if it's something you think you'll use.
I've stopped worrying about remembering details, and just assume that my neural network brain will figure out what is worth retaining. Sometimes it does, sometimes it doesn't, but life goes on. Not having anxiety about forgetting things feels wonderful.
Very often it's a lot easier to remember the journey to a piece of information than the information itself. I remember reading about some Greek scholar who'd imagine his long speeches as walks, which helped him memorize them. I think we are more suited to learning journeys than destinations. Maybe there's simply more concepts to connect neurons to.
Tangentially, I've gotten better at remembering walks and drives as well, not through any conscious effort, but simply cultivating a sense of curiosity about the world around me. The more I learn about how the world fits together, the more interesting tidbits I notice in a location, and when I see the same place again it reminds me of the previous time I was there.
This is actually a really good example of something I've started using GPT-3 or GitHub Copilot for.
If I'm in a Rust program and I don't know Rust, I can type
# set a to the length of the items array
And Copilot or the GPT-3 Playground (if I don't have Copilot handy) will write the next line of code for me, without me having to go and look up how to do length-of-array.
My Dad (also an engineer/scientist) and I joke that our brain is just a massive foreign key relation store to primary keys records on the internet.
Combine these pointers with google and an intuition for finding relevant information, I will nearly always arrive at the correct information/solution.
It's like a muscle, you have to keep training it.
Currently has plugins for Vim (proof of concept) and Emacs (actually good). It'd be lovely if one of you folks would make a VSCode plugin for it. I've thoroughly documented the API and diagrammed the code flow you'd need[2], so that this would be as easy as possible to add to your favorite editor.
Just taking the time to actually learn what the standard library of the programming language you are using provides can dramatically increase your productivity. The same with important libraries you are using. Yes, it is too much to memorize everything but just knowing what actually is there will help you a lot. You can not search for something you don't know exists.
I also really love working on solo projects because I tend to memorize the general shape of the code I am working on. This makes me so much more productive. I might not remember every single line of code in detail but will have a pretty clear idea of the shape of the program. This means I can plan new features or doing refactors in my head and see if they would work. I don't need to sit in front of the computer. I can do the hard mental work while taking a shower or going for a walk and can type in the code afterwards.
Anyone else work like this?
However it isn't always possible. Nowadays a typical dev team uses 10 different frameworks or more, and after a couple of years some of them will have changed.
However at work I am stumped about what to memorize? At work I do a bit of everything from new k8s cluster to front-end optimization.
I should do that more often with whatever I work with. However, my current job is so random (lots of odd jobs left right and center in between ad-hoc meetings and interruptions) I have no need for it.
I will have to disagree. New ideas are rarely totally new, and often build on old ideas. Keeping old ideas in your brain will help you come up with new ideas.
One of the big advantages of having stuff in memory is that you can do background processing.
I can't tell you the number of problems I have figured out while going for a long walk, or driving in my car, or even sleeping (all of a sudden I wake up with the solution).
Keeping old ideas in your brain will help you come up with new ideas.
Agree! Though, "Your mind is for having ideas, not holding them" doesn't mean that you should literally never remember anything. It's just not necessarily clear from context-less quote. I can't tell you the number of problems I have figured out
while going for a long walk
Happily, you are in agreement with the author of the quote - https://en.wikipedia.org/wiki/David_Allen_(author)The point is not to forcibly resist remembering things. The point is to free yourself from the burden of remembering things that can simply be filed away (ie, reference documentation) so that you can be more present, focus on what's important, and free your mind up for more creative/productive thinking - including those walks where many of us get our best thinking done. :)
> Your mind is having ideas, it's just not holding them.
which is still quite accurate, especially given the theme of the article. Augmenting your memory is important, especially since age can make memorization harder, and even impossible.
> While Dash is a $30 Mac app, there’s the free Windows and Linux version called Zeal[1], and a $20 Windows app called Velocity[2]. Of course there’s also at least one Emacs package doing the same thing: helm-dash[3].
[1] https://zealdocs.org/ [2] https://velocity.silverlakesoftware.com/ [3] https://github.com/dash-docs-el/helm-dash
I've installed Zeal and installed the Python documentation. I can see the Python tree (among others that I've installed) in the sidebar. But searching for "python append to list" or "append to list" return 0 results. In the preferences I enabled fuzzy search, still no results.
So I turned to the tree to see how easy it would be able to find the answer. I opened the Python leaf, and from the 15 subleaves I opened Structures. In there I see "PyCompilerFlags", "_frozen", and "_inittab".
I'll be sticking to Google, unfortunately.
Python3 -> Classes -> list
I agree it's not easy enough to use the search for this particular case (it's actually very difficult to find the list "add" -- list.append). I'm going to stick it out with zeal, though, and see if I get better at finding what I want.
Some similar searches in other languages gave back what I wanted, so maybe it's got to do with how the python docs were imported or something.
1. LAMMPS: https://github.com/chazeon/lammps-docset 2. ASE: https://github.com/chazeon/ase-docset 3. VASP: https://github.com/chazeon/vasp-docset 4. QE: https://github.com/chazeon/qe-docset
Old man mode
Remember back when Linux systems just came with huge packages of documentation? You could do anything offline by reading man pages or looking through /usr/doc/, /usr/share/, etc.
We aren't trying to understand minds or caches, we're trying to encourage a practice that has little to do with the inner workings of either.
They are a starting point for encoding new knowledge in a custom symbolic language known only to yourself. It is the basis for all learning.
Some other major tools for learning are mnemonics and spaced repetition.
i could memorize them all but it is less interesting than reading papers about optimizing algorithms for cloud economics or other more valuable (to me) ideas.
Yup... and as the size of L1 cache grows so does its latency. Eventually you reach the latency of main RAM and then... wait I forgot why you increased L1 cache size. Can you just reset it to "fast" please?
https://kids.frontiersin.org/articles/10.3389/frym.2017.0006...:
“Normally, flies will remember very well and will get an A if they are tested a few minutes after learning. But if there is long time between learning and testing, which for a fruit fly is 1 day, they will forget and get an F”
A fruit fly lives for 40 to 50 days, so 1 day isn’t a long time to remember things for them.
https://phys.org/news/2017-03-fruit-flies-memory.html:
“Flies form a memory of locations they are heading for. This memory is retained for approximately four seconds. This means that if a fly, for instance, deviates from its route for about a second, it can still return to its original direction of travel.”
I couldn’t find how many things they can remember simultaneously.
As soon as I hit 40, I can't remeber what I did last week or before that. Did I solve that problem? Which customer complained about it? Where did I put that PDF?
Fortunately I am very organized, and I leave clues for myself all over the place. Readme files in named folders, details in commits, email myself with some information.
But more than often I am asked "was X display compatible with XYZ board?" and I can't give a direct answer anymore.
I've since learned to be slightly more verbose.
Write a short paragraph after each day and a longer one at the end of the week, describing what was done.
Also maybe write down a short paragraph of what you plan on doing the day after each day, and a longer one at the end of the week about what will be focused on next week.
Review each day and you'll keep a good mental map of what has been done.
Then 2 paragraphs in I run face first into a 5 meter high stainless steel marketing pitch. Wtf.
Sometimes I'm not sure what's worse, the not remembering, or the people claiming that you having forgotten means you didn't care enough, or that you're not trying hard enough. grrrr
Has even more similarity between the two halves.
But, according to QuoteInvestigator.com[0], that is not correct:
QI has traced the core of the quotation to the work of an early researcher in artificial intelligence, Anthony Oettinger, who was trying to get a computer to manipulate the English language.
(Lots more detail at the link[0]).
[0] https://quoteinvestigator.com/2010/05/04/time-flies-arrow/
Groucho Marx SHOULD have written that joke. It's good, the punch word is at the end, it's the kind of logic-mocking braintwist Groucho was exceptional at making, and yet he didn't write it. So we ascribe it to him anyway, because that's who SHOULD have written it. Thus is intelligence: we try to make things make sense, and we try to relate things to the big areas of 'sense' we already have in our minds. We have a space for Groucho and for humor, but who's Oettinger and where can we watch his comedy routines? What, he doesn't do comedy routines? And so, sense beats reality…
Once and only once is a requirement for people that can't remember as much but also for good code. Consider it a requirement for code organization and tools:if it takes longer to look up than remember then the tooling needs a bump. Obviously during heads down green fields coding, the cache will get filled efficiently of wjatever is needrd, but by the time v1.1 rolls around a good reread and rethink might be a good idea
Also, the key skill isn't memory but rapidly and well learning new things. My C, perl, and jquery knowledge is well forgotten and replaced with python, go and terraform and i am considered trying out Rust.
Very true. Some programmers have higher tolerance for complex code. Which is often why there are disagreements over refactoring.
Now I want to be honest, I have been trying for a bit over an hour to add ANY documentation that is not included by default. It's way too complicated. Has anyone found any success?
Docs from pyspark, azure data factory, databricks, anything big data related would be a gift for me. But the steps are overtly complicated.
Specifically, in the "Building documentation" part of the article, it's where you are supposed to use doc2dash and use to add new info. If someone or even the author has managed to add any not supported doc, and wants to share, it'd be much appreciated.
Otherwise I am left with the question, how many hours are you supposed to spend in order to save time with each search? Because this programs seem unbalanced for people without the know-how to even start.
There's a gap in the market for a good cross-platform docs browser, preferably supporting user-added notes.
In fact, Programming Interviews directly test for Memory. Do you remember the correct algorithm? If not can you pattern match the correct algorithm? If not can you creatively build the correct algorithm from primitives you remember?
It is my personal belief that if you hold memory to be equal the variability among IQ will significantly be reduced. I will welcome any evidence for or against this belief.
I personally think that there are only two types of stupidity, and both can be adjusted to. You either a) have a bad memory, which means you need to take your time, take a lot of notes, maybe do Anki, or b) you refuse to accept the truth of some basic logical/statistical primitives out of willfulness, which means you need humility.
But perhaps I'm exaggerating? I'm basing these concerns on what we learned about the creator of Dash here, in 2016:
https://news.ycombinator.com/item?id=12684265
(To be clear, the creator of Dash has a different name than the author of this post.)
There was IMO highly manipulative behavior that fooled a large part of HN back then - to the degree that I remembered it just now, six years later.
- the last release was October 2018. The repo shows activity up to 2 months ago, and I'm sure the last release works fine, but it doesn't seem very active. That wouldn't be a problem, I believe software can be done, except...
- the installation documentation is very much lacking. The "make -B build" on the README doesn't work, the build instructions on the wiki is empty except for per distro instructions, the package release for Debian doesn't exist, and the ppa points to an old IP.
That said, I got it working. It seems pretty cool, haven't used it yet, but from what I can see, it is very mouse centric and no keyboard controls, has popup windows and a system tray icon (I loathe them), so for a tiling wm not quite the best. But it does what I need it to do: download docsets, keep them up to date via the RSS feeds for Dash, search them.
I'd love a tool like this but more keyboard oriented, specifically vim like keybinds but any keybind system would work. I know there's an emacs package "helm-dash" but I don't use emacs. I might try it just for this tool.
All in all I think this is going to make my life a lot better. I like to disconnect from the world, but I want to be able to code while disconnected and zeal could enable me to do that.
sudo apt-get install zeal worked fine for me.
Anyway, I found this https://github.com/qwfy/doc-browser that I'm compiling right now to see how it works, looks keyboard focused, simpler and supports DevDocs, and bonus it supports Hoogle if you're a Haskeller.
I'm on bookworm, so it's probably the opposite problem.
Quit smoking pot and things got a lot better for a while, but then things seemed to resume their downward slide. Sleep apnea didn't help, and treating it made things better for a while, but that slide continued.
Now I just try to manage it with lots of notes, scripting everything I can, and trying to simplify my workflow wherever possible. Hopefully things stabilize in my 50s or I don't think I'll be good for much engineering work beyond that. I have a number of dopamine-related movement disorders though so there may be some special issues in my case.
This kind of projects could have a better adoption rate with standards for documentations.
I regulary have to modify code I wrote years ago and I more often that not have completely 100% forgotten ever writing it. So I have to read and learn my own code from scratch. Which taught me how to write code that is easy to read and learn for myself and others.
It's much easier to modify and maintain your own software which is purpose-built to your needs than it is to maintain somebody else's code that was built to their needs, and was probably overcomplicated to make it as a generic library.
For a compiled language, you have the extra complexity of integrating its build system into yours.
So I'd say you should still weigh carefully whether you want to pull an extra dependency or not.
Third party libraries which have a handful or more users are generally speaking relatively bug-free for standard use cases.
Third party libraries are generally a productivity win for languages which have decent build systems, which is most of them notably excluding C and C++.
Every third-party library will feel extremely alien.
What happened? In C ecosystem this is a no brainer. Why can’t my frameworks offer a simple man page I can download into a folder and it works offline. Why does modern documentation require such fancy things like a web browser?
You of course can't use them without permission, but I've found them interesting before.
Takes a couple of times to get the questions right - but worth it. Only complaint is that it is not instant.
Sure! If you just assemble a page with about 500 button or links, and know which one to click on to get the needed information, everything is a click away!
As I understand it:
1. TypeScript is a language (super-set of JS).
2. Code completion is a feature of an IDE or editor.
So how does this language enforce code completion? Or is this something in VSCode? If I were using Notepad++ to code, it won't have code completion, correct? Thanks.
The rate at which you assimilate and synthesize new information on a particular subject is proportional to how much other stuff you already know about that subject.
Its also only really worth learning things that matter or might matter to you in the future, so there is a degree of speculation and risk involved.
The new language versions with things like records take a lot of the tedium away, and Copilot gets rid of most that is left.
Worryingly, while updating my 5 docsets downloaded months ago, I hit a segfault, and trying to debug it in coredumpctl repeatedly hangs gdb and makes it eat all available RAM. I suspect it's a lifetime or multithreading issue.
EDIT: Reported at https://github.com/zealdocs/zeal/issues/1426. QAbstractItemModel is the gift that keeps giving.
A codebase that requires large working memory typically means it's written poorly. Lacking encapsulation, poor modularization/separation of concerns, unnecessarily succinct variable names
It got done earlier and I pushed it out in a bit of a hurry because I was leaving and forgot to update the date.
I have fixed it.
Either that or the programmer forgot what day is today.
Strong agree!