Kernighan's Law: You are not smart enough to debug it
github.com
github.com
Moyer's Law: "When we can finally print from any computer to any printer consistently, there will be no CS problems left to solve."
Since every good law has a corollary, I created one of those too (simply to be funnier):
The Corollary to Moyer's Law: "The printed dates will still be wrong."
While PDFs have brought us closer to "universal printing", I won't claim we're anywhere close to solving all CS problems. Sadly, date conversion and formatting continue to be problems (hint, consider UTC ISO-8601 or RFC3339 dates/times for the JSON representation).
P.S. I don't actually think I'm smart enough to have a law named after me ... nor have I really contributed enough to "our art".
- https://en.wikipedia.org/wiki/Gray_code
- https://www.electronics-notes.com/articles/radio/modulation/...
The reason I stated I'd like the generic version is that there are other types that we use consistently. There's a very nice RFC available for telephone numbers that we've started following and we can marshal/unmarshal pretty easily from strongly typed languages (where we control the code), but wouldn't it be nice if there was a standard way (within the JSON to let systems know it was a telephone number?
If you're going to invent a new language like json-schema, you might as well skip to json2.
People who care about schemas already enforce this at the serialization layer where the objects being mapped to implicitly are the schema.
Pushing the schema down to the protocol just adds complexity and bloat.
Dynamic languages don't have an enforced schema. You can do it manually, but a schema is easier, special-purpose, declarative.
Plus, data schema languages are typically far more expressive than programming languages. e.g. java doesn't even have nonnullable (which C had, as non-typedef structs). They're closer to a database schema.
VSCode even has JSON Schema support built in, which is cool. I use that a lot.
Also, thingo aggressively protected JSON's simplicity, banishing comments because people were using them as parser directives.
a) no mandatory quotes for dicts keys b) date and time intervals in iso 8601 format. c) optional type-specifiers for strings, so we can add e.g. ip4 and ip6 addresses. Eg. { remote: "127.0.0.1"#ip4 }
e.g. { guests: 42, arrival: @2020-02-17T17:22:45+00:00, duration: @@13:47:30 }
Any sufficiently complicated serialization technology contains an ad-hoc, informally-specified, bug-ridden, slow implementation of half of ASN.1.
I see JSON Schemas used all over. Advertising, Medical, Banking, etc.
ISO8601 won't scale to universe-level applications!
:)
JSON isn't meant to be self-describing format. There is JSON schema or the like if that is what you are after.
And yet to a very large extent, it is. Strings, numbers, booleans, arrays, and associative maps made the cut. Timestamps would be a pretty reasonable addition. It would certainly cut out all the controversy here.
Yes :)
That's the beauty of it: time is just a number (of seconds since 1 January 1970).
> Which scalar data type should your parser use?
Integer, since it's an integer number of seconds.
> And is it seconds since epoch or milliseconds since epoch?
Unix time is always seconds.
> That's the beauty of it: time is just a number (of seconds since 1 January 1970).
Having a bare number where the units and zero point rely on out-of-band information is not self-describing.
If you want something self-describing, maybe look into XML?
Wasn't there a BSD that used double instead of long for time_t for a while ages ago?
Javascript time (that 'JS' in JSON) is always milliseconds.
Then it is a complete non-starter for "just use ______". I work on software that requires millisecond precision (honestly it would benefit from even greater precision) at the transport layer. It's not even really doing anything spectacularly complex or unusual. Seconds simply aren't sufficient for tons and tons of use cases.
> Yes :)
> That's the beauty of it: time is just a number (of seconds since 1 January 1970).
> > Which scalar data type should your parser use?
> Integer, since it's an integer number of seconds.
> > And is it seconds since epoch or milliseconds since epoch?
> Unix time is always seconds.
Whilst time can indeed be modelled as "just a number", Unix time spectacularly fails to achieve even that.
* It can go backwards * Simple subtractions do not give accurate intervals
It's not a good measure of time. Don't use it as such.
International Atomic Time (TAI), which differs by UTC by 37 seconds since it doesn't count leap seconds, solves everything I know of. Although the clocks aren't in a single reference frame, the procedure for measuring their differences and averaging them to define TAI is well defined and so sets an objective standard for "what time is it on Earth".
Check out "Falsehoods programmers believe about time": https://infiniteundo.com/post/25326999628/falsehoods-program...
I just spent 2 days programming a timezone selector in a React form that changes the displayed date/time as you switch timezones but the underlying UTC representation wouldn't change.
All this without loading 500k JS timezone libraries. I only used Intl API and tz.js [1].
The trick simply was to temporarily shift the date by the difference between the local time and the edited time zone. :)
If you're near a change in daylight savings then that will go wrong.
Really this time & timezone stuff should be handled by the OS, system calls or libs should be more sophisticated so we aren't "solving" these problems over and over again.
But I'm about to start ranting about Unicode, so I'll shut up now... ;-)
Fortunately, we now have teleconferencing to take our minds off of it.
"Does anyone have an HDMI to DisplayPort adapater on them? What's the phone number for this conference again? What? You only have Slack installed? No, we decided to start using Microsoft Teams for these. Can you speak up? I'm having trouble hearing you. Maybe try dialing in again?"
"Sorry, I was double muted"
"Can you turn mute on? There's a lot of echo on your end"
"Sorry, we're just trying to work out which mic works"
P.S. We're still manually dialing every day while waiting for our ticket to be processed.
It's not that printers are physically unreliable, although they certainly can be. It's the logical complexity of getting an image (or text) from one device via the myriad protocols, connectors (or wireless), page layout languages, etc., correctly onto the printer. And for added fun, even after RMSs long trek, some printers still aren't open enough to drive without secret software.
PR#1
To print?For some reason I've gravitated to printing since my earliest computing days and your law speaks to me deeply. I've recently been working on a source code printing app (yes, I'm nuts) and have been blown away by how bad printers are today. I thought the troubles I had at home were "just me".
Cheers to you!
EDIT: I see that you've submitted an RFC but it misses (what I think is) the point. It's really the corollary of the law that is important as it states how hard it is to do date/time correctly (and you'll notice that the bulk of the discussion here is on the corollary). As I noted above, this was intended to be funny and including it in a list with real laws is just that much funnier!
IMHO there is the biggest gap between the raw intelligence and software engineering skills (including good practices). So, it means that most of the code ends up as code-golf-style contraptions, incomprehensible for anyone else (including their future themselves).
Full disclaimer: I used to be that guy who liked "clever hacks". Now I try to make code readable in the first place (largely thanks to Python philosophy).
Me?
One grad student friend of mine traded me a case of beer for a Saturday of help on his code. It was some code to record spike timings in the olfactory cortex and also control valves (smelly research).
By the 12th nested for-loop, I gave him back the beer.
Less code-golf, more hackathon-style "one more copy-paste and it should work". It has some use cases, just readability or maintainability are not among these.
The problem was that once in a while you'd just get some garbled text on the screen.
Eventually I figured out why: Some of the sentence fragments were malloc'ed, others were returned from the stack. Once the description got over a certain number of bytes he was smashing his own stack. I say 'his' but the problem was introduced by an earlier collaborator.
Buddy, I really want to help you with this but I (much younger me) am not up for unwinding a use-after-free bug of this level of recursion. Good luck, I'm out.
On several occasions, and with lesser crimes, I've told the person something like "I think you are confusing yourself with your own code. I want you to change to meaningful variable names (stop recycling variables) and factor out a couple of child functions here, and here. If you still can't see the problem, come get me and we'll try this again."
Haha, no way. Grad students are poor as is. Drinking a month's beer budget in front of the guy would be bad, leaving him with that code was cruel enough.
Later in college when we started doing more team projects, comprehensible code became much more important. No one had the time to understand someone's doubly-nested list comprehension. We just wanted to finish the project and get on with life. In this sense, I think the moment you read someone else's shitty code is the moment you realize that you need to write good code yourself.
Still, I'd say it was good practice and helped me think about code in different ways.
Of course, Kernighan's law still applies, it's just that the definition of "too clever" recedes with skill and experience.
I might just be dumb but I've been bitten so many times I'm dubious anything can be verified by inspection.
The most brilliant programmer I've ever worked with wrote hardly comprehensible code since they could understand other's code easily, so they didn't feel the need to write readable code because, in their word, "What do you mean you don't understand it? It compiles, it runs, it works, you don't need more than that".
I bemoan lack of nice chaining. Some libraries provide APIs that allow chaining, e.g. Pandas. In other cases, I often find JavaScript (map/filter) and R (dplyr pipe operator) pipelines nicer to read.
However, Python 's Zen of Python can (and IMHO: should) be adopted to other languages.
[1] E.W. Dijkstra - The Humble Programmer, ACM Turing Lecture 1972 (https://www.cs.utexas.edu/~EWD/transcriptions/EWD03xx/EWD340...)
The sign of a great cook is not the ability to make great dishes -- it's being able to fix dishes that someone else screwed up!
Or maybe I just spent too much time in OllyDbg at a young age to notice.
the lower you go, the less context you have about the author's intent. with too-clever code it's easy to miss the forest for the trees.
e.g. it usually works fine to fix spelling errors at word level, but less well to restructure complex sentances without larger context.
If I understand the application really well often I don't need to do the first step, but it can be a handy shortcut.
Yes. Because if you have to use a debugger to understand a piece of code (your own or someone else's), the author has already failed. Good code can be read and understood without any additional tools.
I don't know if you've ever had that experience in your OllyDbg-using days, but did you ever try to analyse software that was using some form of hashing function, or cryptography (ECDSA, for instance) with OllyDbg, without knowing what it was? The idea is kind of like that. While it's easy to get a concrete idea of what that block of code is doing (feed it an ASCII string and some memory structures, it spits gobledygook out), it's not so easy to look at it and conclude "Ha! This implements AES-CBC!" without some strong intuition or experience.
So, yeah, specifying an architecture may be easier than debugging a routine but discovering the "hole" in the architecture can be twice as hard as specifying it.
Tracking down a bug and fixing it with a one-liner does not feel like productive programming (though it brings value to the business).
If I can't debug code on my workstation, e.g. it's running on Kubernetes with service mesh and a cool feature must be added, that's when all hell breaks loose. For example, debugging why Kubernetes deployed on OpenStack will push an image to its internal repository but won't pull from it, that is hell and I'd much rather debug HTTPD than that.
I had this exact issue in an OpenShift system.
I find debugging saves time overall. I don't have to guess about anything, I can see it working or not working, quickly run one liners to confirm. Sometimes getting an application into a certain state takes a bit of wrangling too, and a debugger helps save time by not having to constantly set up that state to see how things have occurred. You can just sit on a breakpoint while you deduce what's going on. A bug that happens at checkout success is a good example. Going through checkout over and over again is time consuming.
Absolutely. I found out very quickly in my career that I can either stare into the code for hours to see the tiny mistake, or spend 10 minutes stepping through the code to let it tell me what's wrong.
But, it is of course more difficult if you have to debug a complex code base that you don't know. It might take you hours to even know where to place debuggers. So, it can be time consuming, but of course it's still orders of magnitude faster than staring into the code, trying to see where the issue is.
If you wrote the code, sure. Then you have the context of each step you took to iterate the code to what's there now, with some insight in to where the issue might lie.
Take someone else's code that's the final, non-working version and it's a different story. Debugging is hard.
This is what RR is for: you only need to be able to reproduce the bug once inside RR.
> ensures some places where it looks like a race are actually not
Uh be careful, if you're writing in a high level language... such as C: there are no safe data races.
https://software.intel.com/en-us/blogs/2013/01/06/benign-dat...
The compiler is free to reorder operations in ways that make what appears to be a safe data race unsafe.
Often it is difficult to prove this to the compiler.
So is the CPU.
And on x86, the rules amount to "the processor will expend herculean effort to make sure that all memory orderings look identical to all readers in the system" (with a few oddball exceptions, read the SDM, yada yada). Really, on x86 you pretty much don't have to worry about this. If two memory operations happen in a given sequence in the disassembly, literally everyone who reads those locations will agree they happened in the order presented, no matter what optimizations the hardware might have done internally.
On most other architectures the CPU can and will change memory read/write order. As a result there is a lot of multi-threaded code that works correctly on x86 that fails when run elsewhere.
(Making simple code for complex problems is hard!)
What I mean is the hard part is stepping back and actually solving the real problem in the best way. That might be utterly trivial code wise, it might be brute force vs optimal. It might be just solving a readily parallelisable problem on a single thread. It might be a unilanguage monolith instead of microservices in the 'right language for every task', etc.
The smartness involved in making code simple (and I mean simple, not elegant) is more about pragmatism, lateral thinking, discipline and focus on end goals than any sort of analytical smartness.
Even in math or science it's harder to create something new than it is to merely imitate it.
It's hard to write a concise example because it really takes a big application to make the point, but consider:
featureFlag(10, false);
processOrder();
What's flag 10? I have to go figure it out, wastes time, oh turns out it was to disable tax, why are we doing that? Okay in the function that called this one we checked if we were processing a subscription order.versus
if(order->isSubscription)
processOrderWithoutTax()https://jonathanfries.net/code-is-still-harder-to-read-than-...
Been reusing and repeating this one quite a lot, since I saw it mentioned in one of this year's RubyConf keynotes: https://blog.jessitron.com/2019/11/05/keynote-collective-pro...
(There is a blog-post version behind this link as well, in case anybody wanted to skim the contents rather than watch video of the entire talk to find out how it fits in.)
Fascinating thought experiments.
Here’s an introduction on YT [1], though it really only scratches the surface.
I would argue that in order to fully and truly understand some code you have to be able to re-create it from scratch. E.g. to explain weird constants, data structures, design decisions, estimations on complexity. Hence, in my eyes, the intelligence required to perform those tasks are the same.
But, there is a backdoor. Solving problems needs immense creativity. Not everyone that easily grasps a piece of code (a domain expert, savant, MIT wizard) could have written it. To my eyes, that doesn't mean that the savant is necessarily of inferior intelligence, but likely less creative in this area.
You still get to be clever, but only you will ever know how clever.
Debugging is twice as hard as writing the code in the first place. Therefore, if you write the code as cleverly as possible, you are, by definition, not smart enough to debug it.
(Brian Kernighan)
The rest of the laws are interesting, so I’d recommend taking a look as well..
Kernighan's Law means that you're continually forcing yourself to level-up in order to debug your own code. It prevents stagnation.
Use caution in multi-developer environments. :-)
[simplicity]: https://github.com/dosyago/dumbass
Personally I find the original law inspiring.
Debugging a K8 cluster with print statements is hopeless, but if the cleverest code you can write in a dumb editor is going through a good test suite and profiler, you might be fine. How many as-clever-as-possible optimization tricks have been made viable by Valgrind?
Gotta love T-SQL, where hard things are easy, easy things are hard, and one must perpetually choose between extreme cleverness and extreme verbosity because there is absolutely no middle ground.
The expression "by definition" is misplaced here. The first proposition "Debugging is twice as hard as writing the code in the first place" is not a definition. At best, we can treat it as an axiom. In any case, the conclusion "if you write the code as cleverly as possible, you are not smart enough to debug it" does not follow by definition.
I believe that the sense in which he meant 'by definition' is that it's tautological to say that more of something is more than less of something.
For n > 0, 2n > n
After a while figured out pretty good ways to manage my interactions with the engineering teams, despite having no coding experience or really any visibility to the code.
How I worked with Development and Continuation teams was DRAMATICALLY different.
When working with "Development Engineering" (they wrote new code) I always asked them:
"What does X, Y, Z do?"
Effectively I always asked them what they think the code "should" do and what the results should be. There was no point in presenting them "OMG it doesn't do the thing" because they would just panic, get defensive / shutdown. Rather I understood they knew what their code does (well what they thought it did) and asking them that was the key to get them talking / sharing.
After I had their words and phrasing I could better present "Hey we see A, B, C under D conditions. I expected to see X, Y, Z." That would get a lot more buy in from those folks.
I also had to avoid some of the more obvious "hey it does G" where G would obviously break the feature entirely... they often didn't understand use cases (one guy I'm pretty sure didn't even know what the product did, but he could program an ASIC for sure..) so you had to be careful about spelling out customer experiences and rope in a program manager if you felt there was a fundamental problem.
When working with Continuation Engineering (bugfixes and etc):
These guys were much more receptive to the customer's story on on how the customer is using the code / equipment, and you could much more quickly present "Seeing A, B, C, under D conditions. So then I changed L and got M..." and so on. Continuation engineering grocked the meaning of why someone would change L... and so it was helpful. Continuation didn't often know what the code "should do" (or they weren't confidant in their understanding) so just saying "Saw A... that's not right" meant nothing to them. In contrast development engineering needed to talk about the happy path first.
Note that much of this occured after I gave them a good writeup / heads up of all the information I had (even if they didn't read it I never kept anyone in the dark).
It got to the point that if a panicked support call came in and it was end times and someone gave engineering a heads up, they would ask it be assigned to me if I was in the office. Technically I was far from the best tech, but I could talk to engineering.
It was telling how much a difference there was between debugging and initial development.
For me I (and the customer) all thought we knew what it should do, but it's always good to spell it out for everyone so we all understand when / or even if we're seeing an exception.
Lots of "Woah hey is that what the protocol really says?" moments too.
For about a day I was terrified that it would never work. Then I found my mistakes (two).
Except in emergencies like this, my time is much better spent writing "inefficient" yet easy to develop code that everyone can understand.
Dealing with that at work right now. We got super clever, changed a couple times, tried to be even more clever, and engineering gets to pick up the mess and try to make it all work.
/s/it/your cleverest code
True but that's too many characters for HN's title field. Which might be why it was submitted with editorialised version.
> /s/it/your cleverest code
massive nitpick but you should be suffixing rather than prefixing the regex with a slash:
s/it/your cleverest code/
As I said, this is a massive nitpick; literally the only reason I pointed this out was because the submission was about debugging code and I liked the irony of debugging someone's posted about debugging code :)(there is another law somewhere about people who nitpick are prone to making mistakes in their corrections thus getting nitpicked themselves -- I'm hoping I'm not exception hehe)
something about needing the hn editor to be a REPL...