Programmers should never trust anyone, not even themselves
carbon-steel.github.io
carbon-steel.github.io
So, good article with a misleading title. Don't be paranoid.
[0] I don't consider 100% test coverage as anywhere near close enough for that.
Even though my entire career has been software, I was an Electrical Engineering major, so I have taken VLSI design (and even designed an 8-bit ALU). My first job was writing embedded software, and would frequently "trust, but verify" the hardware through the use of a logic analyzer.
When I pulled out printouts from the analyzer to show the hardware team that the hardware had a bug, the surprised and incredulous look on their faces was priceless. The blow was somewhat softened by the fact that I'd also found a software bug.
The constant blame shifting between SW and HW teams is one reason I left that job after less than a year.
Going one step further having a culture where finding a bug or defect in your own code/design is rewarded makes it so people aren't afraid but excited to talk about them.
100%. I'm very much a fan of asking _what_ happened and then figuring out how to prevent it from happening ahead. Look back to inform the future, not to find blame. Ultimately, if 1 person can blow things up, it's a bigger issue at play.
It really means something like "act like you trust the person, but secretly, don't trust them and double check that they've fulfilled their commitments". Putting on a smiling face and acting as if everybody was acting in good faith, while secretly expecting the stab in the back because you would do the same, was key to diplomacy in the USSR, even between departments of the same government.
It rolls off the tounge so much better in Russian.
It's interesting to note that English wiki has an article about the phrase, but Russian doesn't.
Not actually. You're describing an abstraction: your opinion/perception of what must/can be done.
It is possible to be comfortable with uncertainty and the unknown (everyone already is, but only in certain, intuitive (in large part due to cultural conditioning, which comes in a variety of forms) ways), it's mainly just counter-culture and counter-intuitive, thus needs strategies, and practice(!) (plus some non-trivial multi-level, multi-dimensional recursion....this is what us HN folks are good at, and love though, right? Right?[1]). We've all been through the hard work at least once, in a certain (mostly) shared way. There are other ways though.
> "Trust, but verify" is much more practical
How do you verify your verification in complex scenarios though? I bet I know: trust/contentment (in your verification skills), though this layer typically is not revealed to us, so causes no psychological unrest ("all is well"), because it does not exist.
> Don't be paranoid.
What do you think your reaction would be if you discovered this is not just wrong, but backwards?
[1] Alternatively: maybe we are only good at it, and only love it, sometimes? But then, "we" is a complex and deep set, into which we have little insight, but also plenty of hallucinated "insight".
If you have to verify, you don't trust. Google "trust definition". Here is the first result:
"firm belief in the reliability, truth, ability, or strength of someone or something".
There is no reasonably sized body of code I ever wrote where I'd have a "firm belief" it was error free.
There are many situations where we shouldn't believe someone's work is error free. It's fine. Anyone who has ever worked in a field where it can be shown that some work has an error knows how many errors humans make. Anyone who is honest with themselves in the software business knows just how easy it is to make an error.
If you need a pithy phrase: "Assume good intent and capability, but verify work".
How do you mean? I think the article (and my experience) suggests that you do have to be paranoid. [I looked up paranoid, just to be sure I knew the exact definition, and I didn't. It's an "extreme and irrational" fear. Is looking it up parnoid? Hahaha.] Colloquially, paranoia is extreme and not necessarily irrational. Think of Andy Grove, "Only the paranoid survive." Or Kurt Kobain, "Just because you are paranoid, it doesn't mean they're not after you."
Anyway, the way I frame the issue of software quality, is to hold the view that there are always errors, and the best you can do to apply extreme vigilance in attempting to ensure errors occur rarely.
But trust isn't a binary, all-or-nothing sort of thing. There are always degrees. "Trust but verify" makes that explicit.
1. type checking, data marshaling, sanity checks, and object signatures
2. user rate-limits and quota enforcement for access, actions, and API interfaces
3. expected runtime limit-check with watchdog timers (every thread has a time limit check, and failure mode handler)
4. controlled runtime periodic restarts (prevents slow leaks from shared libs, or python pinning all your cores because reasons etc.)
5. regression testing boundary conditions becomes the system auditor post-deployment
6. disable multi-core support in favor of n core-bound instances of programs consuming the same queue/channel (there is a long explanation why this makes sense for our use-cases)
7. Documentation is often out of date, but if the v.r.x.y API is still permuting on x or y than avoid the project like old fish left in the hot sun. Bloat is one thing, but chaotic interfaces are a huge warning sign to avoid the chaos.
8. The "small modular programs that do one thing well" advice from the *nix crowd also makes absolute sense for large infrastructure. Sure a monolith will be easier in the beginning, but no one person can keep track of millions of lines of commits.
9. Never trust the user (including yourself), and automate as much as possible.
10. "Dead man's switch" that temporarily locks interfaces if certain rules are violated (i.e. host health, code health, or unexpected reboot in a colo.)
As a side note, assuming one could cover the ecosystem of library changes in a large monolith is silly.
Good code in my opinion, is something so reliable you don't have to touch it again for 5 years. Such designs should not require human maintenance to remain operational.
There is a strange beauty in simple efficient designs. Rather than staring at something that obviously forgot its original purpose:
https://en.wikipedia.org/wiki/File:Giant_Knife_1.jpg
https://en.wikipedia.org/wiki/Second-system_effect
Good luck, and have a wonderful day =3
This is true, but incomplete. All unicode encodings take linear time, not just utf-8.
That's because a character can contain multiple code points. Utf-32 allows for random access of code points in constant time, but not characters.
Utf-16 has variable length code points so it is the same as utf-8 in that regard.
ASCII does indeed take constant time - which I agreed with.
It's probably safe, maybe even close enough to optimal for typical use, to have an array or list of bytestrings for each line. Or maybe more complex Line (of text) objects that record the byte-length, 'codepoint'-length, and '(display)character'-length. There might even be special cases built in for typical and massive documents (number of lines) and overly long lines.
Of course different computers (CPU, memory configuration... they all matter) have different characteristics so where N is large enough is different for each. However in general linear search is fast enough these days. Where it isn't you can look at a profile to verify you hotspot.
I tried to Google for some blog posts, but I could find anything. Does anyone know of any (amateur) research on this topic? Example: Running on Linux, written in C, what is the threshold for N to switch from binary search to linear search?
EDIT: Fix typos
UTF-8 doesn't suffer from this design flaw.
Abstractions, in the mathematical sense, always hold (unless there is a flaw in the definition itself). Axioms in any sense are always going to throw a wrench in things. Thank Godel. But that shouldn't mean we cannot make progress.
Do the work, show your proof! Think hard!
Although sometimes all you need are a few unit tests.
They key is to develop the wisdom to know when unit tests aren't sufficient for the task at hand.
CS seems mostly good for proving that everything is hopeless in the rigorous general case and, hey, here are some heuristics to give you hope again.
But this seems like an argument against formal verification. Formal verification is 100x harder than writing tests...and still doesn't guarantee correctness? Those axiom wrenches are still there, those flaws in the definition are still there, not to mention the flaws in the proof writing. All that extra effort for what gain?
A formal proof that an algorithm makes progress, doesn't require a lock, or whatever property the proof is arguing is 100% guaranteed for every case.
For example, a simple function over the set of integers. A unit test can only test individual elements of the set. A proof demands more: the property must hold over all elements of the set.
The Incompleteness Theorem puts a limit on the provability. That hasn't stopped mathematicians from pursuing the formalization of mathematics. It shouldn't stop computer scientists and programmers either. In fact it tends to make us more honest about the limits and capabilities of our systems.
A good set of unit tests partitions the input space into parts, and then provides an existence proof that the program is correct for an example chosen from each part.
A good formal proof does the same, but provides a universal proof that the program is correct for all examples chosen from each part.
Both strategies can fail if they miss an important part of the input set. Unit tests can also fail if a program is “adversarial” and fails in a specific input that isn’t the particular example.
In practice, achieving “a good set of unit tests” requires you to mentally work out how to partition the input set in a way that matches your program, and at that point you’re most of the way to proving it correct, so you might as well do that. It still might make sense to write unit tests if you don’t have the tooling to enforce a mechanical proof.
It's nice to have when it is though.
The cost is coming down. Proof automation is incredible today and rapidly improving. This is the part that does the tedious parts of a proof for you so that you can focus on the theorems that matter. The languages and proof systems themselves are easier than ever to pick up and use which is bringing the skill cost down.
I don't think most software projects need a huge, dedicated team of specialists to benefit from formal software verification.
"If you feel very smart after writing a particularly intricate piece of code, it's time to rewrite it to be more clear."
Speaking of not trusting yourself.
> Read more documentation than just the bare minimum you need
I wish I practiced this before, I'd be as good and quick as some of my brilliant colleagues.
This left me wondering if reading the docs thoroughly is still a good investment.
On the other hand, having some level of knowledge of a library saves me some time since I can often write things faster by myself than explaining gpt4 in more and more details what I need.
If I was to point out this is an approximation, or a tautology (it is only true to the degree that it is true, which is not (necessarily[1]) 100% of the time), would it make you anxious? And if so, do you think it wouldn't be possible for you to learn [1] a new approach so it does not make you anxious?
As somebody who primarily lives on the testing side of the house, I've definitely run into cases where the developer promises that their unit tests will make a new feature less buggy, then about 5 minutes later I either find a mistake in the test or I find a bug in something that the developer didn't think to test at all.
I've also seen instances where tests are written too early, using a data structure that gets changed in development, and then causes churn in the unit tests since now they have to be fixed too.
I've generally come to think that unit tests should be used to baseline something after it ships, but aren't that useful before that point (and could even be a waste of time if they take a long time to write). I don't think I'll ever be able to convince anybody at my company about this though lol
especially if the company uses 100% code coverage as a metric for success
Unit tests main purpose is not to improve code or reduce bug, it's main purpose is to verify the code to work against the contract that's defined in unit tests. Code improvement or bug reduction are added benefit, if any.
You’re talking about integration tests or e2e tests.
Those don’t sound like unit tests.
Or to put it a different way, your unit tests should cover units that have a good boundary to the rest of the system. This should sound like a module, but there is reason to have module as a larger thing than your unit (most of the time there shouldn't be, but once in a while this is useful), and so while there is overlap it is often useful to consider them different.
Integration tests cover the API, but they do not test the API (well they often use some API as well, but the won't cover all your internal APIs.)
When a codebase gets too big, and devs gets too clever with their tests, the whole test suite becomes complicated.
If your test suite is approaching the complexity of the actual codebase (what with layers of mocks and fixtures that are subtly interdependent,) how could you be expected to trust a test you wrote more than the code you wrote.
All those mocks, and other Jest code, all seem overly complicated but I don't know of anything "better".
omg yes. However, reading this makes me wonder how the TDD people handle this.
Also if the interface doesn't change but your unit tests fail on a data structure change then perhaps your tests are too coupled.
The GP diagnostic isn't good here. Those are bad tests, not tests written too early. What doesn't mean you shouldn't wait for your functionality to be accepted before testing it; some times you should; but this happens for completely different reasons.
Raising hand, guilty as charged. I test for things based in my concept of how the system works, but those darn users may have other ideas!
Actually, when I found a user who seemed to have a knack for finding bugs, they were gold and I let them know I appreciated their efforts.
> unit tests should be used to baseline something after it ships
That has not been my experience. I found that unit tests let me get the pieces working properly so that when assembled, the chances that everything worked as expected were much improved.
Getting those ideas in front of you is the job of the product team.
It isn't your fault if the guy writing the ticket refuses to talk to the users/and or/you.
"Ticket meets acceptance criteria, please submit a new request for this new feature."
I reviewed hundreds of them, tried rewriting dozens of them. Eventually, realize that essentially all of the tests were just testing that mock data being manually manipulated by the test gave the result of test expected.
Absolutely nothing useful was actually being tested.
Some team spent a couple of years writing an unholy number of tests as a complete waste of time. Basically just checking off a box that code had tests.
These tests are easy to write - your mock returns something, and then you verify that the API does nothing (thus the test fails), then returns whatever the mock does and the test passes. These tests are easy to write and they do fail until code is written. However they are of negative value - you cannot refactor anything as the code only calls a mock, and returns some data from the mock.
So . . . is the team of devs who spent years writing and maintaining those tests incompetent or is it the new dev with the complaint? If it was the whole team, how did that happen?
Your "public" tests should document the API for future programmers. This is the concrete contract that should never change, no matter what happens to the implementation. If these tests break, you've done something wrong.
Your "private" tests are experiments that future programmers know can be removed if they no longer fit the direction of the application.
How is this formalized in Go?
In theory, but I've never seen anyone successfully write tests after something ships. By that time much of the context that should be documented in tests is forgotten.
The above those is a very rare exception. The general rule is once code is shipped management doesn't allow you time to make it better.
I disagree, but not entirely. I think there's a balance. Tests can be a great way to execute code in isolation with expected inputs/outputs and can help dramatically in absence of other ways to execute the code. But in general, I mostly agree. Tests are mostly valuable as a way of ensuring you don't break something that was previously working, but they are still valuable for validating assumptions as you go.
If every test requires copy pasting a bunch of sql statements and creating a new user and data, my experience is the team will have 3-4 of these kinds of tests. But if the test set-up code looks like `newUser().withFriends(3).withTextPost(“foo”).withMediaPost().sharingDisabled().with…` then the team is enabled to make a new integration test any time they think of an edge case.
Downside is that test setup can be a bit slow: each `withXXX` can create more data than really necessary (eg a "withPost()" might update some "Timeline" objects, even though you really don't care about timelines for your test). Upside is that it's a lot closer to what happens in reality, regularly finding bugs as side-effect. And also you align incentives: you make your tests faster my making your application faster.
That's as clear a signal that they are testing the wrong interface as you can get.
Unfortunately, developers think of tests as testing code, not interfaces. As a natural consequence, they migrate towards testing the most complex break-down of their code as they can; it increases the ratio of code coverage / number of tests... at the cost of functionality coverage.
I think starting with a test that the code does what the code does is actually a pretty good starting point because it's mechanical. And if you never end up revising that code, it can just live there forever, but if you do end up revising the code, those tests will slowly morph overtime into testing the interface. When you actually do your revision, you get a very clear signal that like hey this test that is now failing as a result of the change, clearly that can't be the important part.
Waiting for the change to see what stays the same I think is often more accurate than trying to guess what the invariants are ahead of time.
Well, the invariants you want to test are what people want the software to do. If you create those test from the beginning, there's nothing to guess. But of course, if people write a bunch of code and throw that knowledge away, you have to recover it somehow, and it will necessarily involve a lot of guessing.
Most tests I've seen aren't aids to diagnosability, so if there is a bug, the developer is still needed to find it.
> tests are written too early, using a data structure that gets changed in development
I wouldn't call this "too early," but "testing the wrong thing" if they were testing internal particulars instead of behaviors.
> Tests don't find bugs.
I want to be make sure I understand your message here. When I write unit tests, I frequently find bugs in my own code. Are you talking about something different?The "A Python script should be able to run on any machine with a Python interpreter." remark is amusing. Recently ended up installing a whole new distribution version just to get Python 3.11 and the new library versions that the script I was running depended upon.
Giving a program finish and polish does have this kind of unlimited depth, but it is critical to remember that comes after the initial coding which is always rough with many gaps since that is always where things start. And then we make many choices. Hearing people commit to strong typing because docs are always out of date is another chuckle. Maybe if the docs were kept right from the start, maybe with some use of automatically generated reference pages, then there wouldn't be such frequent problems getting types right in the first place? To each their own, but strong typing hype is just another currently popular method among many. Strong typing has its value, but like every other methodology cannot be absolutely trusted to save programmers from error.
As more and more stuff™ is moving from hardware into software, for totally understandable reasons, the absolute number of software that could absolutely ruin human lifes is growing. This calls for higher standards when it comes to the whole field of software engineering.
You can apply the same principle to programming (in fact, many programmers do). If your program falls completely fall and explode into pieces once any peripheral, call, timing, or whatever isn't as expected, then it isn't a resilient program. If it fails in unexpected ways while still writing data you might even have a dangerous program.
So many programmers are happy with writing a program that works. That would be as if a car manufacturer would be happy to call it a day if the car manages to accelerate.
The post could also talk about yet another desirable paranoia, namely "Am I building the right thing?" - so talking to customers/clients/stakeholders/users is arguably one of the most important steps when seeking re-affirmation and controlling one's work. Nothing worse than people who think they understand a technical problem, but it's not the one that the customer wants solved...
Yep. Can confirm, the best programmers are often paranoid. I guess over time your brain becomes more logical and you start to notice inconsistencies everywhere around you... Then before you know it, it feels like you're living on planet of the apes.
Demonstrating that language and colloquial "logic" are also abstractions.
> It’s turtles all the way down.
Memes and catch phrases are abstractions.
> These layers of abstractions go down until we hit our most basic axioms about logic and reality.
Reality too is an abstraction. Luckily all humans run Faith, and it runs invisibly, otherwise I suspect we'd have not made it this far. Though, it now seems like that which saved us may now take us down (climate change, nuclear weapons, other/unknown).
> Trust, but verify.
Haha...of course, just use logic and critical thinking!
If the test doesn't do that - if it still passes even if you revert the implementation - then the test isn't doing its job.
The second one largely has to do with ego and is the one that creates insecure people who usually hold others to different standards than themselves, the first is just a realistic view of the world and thus very useful
Somehow, taking into account the state of our industry, yes. But this is not an absolute truth.
I mean, we do have the theoretical frameworks and even tools to come with solutions that allow to proof that code is correct. It’s just that mapping this "know how" with the "how to deal with the expected flow rate feature" is very uncommon.
The undecidability of Turing machines in general doesn't say you can't reason about any programs, as you seem to claim, it says there exists programs which have unverifiable behaviour.
(Although the proof provides unnatural examples, in practice simple recursive mathematical functions such as the Collatz conjecture will probably always be beyond automated analysis.)
Of course software that was not written with verification in mind can be verified correct, provided it doesn't contains things such as an interpreter or complex conditional recursion. There are static analysis tools for analysing programs in languages such as C++ which make use of ATPs (automated theorem provers). For example, most of zlib has been proven correct using Coq. Yes, verification of the entirety of complex programs is in general impossible, but parts of them can still be verified. It's the most common use of ATPs, and an active area of research!
I may be misunderstanding, but isn't part of the problem that these tools are themselves written in code and therefore subject to bugs?
perfectly readable as its Markdown on regular Github
Also performing a critical self review and ensuring you remain skeptical
Trust that a person will do the right thing, but verify that the right thing has actually been done.
I trust that you will give me the money. I also trust that you counted the right amount. But I still want to make sure you did not make a mistake.
In other words: why is it included?
How can you verify as a precondition instead of a postcondition? You can't verify anything until the act has been completed.
Nope. In reality the money doesn't exist. Amazing how many people think that somewhere a cartload of money is physically moving every month when they get paid. When was the last time you deposited anything in a bank? The abstraction is even higher than that. It's just numbers in a computer system. The abstractions work because banks are "too big to fail".
If this sounds silly it's because it is. I thoroughly recommend anyone who is confused to a) do their own accounts and b) invent a silly currency inside your accounts and start a bank for your imaginary friends. Just add trust and government support to your bank and you'll be like any other bank.
Uhm yeah they do. If they take my money and then just do nothing with it then they won't make any money from it. Banks invest your money.
(I don't think that's the main way they make money - it's probably mostly from credit card interest, but they definitely do it.)
The money you pay into banks absolutely exists in every sense.
Banks can create money when they issue loans, which I suppose you could argue doesn't exist. But they aren't allowed to create unlimited money. I'd say it exists as much as any other money exists.
There's a great intro to how banking really works today here: https://positivemoney.org/how-money-works/banking-101-video-...
Either way it's a symbolic system and it doesn't really matter if you execute it on an abacus, or in a digital computer, or by writing in a book and carting bags of coins around.
I think the parts that are missing from most people's mental model are things like the role of the central bank, and the fact that lending limits are influenced but not determined by deposits.
Yes, because then there would be a physical limit on what banks can do. They would have to invent cloning machines or alchemy or whatever to do what they currently do. In other words, it would force bank to actually do the model most people (including the author) have in their heads, namely a bank stores cash and lends it out (called fractional reserve banking).
In reality, without any such physical constraints, the banks are essentially free to "create" money out of thin air by issuing loans. That's how the money supply became 99% "bank money" and only ~1% physical money.
I guess people have hard time accepting this because it seems too absurd to be true. Also it's common to confuse money to be a scarce resource/commodity because that's how it looks like for most individuals (and that's the story usually pushed for us plebs).
If only :)
Far too often I find myself working with tests that patch one too many implementation details, putting me in a refactoring pickle
I’ve written plenty of do nothing tests in my time to be sure that management regularly got a report of tests being added.
- test that were useless, are still useless and will always be useless
- tests that are currently useless but were used in the "wtf should i write" phase of coding (templating/TDD/ whatever you want to call it).
I'm partial towards the seconds, and i like when they're not removed, because often you understand how the API/algorithm was coded thanks to them (and its often unit tests). But ideally, both should be out of a codebase.
But look! Thousands of tests and they all pass! Taste the quality!
I've also seen plenty of tests that test if a template was rendered rather than if whatever thing it actually outputs was in the output. It is just calcifying the impementation making it hard to test.
But it is a tradeoff, and a hard one as well, because if you do all things all the time, combining all variations of database with all variations of the views, then you end up with a test suite that take forever to run. Finding the right tradeoff there has not shown itself to be very obvious, sadly.
When done with that phase and my API looks relatively functional, I remove all relatively trivial tests and I write bigger ones, often randomized and property-based.
This works decently well and you do not have an army of useless tests there hanging after the process is done.
Been there. Change one tiny thing, and 20 tests fail all over the place. But hey, at least we had ~95% test coverage! /s
The more time some piece of code has survived in production, the more "trusted" it becomes, approaching but never reaching 100% "trust" (I can't think of a more precise word at the moment).
For tests it's similar; the longer they had remained unchanged while also proving useful (e.g. catching stuff before a merge), the more trusted they become.
So when any code changes, its "trust level" resets to zero at that point, whether it's runtime code or test code. The only exception might be if the test code reads from a list of inputs and expected outputs, and the only change is adding a new input/output to that list, without modifying the test code itself.
Tests that change too frequently can't be trusted, and chances are those tests are at the wrong level of abstraction.
That's how I see it at least.
It just mean that for 10 years, this codepath has not been taken (The conditions for this specific error case was not met for 10 years) :-(
Actually, it would be a good monitoring information to know which path are "hot" (almost always taken since the beginning), "warm" (from time to time) or "cold" (never executed). It could help build a targetd trust. I guess that it might be possible for VM languages (like based on JVM) because the VM could monitor this... but it might be harder for machine code
That's the point of code reviews.
Trust me; I'm an expert on never trusting myself.