Assertions in Production Code? (2008)
drdobbs.com
drdobbs.com
One faction believed you should never intentionally crash the app, while another faction believed the app was going to crash anyways or maybe do something worse.
They decided to test it out and during a sprint that was already mostly about bug fixes where they would do extra testing before release, they enabled asserts, and closely monitored their crash reporter.
What they found was that the number of crashes did not change much, but the cardinality went down, significantly. The learning was code executing past a disabled assertion may be in one of n different bad states, each of which might lead to a different type of crash. They now had better high-level information about what was causing crashes (knowing which asserts were wrong) and it helped them reduce their crash rate much more than raw crashes without asserts (including cases where the crash was in iOS code, not app code).
Anyone else had similar experience ?
One of the first things I did was turn on every possible "throw on error, die on error, assert" in development (could't do it production as it'd just fall over constantly).
The constant crashes in dev meant I rapidly fixed the low hanging fruit (undefined indexes etc) and then slowly as I cleaned an area up I'd leave the asserts in once in production since at that point what I'm asserting is "I think I fixed all the insanity but tell me straight away if my assumption was bad".
Painful very painful but those areas now cause me very few if any issues.
One place I worked had a team that was very adamant about not really having much error checking. Not much of any qc process, either. Wait for someone to complain about bad data and respond. Honestly, this worked really well for small, skunkworks type projects that needed to be nimble. As you would expect, when errors did happen it was because of bad data from further up. You really had to know the system well to be productive (the cynical of us thought the developers liked this because they could look like heroes).
I prefer to error early and clearly and I got a lot of push back. To be fair to them, often times the errors would be irrelevant to the specific thing they needed and it would have been preferable to ignore and carry on. It felt like I was imposing a bunch of bureaucracy and yak shaving. But bringing on new people and scaling the amount of data going through would have been impossible without more structure.
If you're really working on something that HAS to keep running NO MATTER WHAT then you probably want to do the hardcore critical systems stuff like N independently developed systems voting on the outcome. (Ah, but developed to what spec? Now we've just moved the bugs from the code to the spec... At least hopefully there are fewer of them.)
All kinds of bizarre things start happening, usually bad. The cause is simple, most of these are written in C and C programmers rarely check to see if writes to the disk succeed.
For example, a few months ago, my Windows box would crash every time it would auto-update. I did a lot of cursing about this, as how could Microsoft do this? I eventually realized that the disk was nearly full. Cleared out a few gigs, and the auto-update started working.
This is 2018.
It's not specific to Windows, either. It happens with every OS I've ever tried it with, including Linux. No message like "disk full", or "failed to write file". Just erratic random behavior and weird things happening.
I always try to keep at least 10% of disk free.
Back in the bad old DOS days, NULL pointers pointed at the interrupt dispatch table, so running out of memory (awfully common on a 640K machine) meant you trampled all over the operating system. It was so bad that I'd defensively reboot my machine constantly while debugging.
DOS extenders saved the day, because they ran code in protected mode. I never developed code in real mode again, I just ported fully debugged code to it.
but is somewhat more credible.
Safe Systems from Unreliable Parts https://www.digitalmars.com/articles/b39.html
Designing Safe Software Systems Part 2 https://www.digitalmars.com/articles/b40.html
For debug/testing mode I usually pop a toast or dialog so the tester knows to report something to me even if the app is otherwise working fine.
You often say to yourself "Well I know this assert will never trigger but I'll put it here just in case". Then you check your analytics next day after release and get humbled.
That is exactly what asserts are for: checking for things that "cannot happen" - but do sometimes happen.
For embedded development I have often had ASSERTs log out to a serial port with the function name/line # and continue operating - this can then be connected to any logging device. This makes the system a little more robust (it won't crash, right away at least) but less correct/safe.
During development testing we often had complaints of machines freezing due to asserts, which always increased the priority of those bugs! Definitely a good thing in the long run to fix those bugs and make the system more correct. Stopping all executing may not be the correct 'safe state' for an airplane though!
It absolutely is. It's a (literally) disaster to allow code that has entered an unknown state to control aircraft functions. There is NO WAY such a design would EVER be certified by the FAA.
Here's how it's done:
Assertions in Production Code https://www.digitalmars.com/articles/b14.html
Safe Systems from Unreliable Parts https://www.digitalmars.com/articles/b39.html
Designing Safe Software Systems Part 2 https://www.digitalmars.com/articles/b40.html
The pilot often has the option of rebooting the system and trying it again, but that's dangerous. On an episode of "Aviation Disasters", I don't remember the exact details, but the pilot got a warning of a failed system. After a consult with the ground, they told him to reboot it. After a while, it failed again. They told him don't worry about it, reboot it again.
It (the airplane, that is) crashed.
Aircraft systems are designed to be decoupled from each other as much as possible, so failures do not propagate. In particular, they must not propagate to the backup system, or the "safe mode". The engineers do work hard at this, and sometimes they don't get it right, and another bitter lesson is learned.
The "safe mode" system must be physically and electrically decoupled from the normal system.
An example of getting it wrong is I've seen demonstrations on TV of people hacking in through the wireless key locks on a car to take control of the brakes. That's seriously bad engineering - not so much the vulnerability to hacking because happens, but the fact that the door lock system is connected to the brake system. It's cowboy engineering at its worst.
Boil it down very simply without going into detail. * On start-up each module allocates all its memory (it has a fixed bounds) * Each loop loop of each module has a interation bounds * Assert pre-condition and post-condition of routines and leave them in production.
• In some cases, crashing is a better user experience than proceeding in an unknown state. Usually this is on the backend when data integrity is at risk, but could be on the frontend when we're at risk of making a server call with the wrong arguments. But usually, pretty much anything that doesn't crash is better than crashing.
• `assertAndContinue` requires that you have a reasonable fallback. If code is going to crash the next line anyway, there's no point in `assertAndContinue`. In most cases, it's easy, but sometimes it takes real engineering, like building in an error state into the UI component that skips downstream code. When there isn't an obvious fallback, using `assertAndContinue` is a judgement call based on the difficulty of implementing a real fallback, the likelihood of the bug actually happening, and severity if the bug does trigger.
Even when exception handling is enabled, it's a bit silly and heavyweight if you just need to short circuit a single conditional, loop, or function. Why throw when you can break or return? And when you do use nonfatal-but-logged exceptions, what do you put in the exception handler and what makes it different from assertAndContinue besides minor errata like being specialized for exceptions?
In those cases, you definitely cannot use exceptions but I'm still not sure I'd use something like assertAndContinue. If an assertion has failed, something has gone terribly wrong and I can't imagine just continuing.
> it's a bit silly and heavyweight if you just need to short circuit a single conditional, loop, or function. Why throw when you can break or return?
You throw because it's an error condition, you break or return because that's a potential the normal flow of operation. It's not heavyweight because it's for a state that should potentially never happen.
> And when you do use nonfatal-but-logged exceptions, what do you put in the exception handler and what makes it different from assertAndContinue besides minor errata like being specialized for exceptions?
All exceptions are "fatal" to the current operation, however you define it. It could be fatal to the entire application, bringing it down completely. Or it could just be fatal to a single function, request, or even UI button click. In comparison, assertAndContinue() would continue the current operation in, by definition, an invalid state.
For most traditional uses of asserts, sure. There's plenty of edge cases you might not want to ship with but aren't fatal if they do ship however. In a game this might be something like missing leaderboard definitions. Sure, losing player scores sucks, but it's not as bad as crashing the game outright and losing their progress. It's worth crashing the game at QA time to force it to get fixed, it's worth silently logging in production to avoid even bigger losses to the player. I consider those checks a variation on assertions, and "assertAndContinue" sounds like exactly the sort of thing you'd use for exactly this kind of check. Maybe you don't call them assertions...?
> You throw because it's an error condition, you break or return because that's a potential the normal flow of operation. It's not heavyweight because it's for a state that should potentially never happen.
Exceptions take full stack traces in a number of languages. Even in those that don't, compilers usually optimize for the exception-free path and basically ignore the performance of the exception handling path. And just because it "should" never happen doesn't mean it doesn't happen frequently enough to cause perf bugs.
One example: A title I worked on had significant framerate problems if you pawed at the screen just right. The cause? Exceptions thrown from system APIs on invalid touch/finger IDs when querying finger positions. Because they invalidated the IDs before finger up events could be processed.
It could've returned an invalid state, returned the last known finger position for that ID in release builds, or done any number of faster error handling paths and it wouldn't have even been noticed as a problem. Instead, it threw, I doublechecked there wasn't any way to pre-query the finger IDs or process finger up events faster/in time, added a terrible catch statement, and just ate the performance hit.
Exceptions are heavyweight. Sometimes not so terribly so that they're the wrong tool for the job, but sometimes they are.
> All exceptions are "fatal" to the current operation, however you define it. It could be fatal to the entire application, bringing it down completely.
This is the nonfatal-but-logged case I was specifically referring to.
Again: What do you put there?
> In comparison, assertAndContinue() would continue the current operation in, by definition, an invalid state.
Any code invoking assertAndContinue has presumably taken measures to properly handle the continue path, likely by simply aborting the operation being asked of it. Just like any code that uses a try/catch construct presumably would.
For example:
try {
someVar.someMethod(params);
} catch (e) {
throw new WhateverError("More detail here", e);
}
becomes: assert(someVar != null, "Unexpected null someVar");
assertAndContinue(params.val > 0, "Expected greater than 0 val", params.val);
// etc.
someVar.someMethod(params);
It may not look very different but it adds a documentation like effect to the code. It lays out the potential failure cases before the code execution. Even if internally `assert` becomes `throw`.On my side I don't really see its point anyway. If the code is actually ready to handle such failures, this is merely an error trace, so it should just be called that.
I would never allow such a construct in any code I was in charge of.
OTOH, it’s the perfect title for a sequel of Halt and Catch Fire.
It's stricter than an error trace because if it's encountered in a development/testing context, it's a hard error and the engineer must fix it to proceed. Plain error traces should only be used for expected failure scenarios (like a third-party service times out, possibly), not situations that would be considered "unexpected" to the developer.
And if it is, then what becomes the point of the assertion?
Sure, it is possible that the crash was actually caused by bad state in the kernel, or even in the baseband processor, so we can't say for certain if restarting the app actually put us into a good state. However, the approach of isolating failures at the program level seems to have been greatly succesful at improving the reliability of the overall system. It does not seem unreasonable that when a "single" program becomes sufficiently complex, it would not benefit from a similar containment.
For a safety critical system, that isn't good enough. Put it this way - would you bet your life on it?
* Of two arguments to a function, exactly one is supposed to be non-null (since they're different ways of specifying the same thing). If both end up being non-null, assertAndContinue and just use the first one and pretend the second is null.
* When listing all items of type X in a folder, ask the server for all items and do an integrity check that they are indeed in that folder and have the expected type. If any items aren't like that, assertAndContinue and filter those items out and just display the valid ones.
Same with the second case.
Also, to be clear, both of these examples are in frontend code, where mistakes are very unlikely to cause lasting problems, since anything affecting data integrity would already need to go across the client-server boundary.
We had a similar setup. assertAndContinue will add a trace in memory. When the app crashes it will dump the trace to a file and quit. (similar to https://stackoverflow.com/questions/8233388/ios-crash-log-ca...)
On load the app checks if a dump file exists and sends the contents to our server.
So we will get the last N assertions for any crash plus other info gathered in the exit routine.
No fallback thou. Just continue until the crash happens, if it happens.
We end up with a lot of issues where developers ignore the dialogs, so we are slowly moving to more "crash" checks to catch issues more quickly.
Not failing fast and hard when the application is in unexpected state (which is what assert should check) is robust only on the surface.
I think we have the debate because it's not just about adding asserts. If you want to add asserts to existing codebase, essentially making it to fail fast and hard, you need to have a system architecture that is capable of dealing with said failure. If your codebase isn't architected as such, you will probably make things worse for the user by adding extra asserts.
I think that's what the people who are against asserts fear of, and rightfully so. But the answer shouldn't be do not use asserts, the answer should be make the system more resilient to failure first.
For example the NaCl "randombytes" function has the prototype:
void randombytes(unsigned char *buffer, unsigned long long length);
It has no way to report an error -- its contract is that it must fill the buffer with high quality random bytes before returning.And the OpenBSD function "arc4random_buf" has the prototype:
void arc4random_buf(void *buf, size_t nbytes);
So we have a potential type difference between the two interfaces for the amount of data that we can generate.If you wanted to back randombytes() by arc4random_buf() you have several options: 1. Handle the case within randombytes() where the (unsigned long long) value exceeds the maximum value of (size_t) by making repeated calls to arc4random_buf(); 2. Assert, at compile time, that (size_t)'s range falls within (unsigned long long) and there can be no problem (since arc4random_buf() similarly cannot fail); 3. Redefine the contract of randombytes() such that specifying a buffer length that exceeds (size_t)'s maximum value is invalid -- this is really the same as #2 but instead of failing to compile if (SIZE_MAX < ULLONG_MAX) the program will compile and run fine as long as the contract is not violated; The assert in this case is a guard against the violation that should never occur and helps to enforce the contract during development
With regards to production use of assert, catching a contract violation there means a lot of things have gone wrong. How likely things are to go wrong and how much damage they do when they do go wrong can vary from contract to contract. Various asserts may be used based on the such things, from compile-time to run-time debug to run-time production to an inline debugging framework. There are trade-offs and no answer for all cases.
This advice is not obviously applicable to different kinds of programming with different constraints and different environments. As with so many things: one solution does not fit all problems.
Clearly not all software is flight critical but a lot of software could be improved by treating some parts of it as such.
Very good point. That's is exactly the philosophy behind Erlang (the language and its VM, also extends to Elixir and other BEAM languages).
It has isolated process heaps and process hierarchy supervisors. That is the most critical part of the deal because it means crashes and failures are controlled and only a small part of the system would be affected, without it spidering out and putting the rest of the system and put it in an unknown state.
Microservices or just using OS processes instead of threads can kind of emulate that. But you can only have so many OS processes as they can pretty heavyweight. And with Microservices there whole other stack of stuff involved as opposed to just the language environment.
> I find these options far preferable than going into an unknown state and praying.
The idea of avoiding "unknown" state is key. That's also where crashing and restarting comes in - to get back to a known state. Of course you'd also want to log the failure so someone can eventually fix it. But, if a system is well designed, someone won't have to do it at 4am in the morning. I've seen systems crashing and restarting for weeks. Sounds terrible at first, but the alternative would have been having a service that's down completely and waking someone up in the middle of the night to fix the issue.
Some Scala examples, but they're hopefully quite language agnostic:
https://m.youtube.com/watch?v=Csj3lzsr0_I https://m.youtube.com/watch?v=keTId618iOs
$ python -h
usage: python [option] ... [-c cmd | -m mod | file | -] [arg] ...
Options and arguments (and corresponding environment variables):
...
-O : optimize generated bytecode slightly; also PYTHONOPTIMIZE=x
-OO : remove doc-strings in addition to the -O optimizations
...
It does two things (that I know of):1.) Assertions are removed. The assertion code does not appear in the bytecode.
2.) Any code guarded by a "if __debug__:" statement is also removed.
YMMV
e.g. situations where you're writing to a buffer/disk, dealing with raw pointers, etc. (which should hopefully shouldn't happen with good design, but sometimes unavoidable)
Hard assertions, for things that should never ever happen, or where crashing is preferable to continuing.
As for throws/try-catch - beyond assertion use, they also report all problems that are outside your control (like IO failing due to HDD damage, or connection dropping). Those kinds of problems are here to stay.
Not as much elegance as it is about purity. Minimizing side effects is the number one way to reduce bugs.
It should be the number one guiding principle when creating out reliable software. Which means, you simply cannot use the primitive try catch or similar construct. Don’t break flow of control. Guide it to a terminal value instead.
These days, when I see try-catch and if-else constructs (which is in most codebases) it’s clear there will be bugs over the life of the application.
It’s fine, use them, but there is a world of greatness when you ditch these faulty constructs. Just like ditching OOP constructs. All built on false premises.
Monads, not even once.
For example, today I am working on a service that:
- uses a database
- calls external APIs
- publishes and consumes from a message broker
- interacts with local and remote filesystems over a variety of protocols
All of these things entail error conditions, most of which throw exceptions in the corresponding libraries. Sometimes they are converted to Either/Maybe, and sometimes they are wrapped in "native" checked exceptions.
The important bits:
1. The type system and compiler make sure the programmer has to deal with the error conditions at some point. From a programmer's perspective, a checked exception bubbling up the stack is not very different from returning a monadic object up the stack.
2. The core of the application is entirely pure. No exceptions (in both senses of the word). All side effects are pushed to the boundaries.
In this thread, we discuss asserts, not abstractions. Monad is an abstraction. (On the other hand, many people only see tests and forget about two other solutions, which are as important.)
The real world of software engineering is enormously varied and each project faces a different set of legitimate constraints and objectives. It's more valuable to discuss how and when to use "assertions, throws, and try-catch[es]" than to pretend that they're universally superceded by some other construct.