By crashing and having it start into a good state, the system will at least be operational again.
By crashing and having it start into a good state, the system will at least be operational again.
If you running a thousands of programs on hundreds of thousands of machines 24-7, you're bound to run into weird edge cases at the system level.
Instead of worrying about and optimizing all of these edge cases, SOMETIMES it is better to just have a system that is tolerant of these edge cases by design.
In Java you would explicitly handle this case with a catch and take action on a "division by zero exception". In Erlang you would just let the process crash (and log the reason so bug can be fixed) and restart the state to what it was before the erroneous input, no matter what bug/case the wrong input triggered. By having this generic handling you will make a really resilient system since you don't have to handle every possible bug on a case-by-case basis.
What would be the advantage of handling this case in the Java defensive programming approach? Maybe someone will catch the error and return/introduce null to the state? Or BigDecimal.ZERO? Then you might end up in an unexpected state for all subsequent requests.
(edit) formatting
handle_zero(data, 0) -> {:error, "division by zero"},
handle_zero(data, _) -> {:ok, transform(data)}.
handle_input(data) -> handle_zero(data, data.divider).A key point of Erlang systems, however, is that they are really good at reporting the state of the system when it crashes (due to functional programming, you have the state from before the crash happened, and what event lead to the crash).
The restart is a stop-gap measure that gives you service for the system as a whole. You can then look at the logged bug report and fix the problem. But you are in control of how quickly you want to fix the problem. There is a cost to fixing a bug as well.
Bugs in production will happen in every language, and crashing one of many running processes and restarting it to get it back into a known good state sure beats having to scramble to find a fix, code it, build and deploy a new version of the code.
The way the “let it crash” philosophy is done in Erlang is that you crash individual processes first (e.g. the current HTTP request from one browser). If that keeps happening there’s a counter that crashes the subsystem (e.g. the web file listing component). If that too keeps crashing the whole web frontend might crash, but the node might still be up (e.g. serving FTP or whatnot). And lastly the node will shut down completely if the web frontend keeps crashing.
This way, the rest of the system keeps performing and serving requests even if some parts are not working intermittently.
On top of this you of course have (built-in) logging of all these errors so you can investigate them.