Silent Data Corruptions at Scale
muratbuffalo.blogspot.com
muratbuffalo.blogspot.com
Anyone know of a suitable program for such a process?
When a node runs a random program, it doesn’t know if the output is correct. But it could report the seed and result to a central database.
Then, if you had a several nodes, and two nodes ran the same seed and got different output, that would mean something was wrong and needed investigation.
There are also programs for reducing such programs down to a minimum test case. So once a discrepancy is found, it can be reduced to some small program that recreates it.
I once worked on a compiler backend and a CI job generated random C programs and compared x86 output. against the novel cpu simulator. Any discrepancies found were auto reduced by these tools and then a ticket was automatically created. Lots of our bugs were found and fixed this way.
(My memory is we used C-reduce for the reductions. I can’t remember the tool we used for generating the test programs, but there are several.)
Using "idle" cycles is not a great idea though:
- They may seem "free", but in fact you would end up using more power: CPU turns itself off during idle time, and you'd replace that with an intensive process. Power (and, as a consequence, cooling) is one of the main costs of a data center.
- Machines that are properly utilized (close to 100% CPU utilization) would get less coverage, and those are the ones that need it the most.
So it is better to allocate a certain percentage of your CPU budget to self checks, based on risk and sensitivity of the tests. And have some easy way to put a machine under stress testing if it is suspected of having rare memory or CPU errors.
I presume in theory this would be using idle time only but in practice it’ll slow down your system a lot.
I agree that's a good idea, but don't people usually disable asserts in production?