Cores that don't count
muratbuffalo.blogspot.com
muratbuffalo.blogspot.com
Dynamic reconfiguration: Basic building blocks for autonomic computing on IBM pSeries servers
But seriously, though,
> I think fail-silent CEEs is weaker than the adversary Byzantine failure model.
Of course they're weaker than byzantine failures. There's time locality, and the failure in themselves are not particularly hard to detect if some other core checks the results (although that obviously doesn't happen after every single computation).
experiment :
tell a bunch of people that if they form consensus about, say, a color, they all get $5 (or whatever). have them attempt to reach consensus only using the exact mechanisms of a consensus algorithm. (research what they are)
The article even says this is because of reduced feature size, not chip complexity.
This would help catch a large variety of possible errors, including but not limited to cpu bit flips, cpu bugs, memory errors.
For 99% of cases it is fine, but when you run into that last 1% it hurts.
Rad-Tolerant Lockstep Quad Core 8051:
https://www.eetimes.com/rad-tolerant-lockstep-quad-core-8051...
Tolerating rads with cold-war mcu softcores. Glorious.
This is such a unlikely thing to happen that it is likely a fault in software that is used to validate this.
If all of the logic also operates on ECC with the data, chip yields will also be improved. Say an core of the chip only produces the correct result 99% of the time, currently you have to disable that core. With ECC logic, you can still use it, as it doesn't matter if it has an additional 1% chance of a bit flip, as all of your logic is now immune to single bitflips. For mission critical logic/applications, one can scale up the ECC so its immune to more bitflips before an error is introduced.
A, B ⇒ C
Add ECC bits to the inputs, and you want Aa, Bb ⇒ Cc
Now, if you want this to detect errors made by the “⇒” part, you can’t do this as “drop the ECC bits, compute the result, compute the ECC bits of the result”.So, how do you compute the ECC bits from only the Aa and Bb bits without having to compute the C part?
Depending on the ECC logic chosen, that might be doable for bit shifts, but for addition? For multiplication? For IEEE float square roots?
Since its possible with crypto, I don’t think it’s an insurmountable problem to create an ECC code where
Ecc(a+b)=Ecc(a)+Ecc(b)
It will likely not be as bit efficient as ECC without that property. Checking the ECC would also add a bit of overhead, so at first you would want to only check on a store instruction where you have to wait to select the RAM page anyways.
It wouldn’t, but that ECC wouldn’t work for multiplication instructions, square roots, etc. Dropping floating point and multiplication instructions will correct that, at the cost of significant speed.
I also think any such ECC effectively would be a copy of computing a+b modulo some constant you can pick.
If so, we’re effectively back at “compute the result twice, trap if the two results aren’t identical”