Calling NSA to find your encryption key after a few bits were flipped (2010)
astroengineer.wordpress.com
astroengineer.wordpress.com
At a previous job, we had a customer who suddenly couldn’t send us email anymore. When their IT sent us the server logs to “prove” it’s our fault, we saw that the one letter in the cached MX record was wrong. This was puzzling, until I looked at the ASCII table to verify that the difference was exactly one bit.
We never found out where in the name resolution process the bit got flipped. The problem healed itself a few days later when the DNS cache expired, so it wasn’t worth further investigation.
That really gave me pause how often random bits are wrong in other data.
So that's why they should have let all their chips do ECC instead of making it a premium feature, it would have been better for their brand as "Chipzilla" and had no real cost. And it's dangerous! In fact a soft-error at sea level killed an operating system update on me, lost about tens of thousands of data and money. I have standing to sue Intel, until this sentence clause in which I hereby forfeit the suit, together with requesting them to reconsider ECC (error correction codes) in all their chips as a safeguard needed due to Moore's Law, which was their business plan. Just give it a thought, Intel.
However, intel should have made ecc the standard and not just for 1000$+ Xeons.
Agree 100%. IMHO, the choice between "domestic" and "industrial-strength" should not mean choosing between different degrees of risks of failure.
That's my point, a virtuous monopoly wouldn't do that. It would allow at least some way to have both. Especially since soft errors are easier with smaller transistors.
The bitflip could have hit an OS distribution service. Yes, the target machines should check the checksum, but the distribution service could have flipped the bit before the checksum was computed.
But, yeah, ECC should be the standard. Also, Intel should be better about documenting where they've left holes in their online checksums, machine check exception implementations, etc.
No because the updates are signed (integrity).
Have you ever taken picture of anything important with your phone, like a crime, or a car accident? Have you ever called 911 or sent money?
How dare you use an unreliabke system without ECC, what if a random bitflip would cause it to send 10x more money or data woupd be lost without realtime backup!
This disrespect to users and wanky attitude is the cardinal sin of our industry, people's lives are at stake and it's their fault for trusting us.
Most errors like the bitflips we discuss are not correlated with file locations. They're typically medium errors and the design of typical filesystems will result in something closer to uniform distribution of errors. Bus errors could perhaps be correlated with the start of activity but I think that seems uncommon.
The probably of most things randomly failing at any particular instant is approximately zero. The probability goes up as you increase the size of the time window, or push things to their limits. You can trust your phone to store an important photo for a few days, as long as you don't run it through a washing machine or something. A few years? You're taking a risk. Many people take that risk, and it works out fine for them. I've had phones fail, hard drives fail, heart nerves fail. I take reasonable precautions to back up stuff I care about. I also have plenty of data that I would be bummed to lose that I haven't backed up yet. I'll get to it one day, or maybe it will get corrupted and I'll be bummed.
I don't compare my phone to a server (24x7) with 10000$ worth of data, but even if i do i would make a versioned backup every X minutes, especially before "updates".
But yes phones should have ecc too.
He had the chance to have a good solution and decided it's not worth it, that's the difference.
>and your life is worth something
Restart your phone.
Alright, next time my 80 year old grandoa is having a heart attack I am sure he will have time to wait for android to boot, then input the encryption key
Thats why i still keep a dumbass analogue phone wired to the socket - our industry is full of wankers and cant be trusted
Let's go an buy her a Mainframe and three redundant inet connections and power-lines, but don't try to sue intel because of your own stupidity to rely on a cellphone...alone!
If you trust your life to a cellphone.....your problem.
And just for your information...your cellphone is not a server nor reliable...keep that in mind, go climbing in mountains and you see the massive difference from your iphone to a rugged satellite-phone (with buttons), but even then, never ever trust a single device.
I mean, if you're at the level of manipulating the form of the word, that's already what "ambivalent" means. Strong in both directions. Bi- makes no sense as an opposite of ambi- since ambi- means "both", and bi- means "double". There is no negative element present in either word.
Do you have any evidence or reason to believe that memory error was outside the advertised error rate for the device, or due to some other wrongful conduct by the vendor? And how would you connect that to wrongful action by Intel which contributed to your damages?
I'm no lawyer but I suspect your case is not quite the bargaining chip that you seem to believe.
Can you expand on that? Typical coding schemes use redundancy to be able to detect failures. I'm no info theory expert but I presume there's always a cost. And if you can design a smart enough algorithm you can balance the cost with the actual likelihood of failure to find some kind of equilibrium. Of course the algorithm development has its own cost (NRE, time-to-market).
It's not clear whether you're talking about modern computers or dawn-of-PCs when you say "should have." It might be smaller cost now but back in the day it would have been mega expensive.
About 15 years ago, while working on Windows at Microsoft, a test machine sitting at my desk hit a kernel panic (BSOD). As was standard working on the test team, the machine was already setup for kernel debugging, and so I set out to debug it a bit in order to file a decent bug report.
Hours later, I couldn't make sense of it (I wasn't super experienced at this point). A few of the nearby devs couldn't either, and a small troop of us curious enough about the puzzler eventually escalated to the resident wizard, Raymond Chen[1]. Within 15 minutes of checking our work and poking at the machine, he traced the root cause down to a bit flip.
Ray's condition indeed. :)
Interestingly, a blog post of his was up on HN front page the other day: "The x86 architecture is the weirdo, part 2" https://news.ycombinator.com/item?id=31077912
Damn, I wish I had your kind of expertise.
256 (2 * 128) keys with 1 bit different
32,512 (2^2 * 128 choose 2) keys with 2 bits different
2,731,008 (2^3 * 128 choose 3) keys with 3 bits different
170,688,000 (2^4 * 128 choose 4) keys with 4 bits different
8,466,124,800 (2^5 * 128 choose 5) keys with 5 bits different
So to reach billions of combinations you need 5 bitflips, which seems quite high! But I guess space is a pretty rough environment :)
Without the extra factor you need 6 flipped bits to reach a billion combinations (128 choose 6 is 5,423,611,200).
Thanks!
My latest annoyance: https://www.usingenglish.com/forum/threads/three-times-as-mu...
At least we don have such problems in programming langu... Actually, never mind.
https://crypto.2012.rump.cr.yp.to/87d4905b6d2fbc6ad2389debb7...
There is also a video and a few more details here (starting at 47:38) in their longer CCC talk:
https://youtu.be/v_X0gUzGWsA?t=2858
(Slides for that longer talk: https://www.hyperelliptic.org/tanja/vortraege/facthacks-29C3...)
Sure I dont know how long the key length was, I dont know how long the encrypted string was, but surely it wouldnt have taken that long to cycle through a number of flipped bits, or would it?
Assuming no miscommunication or subterfuge, perhaps it can be explained by a large number of bits flipped by a single ray and/or a preponderance of rays, each flipping a small number of bits. If the satellite's shielding design was poor, there could have been a lot of exposure from a single event, such as a solar flare.
Or perhaps just a single bit of code was flipped and it began writing to protected areas of memory.
https://www.space-travel.com/reports/NASA_Fixes_Bug_On_Voyag...
My understanding is that many SSDs do encryption transparently. The ATA protocol even has a “SECURE ERASE” command that instructs the drive to wipe just the encryption key. This allowed even “bad blocks” to be erased securely.
Tapes at rest don't have to worry about that though.
Most filesystems do not do any error correction. ZFS, btrfs, ReFS, bcachefs are a few notable exceptions. And for the most part these schemes are for multiple device resilience - which is architected quite differently from the schemes used in an SSD or tape (yes, even tape).
> Tapes at rest don't have to worry about that though.
This isn't really true. Pretty much every physical medium at the densities used in modern time requires robust error correction because all physical media has flaws either manufactured or acquired from wear and degradation. For instance, modern LTO tapes use relatively robust 2D Reed-Solomon forward error correction similar to DVD/Blu-Ray.
It would be wonderful if they all have the feature, but I thought only ZFS was really that paranoid.
No it doesn’t.
Many journal filesystems use a checksum for log entries, but that is certainly not covering every block of the filesystem with a checksum. And that checksum only comes into play during log recoveries. Once a block is committed to disk there is no checksum in play (unless the fs has special support for it).
Some newer journaling filesystems support metadata checksumming, but that is not some requirement to be journaling. XFS has not always supported metadata checksumming, and it’s a relatively recent addition to ext4 (like last decade). NTFS doesn’t do checksums on even metadata. This is one reason why ReFS is a thing.
False. A limited number of filesystems keep checksums of data - most notably ZFS and btrfs. Some like ext4 and APFS will do it for metadata only. One of the most commonly used filesystems, NTFS, does not for either data or metadata.
> Typically it's done at the filesystem level or higher, including when using self-encrypting drives.
I don’t know where you got this idea from, but it’s basically the opposite of true.
> If single bits are flipped then they can be corrected.
Also false. Most checksums are used for error detection, not correction. CRCs as are typically used for filesystems are not particularly well suited for error correction.
Also, look into TCG Pyrite. Almost all consumer drives with SED features are Pyrite.
At least 2 people have informed you that you’re wrong. Now it’s up to you if you choose to educate yourself on this topic or remain a fool.
> Checkdisk isn't powered by magic.
What is “checkdisk”? If you’re talking about chkdsk, or Check Disk, or fsck like tools - none of those require checksums to do what they do. At a basic level they check the integrity of on disk data structures - the actual connectivity, valid counts, etc. How in your mind does a checksum contribute to this task?
Chkdsk exists for FAT32, which you already seem to admit has no checksums. How do you think it works?
> Also, look into TCG Pyrite. Almost all consumer drives with SED features are Pyrite.
What does this have to do with anything - it certainly isn’t filesystem level encryption.
Modern hardware tries to detect these sorts of things and halt before the corruption is propagated. Sometimes it succeeds, sometimes it does not.
The best checksums can reliably do is point at the software/hardware component that is at fault.
It made me wonder--what are the odds of that? What is the relative exposure area of the encryption key compared to the rest of the onboard assets which could have been mangled?
* Ionizing dose weakens and disrupts crystalline structure. Wears things out / degrades their specs.
* Single, very high energy particles-- e.g. protons-- come in at high speed and change a voltage somewhere. This can have massively bad effects (e.g. it can, for non-radiation hardened parts, cause parts of the chip not meant to be a transistor to become one shorting power and ground-- this is a "single event latchup"). Or, it can affect the operation of one or a few adjacent bits of memory ("single event upset").
How would you catch them?
Also, these spacecraft didn't start off outside the solar system. They weren't always so far away that a lone prankster would have trouble sending them messages.
Wouldn't they be monitoring whatever frequency it's using? I don't know much about how spacecraft worked in the 70s so maybe it's not practical.
> They weren't always so far away that a lone prankster would have trouble sending them messages.
That and what mlindner said is a good point. The start of the mission could've been easy to mess with.
How would they? They have a huge directional dish antenna for communicating with the probe, they can't intercept every signal on 8 GHz, they would need an omnidirectional antenna which would catch a lot of noise.
Generally that's why you ideally don't just have one key, but multiple. Ideally with voting, but even if you just replicate the key into a second HSM at a different physical location, it's going to improve your situation a lot.
At the simplest, write the key ten different places in memory and compare the read values to determine the most probably correct.
Higher level protocols would signal an error, and require a retransmission until no error is detected.
However, "majority rule" does indeed sound like an error checking mechanism designed by a non-engineer (or non-mathematician) commitee.
Unfortunately, humans tend to make similar errors at similar areas of code when given the same specs.
EDIT: Probably some military doctrine. For science you monster.
Edit: I am thinking of memory chips
Updated May 17, 2010 at 5:00 PT.
One flip of a bit in the memory of an onboard computer appears to have caused the change in the science data pattern returning from Voyager 2, engineers at NASA’s Jet Propulsion Laboratory said Monday, May 17. A value in a single memory location was changed from a 0 to a 1.
On May 12, engineers received a full memory readout from the flight data system computer, which formats the data to send back to Earth. They isolated the one bit in the memory that had changed, and they recreated the effect on a computer at JPL. They found the effect agrees with data coming down from the spacecraft. They are planning to reset the bit to its normal state on Wednesday, May 19.