>
write pseudo-random garbage to the entire I/O register space, and check that the hardware never gets into a state that is not recoverable by either a simple power cycle, or, failing that, a factory reset.Anything with fuses or an PROM is inherently vulnerable to this problem. If the EEPROM is for user configuration, I agree, it should always be recoverable from total corruption. Many buggy UEFI firmware implementations are anti-patterns - the BIOS can be bricked by writing or deleting UEFI variables - they should not exist.
> It's totally okay for the hardware to be bricked by a single stray write?
Yes, it is.
Of course, not for user visible configuration, unfortunately, it's always true for low-level hardware configuration. You overlooked an important issue here - many devices have EEPROM for OEM configuration, not user configuration. The EEPROM stores low-level, critical configuration, such as I/O voltage, clock frequency, device serial numbers, etc. and by definition, a misconfiguration is often unrecoverable.
Toggle the wrong fuse bit, and you're dead. Any embedded programmer can recall the moment when they write the wrong data to the One-Time Programmable ROM, or accidentally disable the main crystal oscillator, and in an somewhat amusing case (saw it on a mailing list), turning on chip's Secure Boot feature without writing a public key.
Given an infinite number of random inputs, all OEM EEPROM settings will be corrupted and brick your hardware, inevitably. These OEM EEPROMs are exposed to the external world, just like normal registers. Ultimately, it's the responsibility of the device driver authors to protect them from modifications. Certainly, a hardware manufacturer should also take some responsibility and make such corruptions a bit harder in the field - for example, many hardware provides an explicit lock/unlock feature for protecting low-level configurations and registers. But it's still the responsibility of the device driver to lock them down.
In this case, Intel indeed provides an option to lock down the content of the EEPROM according to the article.
> The good news is that both bugs have been fixed. The e1000e hardware was locked down before 2.6.27 was released
But in previous versions of the Linux kernel, the lock was unused and EEPROM was left open. In this case, I'd say the hardware manufacturer didn't do anything wrong, it's the negligence of driver developers that allowed it to occur. Just like the articles' conclusion,
> the e1000e driver should never have left its hardware configured in a mode where a single stray write could turn it into a brick.
If the hardware only supports one possible mode, where a single stray write could turn it into a brick, then it's the fault of the hardware. But it's not the case here.
The only potential missstep of the hardware I can think of, is leaving the EEPROM in a writable state as its power-on default, instead of requiring an explicit unlock sequence. But I haven't checked the code or datasheets in question, so I can't tell whether it's the case. But still, RW-by-default, RO-by-request is common in many hardware devices (also software), and it's difficult to say it's a fault.
----
Update 1: I examined the e1000e driver from 2008 [0] in question. The situation is a bit different than I previously thought. According to the fix [1], the actual reason behind EEPROM corruption was not simply writing to the I/O memory in itself, but the race condition it creates.
> The EEPROM corruption is triggered by concurrent access of the EEPROM read/write. Putting a lock around it solve the problem.
In other words, if the commit description was correct, in the majority of cases, the EEPROM was corrupted not because of the write itself, because of an inadvertent write in the middle of an EEPROM read, and the solution was adding a spinlock before performing a EEPROM read/write.
If this is the failure mechanism, I'd say both Intel and the device driver are innocent, it's just a side-effect of an unforeseeable accident.
----
Update 2: Further analysis of the EEPROM read/write sequence in e1000e.
Upon device initialization, all registers and its EEPROM is mapped into the I/O memory by ioremap(). And to write data into the EEPROM, four different methods are used according to hardware types, this includes e1000_write_eeprom_eewr(), e1000_write_eeprom_ich8(), e1000_write_eeprom_microwire(), and e1000_write_eeprom_spi() [2].
Before writing to the ich8, microwire or spi variants, e1000_acquire_eeprom() must be called to explicitly unlock EEPROM access. And interestingly, in ich8, the EEPROM was first written to a shadow RAM before it's checksumed and committed to EEPROM - a good practice and robust design. Also, after write is completed, access is revoked by the device driver.
In other words, both Intel and the device driver are careful enough, most of the time, there's nothing wrong with the device driver.
Unfortunately, for Intel 82573, the only write to write EEPROM is by e1000_write_eeprom_eewr(), which writes straight into the EEPROM from the EEWR register, without any initialization sequence.
Verdict:
* The e1000 driver in Linux (2008) was well-written and contained no mistakes. My initial assumption was incorrect.
* On other e1000e devices, Intel offered reasonably robust protection against inadvertent writes. I'm not sure if these devices were also the victims, but if so, and if the commit message was correct (it's the read/inadvertent write race condition that ultimately corrupts the EEPROM), then neither the device driver nor Intel was at fault, and bricking was purely an unforeseeable incident.
* On Intel 82573, Intel's design flaw of allowing an EEPROM write without any initialization sequences is directly responsible for bricking the hardware.
[0] https://github.com/torvalds/linux/blob/78566fecbb12a7616ae9a...
[1] https://github.com/torvalds/linux/commit/78566fecbb12a7616ae...
[2] https://github.com/torvalds/linux/blob/78566fecbb12a7616ae9a...