Booting Linux using UEFI can brick Samsung laptops
h-online.com
h-online.com
Now I know in this particular case Samsung wrote both the driver and the firmware, so it is easy to point blame here.
But more broadly, should hardware be built/designed so it has a "fail safe" mode where it just won't allow its self to be damaged by OS/software instruction?
PS - Let's just assume BIOS/uEFI firmware updates are off the table for this discussion. Since many modern uEFIs allow you to disable user updates entirely.
I don't think you can expect both to be true nowadays. For example, chances are that your battery charging 'hardware' runs some software. That software, if replaced with faulty software, can destroy your batteries and with it, maybe even your motherboard (through fire, acid leaks, and the like)
This applies elsewhere, too. Historically, we had the 'killer poke' (http://www.6502.org/users/andre/petindex/poke/index.html; variants at http://en.wikipedia.org/wiki/Killer_poke).
Nowadays, it is rumored that buggy baseband firmware for mobile phones can fry the hardware.
I do not think it is feasible to prevent all of these in hardware. Because of that, you must accept "if you can update all software through software, you can brick your device through software".
So, to get to "provide a way to recover from any failure mode that can result from software", you will need to have some unmodifiable software on the device. You also will need that software to allow updating of some firmware and to be free of bugs. I think that is possible, but not economically feasible. Why would anyone spend even a week on bug-checking that earliest running code on a device that will be sold for only six months? That would be giving up the bestselling 4% of the sales cycle.
I'd put that in the same category as writing robot arm control software that throws the arm off the table, for instance.
There is an inherent level of "trust" there. Which makes sense. But the question is: should that "trust" extend to permanent damage to the hardware/firmware?
For example, a lot of GPUs allow the OS to control fan speed. You can literally set the fan speed so low the GPU will over-heat and damage its self. That isn't a "bug" that is a "feature."
Again, we come back the expectations/relationship.
So what can be done?
There are systems with firmware that can automatically detect failed updates or corrupted firmware, or where a failsafe firmware loader can be triggered by a jumper or related request, and that can then perform a reset and (re)load of replacement firmware.
Without requiring a test harness or JTAG access or other equipment.
In various of these cases, there are two copies of the firmware, meaning the old firmware can be immediately accessed, or — pending successful completion — a second copy of working firmware can be generated.
In one case, a system had its firmware mostly in ROM, and had NVRAM that could hot-patch routines via an NVRAM-based vector table, and with space for replacement routines in the NVRAM. This meant that the box would always boot, and bad vectors could be detected by checksum, and firmware bugs could still be patched up to the limit of the available NVRAM.
Put another way, we know how to avoid this mess. It just costs some time and effort and money, and that can get this capability cut.
This stuff is not rocket science.
Why exactly do you blame C's memory model for this issue? The article is not overly specific about the exact specifics of the problem. Do you have a more technical source? I'd be curious to learn how exactly this driver can completely brick the mobo.
There will usually be some kind of firmware reset in the hardware, but it might involve processes that only the manufacturer can do.
I would say C's memory model can contribute towards this because if you have a buffer overflow, or change the wrong bits of memory, you might accidentally change the code, or execute data. Executing data is why buffer overflows are so serious, and allow malicious code to essentially do anything.
EDIT: looks like somebody reporting the bug did this the old-fashioned way: "Just to add, on UEFI machines that got bricked like this I removed the battery and disconnected the CMOS NVRAM battery and this restored the machine to the factory default and fixed the issue for me."
Getting Code Right And Secure, The OpenBSD Way http://www.bsdcan.org/2010/schedule/events/172.en.html
https://plus.google.com/u/0/101093310520661581786/posts/Drq9...
At least in their darkest hours, Linux users can still put on a smile:
"[...] they changed the motherboard and it's working again now. I won't try to install Ubuntu again though. The whole process took about 2 weeks."
To start, it should not possible to brick a hardware in any way...
Interestingly, the UEFI looks to me sufficiently complex to control the device graphics and input before the OS boots, I wonder if that can be used to abuse the system and do again some "on the metal" coding for high performance stuff (ie: games... and scientific things).
I think maybe that cannot be done because probably no UEFI comes with drivers for video hardware acceleration.
I've dealt with kernel dev for BIOS systems, CSM development for UEFI, etc etc. I'll stick with UEFI, even if it does still have some growing pains.
Real mode (or the lack of it in Intel VT-x) still wakes me up in a cold sweat at night.
Isn't it the case that once the kernel is fully booted, up and running, it bypasses BIOS entirely and talks directly to the hardware? Have all those bugs you mention been related to the booting process itself (constituting a relatively tiny part of the kernel)?
What we are seeing here is solid evidence of the fact that software is hard - and a lot of companies just don't have the chops to do a good job. Somehow I doubt the Surface Pro would have bugs anything like this, for example.
QUOTATION START
18. Mandatory. Enable/Disable Secure Boot. On non-ARM systems, it is required to implement the ability to disable Secure Boot via firmware setup. A physically present user must be allowed to disable Secure Boot via firmware setup without possession of PKpriv. A Windows Server may also disable Secure Boot remotely using a strongly authenticated (preferably public-key based) out-of-band management connection, such as to a baseboard management controller or service processor. Programmatic disabling of Secure Boot either during Boot Services or after exiting EFI Boot Services MUST NOT be possible. Disabling Secure Boot must not be possible on ARM systems.
QUOTATION END
0. http://msdn.microsoft.com/en-US/library/windows/hardware/jj1...
Contrast this closed hw architectures like Nintendo and Apple produce. No consumer freedom but incredible profit margins.
Note that modern macs are in fact PC's with just minor modifications so they profit from the PC economies of scale while still locking their customers into their hardware platform.
In that case it would be a security feature -- another line of defense against bootloader malware and/or adversaries in physical possession of your machine.
(I don't know how technically feasible that is; I know Canonical and others are looking at having their own cert so at least their unmodified kernels can run, but I don't know the mechanism for how that interacts with already-released UEFI machines.)
The point is that technologies like this are a double-edged sword, not evil in themselves. A similar argument is made by Linus himself for sticking with GPL v2 instead of moving to GPL v3, which outlaws certain DRM-related uses; he's more interested in providing a functioning mechanism, and leaving the policy-setting to others.
It's the result of computing following designs mandated by the MPAA/RIAA. Who on earth thought that could turn out OK?
1.The BIOS isn't portable, you can compile UEFI for any platform by porting a small base module. Everything uses this module, so the compiler will take care of the rest.
2. BIOS was a heap of 16bit assembly code, with a small memory space. It was quite hard to add any kind of complex functionality.
3. You couldn't use a boot volume greater than 2TB due to MBR. A new one wasn't added because of number 2.
4. UEFI is more like a micro kernel then a generic BIOS, and provides the functionality for vendors to write 'better' interfaces. Here the interfaces aren't actually provided in the official UEFI sample implementation, and most can be arguably called worse then an average BIOS interface. However, you have to acknowledge the capability is there.
5. It added a framework for kernel verification (unlike what some people think, it only verifies the UEFI firmware and the kernel/boot loader it loads directly). The direction that Microsoft is taking it is quite unfortunate, but it's actually a good feature.
6. It has a limited capability as a boot loader, allowing multiple operating systems to be started directly.
These are all features that aren't present in the BIOS, that UEFI has fixed. Unfortunately the current implementation has many bugs, and many features are arguably implemented badly. However is was created to fix real problems, and has been undeniably successful at that. At a linux conference there was a talk about all this, from someone quite involved with the UEFI creation process. If you are interested, the recording will probably be released in a few weeks. There is also a talk from the previous year by Matthew Garrett (the guy who does UEFI stuff in Linux), talking about all the bugs present in UEFI, which is an entertaining watch, https://www.youtube.com/watch?v=V2aq5M3Q76U.
https://bugzilla.kernel.org/buglist.cgi?quicksearch=samsung-... # but nothing else in there looks to be any closer either
But, Intel went out and invented another operating system? Why, when embedded Linux (&coreboot) already existed? with a full stack of software? Would be nice to have a web browser available from the firmware for downloading drivers, etc. I ran Linux on an Itanium at work circa 2003 so it was definitely doable.
Is the answer that MS wouldn't allow it? Or is there another reason?