Booting Modern Intel CPUs
mjg59.dreamwidth.org
mjg59.dreamwidth.org
Except that you don't need to wait 10 ms for each AP -- you can start up the APs in parallel. There's just one small problem: All of the APs start up in the same state -- executing from the same CS:IP, and also the same stack pointer. Good luck having hundreds of CPUs stomping over each other's stacks.
Except that if you're careful, it doesn't matter -- you can even make a function call if you want because all of the CPUs will push the same return address onto the stack.
Implementing this in FreeBSD is on my "speeding up the boot" to-do list. I know it's possible though, because someone told me that they had already done exactly this in a different (non open source) system.
Lexicon for the non-x86 people: SMP = Symmetric MultiProcessing, aka more than one "virtual CPU". AP = Auxiliary Processor, any CPU other than the one which the BIOS starts up for you. IPI = InterProcessor Interrupt, how CPUs wake each other up. CS:IP = Code Segment + Instruction Pointer, where the CPU is reading instructions from.
For the startup code, you shouldn't need to make a function call. A few lines of memory-less stack-less assembly could get the CPU number and then change the stack, assuming you have a global value that gives the base of a preallocated array of stacks.
> For newer CPUs (P6, Pentium 4) one SIPI is enough, but I'm not sure if older Intel CPUs (Pentium) or CPUs from other manufacturers need a second SIPI or not.
When (and if) that became officially sanctioned behaviour is another question.
[1] https://wiki.osdev.org/Symmetric_Multiprocessing#Initialisat...
Except for that TSC_AUX (MSR that stores CPU id number) is going to be 0 for all of the cores. Unless you know some other way to get CPU number?
On the other hand, as an OS on PC-like platform you know how many cores there are supposed to be and what are their APIC IDs before-hand because they were already enumerated by firmware (which is the reason why you can do the one by one AP startup sequence in the first place).
I'm away from my hobby OS to double check, but isn't it the cae that the Start IPI includes a page number which drives CS? If you send those out one by one, you can give each AP its own code page and set the stack page based on that (either using the CS value to index into something, or as an immediate value in the code, that you modify as you copy to the page). Of course, if you do a broadcast SIPI, then all of those are going to have the same CS. Depending on how much early boot code you fancy writing in assembler, you could maybe jump into protected/long mode, find the current cpu id, and lookup the proper stack pointer without using the stack at all, and only then jump into C code? Of course, one probably has nice C functions for some of those things, so it doesn't seem nice to also have it in assembly.
mov rsp, STACKSIZE
lock xadd [currstack], rsp
Of course, the contention on that xadd is going to cost you (if not 10ms per CPU... probably?), and this presumes you aren’t using the kernel stack pointer for anything (like a stable CPU number). To fix that, you probably will need to traverse a CPU -> startup data map in assembly. But it’s a start (no pun intended), and is not as horrendous a hack as having multiple CPUs push the same return address onto the same stack.But even if you allow an entire extra order of magnitude, at one per microsecond, that's still 10000 over the course of 10 milliseconds which is plenty for this usecase, at least for now.
I double checked, and as I understand it, with traditional APIC start IPI, you get to pick the CS address to be (0-255) * 0x1000; although how much of the first 1MB of physical address space is available within that depends on the system memory map. I just start one AP at a time, and use the top of the code page as stack space until it switches to the intended kernel stack, and then that AP starts the next one. That's not time efficient though; you could pretty easily start as many APs as you've got low pages available; although the option where everybody starts from the same place and they figure it out among themselves is probably simpler; because there's never a need to wait for an AP to finish starting before starting more APs; just saying, you've got options, they don't all have to start at the same CS:IP.
It absolutely did not have the hindsight from pre-existing platforms that the platform spec effort in RISC-V has.
With regard to this:
Intel CPUs ship with built-in microcode, but it's frequently old and buggy and it's up to the system firmware to include a copy that's new enough that it's actually expected to work reliably.
...I wish companies put more effort into releasing finished, polished products. Not just Intel. I played with an ESP32-S3 recently and learned not only was the entire digi controller for ADC2 'dropped' from support due to bad silicon, but using ADC1 with DMA was fraught with gotchas and blatently incorrect info in the documentation (reported at least a dozen to EspressIf and they acknowledged).
Pretty much every hobby project I look at, people complain about how bad the ADC is, and then just proceed to take N samples and average them. It's not that the ADC is bad, it's that the underlying firmware is buggy. I suspect there's a race condition with the main CPU and a coprocessor that's using the ADC for some internal wifi stuff.
On the ESP32-S3 ADC2 is shared with Wifi, but I don't think that's the case for ADC1. I made headway by throwing away all their boilerplate code and setting up the raw peripheral registers myself. I actually prefer it that way without all the abstraction underneath. Was getting nice results at 80kHz sample rate.
Would love to turn it into a write-up as I'm not sure anyone has done this before with that chip (at least I couldn't find any sample code around). It's actually all in MicroPython at the moment to boot... but wouldn't be hard to convert to C (and might even clean things up due to the more natural fit of that language working with bitfields). What I'd really like is to make it into a DMA-based ADC module for MicroPython but I'm not familiar enough with compiling that platform from scratch, particularly on Windows.
They tend to usually test their designs a lot more thoroughly, which adds to the final cost, and even when their products do come with flaws, they're more likely to be transparent about it and assist you with workarounds and sometimes send over field engineers on-site to help you.
They also tend to be more honest about the specs, capabilities and limitations of their products in the datasheets. Finding out half-way through the design phase that some claims in the datasheet are bogus is a no-go for most companies.
If you're just doing hobby work, or tinkering, or planning to ship millions of bottom of the barrel products on AliExpress with no intention of providing any warranty or customer support for them, then it's fine to go with whatever's the absolute cheapest, but serious companies like Apple, Sony, etc. who care about the customer experience, won't risk delaying a product launch because they wanted to save 50 cents on a new unproven cheap microcontroller who's ADCs don't work right.
For ST it could be that they seem to be the most popular cheap ARM microcontrollers for hobbyists and consumer products, so it's easier to find faults in them since so many companies use them, similar how the most used pieces of software also have the most vulnerabilities reported on them.
Also, maybe ST was not a great example on my end, as they seem to have an obsession lately with outsourcing and farming out everything to the cheapest offshore location possible and penny pinching to the extreme for their consumer oriented parts. I don't blame them though, competition in the generic ARM microcontroller market is cutthroats and margins are slim and salaries are low and your major customers (Sony, Apple, Samsung, LG, etc) keep putting the pressure on you to lower your prices or threaten to look somewhere else.
One fun one was that when accessing the External Memory Interface module (think: parallel ROM) and switching from Bank 1 to Bank 0, the HDMI controller would get reset, but only when configured for two banks of 32 MB.
If I did one bank of 64 MB and just used the extra pin as CS, it worked just fine.
There are lots of similar quirks in every chip.
Are you some hobbyist or small company who sent your issue report to some generic company email address? Then your report most likely neve reached anyone on the team working on that chip, but probably reached some clueless jobsworth who didn't know what to do with it because large semi companies are highly siloed and there's no centralized management for such things, so you need direct contact with the team responsible for that chip.
If you're large customer, you then have direct email addresses of support engineers, application engineers and the engineers who worked on the chip itself, and so your reports will definitely taken seriously.
There's also another issue. There's public datasheets and eratas which are often not updated or fully transparent, and then there's confidential datasheets and eratas, which are updated and issued under NDA to industrial customers on a need to know basis. As a hobbyist you rarely get all the truthful info.
It's not ideal, but the semi industry mostly focuses on large customers who buy in large volumes of product, not on hobbyists and tinkerers.
Some die photos here: https://mecrisp-stellaris-folkdoc.sourceforge.io/clones-stm3...
The big guys still screw up, and over the journey I've noticed quite a few MCU subfamilies go EOL far before they should and it's usually due to silicon that has too many bugs in it. Maybe the big guys told them 'no' so there weren't any decent volumes on them anymore and they were forced to adapt.
Sometimes they're a bit more subtle. You've probably seen quite a few 'A' revision part numbers recently where they clearly keep the same MCU but fix the bugs. See this on other IC's as well.
For us, logistics (supply chain) and a solid support team are the highest importance. It's rare that we're locked into a single vendor due to a must-have-feature. These requirements narrow down our choices very quickly, and I'm sure it varies region-by-region (and how much $$$ you spend).
That's more or less the history of the x86, from the 8086 and on. I'd add a couple extra expletives as well. I can imagine a number of better ways seeding CPU state prior to starting it in a way they could all be neatly started in parallel.
> I wish companies put more effort into releasing finished, polished products.
Some do. It's just that their hardware is expensive and not very available outside some niches.
I encourage developers that work for me to just read the code and documentation. My advice to them is usually something along the lines of...
Think of how you would have done it as a sophomore in college completing an overdue assignment... chances are it works just like that
[1] https://news.ycombinator.com/item?id=29837884
Protected mode still has segmentation except that segment registers now index into either a global or local descriptor table (DT). Linux tries to make this transparent by setting up descriptors where logical addresses coincide with linear addresses (at least for CS and DS).
These days dreamwidth.org is harder to view and interact with than facebook.com if you don't have an account. We really need to stop linking to it and link instead to an archive.is copy or the like.
After the run-around to the archived copy I see mgj is still complaining about people being able to boot in modes other than UEFI. I'm glad these options still exist. Throwing away all the legacy computing options would remove many abilities no longer possible on modern hardware and software stacks.
I've made the suggestion to Dang in the past to just autogenerate an archive.is link for every story/link posted to HN and have that be an "alternate link" after the main one. I think it's kind of silly some people just post an archive link for every post and gets a buttload of points for it as everyone upvotes it which games the system. I think it would also be good in general to preserve the article as it appeared at the time when initially discussed and hasn't been edited, or removed, since.
I'm not sure how you get that impression, since I'm mostly talking about what happens before you get to that point. Having the CPU hand off control to the firmware in protected mode doesn't preclude the firmware switching back to real mode.
It just opened like a static HTML page for me, no problem at all.
In 2023, not so much.
It's baffling to me this is still supported, when all the other changes in the hardware basically make it impossible to run anything that old on the modern hardware (you can do that in QEMU, but you can then emulate x86 in software fast enough anyway).
Real mode exists today as a gateway to setting up all those housekeeping data tables so that once the "protected mode" switch is flipped on, the CPU will actually find code to execute.
- There might be even a cost in removing it.
- Complexity is a barrier to entry to any competitor that want to produce compatible CPUs.
In today's world, you could probably release an UEFI only cpu and few would notice. But I doubt it would save enough space to make a difference. And you'd open yourself to criticism from those few that still use real mode: this processor is fake, they'd say, because it has no real mode.
They could also claim that it’s not PC compatible. Because it literally wouldn’t be.
Apple also achieved their dream of the Mac no longer being a PC with the release of M1.
TBH (and I'm wildly speculating here, I'm not involved in embedded dev at all), if there were no other options, I believe the SW emulation would rise to the challenge to make it workable.
If people can faithfully reproduce behaviours of ages old consoles to make sure the old games' bugs are still preserved, and for free, I'm guessing someone would step up in case of x86 if there was a business need.
But since you can still use actual HW to do that, there isn't any.
Why? As a fun project with some old motherboard for example :)
It'd likely be a substantial effort to port Linux.
Page tables I am not sure either. Could you still do it like in Linux 1.0? No idea what things looked like then, but I assume much less dedicated hardware support.
outb to 0x3f8 might work.
I know there were once versions of Linux that supported running without an MMU. Those versions, with very limited hardware support, might work in this mode.
And apparently, good enough to run DOOM: https://hackaday.com/2022/12/07/a-tiny-risc-v-emulator-runs-...
I believe that it is not supported in Core CPUs, not even in those of them which were branded as Xeon E or Xeon W.
>Cache-as-RAM (CAR) is no longer a supportable feature in AMD hardware.
https://git.furworks.de/coreboot-mirror/coreboot/commit/a245...
> I'm also missing out the fact that this entire process only kicks off after the Management Engine says it can, which means we're waiting for an entirely independent x86 to boot an entire OS before our CPU even starts pretending to execute the system firmware.
I take it that that OS is Minix?
> But what verifies the first component in the boot chain? You can't simply ask the BIOS to verify itself - if an attacker can replace the BIOS, they can replace it with one that simply lies about having done so. Intel's solution to this is called Boot Guard.
Wait... How can an attacker replace the BIOS? Aren't motherboard nowadays protected from unwarranted BIOS flashing?
Say I'm an attacker and I got root on some PC (Intel or AMD) running Linux, how do I replace the BIOS with a backdoored BIOS without the user noticing?
It's the Minix kernel, I don't think the userland contains much Minix.
> Wait... How can an attacker replace the BIOS?
If you have physical access you can just attach to the flash directly and reprogram it. This is very much in-scope for various people.
Not to mention just replacing the motherboard since the CPU is socketed and could go anywhere.
Not really. You have to come up in it to boot software that expects to already be in it.
This assumes that "firmware" doesn't count against backwards compatibility, which isn't necessarily the case. Maybe Intel (or AMD) doesn't have a 100% monopoly on firmware to be confident enough that the CPU mode is an implementation detail. Or maybe some customers do indeed run their own "firmware" (maybe embedded?). No way to be sure.
I have objections to changing the way the CPU observably starts (i.e. mode in which the BIOS or bootloader starts in).
Then that's what I thought, yeah.
I don't see why you're explaining how your idea would be implemented; I'm rather saying that implementing it in that way might be prohibitive if Intel or AMD still have customers that expect the CPU to act a certain way. And these customers aren't necessarily standard desktop/laptop motherboards.
In other words, changing "what mode the CPU starts in" would be a big and observable breaking change and not necessarily just an implementation detail that can be magically worked around by firmware updates like you describe.
If I remember correctly, this required hacking around a lot of assumptions in the Linux kernel. I imagine the Windows kernel won't be that different.
If Intel or AMD bring out a CPU that doesn't support any operating system in use today (or any UEFI firmware/BIOS implementation for that matter) they wouldn't be selling many chips. Many vendors outsource their driver update tools to third parties, which in turn use tiny operating systems like FreeDOS to flash firmware onto devices; it'd suck for them to end up needing to rebuild their operating systems.
Likewise, a dedicated GPU also plays a role in the boot process, and taking away the legacy assumptions of the GPU boot ROM will probably also require flashing any consumer graphics card with new firmware as well. Then there's PXE network boot, which often still relies on separate firmware, which also brings its own expectations about the state of the CPU.
Bringing up a modern CPU is going to be a terrible hack whatever you do. I don't see why Intel would need to re-engineer their entire boot process. The current system is hacky as hell but it works and it doesn't require much more work than putting the firmware and microcode images in the right place.
I seriously doubt that redoing their entire boot process and guiding everyone from motherboard manufacturers to driver programmers on how to use the new system (and to iron out the bugs in the new process) will be more cost effective than letting all the old code work like it does today. Very rarely do complete rewrites make any business sense.
No OS assumes real-mode for the boot processor at this point. If you boot Linux on a UEFI system you'll jump straight into the kernel in 64-bit protected mode. The only time real mode comes into play is in the bringup of other CPUs (which is something that can be ignored now that ACPI specifies an alternative) and ACPI resume (which isn't relevant on systems that use S0ix rather than S3), so you could absolutely ship an x86 CPU that didn't support real mode and all you'd have to do is modify the firmware entry code. Modern operating systems would Just Work, as would hardware option ROMs.
(And enough systems no longer ship with CSMs that people aren't using FreeDOS for firmware updates any more - it's either Linux or a UEFI executable)
Is this only for OEM systems? I'm used to seeing these happen from inside of Windows, with the exception of motherboard firmware all happening inside of the setup program. It would make sense for much of that to be EFI applications nowadays, although there's not much in the way of context around these GUI wrappers to really indicate what's going on under the hood.
I'd think most of the referece documents can be discovered from that code base.
Relatedly, from the perspective of hands-on programming, the System Programmer's guide is the manual to start with: https://developer.arm.com/documentation/den0024/a/.
(In theory, with source for the GPU 2nd stage bootloader one could change things, but RPI foundation does not provide access to that source).
Should all be performed by ME/PSP now, but the Intel ME/AMD PSP can't be updated (even with new microcode) and older versions are known to contain serious vulnerabilities. Thus we can reduce all of our options down to either (a) not running with Secure Boot or (b) running with Secure Boot, but only on the very latest Intel/AMD processors. So if we have the very latest processors, then we only need to worry about validating "critical" microcode updates to them and that doesn't need to happen at power-on at all, it can happen in the ME/PSP, but we do need to know about them at the C-suite level, since it will impact purchasing of replacement gear, financial statements and risk analysis decisions.
That means we should power-on directly into the on-die ME/PSP and it, in turn, can bring up the RAM, PCI, etc. and check (online or on attached storage) for "critical" updates and, if found, emit a "distress beacon" on the network while also checking for a signal (from a resistor, the network or attached storage) that will inform it to download and update (rather than simply halting) the processor. This allows Management (the ones who have to sign SarBox, not just the system admins) to be informed of emerging vulnerabilities in their infrastructure (because of the "distress beacons" triggering indicators on their enterprise dashboards) so that they may choose to take explicit action to allow processing to continue (by emitting the signal that all ME/PSP would look for to enable critical microcode updates instead of halting).
Likewise, we need ME/PSP to also start verifying all connected communication interfaces (before and after boot) by checking digital signatures that chip manufacturers obtain from Intel/AMD --much like what developers have to do to get their code to run in kernel space on an Apple Mac. I mean, CVE-2022-21742, et. al. --not just Thunderbolt and USB attacks.
The only other rational solution really is to run with Secure Boot disabled (because older Intel/AMD processors cannot be trusted anyway --thanks to their non-updatable components like ME/PSP-- even with post-production microcode updates, even with Secure Boot).
Anyway, 11 years later, will leave this here .. https://www.extremetech.com/defense/133773-rakshasa-the-hard...
Yes they can, their firmware is in the same flash as the system firmware.
How do the hardcoded parts read anything with just an electrical current?
At the very beginning, what is keeping it in 'reset' and initiating the execution of the microcode? Through what means?
From what I understand, in the first few hundred nanoseconds it's just an electrical current that's flowing through the CPU and nothing else.
these days its more likely that there is a system controller orchestrating the bringup. potentially waiting for all the voltage converters and the clock generator to settle before raising the reset line. depending on the architecture it may not really be a simple pin, but addressed by the system board controller through the scan network.
I'm not quite sure of this, how do these controllers 'come up on their own' without any reliance on anything else?
maybe this will give you some context, its about a small controller booting a larger SOC ARM part