Modern CPUs have a backstage cast
devever.net
devever.net
Not quite correct; the OpenSPARC T1 and T2 were publicly released and available by 2008.
https://www.oracle.com/servers/technologies/opensparc.html
"Large parts of this process are handled by vendor-supplied mystery firmware blobs, which may as well be boxes with “???” written in them.
The maintainers of the me_cleaner script likely have the clearest view of what is known.
Points for mentioning this! But things have come a long way since 2008. You can get Intel ME-less machines from the 2008 era. Not sure if OpenSPARC T2 has any management cores.
>The maintainers of the me_cleaner script likely have the clearest view of what is known.
Yep, absolutely. Much of what we know is thanks to the efforts of researchers like these. See also the talks on finding the 'Red Unlock' mode of modern Intel CPUs.
The somewhat surprising but true implication is that on boot, the CPU is initialized before the RAM is initialized. So there is a window of time during boot when the main core on the CPU is running instructions that cannot access the RAM. Even on register-starved x86 it is possible to write code without using RAM, but it certainly seems more convenient to me to treat the cache as RAM.
Documentation for a special compiler that compiles to code that doesn't use RAM: https://github.com/wt/coreboot/blob/master/util/romcc/romcc....
This kind of low-level work is significantly more complicated than even most kernel developers realize - hence the need for articles like OP. Ditto for anything on large (more than single-board) systems. The intersection of the two was, frankly, a bit exhausting. Just keeping track of all the moving parts and their respective states induced a cognitive load that made debugging other already-hard problems that much more difficult. My hat's off for anyone who has kept on doing that stuff longer than I did, or who has to do it in an environment where vendors are keeping so many secrets.
My university (University of Colorado Boulder) bought one of a very few SiCortex systems ever sold. As an undergraduate I competed at SC07 and SC08 in the Cluster Challenge competition.
Our coach was a CU facility member, Doug, who also was responsible for the SiCortex box we bought. At SC08 he told us that another one of the teams was competing with a SiCortex box. We knew that they would win the LINPACK part of the challenge, but we didn't know that LINPACK was basically the only thing they managed to get working.
We had also heard rumors that SiCortex was in trouble financially at that time. When we were walking the show floor at SC08, we came across the huge SiCortex booth, which had 10 or so machines of different sizes (I believe the smallest was a 64-core workstation and the largest was a ~5000 whole rack system).
I remarked to Doug that SiCortex didn't look to be in such bad shape.
Doug turned to me and said, "25% of the machines SiCortex has ever made are in that booth".
The SiCortex idea was like VLIW. On paper the numbers look great. On highly optimized synthetic benchmarks it looks good. On real world code you find out how hard it is to get good performance.
Also the machine positively sucked for linear integer code - like, say, compilers or OS kernels. One of the first things customers would do, naturally, was compile. Bad first impression. Also, it was nearly impossible to get a Lustre MDS for a thousand-node cluster (which is what the biggest machine was) to run for any length of time without falling over, because Lustre was designed around the assumption that the MDS would be bigger and beefier than anything else and have "poor man's flow control" in the form of a relatively slow network. In our case it was exactly the same and completely unprotected because the interconnect was the fastest part of the system by quite a margin. That was my nightmare for those two years. I've heard that Lustre has since added some flow control ("network request scheduler") but I was never able to benefit from that. PVFS2 worked better, and Gluster (which I worked on for nearly a decade afterward) would probably have been better still because it's more fully distributed and less CPU-hungry.
The reason I mention all this is that there's an important lesson: building a system with a very unusual set of performance characteristics is a terrible idea business-wise, because people won't be able to realize its potential. Not even in a fairly specialized market. They'll just think it's slow. Unless it's truly bespoke, literally a one-off or close to it, nobody will want it.
P.S. I actually had to visit CU-Boulder to debug something on that machine, with the aforementioned Doug. It became one of my favorite "war stories" from a 30-year career, but this has gone on long enough so I'll skip it.
TL;DR, the last line is "Once all of the above is completed, the processor will be able to successfully fetch instructions from a boot source. You are now effectively at the same point you would have been 5 months ago, had this been a standard 750 bringup... Board bringup from this point should be very straightforward and follow established methods."
(And that being said, now I’m wondering whether you could force eviction and retainment into L3 cache on demand, to achieve something like memory bank switching…)
It's not really clear to me from the limited bits of info that I've read whether or not L3 is guaranteed to be accessible when doing CAR, but, if it is, you've got enough memory available to do a lot of stuff. (And even the L2 cache is starting to get pretty big on the higher-end current-gen chips.)
(This is why I compared to the SNES: if you have to map the SNES's RAM and [every bank of] the game's ROM, then you're looking at 4–16MB depending on the game. The SNES is pretty much the newest console whose games would entirely fit, I think.)
NB L3 is unified most of the time, but but with L2 you still need to distinguish between data/code.
But I found out what Itanium 2 did had a split L2 cache.
So I can say with some authority that it is possible to have a complete running installation of Windows 95 in 16 megs of disk space; however I had to trim it pretty brutally to fit, so, for example, I think it had 2 fonts left, one of which was used for the widgets that display window-close boxes and so on.
The problem was that the object of the exercise was the benchmark how much quicker it was — but there wasn't enough space left to install any applications, and our benchmark tool used real applications. So, although I did several days work and I did get paid for it, the effort was at heart in vain. But in principle yes you could run Windows 95 entirely out of 16 MB of cache memory, and if you had say 32 MB of cache memory it would be no problem at all.
The Zen 4-based Genoa-X will have up to 1 GB of L3 cache! You could comfortably fit Windows 95 there...
That's with them holding back and not adding any v-cache. If you stacked an extra 64MB on each of the 12 compute dies, you'd have 1152MB.
Think about that. The motherboard "knows" how to read a FAT file system from a USB mass storage device, verify it's digital signature and flash it with no main CPU or memory.
Intel has since replaced this with an 80486 in modern designs; perhaps it also is implemented in the northbridge.
After FW update binaries are located it's not uncommon to write them to a scratch flash and then reset the system. On reset somewhere in the flow the scratch flash is checked for an update and then the hardware sequenced flashing registers in the chipset are utilized to actually flash the firmware. Another reset is performed to boot from the freshly flashed firmware.
There are variations on this flow depending on the firmware implementation and platform/vendor which can simplify it but that is usually the basic idea. Various microcontrollers are definitely employed for other platforms (even on x86, such as the embedded controller though these perform auxiliary tasks to the host firmware rather than the whole thing).
Updating without a CPU installed usually is on the embedded controller itself but that's not a normal update flow IIRC.
https://www.usenix.org/conference/osdi21/presentation/fri-ke...
Four particular ones come to mind:
* The DPT range of SCSI host bus adapter cards, many years ago, had an full blown MC680x0 processor on the card.
* Connor Krukosky, who famously installed a mainframe in his basement with a console front-end processor that was a PC machine running OS/2.
* PC/AT keyboards had on-board microcontrollers running programs.
* And of course who can forget the BBC Micro's Tube?
It's the short period in history where people thought that computers came with only one processor that is the real oddity. (-:
The elegant and well adhered to OS calls made this a straightforward process, if your program ran on the BBC standalone it would work across the Tube for the 65(C)02, but for other coprocessors you had to at a minimum recompile and probably rewrite quite a bit of your code.
https://sites.google.com/site/jamesskingdom/Home/computers-e...
In a typical PC there are > 10 actual processors in the various peripheral and controller chips, and then there is the management engine (a full blown computer in its own right) or equivalent and usually almost every peripheral will have one or more processors as well.
Moreover Intel is just this week actually finally proposing removing real mode. [2] I'm a bit worried for what this means for emulation of old 16-bit Windows and DOS software under Wine (one of the great ironies that Wine can still run Win16 programs on an x64 host OS when Windows can't) - though I suspect the performance requirements of such software is so low by modern standards that emulating such programs wouldn't pose any challenge.
[1] https://www.devever.net/~hl/ortega [2] https://www.phoronix.com/news/Intel-X86-S-64-bit-Only
I did this as a sort of 'just in case' setup, was planning to put OpenSolaris on it and run things under Zones or LX zones and to run it as a backup server. Fast enough to get some work done and possibly more secure if the PSP is ever used/broken maliciously...
What I’m curious about is how does the ‘Self Boot Engine’ initialize itself in the first few miliseconds?
Maybe, a motherboard chip does the actual work, but that just invites the same question, at some point something must be initializing itself, how?
As to how that is accomplished, well, the core's logic is designed so that it reads the instruction from the reset vector in ROM on the first clock cycle after reset is released. Software has control after that.
What is doing the reading?
Are there any 'how-to-make TTL logic chips from scratch' resources that you know of?
For example, the die images and schematic of the most basic TTL chip, the SN7400:
https://en.wikipedia.org/wiki/File:SN7400_1965.jpg
https://en.wikipedia.org/wiki/File:TTL-00-die-schema.jpg
Show odd patterns and I can't figure out how pushing electricity across it will make it do anything other than heat up.
What you see on these dies can also be recreated manually with transistors, diodes, resistors, and maybe the odd capacitor. https://en.wikipedia.org/wiki/7400-series_integrated_circuit... shows further down the on die circuit. There is https://en.wikipedia.org/wiki/NAND_gate to give you further examples, plus there are links to TTL logic in general and all sorts of other gates. It should be easy to find pages that show how to recreate this stuff with actual transistors and resistors (or youtube videos for that matter).
The startup is pretty much just powering these chips and maybe having the input lines defaulting to for example "low" with resistors for a defined input and then output state (if the chips don't offer that themselves already -> chip manual). However, all that doesn't happen in zero time, same as signal propagation which needs time, but this is also part of such chip manuals to tell you about the timings, max frequencies etc. they can be operated at.
Well I haven't found any, even a how-to guide on linking together the gates on a breadboard and getting it to perform any operation that a TTL performs is way beyond the literature I've seen on the internet.
Do you know of some reference source?
or https://www.youtube.com/watch?v=sTu3LwpF6XI
and there are truckloads more of these.
There used to be electronics sets for kids with breadboards (some of them larger scale for easier physical handling then standard breadboards), and they'd had manuals with them explaining all the things. I remember that they also contained descriptions of basic gates like (N)AND, (N)OR, NOT, plus flip flops for storing 1 bit of data.
There's a huge gap between making a single gate or a few gates, and making something even as basic as the SN7400 from 1965.
I don't think I'm the first person in history to have noticed this, but the lack of learning materials is very odd.
The difference between the SN7400 and gates on a breadboard seems to be a black-box. At least there is not any published material in the whole world that actually explains this gap, that I could find.
Have you managed to find any yet?
As an aside, this surprised me, there's not a record of anyone, prior to this conversation, even raising the question, across all the major search engines, academic libraries, published books, etc...
People have been recreating CPU parts and I think also full blown CPUs in Minecraft with this redstone stuff (and had to "de-pig" their machines). That runs at about the same level as the above simulation, sort of.
There is also no "black magic" to a 7400, but with just superficial knowledge there is no way to break through that wall, i.e. you will have to go down that rabbit hole for quite some time before all the puzzle pieces fall into place. Then you should also be able to recreate a full 7400 on a breadboard, however, it may not look 1:1, but by then you are able to say why.
What I wonder though: What do you hope to gain from that? These are "just" transistor circuits, you provide power, they are "ready". Yes, there are timings involved: Nothing is instant, power is ramped up, signal distribution takes time. There are many 8 bit CPUs/SoCs around, just grab one of the manuals, often they are split into 2, one aiming at the hardware side of things (voltages, currents, frequencies / timings, etc.), and one at the "logic" and how to program all the parts and pieces. One thing that is shared by all of them: A reset line or at least some reset timings for things to settle in the chip to a defined state, because, as noted, things are not instant. Afterwards the initial piece of software takes over.
Can you explain how you know they are correct approximations?
From my perspective it seems to still require knowledge, of the actual SN7400 for example, to verify.
That aside, I'm still wondering what you are on about. The "internal logic" of something like a 7400 is fully defined by the transistors, so to understand that, you have to understand how transistors work. Luckily, we are not dealing with analogue computers, but digital/binary ones, so things can be simplified to on/off. Transistors are something like switchable resistors, ideally between zero resistance ("on") and infinite one ("off"). This state is controlled by the "base" (BJT) or "gate" (FET). Which means if this base/gate has a defined state, the output will have a defined state, and that is where the specifics of the respective 74XX chip are important -> chip manual will give all the required details. But as such there is no "start-up procedure" or anything with these things as long as they have no internal state, and NAND gates don't have that. Once we talk about flip-flops, latches, registers, and similar things, then we have state, and usually also some way to reset that state. 8 bit CPUs all come with a reset line for some reason, or handle this automatically as part of their power-up.
And that is where these "we build stuff from 74XX chips to a 8 bit CPU" video series come in, because they show all the CPU components, their wiring, and potentially also give access to a full schematic so you can follow the traces and see + understand what happens as part of the power-up.
> That aside, I'm still wondering what you are on about. The "internal logic" of something like a 7400 is fully defined by the transistors, so to understand that, you have to understand how transistors work.
What I'm on about is precisely that the 'internal logic' of the 7400 is unknown and not fully defined.
You make an assumption here, that does not seem to be backed up by any published source, "it's fully defined by the transistors".
For example, it might very well be reliant on other aspects, such as the length of the circuit traces to function correctly.
Datasheet for a 7400 type chip (or actually multiple). Go check it out.
Are you getting confused as to what a bipolar NAND gate is vs. what we were discussing (the SN7400 TTL chip)?
Might be useful to know what BJT and FET means and then you will see that standard / original TTL is always "bipolar", for example.
They are even manufactured using very different technology.
Like I said, just because it shares the same part number and is labelled "bipolar NAND gate" does not mean that it's approximately the same.
That we have an incredibly fancy chain of smaller CPUs booting bigger CPU is more surprising! Things should be initializing themselves. Those are 'just' voltage levels in the chip.
I would 'simply' have the main core of the main CPU come out of power-on-reset in a vaguely usable state, and then let regular firmware on a NAND spin up the RAM and setup everything else!
It's still not clear to me why is a second core needed to setup the first core the right way, instead of coming out of reset the right way
For example, there is some initialization step in the way between the big core and the instructions, and it is deemed easier/less risky to let software perform this initialization rather than do it in hardware and not have a fallback in case of hardware bugs. Or, if you have the fallback, then you've just reinvented the simple-core boot mechanism so you might as well ditch the hardware work.
I hate this trade-off because it means so much irritating work for the software engineer and frequently has downstream impact on other parts of the system (e.g. does PCIe link come up in time?) but the risk of silicon problems drives this design decision.
The very first instructions executed by the SBE are from OTPROM (one-time programmable ROM, set at the factory). This is just a few instructions: https://github.com/open-power/sbe/blob/master/src/boot/otpro...
The first SBE code embodied in a mutable nonvolatile storage device executed is here: https://github.com/open-power/sbe/blob/master/src/boot/loade...
Interestingly the SBE code is actually stored on a serial I2C EEPROM on the CPU module itself, not on the host. This is quite unusual from an x86 perspective where there is usually not any mutable nonvolatile state on the CPU itself.
POWER9 is also a little unusual in that just powering on the CPU doesn't cause it to do anything by itself. You have to talk to the CPU over FSI (Flexible Service Interface), a custom IBM clock+data control protocol which is mastered by the BMC and used by the BMC to send a 'Go' signal after powering on the system's power rails. This FSI interface is basically the CPU's "front port" - think of an operator front panel on an old mainframe. It's kind of fascinating how IBM's system designs still reflect this kind of past. Indeed, until very recently (I think this only changed with POWER8, but it may have been POWER7) IBM's own POWER server CPUs weren't designed to boot, in the sense of being brought up in a self-sufficient way - they were externally IPLed by the FSP (IBM's version of a BMC).
Basically, what we think of as boot firmware would actually execute on the FSP - which is a PPC476 running Linux - so you have the boot firmware executing as a userspace Linux program on the FSP, doing register pokes to get the CPU initialised, do memory training, etc. all remotely via the FSI master, rather than it being done on the CPU itself. It would even load the hypervisor, PowerVM. The CPU itself traditionally was essentially held in reset until the service processor had completed all this init.
So the POWER CPUs have traditionally all been initialised externally, rather than via 'bootstrapping' in the literal sense. Even memory training was done this way, all via the CPU's front port. What's really amazing is that when IBM wanted to switch to a self-boot model around the time of POWER8 (I think), they looked at their boot/IPL code, which was designed under the assumption it would be running on an external processor and initialise the CPU by probing it. Their response was to write a C++-level emulation environment for all the internal API calls this firmware uses to access the hardware. Basically, you have hardware initialisation procedures written as though running on a separate machine, and then a C++ framework all of this is hosted in which pretends that this is still the case, but allows all of these hardware initialisation procedures to be reused without having to rewrite them all. Amazingly it works. Though the size of this boot firmware is so large, it even has to have a paging mechanism for its own code while running in CAR mode(!). http://github.com/open-power/hostboot
With POWER8/9 since IBM moved to a self-boot model, the FSI interface is often just used to send the 'Go' signal, and I believe by the BMC to query temperature information after boot. But the CPU won't do anything until you send that signal. It's kind of cute, really. Sending the 'Go' signal starts the SBE running. And yes, this means that all of the hardware initialisation procedures which were written to be used from a separate service processor are now never used that way.
This kind of 'front port' based design to system 'IPL' in which a service processor just pokes the CPU remotely to initialise it is pretty fascinating. It's like the idea of automating what a human operator would have done to initialise a mainframe long, long, ago (makes me think of the autopilot in Airplane). Though of course the amount of stuff you'd have to do flipping switches on an operator panel to initialise one of these CPUs may take a lifetime if done manually...
So this is an example of how IBM's technical history very definitely lives on and is reflected in their systems engineering today.
Worth noting an advantage of this FSI interface is that it's also effectively a hardware debugger interface. In fact I believe the Talos/Blackbird systems ship with the 'pdbg' hardware debugger tool accessible via BMC SSH shell. So these systems effectively have a hardware debugger built in, which is just a Linux userspace tool which pokes the CPU's debug registers over FSI. https://github.com/open-power/pdbg
The idea of 'front ports' seems pretty rare nowadays in most SoCs. Or rather, it actually does exist: it's called JTAG. In that regard IBM's traditional boot-via-FSI model isn't so different from the idea of initialising a chip at boot using JTAG. ...I'm fairly sure I've heard some horror stories in which some systems actually do this.
The reset pin "master core release" appeared with POWER9 to support BMC-less (IBM docs lingo: SPless) boot process, and depends on tying JTAG_TMS pin high to enable autostart mode (OpenPower and FSP tie that pin low and use FSI magic write to trigger SBE)
I believe the actual "go" signal is a pin separate from FSI since POWER8. Even in the BMC-less bootstrap (documented for power9 but afaik not used by anyone) you essentially need to have a bit of logic on motherboard responsible for sequencing startup and only releasing the cpu reset pins once power is stable etc.
https://github.com/openbmc/openpower-proc-control/blob/maste...
> The very first instructions executed by the SBE are from OTPROM...
How is the very first instruction executed by the SBE in an un-initialized state?
Thanks, that's what I was trying to ask about.
How does the 'hardcoded state machine in the RTL logic' in the SBE execute the very first instruction?
PCIe in particular is literally a packet-switched computer network - it has a physical layer, data link layer, and a transaction layer which is basically packet switched. There are even proprietary solutions for tunnelling PCIe over Ethernet.
Each EV7 computer had, instead of normal BMC, a bigger management node connected to 10MBit ethernet hub (twisted ethernet, fortunately :P), and this network was then connected to things like I/O boards, power control, system boards... including to each individual EV7 CPU. Each so connected component had a small CPU with ethernet that was responsible for interfacing their specific component to the network, and when the system booted part of it involved prodding the CPUs over ethernet to put them into appropriate halt state from which they could start booting.
As someone that used to do embedded, there is a reason i felt most at home in erlang and elixir.
Their processes that share nothing and use message passing was really close to how it looks to build and code for an embedded platform.