Modern Microprocessors – A 90-Minute Guide (2001-2016)
lighterra.com
lighterra.com
https://speakerdeck.com/alblue/understanding-cpu-microarchit...
The presentation was recorded at the London Java Community meeting in April 2020, and a recording is available here: https://youtu.be/C4HEoBYL0yk
2018, 87 comments: https://news.ycombinator.com/item?id=18230383
2016, 12 comments: https://news.ycombinator.com/item?id=11116211
2014, 37 comments: https://news.ycombinator.com/item?id=7174513
2011, 30 comments: https://news.ycombinator.com/item?id=2428403
This link has also appeared in 9 comments on HN, featuring threads on "Computer Architecture for Network Engineers", "X86 versus other architectures" by Linus Torvalds, and "I don't know how CPUs work so I simulated one in code", also recommending a udacity course on how modern processors work (https://www.udacity.com/course/high-performance-computer-arc...): https://ampie.app/url-context?url=lighterra.com/papers/moder...
Jason also has a couple of other interesting articles on his website, like intro to instruction scheduling and software pipelining (http://www.lighterra.com/papers/basicinstructionscheduling/) and the one I liked a lot and agree with called "exception handling considered harmful" (http://www.lighterra.com/papers/exceptionsharmful/).
PDS: Why do I bet that the Transmeta Crusoe didn't suffer from Spectre -- or any other other x86 cache-based or microcode-based security vulnerabilities that are so prevalent today?
Observation: Intentional hardware backdoors -- would have been difficult to place in Transmeta VLIW processors -- at least in the software-based x86 translation portions of it... Now, are there intentional hardware backdoors in its lower-level VLIW instructions?
I don't know and can't speculate on that...
Nor do I know if the Transmeta Crusoes contained secret deeply embedded "security" cores/processors -- or not...
But secret deeply embedded "security" cores/processors and backdoored VLIW instructions aside -- it would sure be hard as heck for the usual "powers-that-be" -- to be able to create secret/undocumented x86 instructions with side effects/covert communication to lower/secret levels -- and run that code from the Transmeta Crusoe's x86 software interpreter/translator -- especially if the code for the x86 software interpreter/translator -- is open source and throughly reviewed...
In other words, from a pro-security perspective -- there's a lot to be said about architecturally simpler CPU's -- regardless of how slow they might be compared to some of today's super-complex (and, ahem, less secure...) CPU's...
See if Transmeta Crusoe is vulnerable to Spectre?
But, even if it is... keep in mind that when running x86 instructions, you still have the x86 translation software proxy layer... that means that could could grab any given offending / problem-causing x86 instruction -- when you encountered it, and recode the VLIW output from it to output a different set of native VLIW instructions -- that you knew were safe...
In other words, with a Transmeta Crusoe -- if the x86 translation layer is open source and you possess it (and can code / understand things) -- then you'll have some options there.
Which is unlike a regular x86 CPU -- where the way it decodes and executes instructions -- cannot be changed in any way by the user...
An open system with such a design would indeed be fascinating (the original wasn't open, and Transmeta was big on their patents on this stuff). More flexibility than microcode patches too.
Agreed completely!
>More flexibility than microcode patches too
Agreed completely!
The point is to leak privileged code flow.
In the latter case -- you have absolutely no control whatsoever over how the processor interprets and dispatches its x86 instructions...
I'm mainly saying that it's a problem space that both has actively shipping implemetnations (Nvidia Denver), has new levels of cache which affect performance based on previous codeflows, and hasn't been fully explored publicly AFAIK. There's probs some dragons in there in at least the pre spectre versions of that software.
This is in contrast to Transmeta where the whole system more or less ran out of the one translation cache.
Now Nvidia Denver on the other hand...
Here, let's put a link for posterity:
https://en.wikipedia.org/wiki/Elbrus_2000
PDS: Opinion: Transmeta/Elbrus/VLIW designs in general -- worthy of future study...
You can either separate your hardware (physically or virtually) or do no speculation.
Edit: originally said "outdated".
I said "outdated" in a wrong way, because this means is not valid today, which is obviously not the case.
It's interesting to see how many CPU architecture ideas that we consider modern were first developed in the 1960s, and how they took a long time to move into microprocessor.
My point is that there's still a lot of advancement made.
The wiki page on x86 has a pretty good summary of the addition of things https://en.wikipedia.org/wiki/X86 .
>"From a hardware point of view, implementing SMT requires duplicating all of the parts of the processor which store the "execution state" of each thread – things like the program counter, the architecturally-visible registers (but not the rename registers), the memory mappings held in the TLB, and so on. Luckily, these parts only constitute a tiny fraction of the overall processor's hardware."
Is each "SMT core" then just one additional PC and TLB then? I'm not sure if "SMT core" is the correct term or just "SMT" is but it seems like generally with Hyper Threading there is generally 1 hyper thread available for each core effectively doubling the total core count. It seems like it's been that way for a very long time. Is there not much benefit beyond offering single hyper thread/SMT for each core? Or is just prohibitively expensive?
>"The key question is how the processor should make the guess. Two alternatives spring to mind. First, the compiler might be able to mark the branch to tell the processor which way to go. This is called static branch prediction. It would be ideal if there was a bit in the instruction format in which to encode the prediction, but for older architectures this is not an option, so a convention can be used instead, such as backward branches are predicted to be taken while forward branches are predicted not-taken.
Could someone say what the definition of "backward" vs a "forward" is? Is backward the loop continues and forward a jump or return from a loop?
Also are there any examples of "static branch prediction" CPU architectures?
It must be noted this article discusses von Neumann architecture and alike (Harvard).