Pentium-III autopsy
sciencystuff.com
sciencystuff.com
The electron microscope images are all cross-sectional, because it appears he doesn't have the equipment to do surface etching, and just cleaved the chip. I've not generally seen good sectional images around though, so it's definitely an interesting look.
http://www.flylogic.net/blog/ has a lot of stuff about depackaging and reverse-engineering chips, as does "Dr Decapitator" (http://decap.mameworld.info/), who decaps old arcade ROMs, and then extracts their actual data from micrograph images to produce romfiles for emulators.
Edit:
The Sparkfun Saga of the Fake MCUs:
Part 1: https://webcache.googleusercontent.com/search?q=cache:kMgE8B...
Part 2: https://webcache.googleusercontent.com/search?q=cache:mEZ-8g...
Part 3: https://webcache.googleusercontent.com/search?q=cache:3Tlcu2...
(Links via google cache because they seem to have broken their old news URL structure)
Edit^2: I forgot I had this old image of a System-in-Package radio module that I made myself (Digital camera through optical microscope at, iirc, 20x)
http://metavore.org/faff/chip.jpg
The thick black lines at the bottom are millimetre markings on a ruler. The processor is at the centre, and the various other modules are SAW filters (https://secure.wikimedia.org/wikipedia/en/wiki/SAW_filter#SA...)
Chipworks is a company that reverse engineers chips for analysis by competitors. Their websites has lots of cool SEM cross-sections: http://www.google.com/search?cx=w&q=chipworks+SEM&tb...
Even though I design circuits using these devices, I still find the SEMs mindblowing.
Instead, drop the whole chip in a test tube of fuming nitric acid at room temp, and let it work slowly overnight. Anything with fuming nitric acid is dangerous, but this is much safer than the hot plate method. See this tutorial by Travis Goodspeed:
http://travisgoodspeed.blogspot.com/2009/06/cold-labless-hno...
The pic of mine had a simple brazed metal 'lid', which I managed to separate by scraping out the solder with a scalpel, and then finally a hot-air gun to melt the rest and pop the top off.
I'm sure I've seen references to HF (Hydrofluoric acid) used in depackaging, but since it also etches silicon, I'd imagine you'd need to be pretty amazingly accurate with it.
You need the HF to etch the top layers of metal/silicon from the die. These layers are above the actual circuitry, and help increase security so that an attacker can't steal IP by using nitric acid and microprobes or UV light to modify the operation of the IC. Flylogic has a special technique by which they selectively etch areas of the metal layers, but not the circuitry below.
Edit: if the torrents aren't being seeded anymore, you can watch the video here http://www.podcast.tv/video-episodes/24c3-2378-mifare-282189... or download from ftp://media.ccc.de/congress/24C3/matroska/24c3-2378-en-mifare_security.mkv
IIRC there's a chapter on it in "Security Engineering" by Ross Anderson, too: https://www.cl.cam.ac.uk/~rja14/book.html (whole 1st edition is at the bottom of the page)
Edit: The story of Chris Tarnovsky and his work in satellite TV smartcards is also a pretty good read - http://www.wired.com/politics/security/news/2008/05/tarnovsk...
Register allocation is one of the more expensive phases in a compiler, and register allocation on the Intel instruction set is particularly hard because it has some few registers. It's kinda ironic that internally modern chips have zillions of registers. There's a fat chunk of software that squeezes a program in 8 registers and then a fat chunk of silicon that expands into however many registers the chip has. Not only is this wasted effort, the extra silicon costs Intel in terms of power consumption and it one reason at ARM are pWning them on low power platforms.
But it gets a bit more complicated. I'm talking about (and your references to compilers suggests that you are as well) architectural registers; the things assembly programmers and compiler writers see.
The actual register file in the CPU has many more registers on out-of-order CPUs like the Pentium Pro and later; these are physical registers, and are mapped to the architectural registers dynamically by the register renamer. This is done to prevent false dependencies from preventing parallel execution of instructions in the same instruction stream.
The work done by the compiler to map live variables into architectural registers is mostly orthogonal to the work the CPU's instruction scheduler does to map architectural registers to physical ones. In some sense, more architectural registers are better, but there are diminishing returns, and eventually other considerations make more worse. x86-64's 16 architectural registers is a pretty good compromise for current generation CPUs.
Your comparison between ARM CPUs and x86 is somewhat spurious, since the instruction set defines the architectural registers: 16 for x86-64, 32 for ARM. A given out-of-order implementation can choose whatever number of physical registers gives optimal performance, and given the stakes and available engineering resources, both ARM and Intel/AMD will choose optimal (or exceptionally close to optimal) numbers for this.
Long story short, the number of architectural registers is an extremely small part of why ARM is or isn't "pWning" Intel on low power platforms.
Intel was quite correct in predicting that we would continue to get faster transistors, and the Pentium 4 was absolutely the right design to take advantage of faster transistors. The story of the past decade of CPU architecture is all about heat becoming a primary design problem for the first time in history.
> Pentium 4 was absolutely the right design to take advantage of faster transistors
No, it was a horrible design in many ways, irrespective of those constraints and any wishful thinking about the power wall, and it gave the advantage to AMD for years. Not a big surprise as it was designed by Intel's B team. The Pentium M which has been the microarchitectural basis for their subsequent desktop processors was designed by the A team.
To be fair, nobody has ever designed a CPU based on current process technologies. Design teams are always saying "we think the process technology we'll end up working with will look something like this...", and based on past history there was no reason to think that the power issue wouldn't get solved like it had the past N times.
Historically (pre-P4) this was a great tradeoff. The Alpha 21064 was a barn-stormer; it ran at nearly twice the Mhz of its brainiac competitors at the time. This more than made up for a deeper pipeline and simpler execution units.
What killed the P4 was that they hit the power wall. Moore's law scaling of gate size no longer resulted in a corresponding scaling of power usage.
It is hard to blame the P4 designers too much for this, since people have been prophesying the end of Moore's law since the mid-80s at least. They happened to be designing when it actually happened on power. Current designers should look to that lesson when relying on increasing density...
Of course, these lessons are pretty useless, because much like a stock market bubble, Moore's law continues on a given metric until it doesn't. And it is always obvious in retrospect.