Standard cells: Looking at individual gates in the Pentium processor
righto.com
righto.com
The grad student was Carl Sechen, advised by Alberto Sangiovanni-Vincentelli.
https://archive.computerhistory.org/resources/text/Oral_Hist...
"...we finally made the decision that we should go with automatic place and route. Neither one of those things existed at Intel and the concern was could we get it done in time and would it blow up the areas of the chip so that they wouldn’t fit and then it would all fall apart and we’d have to do it by hand. So what we did, we got an automatic placement program from a grad student at Berkeley, it was called Timberwolf and we checked it out and it seemed to do an adequate job so we had his software. He moved to MIT to work on another project and we actually had a terminal set up in his campus room where he’d fix bugs in the auto placement program as they came up. But luckily the whole thing came together and worked. There are several points in time where we’d get stuck and have to be waiting for him to fix his program. So that would take the individual cells and put them within a rectangle in an optimal situation for speed.
"...I was just going point out that if management had known that we were using a tool by some grad student as the key part of the methodology, they would never have let us use it."
EDIT: I didn't realize that Right-o had an article on i386 place and route with standard cells that also links to the panel interview. The specific areas of the i386 die that used standard cells are identified.
https://www.righto.com/2024/01/intel-386-standard-cells.html
"Bart, don't make fun of grad students. They just made a terrible life choice."
This is because of CloudFlare.
When I go to the page, I get the CF "are you human" check, which I complete.
However, every image load is also getting that check, but those checks are not presented to me - just the image doesn't load because a HTML page is being returned.
Almost as though they had rejected me before the captcha and were merely torturing me for their amusement?
Even more bizarrely, VirusTotal presented a second upload form on the captcha page... which itself is captcha-free...
So, can we all take up a collection so Ken can get a nice electron microscope, or what?
Leading edge nodes are basically black magic and are right on the edge of working vs producing broken chips.
You as a customer would never want to be in a position where you are solely responsible for yields.
Nobody wants to give away trade secrets, so everything remains proprietary and behind an NDA until it has become completely obsolete.
[0]: https://www.skywatertechnology.com/sky130-open-source-pdk/
https://www.epfl.ch/research/facilities/cmi/equipment/photol...
NB: this is obviously a simplified explanation.
1: https://semiengineering.com/knowledge_centers/materials/fill...
As an aside, few analog layouts will use the "line of diffusion" style that's common in the standard cells. And in analog one can find more exotic transistor patterns, like waffle [1], that aren't used in digital. Many things are possible, but it has to pass DRC.
[1] https://www.radioeng.cz/fulltexts/2019/19_03_0598_0609.pdf
What I can imagine, is that the foundry only tests their magic OPC algorithm with DRC clean inputs. If your mask isn't DRC clean, who knows what's coming out the other side.
If you haven't watched this video, I highly recommend it: Indistinguishable From Magic: Manufacturing Modern Computer Chips https://www.youtube.com/watch?v=NGFhc8R_uO4 It's a little outdated now, but very comprehensive and gives you an idea of how totally nuts the chip business is.
You can't build a high-end fab for a couple of billion dollars.
My project has been to design (and create) better EDA software that will simulate, optimize and therefore can form and place each individual transistor optimally to achieve lower power, higher speed and lower cost. There is only one drawback over all existing EDA software: my EDA tools must run on a ($100k) small supercomputer or FPGA cluster because it deals with a billionfold more transistors than existing EDA software and that takes more compute. It means my software is much cheaper than existing EDA software but will yield much better chips and wafers with much faster, better, cheaper and fewer transistors.
A high level overview of my software is mentioned indirectly in my talk https://vimeo.com/731037615
I'm eager to give a talk on my EDA software as well, please consider inviting me to give it?
Other researchers and companies have proven that optimizing transistor design and placement over standard cell libraries and PDKs can be done, for example:
https://www.micromagic.com/news/Ultra-Low-Power_PressRelease... was done with their own EDA software.
I am very certain (but have no hard proof) that this is what Apple did on their M1, M2, M3. M4 and M5 processors, especially their high end M2 and M5 Ultra chips.
What I'm claiming here is that humanity can design three to four orders of magnitude faster computer chips using at least two orders of magnitudes less energy making chips orders of magnitude cheaper if we only used better EDA software (CAD=> SYM=> FAB) that we use today. Moore's law is not at an end. I'd be happy to provide proof of this, but that takes a bit more effort than a HN comment.
With the approach you mention, would it involve creating "custom standard cells", or would the software allow placement of every transistor outside of even a standard cell grid? If the latter, I would have trouble believing it could be feasible with the order of magnitude of computing power we have available to us today.
The trial and error you do mostly by simulating your transistors which you than validate by making the wafers. You can simulate with mathematical models (for example in SPICE) but you should eventually try to simulate at the molecular, the atom/electron/photon and even at the quantum level, but each finer grained simulation level will take orders of magnitude more compute resources.
Chip quality is indeed limited by the magnitude of computing power and software: to design better (super)computer chips you need supercomputers.
We designed a WSI (wafer scale integration) with a million core processors and terabytes of SRAM on a wafer with 45 trillion transistors that we won't chip into chips. It would cost roughly $20K in mass production and would be the fastest cheapest desktop supercomputer to run my EDA software on so you could design even better transistors for the next step.
We also designed a $800 WSI 180nm version with 16000 cores with the same transitors as the Pentium chip in the RightTo article.
I taped out 9 mm2 test chips to test transistors, the processors, programmable Morphle Logic and interconnects.
The ultra-low power 3nm WSI with trillions of transistors anda Terabyte SRAM will draw a megaWatt and would melt the transistors. So we need to simulate the transitors better and lower to power to 2 to 3 terawatt.
There is a youtube video of a teardown of the Cerebras WSI cooling system where they mention the cooling and power numbers. They also mention that they also modeled their WSI on their own supercomputer, their previous WSI.
Also, have you checked out the OpenROAD[1] project? It’s a pretty impressive open source RTL to GDSII flow.
I went to their most recent meetup at DAC’24 and there’s a great community around the project.
I'd love to but what do you want me to elaborate on?
We started making EDA tools and simulators (CAD, SYM FAB as Alan Kay says) and designing a wafer scale integration to run parallel Squeak Smalltalk (David Ungar's ROARVM) in 2007 and we are still working on it in 2024 so I estimate 30,000 hours now. I call that very ambitious too.
>Also, have you checked out the OpenROAD[1] project? It’s a pretty impressive
No it is not pretty impressive EDA software, OpenROAD software quality is like Linux, " a budget of bad ideas" as Alan Kay typifies it. Openroad is decades old sequential program code, millions of lines of ancient C, C++ and bits of Python programs written in the very low level C language, riddled with bugs and pathes. The tools are bolted together with primitive scripts and very finicky configurations and parametric rules. Not that the commercial proprietary EDA software is any better, that usually is even worse but because you don't see the source code you can't see the underlying mess.
Good EDA tools should be written by just a few expert programmers and scientists in just a few thousand lines of code and run on a supercomputer.
So the first ambitous goal is to learn how to write better software (than the current dozens of millions of lines of EDA software code). Alan Kay explains how [1-3]:
[1] https://www.youtube.com/watch?v=ubaX1Smg6pY
[2] https://www.youtube.com/watch?v=Kj4fLRm2UC4
[3] https://www.youtube.com/watch?v=1e8VZlPBx_0
The second ambitous goal is to (learn to) design ultra low power transistor and free space optics. Learn from the best quantum physicists: [4].
https://www.youtube.com/watch?v=-dQoImLNgWs
The biggest problem is you need to get a couple of million investment just to test your software with a few tape-outs.
I only managed to invest the money for the 30,000 hours of labour so far, you can guess how many millions that's worth.
My transistor and atomic simulation software is extremely parallel but not in the limited SIMD way that GPU's are.
There are other processors, such as Ambiq Micro, but they are Cortex M4 and M55: https://www.top-electronics.com/en/apollo510-soc-250mhz-3-75...
I would argue that a solar powered computer would benefit from a 2 megabyte (could be as low as 128KB) SRAM operating system with GUI like Squeak or Smalltalk-80 instead of a Linux as you propose. We've learned a lot from the low power OLPC designs.
Thanks for the invite, I'm eager to collaborate on your solar powered computers but I'm having trouble finding your email in your githubs. Could you email us morphle73 at g mail dot com?
Also, Ambiq Micro has a 2MB MRAM (+2.75MB SRAM) processor https://www.top-electronics.com/en/apollo4-blue-plus-192-mhz... that can run on solar power (it uses ~5uA/MHz, Cortex M4) and Andreas Eriksen got LISP text editor to run on 384K RAM with the Apollo3 https://hackaday.io/project/184340-potatop and 4'4" Memory In Pixel Display.
I read that lithium-ion capacitors have much faster charging than regular li-on. https://www.tindie.com/products/jaspersikken/solar-harvestin... (and longer lifespan)
> Typically, the energy required for a full switch on an E-Ink display is about 7 to 8mJ/cm2.
>The most common eInk screen takes 750 - 1800 mW during an active update
The Smalltalk-80 Alto, the Lisa and the 128K Mac had full window GUIs in black and white and desk top publishing.
The One Laptop Per Child (OLPC) had low power LCD color screens especially made for use in sunlight and would combine nicely with solar panels.
the particular memory lcd i have is 35mm × 58mm, which is 20cm², so at 7½ millijoules per square cm, updating the same area of epaper would require 150 millijoules to update if it were epaper. the lcd in fact requires 50 microwatts to maintain the display. so, if it updates more than once every 50 minutes, it will use less power than the epaper display, by your numbers. (my previous estimate was 20 minutes, based on much less precise numbers.) at one frame per second it would use about a thousand times less power than epaper
so in this context epaper is ultra high power rather than ultra low power. and the olpc pixel qi lcds, excellent as they are, are even more power-hungry
pixel qi and epaper both have the advantage over the memory lcd that they support grayscale (and pixel qi supports color when the backlight is on)
I just googled "low power epaper" and read the summaries for mentions of mW and J
https://www.anandtech.com/show/5555/intel-at-isscc-12-more-r...
>At IDF last year Intel's Justin Rattner demonstrated a 32nm test chip based on Intel's original Pentium architecture that could operate near its threshold voltage. The power consumption of the test chip was so low that the demo was powered by a small solar panel. A transistor's threshold voltage is the minimum voltage applied to the gate for current to flow. The logical on state is typically mapped to a voltage much higher than the threshold voltage to ensure reliable and predictable operation. The non-linear relationship between power and voltage makes operating at lower voltages, especially those near the threshold very interesting.
Quickturn was bought by Cadence and now seems to be gone.
Source: I work in EDA
With the foundry provided library you get what you pay for...
Although I suppose if the problem is embarrassingly parallel, the SpecINT x #cores curves might just about reach the #transistors curve.
[1] https://substackcdn.com/image/fetch/w_1272,c_limit,f_webp,q_... via https://www.semianalysis.com/p/a-century-of-moores-law figure 1
your problem doesn't have to be ep to scale to 10² cores
i suspect it's true that compute power per transistor is dropping because thermal limits require dark silicon, but that plot doesn't show it
The logic is built out of standard gates and logic blocks like flip-flops anyway, so the overhead of using standard cells that implement those building blocks likely isn't too great.
This scheme works even with just poly and one level of metal, but if you have enough metal layers than you can run them through the cells themselves. You just have to avoid the vias that take the inputs and outputs down to the transistors. You have an additional gain if you flip every other row of cells so that the PMOS of two rows have the Vdd rail overlap and the NMOS of two rows have the ground rail overlap.
*"Below" meaning "on the opposite side from the objective" - you illuminate _through_ the sample.
Biological microscopes, on the other hand illuminate the sample from the back side (which doesn't work for fully opaque objects).
An external light works for something like an inspection microscope. But as you increase the magnification, you need something like a metallurgical microscope that focuses the light where you are looking. Otherwise, the image gets dimmer and dimmer as you zoom in.
That said: we have video players in our pockets. Sure, dissecting one frog might be a more educational experience than watching somebody else dissect a frog, but is it more educational than watching 20 well-narrated dissections? I suspect not. I don't think we need to do either.
https://www.vlsitechnology.org/html/libraries.html
https://opensource.googleblog.com/2022/07/SkyWater-and-Googl...