Automated CPU Design with AI
arxiv.org
arxiv.org
The toy example in the paper is an 8-bit adder. So given many examples of 8-bit addition, their procedure constructs a possible logic circuit for it. Then they still have to pass it onto standard logic synthesis tools.
The RISC CPU example is the same, which means the resulting logic is not even a pipelined design. There are NO microarchitectural design optimizations (the authors propose this as future work, somehow).
So given the C++ simulation I/O test traces of some CPU design, it is possible to extract a logic circuit for it, one that is not blowing up exponentially--mainly because the proposed algorithm is enough to separate a control and data path (ALU)--but yet it does not give any consideration to architectural optimizations like pipelining, caches, or anything else in the standard CPU architecture book by Patterson & Hennessy.
The main point is that something like this could be used to bridge a design specification to an unoptimized logic function, so this is purely "functional" as opposed to architectural. It's a step but does not reflect the rather grandiose title of the paper.
> We verify our output netlist on the FPGAs and tape out the chip with 65nm technology. The automatically designed CPU was sent to the manufacturer in December 2021.
They compared performance of 80486 with their tapeout at 65nm (20 years of technology betterment). Judging from this, I suspect they may introduced long delays into some of the circuitry.
They also "cleverly" avoided exponential blow up of any BDD representation by not implementing any multiplier or divider. Their ISA is RISC-V32IA, which does not include multipler.
AMD used ACL2 for complete symbolic verification of their FPU and discovered that their 50 million tests test suite still allowed for some bugs to pass.
Here's the same problem - in their approach, a function is specified by some of input-output pairs and their algorithm discovers BDD representation of such function. These input-output pairs may be incomplete to specify function fully, giving the opportunity for something akin to FDIV bug.
> [...] This approach generates the circuit logic, which is represented by a graph structure called Binary Speculation Diagram (BSD), of the CPU design from only external input-output observations instead of formal program code. During the generation of BSD, Monte Carlo-based expansion and the distance of Boolean functions are used to guarantee accuracy and efficiency, respectively. By efficiently exploring a search space of unprecedented size 10^{10^{540}}, which is the largest one of all machine-designed objects to our best knowledge, and thus pushing the limits of machine design, our approach generates an industrial-scale RISC-V CPU within only 5 hours. The taped-out CPU successfully runs the Linux operating system and performs comparably against the human-designed Intel 80486SX CPU. In addition to learning the world's first CPU only from input-output observations, which may reform the semiconductor industry by significantly reducing the design cycle, our approach even autonomously discovers human knowledge of the von Neumann architecture.
The von Neumann (and Mark) architectures have an instruction pipeline bottleneck maybe by design for serial debuggability; as compared with IDK in-RAM computing with existing RAM geometries? (See also: "Rowhammer for qubits")
(Edit: High-Bandwidth Memory; hbm2e vs gddr6x (2023) https://en.wikipedia.org/wiki/High_Bandwidth_Memory )
Hopefully part of the fitness function is determined by the presence and severity of hardware side channels and electron tunneling; does it filter out candidate designs with side-channel vulnerabilities (that are presumed undetectable with TLA+)?
I.e., they manually designed the CPU but left a (large) number of parameters open, then used AI to find an optimum for those parameters.
So anything the AI did was completely correctness-preserving.
Note that this may sound like it's a small achievement, but keep in mind that for modern CPUs the search of the design space is hugely important, and probably the reason for the success of e.g. Apple's M1.
As I understood it, component placement and routing was already done somewhat automatically?
Just wondering how far from human design-space you could end up with this.
If this actually worked, it should be able to cough up a 6502, 6809, 8051, etc. as well since they are so much simpler--especially since they even mention a Commodore 64.
The fact that they don't do this stinks very strongly. There are other concerning signs in the paper as well.
I think of lots of reasons to do it with a riscv
* lots of excellent simulators and emulators
* great tool chains
* both software (Verilog, VHDL) implementations as well as hardware
* regular, compact instruction set (no condition codes)
Using anything besides RISC-V would have been an order of magnitude harder.Yes. 6502 was quite cheap for the day so is much more optimal for cost than most designs. The 6809 was done fixing the mistakes of the 6800 and it's implementation is much more orthogonal. The 6800 and 8051 are probably the best documented. All of them have extremely long lived tool chains and support. Pick your optimality.
In addition, then "Why should it produce a RISC-V design?" RISC-V is definitely sub-optimal on quite a few fronts.
If a system is doing actual CPU design, as claimed by the paper, those designs (6502, 6809, 8051) are a simple sanity check. The designs are extensively documented to the point that we have web pages that simulate them down to the transisitor. You should be able to provide a "relatively" small input and get back a compatible design as an output. A 6502 has only 3500 or so transistors. That's on the order of the complexity they claim in the paper.
This would prevent someone like me from saying: "You basically stuffed a RISC-V design into the training set, managed to launder it through ML/AI to get the computer to cough it back up, then deployed a legion of humans to patch the result suffciently that it could be called "Linux compatible", and finally barfed out a publication with 6 pages of link references in a 12 page paper."
Here's the touchstone for whether AI is doing chip design: "When AI can distinguish between control plane and datapath and synthesize and place them differently, AI is doing actual design."
The abstract is also misleading, their design is not even pipelined (in the paper they claim only preliminary results for a 2-stage pipeline).
This is unfortunate because their idea of using BDDs--which are studied extensively in hardware design automation--is a fair paper on its own.
Nevertheless, the claims about design are inexcusable and any of the cited authors, many of whom are experts in CAD, would object to such abuse of accepted terminology.
Reproducible? Good luck. Fraudulent? Very likely until someone does something similar, especially considering the history of fraudulent papers coming from China.
Does anyone know if the institutions named have a good reputation outside of China?