Why It’s So Hard to Create New Processors
semiengineering.com
semiengineering.com
I generally see a lot of more effort and good quality tools in hardware verification than in software validation. The hardware design and verification industry is not really 30 years behind the curve compared to software development.
Hardware and software are two very different problems with their own set of constraints. It's actually more often that you see software barely tested being deployed in the wild to millions of users rather than hardware. Hardware bugs tend to get more attention because you can't really deploy a patch to fix the issue.
but this is NOT a misconception. Show them actual tools that constantly crash and lag on you, that demand checking out licence for a runtime which costs 5-100k/Year, show them environment in which we have to write our code - they will laugh at you. This was exactly the state of the software industry 30 years ago and this industry is long past that stage.
EDA tools are expensive, yes. But you are not just paying for the tools, you are paying for all the engineering support being those tools. Bugs and problems happen in any development tool and those companies are ready to pick up the phone at any time of the day to help you with anything you might face, whether it is a problem with their tool or a problem with their setup, to the point that they can deploy a new version of their software on the spot just for you.
And they have actual engineers on the other end of the line who know the product like the palm of their hand.
While this might sound superficial, it is extremely important when you are racing towards a $100M dollar deadline for your project.
Things being bleeding edge and fast evolving in certainly true for some EDA tools or parts of them but there's lots of more bread and butter stuff that never feels quite right.
Of course, I hope you are not referring to Windows or any day-to-day software by "important".
I do not know what SW tools are you using but i'm still looking forward for a SW tool that doesn"t suck.
Ignoring text editors here, because text editors are not programming environments (no debugger, no compiler, no source control, etc...)
So I can buy the tool without the support contract?
> Bugs and problems happen in any development tool [...]
So it's automatically a wash? Why bother trying to improve things if there will always be bugs, right?
> [...] whether it is a problem with their tool or a problem with their setup, to the point that they can deploy a new version of their software on the spot just for you.
I'm not impressed. This is a kludge.
> And they have actual engineers on the other end of the line who know the product like the palm of their hand.
Okay, that actually sounds pretty nice.
Probably not but it's not something you would like to. Cutting corners is a really bad idea when it comes to hardware development because once it is out there, there is not going back.
Having access to a professional engineering work force that knows how the tool works and know it will ready and interpret the Verilog and process it can make a difference between a 5 minutes delay and 2 week delay in the project.
> So it's automatically a wash? Why bother trying to improve things if there will always be bugs, right?
You are distorting my words. I'm not saying that we should be complacent and just accept incompetence.
> I'm not impressed. This is a kludge.
So how would you handle it? Tell the customer you'd love to fix their problem but they have to wait for the next formal release?
You can still get proper support for, say, IntelliJ if you want it, but tools and ecosystem have improved so much that it’s just not as important as it used to be.
To be clear, I’m not claiming that hardware development is the same as software development. It’s not. It has different problems. But the tooling _does_ seem well behind.
How many people are doing hardware verification work vs. cranking out webapps with IntelliJ?
JetBrains can invest a lot in software ergonomics because they're basically selling mass-market software. If you have sophisticated but niche software, your probably going to invest your engineering effort into the it's capability rather than its ergonomics.
Given that I’d suggest that the reason that Xilinx’s development products have such a poor user experience vs IntelliJ is NOT that Xilinx is impoverished, but rather that the market has extremely low expectations so they don’t feel the need to spend on fixing them.
This isn’t a problem unique to hardware-land; lots of verticals have incumbents who produce low quality products on large profits because, frankly, the competition is just as bad and there’s a high barrier to entry. I’m not sure what would fix this for hardware-land, though. Possibly more pressure from open source; that’s largely what did it for software development (though even before that, companies like Borland put more than the minimum effort in).
As for open source, SystemVerilog is an open standard, yet can you show me an open source simulator that competes with the free one that comes with Vivado?
Zero support for source control? SoC have no examples, why can't they systematically provide an hello world example (blink a LED) with their development boards? The IDE itself has close to zero doc and there is no help on the internet, it takes 3 freaking days for an experienced developer to figure out how to create a project.
I don't doubt that the hardware verification tools are expensive, buggy, and probably have terrible UX. However, all those things are also true of software verification tools, but where license cost is replaced by esoteria and incompleteness cost.
I've done VHDL/Verilog and tons of software, the HW tools do lag significantly in my experience.
Not everything gets verified 100% in HW, otherwise we wouldn't have silicon errata. It's all a question of how deep you want to go in verification and what your costs are for screwing it up.
But that doesn't really compare to state of the art of what's possible in software. Property based testing, symbolic execution etc all have incredible power that afaik is not available in the hardware world. I'm just a hobbyist who uses FPGA's but it's way way easier for me to ensure a function has some properties compared to a VHDL module or whatever.
But the comment came from a hardware designer, at least the grandparent one (from artemonster)?
I really hoped you were referring to the secrecy
It’s a space with a high barrier to entry, which is probably the biggest problem. Software engineering tools got good, and affordable, due to a combination of lowish barrier to entry and open source.
capitalism is at its best when optimizing industrial production (manufacturing materials)
knowledge creation is not quite an industrial process.
I would highlight the importance of open source as a key component for better tooling.
But knowledge does not quite play well with capitalist markets. there must be a better way, but finding it requires revising several 'fundamental' assumptions about how this (market) society works.
Our VDI development environment runs an OS from 2 decades ago. It's ugly and the latency is mildly infuriating. And yet, my co-workers are totally OK with it all.
Well, microcircuits are also subject to physical constraints that have been around for almost 14 billion years...
These days I design Verilog for FPGAs without a verification department. I've learned to code defensively, to use what I know is robust without creating complex corner cases. This often involves clear handshaking and dataflow. Anyway, it's more fun than fixing bugs deep in the weeds of a complex system.
https://alastairreid.github.io/CPUs are arguably the most crucial mechanical invention in history and yet the practical art of making them actually work right is shrouded in secrecy in service of greed.
On the other hand, if someone were to murder all of the currently working or retired processor designers, we'd probably have new i7s in a few decades.
From what I understand, the same can't be said of real microprocessor design - the complexities of aligning logical requirements with EE complexities, analog responses, and manufacturing process limitations are not at all captured in the 'code', not at the same level. We can hope that they are documented, at least to some extent, but we all know the general priority of internal documentation against other concerns, and the difficulty of documenting design processes.
Brings me back to 1995 when I was working on the C-Cube CL4010 MPEG-2 encoder processor. A fully custom 32-bit (with some 36-bit registers) RISC architecture.
Are we sure about that? They start out talking about how hard it is to write correct state machines in Verilog, and they follow that with a lot of skepticism about RISC-V processor designs. But most RISC-V designs use Chisel, not Verilog. Chisel is much higher level than Verilog and using it should also make it quite a bit easier to prevent many design-level errors.
But Chisel will not help you when you have some errata like "the BRAMs on these devices take 67 cycles to initialize after coming out of reset" tucked away in a manual somewhere that you forgot to read. That's just an FPGA example -- you can run a hundred simulations and your design will still fail immediately in real hardware because of such things. (Even "cycle accurate simulations" from your vendor may not account for these things.) You're going to encounter a lot of problems like this, before you even move to ASIC level flows.
Ultimately the hardware industry does need much better RTLs, because the current ones mostly suck from a language POV in a huge number of ways. But verification is still a much bigger problem in general no matter what RTL you've chosen.
tucked away in a manual
With a better eco-system, the (temporal) specification (such as "take 67 cycles to initialize after coming out of reset") such be given formally, so a tool can ensure that it's not forgotten. In principle this is possible. much better RTLs
Have you got any concrete ideas and proposals? I'm asking because (A) I agree with you, and (B) I'm very much in a position where I can influence research in this direction.Could you elaborate on what you do?
> Have you got any concrete ideas and proposals?
My recommendation is: try to build something complex, or take something from open-source that is very complex and extend it (test it, verify it, improve performance, etc.). Notice the problems you encounter, and then try to solve those problems.
I could give you a laundry list of pain points myself (why do I have to pay money for lint, and why does lint still suck???), but I think the best research is done by people trying to solve problem X, but along the way ended up having to solve Y and Z just to get to X.
I think there is an interesting C. P. Snow-like Two Cultures thing at play between SW and HW people. I think the latter don't quite get the abstraction power of modern programming languages, they don't understand that syntax, types, modularity matter. They think PLs are something like C and C++: a mess that we simply accept and get on with life. The former really underestimate the extreme scale and complexities involved in processor design and verification. Never the twain shall meet?
laundry list of pain points
Think bigger: what would you do if you have a team of 3 engineers, or 30 engineers to work on a better tooling ecosystem?You want to make sure that the hardware you are designing never ends up in a situation that you didn't foresee. So you want to make sure that during hardware verification you go through every single possible combination of inputs and you want to make sure you exercise every single possible outcome.
The quality of the hardware is not about the language you write your hardware in. It's about the tools, the methodology and the infrastructure you use to exhaustively test your hardware.
And this is the point of the article, current processors are so complex and do so many things that the hardware verification suffers from an explosion of states to exercise.
Why can't we reuse that knowledge, that code ?
>Chisel is much higher level than Verilog
have to disagree here. the level of abstraction is exactly the same in chisel as in verilog, its still RTL level. High-level abstraction in hardware is a long-chased holy grail that is revived in the industry every now-and-then and then immdediately dismissed and forgotten, when faced with real-world industry challenges.
Doing a complete verification of a CPU in a strict sense is impossible.
These days we are relying more and more on computational proofs for math and science research.
Therefore - what can we say about the solidity of our scientific knowledge going forward?
Metaphorically, it's like the rush to build auto-autos (self-driving cars): the problem is too hard. If they had started by trying to make a self-driving golf cart Elaine Herzberg might still be alive.
If "verifying hardware is a much more complex task than designing it" (and I don't doubt it) then that is the limiting factor. (Or should be IMO. The liberties taken by software cowboys are bad enough w/o the hardware getting all squirrelly too.)
The hard part is creating something that is competitive with top of the line commercial processors that have thousands of man years of R&D poured into them. Its not just verification, but the huge effort that goes into eaking out another couple percent on something like a branch predictor, or optimizing some "edge case" that turns out to be a significant portion of a benchmark if its not done correctly. Then there are all the general optimizations that give you a 10% uplift here and there. Worse, yet if you go with something that doesn't have a large installed software base (x86/arm/power?) because your going to be spending crazy amounts of effort doing compiler+application optimizations as well.
They show a lot of the verification process throughout the film, including an exciting moment when the chip boots Windows for the first time.
Therefore most formal methods blow up on cpu designs and random coverage is really hard to define and even harder to reach.
I'm coming from a physics background so I never really know where to start when I inevitably start looking this stuff up at 4am.
The goal of UVM is to try to catch all the edge cases in simulation before you fab. The basic idea is to make a model of your design and compare it to your implementation by making a set of random inputs and checking the outputs.
You'll first have to become 'fluent' in EE, but for a physicist, it's just spending the time and getting used to things. Not terrible, long, but straightforward.
As towards what the article is talking about, you need to be trained in it. Honestly, you have to apprentice with the Greybeards (they are mostly men, but not always). There are other ways, like reading through Intel docs or the manuals for ICs or digging through forum posts from 2003. But those guys in the basement with funny newspaper clippings from the 80s or old xkcd printouts are a much better return on your time. They have tons of knowledge about specific chips and machines, stuff that is nearly impossible to recite unless prompted. You just got to spend long lunches blabbering with them, despite their strange political and societal views. Just listen to them, then write down every little thing they said. They are gold in terms of hardware.
Both true, but also an incomplete picture.
It's not just the mental models and the language but also the culture that is very different.
H/W guys are bred, born and raised in an environment that thrives on secrecy and where nothing is ever free.
The way they transact with one another, the tools they use, and in the end, the very thing they produce all exude that culture.
It is exactly the software industry 40 years ago.
Networking for example is the same. If you want to test high scale network equipment OR virtualized network functions, you will buy hundreds of thousands of dollars worth of closed-source testing hardware, software and/or professional services from one of a few big vendors. You will not let anything about your algorithms and designs slip to the outside world, and neither will your test vendors.
Edit: the same is true of most of the software world in general. Sure, you have Microsoft and Google and many others collaborating on Linux, or releasing Kubernetes, VS Code, Go and so on. But the core IP that is key to their business? That is staying in-house, fiercely guarded, developped and tested by an army of engineers.
The main difference is that there are far fewer well-defined software classes that can be tested generally, so it doesn't make too much sense to look for a 'software testing' industry, like you can for hardware. There are some tool vendors, but they offer far fewer guarantees, since it's hard to imagine a product that could find a large proportion of the bugs in both the Haskell compiler and World of Warcraft.
Not sure I understand this. An SoC is a processor, plus more stuff (memory, I/O), right? Is the idea that it might be easier because the "more stuff" abstracts away some inner details?
Using the raspberry pi example from another thread: broadcom designed the BCM2835 SoC, which included an ARM1176 core. Broadcom probably didn't do a ton of verification for the ARM1176 core itself, since ARM already verified it.
If by dependent types you mean theorem provers, then that is used, but rarely -- hand-verification doesn't scale to modern processors, usually you model check against some temporal logic formulas that the processor meets its specification. If OTOH you mean using HDLs (= hardware description languages) that use dependent types, then mostly not. Arm's ASL (= Architecture Specification Language) has a tiny bit of dependency build in to reason about length of bit vectors.
[Citation needed] There's literally 0 upfront costs for a Cortex-M3 [0]
Spawning new ecosystem is not about making something completely from scratch like processor. One have to align a lot of stars in the sky to make that happen. They had a specific goal and niche where they planted the seed for RPi.
Or is that view to simple?
And you have to check/debug those 10 million compiler passes at various stages, and each design change may require developing a new debugger or disassembler from scratch to plug into the compiler at each stage of compilation.
What I'm saying is that CPU designs aren't programs, because you can generally trust the compiler to be infallible (and compiler bugs are there, but they're rare). In a CPU process you have to consider the physical impact of the design on manufacturing, what yields you get, how the product is binned, and so on. There are feedback loops between the packaging, testing, and design teams to alter the silicon before production ramps up to go to market. There are tons of moving parts to the actual design process itself, let alone what is being designed.
I guess everyone underestimates how long it takes to write software - even hardware designers.
If you count semi-custom cores derived from ARM designs, then add Ampere Computing and Qualcomm as well.
Perhaps a fuzzy/AI type approach?
Yeah does seem like a intractable problem for sure
Why can't you simply upload the design to an FPGA, and then check that it can:
1. Boot all available operating systems (Linux, *BSD, Windows, etc.)
2. Successfully compile and run the testsuites for a bunch of open-source software (several languages like Rust have a standardized repository and method to build and run tests, so this is very easy)
3. Correctly run stress testing software (Prime95, etc.)
4. Correctly run several software unit tests that you write to exercise instructions that may not be produced by LLVM/GCC
5. Correctly run tests you write to exercise specific processor/cache states
6. Properly handling fuzzed code without freezing the whole CPU (using afl-fuzz)
Start with the simplest possible in-order core so that you get it working very easily, and then evolve to your desired end-state with a series of small commits, and if the verification fails use `git bisect` if needed to find the offending commit, insert any instrumentation you might need to detect the issue and fix it.
I don't see why you would need a specialized tool for that, or even what a specialized tool could possibly do.
While you could never release a CPU that didn’t pass the tests you describe, they don’t even begin to exercise all the corner cases for a chip. Multiplying two specific numbers together, while the instruction crosses two memory pages, when an interrupt arrives? How do you even test for that kind of thing?
We know about techniques to reduce large classes of errors. Data races in particular can be prevented by some languages statically. Other types of “once in a blue-moon” errors that happen as a result of two coupled systems doing something in tandem can be reduced by introducing stronger boundaries between the systems, and then you can test each system independently and make sure it works regardless of what the other system does (I.e. dependency testing, or maybe even fuzzing).
These approaches aren’t bulletproof, but I think they do illustrate a point: that there are techniques to reduce the likelihood of the errors you highlight. Whether they do it at a competitive cost to existing industry practices or not, I have no idea.
I.e. trying to despell the idea that large systems are intrinsically difficult to test.
Once you do all that, you can try carrying out your program as described above. But it guaranteed that the first time you try it, it won't work, because your design will have bugs. Then what?
You need the ability to debug. This means you need to have probes, you need to be able to extract the data, and you need very high bandwidth. You need to have testbenches that are partly in software and partly on the FPGA hardware. Again, that's what the EDA industry will sell you: the hardware and software to do it, as well as the expert consultants to walk you through the process.
And your device needs to interact with the environment. Some of the hardest verification problems have to do with the timing of interrupts; if one comes when the processor is just at the right point, and that case isn't handled in the design, it could lock up. THose cases have to be covered.
Now, for your example of a small core, perhaps it's small enough that you could get it to synthesize and fit into one large FPGA chip and avoid some of these issues. Good luck doing something that can boot Android in that size.
Not that this makes a difference for the core of the point you're making, it's more an aside.
At the end of the day, I don't care (or know) how it's implemented. What I can say is a) they aren't super fast, only ~1Mhz b) they take forever to compile down to, c) they have some, but not great visibility to what went wrong, and d) they cost a stupid amount of money.
> Start with the simplest possible in-order core so that you get it working very easily, and then evolve to your desired end-state with a series of small commits
You can't just trivially evolve a simple design into something more complex, much in the same way when Linux does a new major release they haven't started with some stripped down basic *nix and worked their way up from there.
If you can actually test on an FPGA (which is usually not possible, and if it is, it's not representative of the actual silicon/analog system you're building anyways), you can get 1T instructions in roughly ~6 hours at 50 MHz. But most hardware emulations are ~1 MHz, so now you're looking at weeks to hit 1T instructions (SPECint alone is 20T).
And what happens when you hit a bug, 2 weeks in? It may not be because you actually wrote new, buggy code, but because a new, higher performance branch predictor uncovered existing bugs. But you'll never know, because the FPGA historically gives you terrible visibility.
But in simulation (where testing is actually done), you're looking at ~1 Hz for a cpu core. Ouch. Obviously a better approach is required (unit-tests against models, formal, etc.), since at the level of detail you can't test much of anything.
> 4. Correctly run several software unit tests that you write to exercise instructions that may not be produced by LLVM/GCC
and
> 5. Correctly run tests you write to exercise specific processor/cache states
These two alone seem like they could be really quite complicated.
Also, "run the world" is pretty slow when you have to do it in simulation (or emulation if you wait to have a netlist to find out how broken it is).
I suspect the coverage from your list is substantially lower than you might expect. Would this have caught F00F? FDIV? AMD Phenom's TLB bug?
People are already doing every step you mentioned but there are three problems:
- Processors nowadays are so advanced and complex that you can't simply approach them as if it was one single block. You need to divide the processor into smaller blocks, develop those smaller blocks and put them together by the end of the development. Like in any engineering problem. The main problem is that it takes a lot of time for all sub blocks to be mature and stable enough for top level integration.
- Once you have all your blocks ready, you can start integration and bring up the system with FPGAs. But now you face the problem that FPGAs are really slow and are not really usable as a normal system. You can run some preliminary tests, the short ones but it would take months to properly execute a normal benchmark.
- You could create and tape out test chips but then you would also need to create all the infrastructure needed for the processor to work like the memory system, memory RAM, communication buses, firmware and etc.
Getting from an initial logic design to a manufactured chip is a big adventure. This requires a lot of layers of software, a lot of it highly specified. And yes, you can buy very good simulators for chips, just search for Cadence Palladium for example. They are huge monsters.