A Case for Asynchronous Computer Architecture (2000) [pdf]
avlsi.csl.yale.edu
avlsi.csl.yale.edu
If someone could convert synchronous verilog to async circuits under the hood, they may see huge gains in speed and power use for their circuits, but that is a huge uphill climb.
The Achronix folks are going strong today (still independent), but with a much more conventional FPGA. The on-chip network may be async, but they hide it well. I hope they have lots of success in the future.
It's one of those things from the department of "we can make it a little faster at the cost of much greater complexity, higher cost, and lower reliability". That's appropriate to weapons systems and auto racing.
Even for low power applications, you would probably use less battery getting the work done quickly in a clocked CPU and then falling back to a lower power state ASAP. Allowing the pipeline effects to take hold in a modern clocked CPU should quickly offset any relative overhead. Heterogenous compute architecture is also an excellent and proven approach.
Certainly, there are many things that happen in a CPU that should not necessarily be bound by a synchronous clock domain (e.g. ripple adder). But, for these areas where async cpu a clear win, would we actually see any gains in practice using real software? Feels like there's a lot of other strategic factors that wash out any specific wins.
* An async add operation takes variable time based on the number of carries, whereas a sync one is set to the worst-case.
* The clock for an ALU is set for the worst-case even when doing something faster (e.g. an ADD rather than a NAND)
* If you have multiple logic stages handled in one clock cycle, the problem is compounded. The clock is set by the slowest stage for all components in the system.
* If your system is doing nothing, you're still clocking it. Clocks are adjusted, but not at a nanosecond-by-nanosecond level.
All-in-all async gives a nice power boost and a nice performance boost (not enough of a boost to displace an entrenched ecosystem, mind you, but a nice boost nonetheless).
The reality is that doing clockless logic introduces a lot of overhead at every state, both area and timing. There is different styles and the issue are different for them, but the bottom line is that nobody has been able to realize the theoretically wins in production (note1). And that's not even addressing the lack of tooling.
note1: the closet IMO is Ivan Sutherlands group which have some very impressive claims, but still nothing you can run out and buy.
1) I think an architecture change -- any architecture change -- is expensive. Intel and AMD dumped many billions of dollars into R&D around existing architectures, and an asynchronous one starts without a lot of that benefit.
2) There's a ton of stuff -- chipsets, RAM, software, etc. -- built up around synchronous. The engineering cost go up astronomically.
3) That's not to mention baseline engineering costs.
Async won't give a 2x boost to performance. I would guess it'd be 10%, maybe even 30%. That's not nearly enough to justify the investment.
Ivan Sutherlands' work certainly won't compete with teams 10+x times that size and investment.
ADDED:
With the millions spent on getting just a minor single digit improvement, you think the big players wouldn't jump on clock-less immediately if they could? Note, Intel did use (does?) use domino logic in the ALU and FPU. The benefit just isn't there for clock-less. I personally know of companies that tried and gave up on it.
To your other points, you wouldn't boil the ocean; you keep everything else clocked as usual and bridge between them.
I agree it's not too expensive to prototype, but it's super-expensive to do *right*.
> If you have to poll or await some other component arbitrarily, there will necessarily be extra overhead and delays in these areas.
You don't poll. You have a lot of small input-clocked domains which work at a speed with which data comes.
It is very difficult to estimate which of the 2 approaches will need less area and power for some given requirements.
It is likely that for a sufficiently complex device an asynchronous implementation will use less power, but the effort to design bug-free complex asynchronous logic is much higher than for synchronous designs, which is probably the main reason why very few commercial asynchronous devices have existed.
https://authors.library.caltech.edu/43698/1/25YearsAgo.pdf
It was the original paper for this that got me interested in building silicon tools
It's like:
* having ECC everywhere
* having a single display standard (as opposed to HDMI/DisplayPort/USB-C/DVI/VGA/...)
* some kind of architecture where a single bad expansion card (USB, PCIe, etc.) can't crash a whole computer
... and so on
On one hand, no brainer. On the other hand, it hasn't happened.
NVidia is breaking ground on the move to SIMD/MIMD-style architectures, as predicted at the same time, and only because it gives a 30x boost in performance. Async will probably net us a 50% performance boost or something.
If you mean IOMMU, we do have that. It doesn't seem completely doable because someone could still plug an etherkiller into the card.
EDIT: talk begins around 7 minutes.
But you need to add the ability to switch things off dynamically, meaning cores on CPU/GPU; so far the industry has solved this with little.big but that requires all software to change, it's going to take time that we unfortunately do not have as hardware is closing the ownership model.
There can be dynamic synchronous logic, and vice versa.
Dynamic vs. static determines whether the circuit as such needs to be driven by any constant pacing input, whether embedded clock, or external clock, vs. not needing it to arrive to a settled state (to latch.)
If you are to speak strictly, asynchronous vs. synchronous determines whether that pacing input is external, or recovered from input.
However, the MiniMIPS pipeline structure can execute instructions out-of-order with respect to each other because instructions that take different times to execute are not artificially synchronized by a clock signal.One way to do this is to have each component have an output clock, which raises when it's output is known stable. If an adder has no carries, that takes 1ns. If it has each possible carry, it takes 2ns. You have a second clock propagating backwards to know when the next stage is ready for it's next input.
You still have timing. It's just set to when a component is ready with output, or ready to receive input.
Everything goes faster and uses less power.
Async CPU solved a problem that would have marginal benefit in a metric we care about
Also, I imagine, they would need to be implemented assuming the worst timing delay from the processes. They can't be binned like modern CPUs.
Am I missing something?
So it must consume power to retain its value.
For start, it doesn't make sense to power gate a SRAM. So they are always leaking power. And despite writes not being common, reads are. Most application with SRAM reads all the metadata in parallel looking for a match (and often the data too due timing constraints and increased size of control logic because the extra complexity). And reading uses power.
In some sense they're already asynchronous, despite clocked.
For instance, on my CPU which is AMD Zen 3, the idiv instruction (it computes integer division and modulo) takes between 9 and 19 cycles for 64-bit version: https://www.uops.info/html-instr/IDIV_R64.html#ZEN3 That’s for the operand already in a register i.e. no RAM access involved.
Whether it takes 9 cycles, 19 cycles, or something in between, depends on the arguments of the instruction, i.e. on the numbers being divided.
Same applies to quite a few other instructions: floating point divisions (divps, divpd), floating point square root (sqrtps, sqrtpd), even 64-bit integer multiplication (imul).
It’s not just the math. Jumps, branches and function calls take very different count of cycles depending mostly on two things: predicted or not, and the state of micro-ops cache at the target address. Albeit these effects are very hard to measure reliably, depends on the code too much, probably for this reason uops.info doesn’t have latency figures for jmp/call/etc.
Think of the extent to which a clock can propagate as a lightcone, within that domain everything is synchronized, but it need not be synchronized with things on the outside. The smaller the domains the more asynchronous a design gets. But you'll never be async all the way down, at some point you will have to worry about stabilizing your outputs and passing them on to the next stage in that stable situation.
Compare synchronous serial lines with an asynchronous interface such as a centronics printer interface. The former will happily send zeros + clocks all day long absent a signal, the latter will strobe it's 'ready' output only when there actually is data, but it will still have to output that pulse, which serves as a very local clock.
Integer addition is too easy. On all modern computers, add instructions take at most 1 cycle. Even vector ones like vpaddq AVX2 which adds four 64-bit numbers to another four numbers.
The pipeline takes variable count of clocks to complete an instruction. The number depends on the instruction, input data of the instruction, and quite a few other things. In some exotic cases it even depends on power state, e.g. some Intel CPUs took ~20k cycles to power on their AVX pieces, during that window AVX instructions are much slower.
If for any reason the pipeline is unable to deliver the result by the end of the clock, CPUs don’t delay the clock, they continue running the clock. You simply gonna get the result on some later clock cycle.