100-GHz Single-Flux-Quantum Bit-Serial Adder Based on 10-KA/Cm2 Niobium Process
ieeexplore.ieee.org
ieeexplore.ieee.org
But likely not when the energy cost of manufacturing the cryocooler (which has a limited working life) is included.
If you count dollars instead of joules the answer is a definite "no". The manufacturing cost per joule of the cryocooler, amortized over its usable life, far exceeds the value of the energy saved.
While your actual point is true, your example is too loosely specified - the energy used by a perfect (or even imperfect) refrigerator depends on the temperature of the environment it's dumping the heat into. In the extreme (and also engineering-wise preferred) case, with a heat sink at cosmic microwave background temperatures of ~2.7K, a perfect refrigerator (rather, heat exchanger) would actually gain energy. (Of course, since the literally-glowing-hot sun takes up some portion of the sky, actually getting a <4.2K environment would likely require siting your computer in the Oort Cloud, if not outside the galaxy entirely, hence why your actual point is true.)
If you want to go down that road, I believe you’ll find that the maximum amount of heat that any machine can exhaust into the CMB (or, more generally, radiatively transfer into a medium at an effective temperature T_h) is precisely the amount of heat it would emit as a blackbody at T_h. IOW sending entropy into space requires generating a uniformly distributed random distribution of something and sending it into space and, if you want to use light, you end up with delta S = Q / T, where T is the temperature of the light. So your device will need a silly surface area, in addition to needing to be very well shielded from the sun. Now maybe you can cheat because, at the wavelengths in question, the effective surface area of an antenna is quite large, but I doubt this buys you much.
But in some moral sense, you are entirely correct. If you put your computer on a cold moon with no view of the sun, you can sink quite a lot of heat at a low temperature. (A really good spectrally and directionally selective filter means you may not need to be quite as far from the sun as the Oort Cloud. Even just pointing an IR thermometer into a clear night sky gives a nice low temperature.)
Well actually you can just use multilayer sun shades made of sheets of aluminized kapton like the James Webb Telescope. It wouldn't get it down to CMB temperature but supposedly some of the Oort cloud is already down to 4K so it should be possible to reach <4.2K with enough attenuation by a sun shade.
So as a corollary you can transfer 1J of heat using less than 1J of work. Since the temperature difference is so large the theoretical limit is really close to 1J of work though.
An ideal heat engine operating between reservoirs at T_h and T_c has efficiency 1 - T_c/T_h. This means, when operating reversibly, it will send Q_c heat to the cold side, remove Q_h from the hot side, and do W work, where Q_c = Q_h - W and W = (1 - T_c/T_h)Q_h. If T_c = T_h/10, then 1 - T_c/T_h = 90% (very efficient!) and Q_c = 0.1*Q_h (not much waste heat!). Now turn this around, because it’s reversible. A heat pump removes Q_c from the cold side and exhausts Q_h to the hot side. If you remove 1J from the cold side, you exhaust 10J to the hot side, and that 9J difference is the work done. That is, it costs 9J to remove 1J of heat from a 27.3K refrigerator if the exhaust is at 273K. It’s worse at 4.2K.
(You can do an equivalent calculation a little more tidily by considering the entropy change of each side. To remove a given amount of entropy from the cold side, you must add at least as much entropy to the hot side. Plug in the usual formula for isothermal entropy change, and you get the answer.)
I suppose that means that the cooling requirement will increase the energy requirement by nearly 2 orders of magnitude.
For refrigeration the situation is different: we are interested in the amount of heat _removed_ from the cold side. In this case, for a Carnot cycle COP_cooling = T_c / (T_h - T_c) and is NOT always greater than 1.
The "100GHz" headline tempts comparison to CPU clock speeds but those are absolutely not the right point of comparison. GHz doing one thing != GHz doing another. CPU clock cycles are very deep compared to a bit-serial adders and operate under heavy thermal constraints. To compare apples to apples, we need to know: what is the speed of a bit-serial adder in near-future CMOS? Silicon bipolar? 3-5 bipolar?
My guess: you could implement a 100GHz bit serial adder in CMOS today and get several hundred GHz if you dropped thermal constraints, went bipolar, etc. That doesn't invalidate the research -- we need more public cutting-edge device research, not less -- but the posts assuming this result translates to 100GHz CPUs are getting a bit ahead of themselves.
Comparing a PLL to an adder is silly.
> My guess: you could implement a 100GHz bit serial adder in CMOS today
7nm FINFET CMOS will get you to roughly the same performance point. It's also quite mature compared to the technique in the article.
Yeah, but comparing a bit adder to a 64 bit full adder integrated into a modern CPU clock cycle is really silly :)
Glad to hear my intuition about modern CMOS was in the ballpark though. I've been out of this scene for the better part of a decade.
This is doing quite well for early rounds on this technique.
I did a cursory search but couldn't find them so I went with slightly silly clickbait on the rationale that it was an order of magnitude less silly than the other comparisons being made in the thread. I stand by that: as "benchmarks," a bit adder is much closer to a PLL divider than a 64 bit full adder in a CPU cycle.
I fully agree that there's another order of magnitude before the comparison starts to become apples-to-apples, or bit-adders to bit-adders, but you're going to have to help with the search if you want to see it happen.
Here, we're discussing a paper characterizing the current version of the process first with ring oscillators and measured delays through a single cell, and then with a serial bit adder, so.. if you mean those, yes?
I think it's exciting here that they're dicking around with new technologies in their infancy, and in a few coarse design changes of cells went from 60% of CMOS leading edge to about the same numbers.
Of course, total system power, etc, is atrocious, considering that this is cryogenic, etc. You need a lot of logic before this could come ahead in efficiency.
Yeah, that's what I thought and why I settled for a less accurate comparison.
> I think it's exciting
I think it's exciting, too. The extent of my claims is "no, this does not mean 100GHz LN2 CPUs in a few years." I did not intend these remarks to be meaningful to chip designers, only to CPU consumers who see "GHz" and think "CPU clock speeds."
The leverage of having a slightly better fundamental device is extraordinary. We should spare no expense looking for them and leave no stone unturned.
RSFQ (Rapid Single Flux Quantum Logic, or as some would jokingly say Russian Single Flux Quantum Logic) was quite hyped up in the nineties, Prof. Likharev at Stony Brook had a collaboration with IBM working on RSFQ circuit elements to replace conventional semiconductor logic. At the time the achievable speed was fantastic as compared to regular circuits, (un)fortunately semiconductor processes kept evolving and today RSFQ is only interesting for some niche applications like fast microwave circuits at cryogenic temperatures (and even there HEMT transistors are often a better solution nowadays).
Also, getting circuits with more than 10.000 junctions to work was quite tricky as the fabrication processes weren't very reliable and transferring flux quanta is a bit more noisy than storing charges on an FET, so I'm doubtful whether we could even have large-scale RSFQ circuits without extensive error correction.
Well, it's still an amazingly fun and fascinating field, really hope we might see a revival of it one day (maybe if we get room-temperature superconductors).
Since you sound like the appropriate person to ask, what is the current state of research towards room-temperature superconductors?
BTW it would already open up many new opportunities if we had a superconductor that could be patterned as easily as a regular semiconductor and that would work at "high" cryogenic temperatures (e.g. -40 or - 80 °C), as those temperatures can be easily attained with conventional coolers that don't rely on liquid nitrogen or (much more expensive) liquid hydrogen or (super expensive) a mixture of different Helium isotopes.
So far there's no fundamental reason that superconductivity can't work at higher temperatures, but there's also no clear path to room-temperature superconductors as far as I'm aware.
[1]: https://phys.org/tags/high+temperature+superconductors/
"Automatic Single-Flux-Quantum (SFQ) Logic Synthesis Method for Top-Down Circuit Design"
https://iopscience.iop.org/article/10.1088/1742-6596/43/1/28...
>"Abstract. Single-flux-quantum (SFQ) logic circuits provide faster operations with lower power consumption, using Josephson junctions as the switching devices. In the top-down flow of SFQ circuit design, we have already developed a place-and-route tool that covers backend circuit design. In this paper, we present an automatic SFQ logic synthesis method that covers front-end circuit design. The logic synthesis is a process that generates a gate-level logic circuit from a functional specification written in hardware description languages. In our SFQ synthesis method, after we generate an intermediate circuit with the help of a synthesis tool for semiconductor circuits, we convert it into a gate-level pipelined SFQ circuit. To do this, an automatic synthesis tool was implemented."
PDS: Phrased another way, the idea here can basically be boiled down to:
For the most time-critical parts of a conventional CPU, i.e., the Adder/ALU -- instead of using conventional circuitry for that, let's use circuits specifically designed for Quantum Computers because of their fast switching speed...
The end result is that you still get a conventional digital CPU -- albeit, in theory, a much, much faster one...
The "magic phrase" (for research) for all of this -- is Single-flux-quantum (SFQ) logic circuits...
What I want to know is, what were the problems that prevented them succeeding before, and have they been overcome now?
So to run at those speeds, you need extremely compact circuits, and you need to bring memory closer too.
Maybe move to 3D, another way to increase reach.
What is the current thinking regarding this problem?
Bit-serial adders are extremely simple circuits with a single carry register, which can run at essentially the full speed of the underlying switching devices. Ripple-carry adders have a circuit depth proportional to the number of bits in the adder -- they're only stable when the clock is substantially slower than the switching time (proportional to the circuit depth), because the carry signal needs to propagate through all of the 1-bit adders.
That's not what I said. I'm just saying that it's obvious that if you wanted a fast 64 bit adder on the process, you'd build a carry lookahead adder or other fast adder, not stack a single bit adder 64 times. Saying "YOU NEED TO MULTIPLY PROP TIME BY 64" seems kinda dishonest.
I said:
> > You could obviously use the process to make wide adders with faster clock speeds than stacking serial adders in front of each other.
Do you really think the optimum thing is going to be stacking 64 of these in a row?
Propagation delay problems are real, but for caches IIRC the biggest problems are transmission lines being lossy and having propagation speeds much slower than the speed of light in a vacuum, and for DRAM the biggest problems are waiting for precharge and waiting for the sense amplifiers to stabilize. Waiting for the precision required to read obscenely small capacitors, in other words.
If DRAM is 1T1C and SRAM is 6T, I always wondered why we couldn't just have DIMMs of SRAM at 1/3rd the capacity or whatever. Overprovisioning DRAM by that factor is already common, wouldn't eliminating all the downtime be a way to put that slack to work? Like the HDD -> SSD transition? Ah well, I'm sure there's a constraint I'm just not thinking of.
SRAM costs way more, but afaik that's an artifact of market size, not fundamental. Hence the 1T1C vs 6T comparison to tease out the fundamental cost difference, which looks to be not worse than 3x-6x more expensive, which would put SRAM DIMMs will within reach. I lean on heinously expensive SRAM in my own embedded designs and would really like to see some market scale drive down those costs!
Where did that go?
I'm imagining a small in-order 32-bit core with ~256KiB static RAM running at 100ghz. Call it the Serial Killer.
Having a "100-GHz Single-Flux-Quantum Bit-Serial Adder" means we can finally have hope for classical computations at hundreds of GigaHertz?
Quantum computers are very different in that they maintain a non-classical state throughout many components, allowing certain kinds of matrix calculations to be solved significantly faster than is possible for any classical machine, analog or digital. (Although quantum solutions have a stochastic nature and there are other caveats to quantumness that potentially stand in the way of actually exceeding classical computers).
I can't speak for all architectures, but the ones I'm aware of absolutely are analog computers.
Source: building a quantum computer.
They're closer to analog computers than digital, but really they're neither and share some similarities of both. In the sense that the qubits usually only have 2 measurable states in computation, they're similar to digital computers, but that only goes for when you measure. On the other side, the 'programs' going in are not really digital information (though they can be represented as such), and because qubits are not only in one of 2 states (unlike bits) the way in which you interact/program them is definitely not digital, it is analog.
If you're interested in learning about writing quantum programs, I recommend checking out IBM's Qiskit toolkit[0]. I found the tutorials to be helpful for helping grasp the fundamentals of quantum programming.
Are there any resources you particularly recommend if I wanted to learn more about quantum computing and how quantum computers work on the hardware level?