Path to 2 nm May Not Be Worth It
eetimes.com
eetimes.com
Qualcomm are comparing 10nm and 7nm, and those figure are not as high as a generation node should be, and that is perfectly fine because 10nm from TSMC are half node. They should be comparing 14/12nm with 7nm, and there are multiple iteration of 7nm. Does it scale as perfectly as before? No.
It isn't about the complexity and design rules. It is about all about cost, and it only means chips are going to be more expensive going forward and only few companies will be able to afford it, but there are nothing on the roadmap suggest 2nm isn't worth it.
This article reminds me of the old days 10 to 15 years ago when they keep saying the end of Moores law.
Basically, the stars are aligning to spell the end of traditional CMOS scaling. 7nm nodes are in early production now; 5nm is on the way; 3nm is likely to happen; 2nm may or may not depending on the economics; 1nm is unlikely but you never know; and sub-nm CMOS nodes are probably just not going to happen. (Don't confuse actual nodes with what happens when marketing gets involved. Some twat sticking a "2nm" label on a 20nm-class process doesn't actually make it any better!) This isn't to say that semiconductors will stop improving, just that it sure won't be classic CMOS any more.
One under-explored area I think we're going to see a lot more of is backfill of older processes. Very few things actually need the latest and greatest digital process: main processors (CPUs, APUs, GPUs) are about it. I think there's plenty of room for improving cost on older processes to make them more widely deployable. To some extent we're seeing that now with SOI processes like 22FDX, but I think that trend will continue. (Unfortunately, there's a lot of stuff that just doesn't shrink, period: some analog stuff, all HV I/O cells, and just about all power management.) So the potential upsides aren't universal, but I think there's a big market out there for stuff that can benefit from smaller processes but can't afford the obscene cost of a multipatterning mask set.
Gimme a coprocessor that's just 1000 Pentium I's on the same die.
The bottleneck of the Phi was those non-vector instructions you inevitably needed.
If your threads were tied down executing excruciating slow scalar operations you couldn't schedule the vector operations you needed. There were also problems with memory latency in the ring bus. It was a great architecture if everything worked out perfectly, but then you could say the same about systolic arrays and other exotica.
The GPU model of having every thread execute its own scalar operation in parallel feels like duplicated effort but it ends up working out better in practice for a wider variety of applications. On the Phi they added mask bits to get you closer but it was never really designed to work that way.
That is just one of the many, "obscenely expensive" parts of the most impossible machine in the world that is a bleeding edge fab.
> “The abstractions we worked so hard to maintain are becoming porous as secondary and tertiary effects bubble up through the whole flow … from node to node, they intensify or change character,” he said, expressing optimism that “all of these effects will get solved.”
Might you be able to provide a little more information about which effects are becoming problematic?
I'm not up to speed on the details of why we can't just use X-Rays but apparently its a difficult problem.
At least not yet until 2nm. The TSMC roadmap and its Grand Alliance as Morris Chang likes to call it, has it all till 3nm. That is roughly equivalent or slightly better than Intel's 7nm. We could do 450mm Wafer, when the time comes where it make economics sense. 3nm is likely 2021 / 2022, it is a little hard to tell 5 years down the road what road blocks we have.
I think we are more in agreement than disagreement. I just don't like the article headline and its handpicked statements and content tries to convince whoever is reading it we cant do 2nm. We can and we absolutely will. I would be very much surprised if our 1.3 billion smartphone per year sold and all the money in AI / deep learning as well as GPU can not absorbed the increase in cost. AWS is still expanding at a unbelievable pace, and China and India still has lots of rooms. The bitcoin mining chip has already helped to recoup some of the 16 / 7nm R&D cost ( So bitcoin is not all useless after all as some said ). We have Cars that need lots more transistor for Autonomous Vehicle, I doubt the few thousands dollar increase in BOM would be problem for these car buyers. As long as we can increase the total addressable market, 3nm or even 2nm cost should not be a problem.
>One under-explored area I think we're going to see a lot more of is backfill of older processes.
We already are. 28nm capacity demand has outstripped supply for a while. TSMC is working on further simplifying and lowering cost of 12nm, so it should be ready to replace 28nm by 2020 or later.
In terms of Semiconductor industry, I am much more worry about DRAM cost not coming down.
The funny thing is they might stop being semiconductors actually.
> Digital logic integrated circuits (ICs) implemented with complementary metal-oxide-semiconductor (CMOS) transistors have a fundamental lower limit in energy efficiency because transistors are imperfect electronic switches, having non-zero OFF-state current (IOFF) and finite sub-threshold slope. In contrast, electro-mechanical switches (relays) can achieve zero IOFF and perfectly abrupt switching characteristics.
https://www2.eecs.berkeley.edu/Pubs/TechRpts/2017/EECS-2017-...
And the work on the optical technologies to use instead of transistors are massive recently.
The starting point in the article seems sensible, but unremarkable: 2nm is not inherently valuable. It needs to pay out in smaller chips, lower power usage, faster per-transistor speeds, or lower cost per FLOP. But the article is written like a shrink is only valuable if it does all of those things, when in reality it's still worthwhile to get any of them.
Assuming the article is accurate, per-transistor speeds may stall and cost per FLOP may even move backwards (that's not new though - just a question of how fast cost savings catch up with new processes). Power savings might slow, but look substantial at 7nm and 5nm and there's no pitch for why 2nm would differ.
Meanwhile, area scaling is still a major improvement, just as you'd expect. Perhaps that won't be reason enough for desktops and servers, but that's been an open question before now. Lowering minimum chip size creates qualitative change in smartphones, IoT hardware, sensors, etc. It's in that "world market for five computers" vein of only looking at the use of a development in already-extant devices.
As a last note, it's downright sleazy for the article to write "'...It’s not clear what will remain at 5 nm,' said Penzes, suggesting that 5-nm nodes may only be extensions of 7 nm." From the first half of his quote, Penzes is clearly saying that it's not clear which advances will persist through 5nm, not that it's unclear whether any will persist.
You're more optimistic than the article about power use, but if you agree it makes an okay case about speed and cost, and the only strong improvements are on area, then we're not in a great spot.
In the olden days, poly silicon and metal wires could be thought of as having a square cross section. So if a wire needed to carry a certain current, it had to have a certain area. But in order to manage that current when one dimension shrank (the base) then the other dimension had to grow (the height).
Around the time that chips were getting into the GHz range, the wires had a small footprint but they began to have a taller height relative to that. So instead of square wires, they ended up looking like skyscrapers. All that parallel cross section facing each other makes the wires act like capacitors. So that sets a lower limit on feature size.
Edit: so if the ratio of height to base doubles, the capacitance almost doubles, which decreases the bandwidth by almost half:
https://en.wikipedia.org/wiki/Electrical_impedance#Capacitiv...
https://www.electronics-tutorials.ws/accircuits/ac-capacitan...
We've gotten complacent. Software keeps getting faster, despite bad coding - just because of the improvements in hardware.
Engineers have known for a long time that we aren't as efficient as we could be.
[0] Not attempting to beat up on either of those companies.
At least one aspect of windows10 design will add input latency, which is the forced vsync in the windowing system.
We think we have been optimizing for those things, but I don't know that we actually have. Does Ruby/Python really help you get to market faster than would similar languages with static typing and compilers or JITs? (Kotlin, C#, F#, Nim, etc)
There is research that suggests not, certainly I don't think the effect is large if it even exists. Many people use TypeScript because they feel it is a productivity boon over Javascript. We might one day leverage that and use the static types to compile TypeScript more efficient WASM or something, win win situation.
- Garbage collection make performance unpredictable because of collection spikes.
- If you use a big library, you now have to load the library, which is an extra time-sink.
- Likewise for frameworks.
- Optimizing compilers optimize the 80% of our programs that don't need optimizing, and under-optimize the 20% that do.
(This is DJB's argument.)
That might be what he's saying. He didn't say anything about dynamic vs. static typing.Also, there's a counter-argument to this, that writing non-garbage-collected code might result in security bugs (to do with memory management). And security is a non-functional requirement that some might say is more important than speed. It's hard to judge though; there are loads of requiremenets in software dev, and it's hard to balance them.
That's falacy. Don't mix "fully manual memory management" with "I have no GC". Those are 2 totally unrelated topics. GC is one of the way to manage memory automatically. There is hundreds with their own plus-sides and downsides. Here's the least unpopular:
* Stop the world and incremental GC: Makes it easy to write simple code. Doesn't really make the lifecycle explicit so when the codebase grow larger you start to have boilerplate code to manage/close connections/files/transactions. (all toys languages, also Java and Go).
* Reference counting: Can be made fast enough by using some of the pointer 64bits to store its state. A giant problem is that all your libraries have to support compatible "smart pointer" formats or the whole thing become unmanageable. (Modern C++)
* Contacts/Lending: Does it's job, is safe but makes the code more complex by making a lot of implicit things explicit (Rust, Modern C++)
* Explicit ownership + message passing: works well enough for GUI because just like HTML everything is less or more a tree. Also consuming messages allow "higher level" objects not to have to care about pointers. For backend code it's not very good. (Qt-C++)
* Pure functional: Everything is a copy, they get freed when out of scope. Works well for some domain, but gets in your way for more stateful constructs.
* Copy-on-Write: Pass everything by reference (and count them). Copy when they are modified, otherwise point to the real memory. Only solves half of the problem. Assuming most allocations are containers it works well enough. It doesn't attempt to solve objects management.
* Pipelining: Get input, process them, pass them on or free them. The processing should not have to dynamically allocate at all (Bash, Erlang).
* Pre-allocation: Allocate everything at compile time and use buffers big enough for your use case. (C)
* State machines: Encode the object lifecycle as a series of state changes. Release when it reaches end-of-life. Great for protocols, parsing, transactions. You usually can use code generators to generate proven safe C from an higher level format or proof. However it doesn't scale and only matches a subset of programming projects. It's also hard to learn for non EE programmers because it doesn't map perfectly on the iterative Von Neumann CPU. Easy to read but hard to write.
No-GC != security-bugs because of memory management. If a C program uses pre-allocated buffers, then the security issues is because the language is unsafe, not because it has no GC.
I never dreamed that software developers would manage to make things slower still (by orders of magnitude) by throwing everything onto the 'net and using multiple round-trip web-api calls to do the most trivial things. I just didn't anticipate that this source of inefficiency was there to be harnessed.
(This came to a head for me recently when I was trying out planner apps for my phone. It boggled my mind that, even on a dual 1.8 GHz processor, Asana and MS Planner do not respond instantly.)
But it also means there's some relatively low-hanging fruit these days if you find a niche where performance is much more important than accessibility :)
The Dylan folks were all trained computer scientists, mostly MIT PhDs, almost all with years and even decades of experience in writing garbage collected languages on resource-constrained machines (e.g. MACLISP). By contrast many of the apps you write are written by much less trained folks using a lot of automation and boilerplate.
While I personally prefer to write tight, clean code directly to the underlying system, I am really glad for this democratization of programming which has lead not only to good careers for people who didn't get to go to MIT/CMU/UCB but also lead to a profusion of great apps. Not everybody has to design TCP but everybody gets to use it.
And note that Dylan had little impact on society while Tinder sure has.
Don't know what I was thinking.
I think we can go even stronger than 'probably' - we know that there are people who do work to these constraints, and it's mostly experienced devs from a few starting tracks working on years-long projects.
As far as I know, NASA's coding standards haven't slipped from their famous "0 bugs per 500,000 lines", and that code (along with JPL, Orbital, SpaceX, etc) is often written to harsh physical constraints. Systems like the Mars entry controller have to operate unsupported in environments where slowness, power, and weight are all serious limiting factors. For all that the IoT is a disaster, embedded control software for non-consumer projects is often concise, fast, and reliable.
It would probably be possible to train a lot more high-performance developers than we have now, but the demand certainly hasn't been present and even if we could there'd be a real price exacted in lost development time and features.
It'll be interesting to see what happens as chip speedups become increasingly unavailable; it's possible we'll revisit the old stuff and make it faster, but so far the leading answer appears to be falling back on networking to handle tasks on bigger hardware.
Isn't the problem really memory bandwidth? Aren't most CPU cycles "wasted" on speculative execution while waiting on fetches from memory slower than CPU caches?
Wouldn't a hardware model where all of the RAM (SRAM?) was on the CPU die (fabrication and DRAM clock speed aside) solve this problem and make everything an order of magnitude faster? What am I missing? Is this really the programmers fault?
The CPU literally spends most of its time waiting on my instructions, not choking on them, so why is this my fault for being lazy?
You are missing money...
SRAM and wider data paths are expensive.
You are totally right about that taking more logic into synchronous domain domain is an almost universal solution to getting more performance. The faster the chips become, the more IO gap becomes.
Experienced MCU programmers know quite a number of algorithms that may run faster on "dumb" 100mhz MCUs than on a 20-cored Xeon, and are proud to show them off.
Even when some of MCU cores may take like 6 cycles to do an arithmetic op, their strong side is that it will be taking them these 6 cycles to do that all and every time consistently, and that their memory runs synchronously.
Say, in a 1000000 cycle loop, where every operation results in cache miss, Xeon will spend more time on a loop than a wimpy 100Mhz MCU that simply don't take any penalty on unpredictable, heavy on synchronous ops cope.
Apple A11 is so fast for an ARM just because they choose the simplest, the surest, most direct, but also the most expensive way to scale performance: tons of SRAM and prefetch logic everywhere to reduce IO gap as much as possible.
I have read estimates that go as high as 1/3 of the CPU lost to it (sorry, no reference here), but even those were not usual.
I was reading a thing a while back on here or Reddit where someone asked about an OS that ran entirely in cache, after all in principle a modern processor has more L3 than an original Unix minicomputer.
Since many of the memory subsystems operate on the assumption that changes should be propagated to memory that you would still need to have RAM there even if you weren't really "using" it.
There was a suggestion that since the bringup code begins entirely on-processor (before initializing RAM) that you could hack the Management Engine and keep running in that mode forever.
My iPad Pro 9.7" (one version old, came before the 10.5") drops frames in iOS 11's full screen blur transition animations. Before iOS 11 it was fine, but either the iOS team didn't test animations on anything but their brand new hardware, or they tested and didn't care if it ran smoothly.
Having one of the OS's most frequent animations choke on anything but the brand newest GPU in the $650+ model? Not a great look, especially with Apple's reputation for making devices that last a while.
Fingers crossed that iOS 12 fixes it.
Power consumption scales roughly quadratically in relation to frequency.
Modern cores all have features like power and frequency gating, performance counters - the stuff needed to allocate just enough of computing power to complete computations in time. This is very, very evident in things like music playback - a lot of intermittent io and computational spikes.
It's true that switching power scales with the square of voltage (P=V^2Cf), but that's not really a problem for the race to sleep as switching power even with V as low as you can get it is going to be orders of magnitude higher than your static power.
The "race to sleep" is a misconception long held by software people.
The situation you describe is only true when the part of the chip doing computation is completely off, and is shunted off from power. That's so called deep sleep.
There is a limit to how frequent you can switch the big and complex chip to low power state, as doing so by itself takes some power and computational overhead to keep track of your power use profile.
For many cases, one have to endure leakage and keep working, like when you have to do IO have to keep the chip "on" anyways, and want to utilize your "on" time as much as possible.
A ton of work have been done on power saving strategies like partial device shutdown, shutting down part of registers, trying to operate different parts of the chip at different frequencies to match io wait and processing time, downshifting to more iterative computational paths, and on and on and on
It's very easy to stamp out more ALUs on a chip than we can reasonably fill with data (and, arguably, we've already hit that point with consumer hardware). Improving parallelization nearly always comes at a cost to serial performance, and if you have code that requires a latency-sensitive, unparallelizable critical loop... well, you're basically saying "screw you."
They're working on a new processor architecture, which is very interesting in many, many ways.
And they're working on ways to parallelize existing code, in ways that can't really be done with conventional superscalar architectures.
That's the biggest problem with Moore's law fizzling out - computer architecture isn't actually that hard, most of the work is just figuring out how to apply the extra transistors you're given. So no matter how specialized you get, you aren't going to get more than 1-2 generations of a chip out of a single semiconductor process, as there is only so much that clever architecture can accomplish.
So each 'problem' that's worth it might get a specialized ASIC or two, but all of that could very well run its course in the next 5-10 years, and things look kind of bleak for performance gains after that.
That was sort of my point; however cool the conceptual architecture is, they've definitely reached the point where they ought to be able to get at least rough comparisons to competitive hardware. That such comparisons are lacking should be taken as a sign that the architecture is too immature to make those claims. I just wish people would stop bringing it up as the magic solution to parallelism, since there's insufficient evidence for those claims.
> That's the biggest problem with Moore's law fizzling out - computer architecture isn't actually that hard, most of the work is just figuring out how to apply the extra transistors you're given. So no matter how specialized you get, you aren't going to get more than 1-2 generations of a chip out of a single semiconductor process, as there is only so much that clever architecture can accomplish.
I've been swayed by the idea that the future of processor technology is essentially accelerator land. You're going to have a few big CPU cores because they're good for code that has no real parallelization potential, but you're going to also have clusters of accelerators that are optimized for different kinds of workloads (such as data-heavy, communication-light workloads that are good for GPGPU).
That is possible, their progress has been slow.
> For the most part, the ISA just needs to be 'good enough' to allow things like branch prediction, speculative execution, out-of-order execution, etc. You can do that with x86, ARM, POWER, whatever - the instruction set doesn't matter a whole lot, it's just how many/what type of transistors you can throw at it, and how hard you want to work on timing and power optimization.
If you watch all their videos, then this isn't really correct.
Architecture matters a lot. With modern processors, there is a vast gulf between the instruction set used (the programming model) and how the processor actually works. A lot of the die area on the chip (aside from caches) isn't devoted to actual computation (like adding numbers in an ALU) but moving data around and scheduling.
The Mill Computing approach is to move all this functionality back out of the hardware into the software compilation / loading step. Where it can be done once and more efficiently. They can then use that die area to add in more functional units (because the architecture is wide-issue) and more cache.
There is only so much that clever architecture can accomplish, but that is still an order of magnitude more than what is being done now with existing architectures, so this seems a worthy pursuit.
Why will they succeed where Itanium failed? There's a long, long history where people expect compiler smarts to maximize performance of hardware, and compilers just aren't up to it. It's also very necessary to demonstrate this on larger, more complex code, rather than focusing on things like GEMM kernels.
There's a lot that needs to be justified, and the Mill people just haven't presented the justification.
Ivan Godard mentions Itanium a few times during the talks he has given. So they are at least aware of that architecture and its faults.
I guess it all comes back to the question of where all this instruction scheduling ought to be. Yes, it is being done in hardware now, because that is the path that offers software compatibility at the binary encoding level.
But in a sensible world, is that really where the instruction scheduling should be? And if it can be done in hardware, I don't see any theoretical reason it can't be done in software.
The question, then, becomes how to do this in a practical manner. It remains to be seen if the Mill Computing scheme will be successful. I would argue that we (as an industry) haven't failed enough with the software scheduling approach.
Hardware instructions can take a variable amount of time to complete, depending on their inputs--particularly branches and memory instructions. That means that the set of operations that are ready to execute is a fundamentally dynamic property, and software can only make a static decision.
Everything on the Mill CPU architecture is statically scheduled, because everything has a fixed latency.
Loads / stores from L1 cache are also a fixed latency. If you're going out to DRAM, then you'll have to wait a lot no matter what architecture. Branches are also fixed latency, assuming the target has been prefetched.
So then it is a question of how good is the prefetching, and they've got some techniques to improve that too. They've also got techniques to help with cache pressure and effectively double the I-cache size without the usual power and latency penalty.
Details are in the papers and videos, if you are interested.
For memory accesses yes, that's a big problem which is why there's so much effort in the design devoted to making software memory speculation feasible. The Itanium put some effort into that too but generally at a much higher overhead and the Mill's solution for dealing with memory faults in speculated accesses was frankly brilliant.
However, starting up an entirely new computer architecture isn't easy. They want to retain control of their company, so they aren't accepting tons of VC money. Progress is slow.
The Mill has some sets of innovations that seem to come together, for instance the Spiller, Belt, and memory speculation model seem hard to seperate. But I'm pretty sure you could detach that from having a virtually addressed cache and their protection model. And then there's stuff like the super weird organization structure they have. I feel like they're setting themselves up for failure by trying to make changes in so many ways all at once.
But you were still getting more transistors at the time. You just had to use them in different ways.
Disagree. Much of it keeps getting slower.
I'm sure it solves problems they just aren't my problems.
Which frankly applies to nearly everything AWS offers.
I've seen people design for massive scaleout before they even have a final idea of the product.
Frankly for the kind of systems I work on I have collosal vertical headroom before I even need to think about going lateral with scaling.
I mean the largest database I deal with would fit into the RAM of a single mid-range 2018 server (192GB would do it).
The most important database on our main work system will fit into the RAM on my Thinkpad with enough space to run an IDE and Vagrant (32GB of RAM in the laptop so high but insane).
You can always just add more computers in "the cloud" I suppose. But, at some level, increased functionality depends on better price performance.
We're already seeing this in machine learning. I used to follow the processor space fairly closely and I'd see various companies come out with designs that were specialized for some workload or other. The problem was that you could instead wait for a couple years and x86 would be just as fast.
You also have some other things going on like open source software, centralized computing in clouds, etc. that make having hardware that's optimized for some specific workload much more tenable.
This still holds for the majority of users. Only tier 1 dotcoms can run stuff on GPUs economically. For everybody else, just getting many times, possibly few magnitudes, more commodity x86 machine time still makes sense.
I always remember the saying of my ops unit head once I worked in an advertising network. It was something to the meaning of: "a new coder like you who can find a fancy algo to do thing A and B faster pays off in two years, but a 10 new servers will pay off in just 2 month"
And I am still convinced that fabs will continue mercilessly squeezing CMOS till the last drop of blood for at least a decade. There are tons and tons of avenues to make cells smaller, cooler, and more performant that don't involve further process down scaling or switching to a next gen semiconductor.
Disagree. Anybody running appreciable ML workloads buys and racks their own GPUs.
Moreover, say, a tier 2 player like a major AdNet rents its CPU time at rates that make AWS users cry. For them, it makes even more sense to just "let the code be a horrid contraption written by 10$ per hour outsourcing shop, but we can run in on cheap hardware"
* Uptime guarantees - AWS goes down. Running your own hardware locally, you at least have a recourse.
* Storage and bandwidth - some of these datasets are huge and not amenable to streaming onto a temporary rental.
* Security - the ever-present bogeyman of cloud services. It might be harder to lock down your own machines, but it's the only way to fulfill some contractual requirements.
For commodity processing that doesn't involve these specialties, sure, fire away. That's what it's good at, and what most people need most of the time. But there's always a final TCO number and arriving at it requires accounting for all the risks.
I think this is the key. GPUs are pretty much the only consumer-level specialized co-processors, and those have been selling like gang busters. Plot twist, it's the data centers that are buying them up... to the point that the manufacturers tried to prevent you from buying more than one at a time.
Of course specialized processors need to have proper networking and storage proximity to really deliver value. May be more challenging to offer a specialized service like this than you'd hope. Perhaps the big cloud vendors will dominate again, so chip manufacturers need to just hope for a few big contracts to Amazon / Microsoft / Google.
No, it's the cryptocurrency miners that caused that restriction to be introduced.
AFAIK Intel & luxtera are the two main leaders in silicon photonics so I wouldn't count intel out yet.
Or a tube-like cylindrical thing?
I suppose a long tube might work, but that'd be a weird shape to work with.
The trick is to add 3D features that don't involve making an entirely new device on top of another. This is the case for RAM, MRAM, and flash memory: they share devices in the stack, the only parts that are being scaled vertically are charge/spin carrying parts, but not amplifiers, backend, data lanes, or other devices.
There was a lot of talk about making 3D standard cells (stuff from which normal, non-memory, devices are made of.) The amount of work is immense. Every year there are a dozen cookie cutter PhD work like "3D NAND/XOR/INVERT device that is N percents smaller than before," but it will take years to cover and unify the whole cornucopia of devices in cell libraries. And only once it's done, will major fabs think of switching to that. No fab will try to add much more litho layers just to reduce footprint of only one device or macrocell.
(2) How the heck would you cool it?
Now imagine, thanks to 3D cells, you now magically get double the number of registers. How much efforts do you think you need to spend to double the amount of loops running in parallel?
2) I seem to remember IBM was looking at microfluidics for both power and heat distribution, like our blood.
I would love to see the microfluidics cooling. I wonder what would happen if you reached boiling though--it'd probably explode the chip.
https://en.wikipedia.org/wiki/3D_XPoint https://en.wikipedia.org/wiki/Flash_memory#Vertical_NAND
So I am not sure what to believe.
When do we start using optical interconnects?
im wondering just what your implementation is, are you talking "audiophile" equipment, are you talking digital data transmission... is it a human eye/ear noticing latency or is it an electronic component, that will consequently influence the smooth operation of the circuit?
> is it an electronic component, that will consequently influence the smooth operation of the circuit?
So, it would be this.
for this reason you will have a maximum frequency threshold for statisically valid data...for parallel this is an average of each lane in the bus...for serial this is the overall latency of the junction operation... this is only considering the Tx>>Rx junctions...the signal path between must also be considered...potentially quantum chips would pump individual photons and tunnel individual electrons...or when we get there, we could talk about entanged photons Tx Rx bidirectionally through whatever it is that "tangles" between photonic superpositions...
As a lay person (without EE or physics training), I'm still just looking for a high-level summary, as there's no way I'll learn anything useful, otherwise.
https://en.wikipedia.org/wiki/Transistor#Transistor_as_a_swi...
...latency matters when a chip has to operate so quickly that someportion of the chip is left waiting for another portion to finish its job... if you want information[signal] there is a required latency. if you flick a switch on, you have to wait for it to be flicked off before you can flick it on again...try this with a light switch, flick it on count to one flick it off count to one again flick switch on again...you should observe a light bulb light up 1sec, go dark 1sec, light up again..etc. ...now do this, flick the light switch a rapidly as possible up and down over and over again, if you do it fast enough the light will look very different it will seem to be illuminated or constantly flickering but not going dark...this is due to latency of the lightbulb filament, there is a frequency of switching that will break current and resend current to the light bulb faster than the light bulb can deenergize and darken...
AMD's High Bandwidth memory and Samsungs high density NAND seem to open up another axis to continue doubling transistor count in a given area.
It'd be interesting to see things be engineered this way as devices are starting to become too thin to be comfortable.
3D memory is a lot more tractable because it doesn't have to do nearly as much as a CPU. You don't typically worry about cooling your RAM.
There is a type of logic efficient enough that we can think seriously about monolithic 3D integration; superconducting Josephson junctions at liquid-helium temperatures are really, really efficient. There's some hopes that spintronics may one day give us sufficiently efficient switches at room temperature.
But we've wanted to integrate our chips in 3D for decades - transistor scaling aside, routing latency adds a lot of overhead and the length of the wires you need only gets worse the more transistors you need to connect with them. (It's an n^2.something scaling, IIRC). All else being equal, 3D integration of the exact same number of transistors as a 2D CPU would result in a major performance boost all on its own, by making the routing much more straightforward and direct.
The reason we haven't done it isn't because we haven't thought of it, or haven't tried to do it - it's because we can't. (Also, manufacturing a 3D chip is really tricky)
But then, what is really the gain? You will be trading a larger number of chips per die for a proportionally more expensive process. Besides some faster communication (may be great for L1 cache), I can only see it adding problems.
Edit: This is not a topic of expertise for me so I have no idea whether the content is as valid now as it was when it was published.
I was kinda hoping someone with expertise could inform us, given that it’s a fast-moving field.