TSMC approaching 1 nm with 2D materials breakthrough
edn.com
edn.com
But despite the weird naming scheme, it's clear from transistor density [1] and GPU prices [2] that foundries are still making progress in transistors per dollar. That progress is just barely beginning to make large neural networks (Stable Diffusion, vision and speech systems, language model AIs) deployable in consumer applications.
It might not matter whether your cell phone renders this page in 1ms or 10ms, but the difference between talking to a 20B parameter language model and a 200B net is night and day [3]. If TSMC/Samsung/Intel can squeeze out just one or two more nodes, then by the middle of the century we might have limited general-purpose AI in every home and office.
[1] https://en.wikipedia.org/wiki/Transistor_count#GPUs
Are they?
Home GPU price have been elevated due to crypto.
Latest node CPUs are mostly opaque long-term contracts.
The xxBN transistors chips are priced at $xx,000.
If you have a (better) reference I would love to see it.
Intel and Nvidia just got too lazy or fumbled their node progress.
I understand roughly why this shift is happening (machine language proving to solve a whole raft of hard problems) and how it's happening (specialized chip designs for matrix math). But I don't understand where it's all going, or how I can plug into it.
It feels like a fundamentally different landscape than what I'm used to. There's more alchemy, perhaps. Or maybe it's that the truly important models are trained using tools and data that are out of reach for individuals.
Does anyone else feel this way? Better yet, has anyone felt this way and overcome it by getting up to speed in the ML world?
Don't know how language translation models work? No problem, use one that someone else made to make a web framework that self-internationalizes to the user's browser default language without the site creator even knowing that language exists!
I guess the difference (for me, anyway) is that this change isn't incremental. It's a fundamentally new type of computing-- One that comes with a totally new way of approaching problems. Listening to Andrej Karpathy talk about Software 2.0, for instance, it seems probable that ML has a place in many parts of the stack.
It's possible I'm just projecting my insecurities, here, but my experience has been that changes to computing hardware usually result in changes across the entire industry. And this feels like a pretty meaningful change.
That said, getting hold of the data and the computational resources for training are barriers.
Using ML/AI tools is not terribly difficult, and the principle of how they work is simple enough (feed the model examples of what you want, eventually it can reproduce similar ideas when prompted or recognize new examples as something that it's seen before). Maybe it's ignorance on my part in regards to what it takes to learn these technologies deeply, but right now, I wouldn't even know where to begin to start studying to really learn how this technology works/operates on a fundamental level.
Despite my ignorance of the subject, I could probably work out how to step into an ancient archaic COBOL system or whatever, but ML/AI _feels_ so far out of my reach as a webdev.
This is one reason I think ML as a topic is cool, even if I think the practical uses for the every day person is still really far away from being a realistic.
I avoided many areas of computer science and programming for far too long because I thought they were "hard". Many of them turned out to be way easier than I expected (I often found myself wondering, "why didn't someone tell me how easy this is! I could have done this years ago!!"), or so much fun that working on them felt effortless, despite the increased effort in objective terms.
I guess what I'm saying is, start learning about what interests you today! Make small and consistent progress.
I've found that setting aside time every day (rather than just when I feel like it) to study areas of interest has been extremely helpful in this regard.
ML is just math. We're at a point where you don't even have to completely grok the math to apply ML techniques productively.
If you want somewhere to start and need a project idea, read up on how to build a simple binary classifier. The "hard" work is building good training and validation sets if you use a ML framework.
So I don't think you need to worry about being left behind or that your skills are stagnating if you're not directly developing ML models. There's no existential reason to jump in, unless you're particularly interested in ML.
Aren't ML models kind of the front lines these days? If I'm understanding correctly, it's the models, the training techniques, the curation of data sets-- These are the things that will inform the next generation of products and services.
I agree that there's more to it that just letting ML models loose, but it certainly seems like the core of it.
Interesting. Is Stable Diffusion a product of neural network size then? Is the size of the network a function of chips density? Also is there a Stable Diffusion app that currently works on edge devices?
Even bolder prediction: When we finally understand how the brain actually does it, the algorithmic improvement will be so enormous that the machine learning tasks which run on massive servers today will be able to run on the phone currently in your pocket.
The semiconductor researchers spend a lot of effort to make ever-smaller transistors. What is a transistor? It’s a tiny switch.
The ML researchers meanwhile use the language of linear algebra to define mathematical transformations of real numbers with nice differentiability properties.
The chipmakers are then tasked with reconciling the two. So they use transistors to make gates. And gates to make adders. And adders to make integer multipliers. And integer multipliers to make floating point multipliers. And fp multipliers to implement matrix multiplication. And now you can run your cat diffuser model on those transistors.
But what is the chance that the configuration of transistors in a floating point multiplier is anywhere close to the most efficient transistor configuration for learning?
The only reason we’re using multiplication of real numbers is because the math people said so.
I think what we get wrong is that individual neurons rarely represent anything. They are a medium for the waves. The waves are the currency of thought. A brain is a series of electro-mechanical oscillators that resonates with abstract concepts and patterns.
AFAIK, most research is still using the old "neurons represent single things" paradigm. Someone needs to tell them, there's no such thing as a "grandmother neuron".
If you dig into how neurons work in the brain you'll discover that a single neuron has the complexity of a large neural network internally and it's behavior is not nearly as simple as the typical model explanation. Different ion channels, time-dependent behavior, up/down regulation of neurotransmitter receptors and release, and much more.
It is entirely possible that the brain "does it" by throwing vastly more computing resources at the problem than we previously believed.
I think that even when we understand the brain completely, it will be very difficult disentangling what is useful for artificial neural networks and what doesn't really matter.
[0] https://www.lesswrong.com/posts/K4urTDkBbtNuLivJx/why-i-thin...
[1] https://www.lesswrong.com/posts/6Fpvch8RR29qLEWNH/chinchilla...
[2] https://twitter.com/_akhaliq/status/1479265403142553601
[3] https://twitter.com/TacoCohen/status/1584499066410790912
Scaling diagrams also showed no sign of plateauing AFAIK.
The industry has chosen angstroms as the next unit. Where 1.4nm will becomes 14A. ( Intel for now, but TSMC has uses the term in a few of its presentation )
>If TSMC/Samsung/Intel can squeeze out just one or two more nodes,
We have a very solid roadmap all the way to 1nm, or 10A by the end of this decade. As long as the market is willing to pay for it. At least TSMC 3nm and 2nm is pretty much done.
>limited general-purpose AI in every home
Like the comment below, I am not convinced the brute force approach works. You can already build a 800mm2 NPU today that is equivalent to a chip used in "in every home and office." by the end of this decade. But we are still no where near it.
I am not too fond of this, but obviously thats the easiest way to explain the significance of research results.
Getting a new material into a semiconductor manufacturing process takes years or even decades. Whatever material will be in TSMCs "1nm" process, it is already known.
Another one of those "right around the corner" things that is always coming soon.
Some of it is pretty fun for playing what if with (enforced) pastoralism. Using primitive materials with a modern understanding of them, instead of limiting yourself to historically accurate processes.
So unlike before, compression has stopped being a primary driving factor in the engineering of many types of devices and types of software.
We are still interested in feature extraction and scene composition, which both benefit from more compact representation.
To compete with hardware-assisted decoding in display hardware, fractal rendering would probably need to be coded to run in Vulkan.
edit, see Zeiss page on EUV: https://www.zeiss.com/semiconductor-manufacturing-technology...
The world will not be free of short-term thinking until the last MBA is strangled with the entrails of the last PMP.
> An EUV system uses a high-energy laser that fires on a microscopic droplet of molten tin and turns it into plasma, emitting EUV light, which then is focused into a beam.
is the spectrum naturally pure, or is there some filtering/selection going on?
> The tricky thing with EUV light is that it’s absorbed by everything, even air. That’s why an EUV system has a large high-vacuum chamber in which the light can travel far enough to land on the wafer. The light is guided by a series of ultra-reflective mirrors
The system uses 1MW of power!
1. https://medium.com/@ASMLcompany/a-backgrounder-on-extreme-ul...
https://www.degruyter.com/document/doi/10.1515/aot-2017-0029... see pre-pulse part.
Not entirely sure how you’d avoid voids with tapers facing upward, but you can orient them in five out of six directions. And I’ve seen a few examples of the weird games they have to play with light to get the crisp corners they are after so maybe tapers would be easier?
>The breakthrough relates to a new set of materials that can create monolayer—or two-dimensional (2D)—transistors in a chip to scale the overall density by a factor matching the number of layers. The teams at TSMC and MIT have demonstrated low resistance ohmic contacts with a variety of existing semiconductor materials, including molybdenum disulfide (MoS2), tungsten disulfide (WS2), and tungsten diselenide (WSe2).
So it is a better way to make a contact to a semiconductor material that noone knows how to manufacture reliably on a large area wafer.
Is anyone else using non-silicon materials yet?
Gallium
It’s used in current-dense applications and leakage current is out boogeyman with silicon.
Source: I helped with a research project using Ge. Here's the paper if anyone is curious https://ieeexplore.ieee.org/abstract/document/8653561
I wonder what happened to plans for optical interconnects?
What does matter to most people is transistor density and other performance metrics that actually have an impact on the product. The "Xnm" has always given a rough indication of the improvements in transistor density, and all the foundries has done is continue to name the nodes as if that's what matters... which is indeed the case.
TSMC does sometimes call the nodes eg "N7" now. But nobody that matter cares because everyone understands that "7nm" is a marketing term.
TSMC has continued to improve the process, and we've seen Apple and AMD design performant and efficient chips as the process improves.
What changes if you don't understand the exact measurements involved?
The actual visible effect of this is an economic one: we've gone from spending linearly more marginal CapEx per fab process-node upgrade, to spending geometrically more marginal CapEx. Each process node under 8nm has been twice as costly to design as the one before it. This is untenable.
This is such a HN issue. It's not tricking people because the vast majority of layman don't even care about node size, gate size or the manufacturing process. The number of times I have encountered a person outside of a tech bubble that cared about the manufacturing process of a CPU is exactly zero.
What people do care about: does it work, is it fast, is it efficient and can I afford it. At least for performance and efficiency the marketing term does give a somewhat reliable indication of the generational gains.
Analysts are often not from the exact line of business. They may have a bachelor's in engineering, or something, but they're usually not domain experts beyond having studied the field from a business perspective and gotten used to its lingo. So, these things can be kind of manipulative. That said, it's par for the course. It's better than outright lying and bribing the analysts like in the 2000-2002 telecom bust.
It makes no sense to speak about a transistor that would be smaller than the gate, there is no such thing.
Besides the active part, whose conductance is controlled and variable, the transistor includes parts that are either electrical conductors, like the source , the drain and the gate electrode, or electrical insulators.
While those parts may be less important than the active part, they also have a major influence on the transistor characteristics, by introducing various resistances and capacitances in the equivalent schematic.
What matters is always the complete transistor. The only dimensions in a current transistor that are around 1 nm are vertical dimensions, in the direction perpendicular on the semiconductor, e.g. the thickness of the gate insulator.
The 2D semiconductors that are proposed for the TSMC "1 nm" process are substances that have a structure made of 2D sheets of atoms, like graphite, but which are semiconductors, unlike graphite, which is a 2D electrical conductor.
In this case the thickness of the semiconductor can be reduced to single layer of atoms, which is not possible with semiconductors that are 3D crystals, like silicon, because when they no longer form a complete crystal their electrical properties change a lot and they can become conductors or insulators, instead of remaining semiconductors.
There is little doubt that for reducing the transistor dimensions more than it is possible with a 3D semiconductor like silicon, at some point a transition to 2D semiconductors will be necessary. It remains to be seen when that will be possible at an acceptable cost and whether such smaller transistors can improve the overall performance of a device, because making smaller transistors makes sense only when that allows a smaller price, smaller volume or higher performance of a complete product.
They never state it was the actual gate size in the first place. May be you should blame the media for it?
And just to reply to parents.
> TSMC does sometimes call the nodes eg "N7" now.
It is not sometimes. TSMC has always been very careful and calling it N7 or N10 all the way back to 28nm era. What the media, marketing and other PR decided to call it are entirely different matter. This is especially problematic since ~2015 when they have gotten thousands if not millions times more media coverage.
Hear hear.
Processors are just fundamentally complicated and boiling them down to a couple number is hard. Any modern one is pretty good and saying anything more than that -- depends on your workload I guess.
Nice.
And very consumer-friendly too, congrats.
Like 2 bil transistors/square cm, let's call it 2 btsc or something. Surely they use something like that internally, why not publicly? Maybe there are different kinds of "transistors", some (CPU) bigger than others (RAM)?
Moreover, density is only one metric of improvement, the other major one is power efficiency. Foundries want to market these improvements too.
Interally we might say 5nm to refer generally and N5 or N5P to refer to a specific process. When you need more specific numbers you look it up. Exact design parameters are carefully protected trade secrets.
Commenters on Hacker News care a lot more about this distinction than the engineers working with it. Nobody cares that it is not gate length or half pitch.
https://www.hardwaretimes.com/wp-content/uploads/2021/07/TSM...
A new metric could be more precise, but Xnm has mindshare, and works perfectly fine for the layman.
Hyperbole aside, it's undeniable that TSMC's chips for Apple have a greater computing-power/watt than Intel's chips -- but is that all they bring to the table? I don't mean to disparage power efficiency, it's awesome. But the question stands.
First order things that really matter are: - power dissipation - switching speed
These rarely get mentioned in marketing materials.