ARM Pioneer Sophie Wilson Also Thinks Moore’s Law is Coming to an End
nextplatform.com
nextplatform.com
The article says that 28nm will dominate for another decade, even though 14nm fabs exist. Having to use extreme ultraviolet (really soft X-rays) for lithography runs costs way up. EUV "light sources" are insanely complex, involving heating falling droplets of metal to plasma levels with lasers. It's amazing that works as a production technology. The equipment looks like something from a high energy physics lab.
It's interesting that we hit the limit of photons before the limits of atoms or electrons.
Another problem with all this downsizing is electromigration. Every once in a while, an atom gets pulled out of position by the electric field across a gap. Higher temperatures make it worse. Narrower wires make it more of a problem. This is now a major reason ICs wear out in use.
Getting rid of the heat is another problem. High performance CPUs are already cooling-limited. This is also why 3D IC schemes aren't too useful for active components like CPUs. Getting heat out of the middle of the stack is hard. Memory can be stacked, if it's not used too hard.
There's no problem making lots of CPUs on a chip, if the application can use them. Things look better server-side; you can use vast numbers of CPUs in a server farm, but it's hard to see what 20 or 100 CPUs would do for a laptop.
Drastically different architectures may help on specialized problems. GPUs have turned out to be more generally useful than expected. There will probably be "deep learning" ICs; that's a problem where the basic operation is simple and there's massive parallelism.
For ordinary CPU power per CPU, we're close to done.
It requires a few assumptions, but here it goes:
1) Assume applications increase their CPU demands significantly over the coming decades. Why? Who knows. Maybe (very plausibly) extremely advanced (compared to today) AI use & integration.
2) Move to, essentially, a CPU per process model
I have a couple dozen processes running on my system. Make the CPUs cheap enough and give me 30 of them.
Perhaps the average system will have four or five major AI agents running locally on it, that are particularly good at various things, and those agents will be constantly running processing intensive tasks (I'm assuming here that ~95% of all AI tasks will be performed in the cloud; which is to say I expect the computing power consumed daily by an individual in ... 30 years to be a hundred plus times what the average user is consuming per day now (doing things like watching YouTube or checking Facebook or running Snapchat or WhatsApp)).
Possibly for the same reason it's already increasing today: as software becomes more complicated (or complex, depending on your POV), we are shifting toward languages and abstractions better suited to manage that complexity, often at the cost of CPU cycles.
Examples: dynamic dispatch, interpreted code, libraries on top of libraries on top of libraries, boxing of machine-native types, bridging between old unsafe languages and new safe languages, automatic memory management, automatic bounds/overflow checking.
The problem, though, remains, that for most systems, unless each core is tremendously weak, you can run most of the ongoing tasks on a typical current desktop system on a tiny timeslice of a single core.
OSs certainly can paralleise more. E.g. AmigaOS relied far more on using multitasking as a basic OS primitive 30 years ago than most modern OSs, because message passing was exceedingly cheap (no memory protection, context switched involved moving a few registers). Something as simple as cut and paste could involve nearly a dozen processes (device drivers for mouse and keyboard, input handler, console device and higher level console handler, a process specifically managing the console clipboard (other apps might handle it themselves), the clipboard handler, the clipboard device, a filesystem handler for wherever the clipboard was mapped, the underlying disk device - usually ramdisk but could be anything, and more depending on setup). That was to make latency predictable first and foremost, at the cost of throughput, but the overall model could be adapted.
But part of the challenge is to write applications this way without just ending up being bottlenecked on the message exchanges.
Even if you can make that cheap enough, we still have the problem of the vast number of tasks that are simply very hard to parallelise. For some of them splitting them up and spreading them out certainly could make things like pipelining processing help us cut latency when something is done repeatedly, but whatever we do, we will have big challenges ahead.
[1] https://www.parallella.org [2] http://www.adapteva.com/announcements/epiphany-v-a-1024-core...
I'm willing to bet that "extremely advanced AI use" will not come to applications in the "coming decades". It's another fad (in fact a cyclic fad, it had made the same promises back in the 80s), like all the buzzwords that take over an industry for 3-5 years and then give way to the next.
In 2040, for all the predictions for its transformation or demise, the business and home desktop will look more or less like it does now.
The first time a hardware-side friend explained this to me, I was 100% certain he was BSing with hyperbole to emphasize the difficulties involved.
Now that I know otherwise, I am still amazed that such a complex method is the easiest one we have. It certainly makes some of my complaints regarding the complexity of software feel trivial!
This only if we want to keep the laptop in its current definition: a clunky device requiring lots of cables and holes in the casing, showing flat image at 4K at best and average quality audio, typing every single letter of a every command to get something done. No useful voice recognition/speech synthesis, no gesture tracking etc.
Specialized (sub)systems (media, wireless, smart networking) already use dozens of CPUs. Having generic CPUs to do the same work is not currently justified, mostly because of power and memory consumption. As these will go down, generic CPUs will assume more and more specialized tasks, because they are always easier to program, and probably cheaper to use overall (licenses, software and hardware tooling etc)
Thinking about it, 20 CPUs in a laptop seems like a minimum to me.
The laptop has been somewhat eroded by tablets and hybrids like the Microsoft Surface, but until we get rollup screens the form factor of 10"-17" portable screen with keyboard will continue to exist.
Mind you, I'm old enough that I took all my lecture notes longhand. I'd love to see a dictation system cope with maths notation.
(Thinking back, 20 years ago I went to a lecture on the subject of "what on earth will we do with >4Gb of RAM"...)
No, the advantage of voice is natural language queries. Imagine the following:
"<computer name> open the file I was working on yesterday."
"<CN>Find the email from Bob with the attached document on system requirements."
"<CN>Open the calendar and see when I can schedule a meeting with Alice and Jane."
"<CN>Find the picture of <wife> in front of a waterfall from our trip up north."
This requires smart apps that accumulate meta-data about your files and activities. That data then enables natural language communication with the machine, and that is what's needed to make voice recognition worth while. But with all the pieces I think you'll find it very rewarding.
I've often imagined a photo-realistic office assistant in a small window off in the corner that speaks and understands and takes care of things like this. It may seem like a gimmick but that's because nobody has done it correctly yet.
Those are gimmicks anyway. Unless incapacitated for some reason, nobody wants to talk to their programs -- except if it is combined with very advanced/strong AI to just utter some command and have them figure out all the rest (try talking out aloud for an hour for e.g. writing a post -- it gets old very soon). Making gestures in the air even less so, unless you want to switch on the lights (e.g. something you do very rare) or play a gesture tracking game to get fit. Those Minority Report style UIs is something that goods look in movies, but is not a good idea in practice (just like the translucent glass monitors they use in those movies. What's the idea there, who thought that, outside of a HUD, it makes sense to have anybody/anything on the other side of the glass compete with the on-screen data for your attention?).
>a clunky device requiring lots of cables and holes in the casing, showing flat image at 4K at best and average quality audio, typing every single letter of a every command to get something done. No useful voice recognition/speech synthesis, no gesture tracking etc.
You make the pinnacle of modern technology sound bad. Perhaps born quite recently and lacking some context?
Some of the sneer doesn't even make sense. "4K at best" -- the eye has a certain physiology, and for a laptop screen that's 15 or even 17 inches beyond 4K is marginal returns (and after that, no returns at all).
And "average quality audio"? Not sure what you have in mind, but audio in laptops is constrained by the speakers, not by the soundcard or CPU. It's trivial to have a great DAC in a laptop that can do 192Khz/24bit that's better than what a human can hear (or even tolerate: 24bit dynamic range allows for building-crushing audio).
a) laptop as a back-end for complex measurement system. Connected to a high-voltage signal, sitting in a crowded labs full of noise and cables. Would you like to reach out to type in a command every time, if you can speak or gesture instead?
b) laptop put next to a rendering device (video, audio, whatever), close enough so a wide-band wireless becomes efficient. Would you like to have all rendering devices perform the same computation, closed-sourced, paid-licensing, full of non-addressed CVEs? Or you'd rather concentrate the same processing in one place -- your laptop -- source code open to public inspection etc?
First, wouldn't that be a totally niche use? I'm pretty sure people who can't type (e.g. no hand control due to some accident/condition) are far more than people in "crowded labs full of noise and cables" using their laptop as a "back-end for complex measurement system".
So, I didn't say there are no uses for speech/gesture control (I even pointed to incapacitated users and such). I said it's gimmicky for most people, and this is not really a counter-example to that.
But those things aside, if the lab is "crowded" and "full of noise" wouldn't that make speech commands either perform badly (due to the noise in the room), or further annoy the others in the lab (due to them adding to the noise)?
That leaves gestures. But those wont work for complex commands, but just for stuff like "pick among a few options". Else you'd have to learn a whole gesture based language (like sign language). I'd take typing then -- or some special controller that encodes the options (e.g. a foot switch).
>b) laptop put next to a rendering device (video, audio, whatever), close enough so a wide-band wireless becomes efficient. Would you like to have all rendering devices perform the same computation, closed-sourced, paid-licensing, full of non-addressed CVEs? Or you'd rather concentrate the same processing in one place -- your laptop -- source code open to public inspection etc?
I'd call this out as another totally niche and contrived example. And my laptop is not "source code open to public inspection" anyway, much less the hardware in it. So why would I trust the processors in my laptop more than the processors in my rendering device? I bought both anyway, and have not inspected any of them.
Laptops can be roughly the size of a screen and a battery, and very sleek. They can do all power and I/O through a single USB-C cable. Laptops can do 3D, and if you find a screen that's better than 4K they can do that too. The only real limit is that it's fundamentally one screen, which is pretty minor. Laptops can have great audio. They can do voice recognition and speech synthesis. They can do gesture tracking.
And doing all those at the same time only requires a couple processors.
A couple of general-purpose CPU cores you mean?
How about: at least one CPU in WiFi subsystem, another in the Gigabit Ethernet card, another as high-speed interface host controller, video subsystem probably several (8K resolution / stereo, decoding + transcoding + (de)watermarking + error concealment + post processing), audio one or more (Dolby Volume, spatial sound), demodulation + demuxing of container formats at high speed at least one more maybe?
My point is really -- laptop should be able to do all that great stuff you enumerated, and more yet, all at the same time. As data volumes and our expectations grow, and Moore law comes to end, more cores will be needed, and I say there will be work for more of them than just 4 or 8.
I expect more cores to be used in the future, but all the things you mentioned are either already handled on the main cores, or take utterly trivial amounts of processing. They can all be handled on couple cores. They won't be what pushes growth.
For a counter-example, look at: http://www.trinnov.com/products/home-theater/magnitude32/spe...
Home Cinema Audio Processor. Intel i5 Quad Core (audio alone). High End. Today.
Some applications would benefit a lot from higher single core performance, some applications would benefit from many cores on the same mainboard. For well know reasons (little competition) we are stuck with desktop 4 cores (and 2 on notebook even the i7) with little increased single core performance for years. Now that AMD comes with more competitive CPUs with a good price, Intel finally dusts of some years old designs from their basement and finally releases their first i7 notebook CPUs with more than 2 cores. So AMD and ARM CPU competition to Intel almost-monopoly is very good to bring fresh air to the stagnated desktop/notebook/server market.
You have to be more specific here to support such a claim.
I have to do nothing. You could fire up Wikipedia or Google or Amazon...
Entire articles and books are written about Cray supercomputers and are waiting to be read by curious individuals.
For Cray 2, that's the Wikipedia article: https://en.wikipedia.org/wiki/Cray-2 it just gives an basic overview so don't jump to conclusions based on it. But it's worth to dive further in, really interesting topic.
If I do a search I will start with no knowledge of Cray and no context to find this answer, even the wikipedia page is too general. If I find common fan and water cooling solutions I might not see why they are better because that information might be too subtle for me as a novice. If you don't provide it is unlikely to be found this discussion; this conversion and you will be intellectually poorer for it.
The onus truly is on you to provide interesting information in conversation when you talk.
Getting rid of the heat in the Cray I involved lots of copper with fluid channels, like the stuff overclockers use. The Cray 2 ran the whole CPU in a tank of inert fluorocarbon liquid. This was generally considered a pain. The idea comes back now and then. One Bitcoin mining farm in Japan ran submerged in coolant. Everybody else just used air cooling, sometimes in a cold climate.
The issues revolving around nano-meter scale heat bottlenecks is fundamentally different. You're smaller than airconditioners and fans.
Wrong. Read at least the related Wikipedia article. https://en.wikipedia.org/wiki/Cray-2
Yeah, its nothing more than a fancy air conditioner. http://www.energyquest.ca.gov/how_it_works/images/air_condit...
In any case: I'm curious why you think this cabinet hold any use with regard to nanometer scale engineering tasks:
https://upload.wikimedia.org/wikipedia/commons/4/49/1985-Cra...
There's a big difference between "cooling down a cabinet" and "cooling down a 10x10nm hotspot"
-----------
IBM's "liquid cooled chips" might be a potential solution: where water is pumped through channels on the die itself. (!!!). The water absorbs the heat directly from the die.
https://www.cnet.com/news/how-ibm-is-making-computers-more-l...
But the technological hurdles to accomplish this sort of task are massive! IBM's done it, but no one else has figured out how to do it. So its not a simple task by any stretch of the imagination.
It's not only raw power. GPU programming comes with builtin asynchronism (waiting on promises of computations, so you can hide latency) and a memory hierarchy (local, block, global, host memory) each one being more local and smaller than the next level. This makes caching much more efficient, since a lot of values don't need to escape the thread/warp/block. The problem on CPUs is that memory is flat and considered equal.
CPUs have a similar memory hierarchy, the difference is that you're a context switch away from losing your preciously cached data. Even in GPUs memory is more often than not a constraint on the speed of computation, it is very easy to starve a GPU by accident.
My personal favorite trick to see if I'm using my GPU effectively is to run a benchmark on it to max it out and register the power consumed with an e-meter. As long as I don't hit that power level in my own application there are still improvements to be had (and I almost never manage to reach the maximum but I'm happy with 60 to 80% or so of that). It's one of the dumbest debugging tricks ever but it is surprisingly effective at figuring out if I've laid out memory access patterns properly.
That's an argument for more cores, each with it's own cache. It also suggests that the OS schedule all of the low CPU load tasks on the same core to prevent context switching on the big workloads. I'm not sure if that's how it's done today, but when my dual core is busy the CPU usage graphs tend to switch places regularly.
For instance, in a one particular quad core CPU all cores would have their own L1 cache, but cores 0 and 1 share the same L1 cache, and cores 2 and 3 another L1 cache. Then finally all cores share the same DRAM. So depending on whether your tasks have been 'pinned' to a certain core you might see a smaller or bigger hit when there is a context switch and a migration of a task that is not pinned to a certain core might flush L1 or both L1 and L2.
Of course every architecture has its own quirks so to gage the impact you'd have to know exactly how things are set up in that particular machine.
Another set-up (the one I'm writing this on) has 32 K per core L1 cache, 256K L2 cache and a 15 MB shared L3 cache.
Migrating from one core to another on this machine is cheaper than on the four core example above.
So apart from a generic inspiration to keep going further, this really does nothing for you. In fact you may still have further performance gains possible after this heuristic tells you to stop.
It's not suitable as a detailed profiler.
One way to do so is silicon photonics: it's possible to achieve 280 tbps bandwidth for 168W @ chip area of 4mm2 for 168W. Or you could triple that but you'll need ~1.5KW .
This of course will require better cooling, but DARPA has a project, using microfluidics channels, that can cool 1KW/cm2 and 30KW/cm2.
And as for power delivery - with TSV's you can deliver 150W/cm2.
And sure, such chips may be more expensive both in cost and power than today's are. But infinite memory bandwidth and photonics can make programmer's lives so much easier, and offer great acceleration in memory constrained problems, and make new architectures feasible(network computers with infinite bandwidth?). So maybe there's an initial market ?
And once we're there, and money starts flowing, maybe we'll see another type of "moore's law" focused on reducing the costs of cooling/photonics/power-delivery?
And BTW, some good news: ST(i think) is opening a photonics fab.
links:
https://www.extremetech.com/extreme/224516-microfluidics-dar...
https://pdfs.semanticscholar.org/4f9b/d24ea869f64c4457f69ad6...
http://drum.lib.umd.edu/bitstream/handle/1903/17153/TR%20DRU...
Push the latency problem of pointerful code closer to the data.
Currently DRAM is designed to fetch large pages at a time which are cached in the CPU. Cache is highly beneficial to latency when there is good memory locality in the program being run.
Our data hierarchies create graphs that have lots of internal pointers in them, and we are constantly warned about the cost of pointer indirections. Memory already knows how to deal with addresses. Teaching it to chase pointers should be easier than teaching it Boolean math.
[edit: it appears the real limitation to both our ideas is that virtual memory prevents you from making any decisions on the wrong side of the MMU. You could only ever make conditional fetches from the same page, which could be slightly useful but would be so hard to use I don't know who would bother]
It's all about the speed of the memory itself.
The real problem is the speed of the RAM, not the distance, for now. When RAM gets a request it can take many cycles to respond, sometimes as many as 50 or 60 before the first byte is returned.
If you are smart you can request large sequential chunks and after the first bytes things come down to the the CPU very fast, but that initial hit is still something a lot of people optimizing algorithms are wrestling with.
You can add more channels and more ranks into the system, which gives some small linear increase w.r.t. to concurrency, but brings a a big (channels) or minor (ranks) power increase. This doesn't change non-concurrent latency, so does not help sequential workloads.
ps: also mostly clockless and microwatt range
That's the thing with these guys, you're never sure if they're insane or just that the world doesn't tune enough to their ideas.
Parallela has managed to conduct a very succesful KickStarter. They've been under the radar since but still active, they published new plans for 1000 core coprocessor array. And I suppose, now they've proved themselves alone, they won't stop soon.
Systolic array/vector processors have been used since the late 60's for neural and other gemm heavy workloads. The Warp architecture from CMU is an excellent example.
https://www.dnd.utwente.nl/~tim/colorforth/PTSC/IGNITE_Proce...
Faster than clock processing before that was a thing.
But...
One has to embrace parallel languages instead of single-threaded scripting languages or the music is going to stop playing RSN, no?
Prophet Moore predicted the future and now engineers start breaking the law?
Isn't it the other way around that Moore made an observation about some effect that arose naturally? The formula was then called Moores law and its extrapolation had great predictive power for a long time.
Similar effects occur all through industries when you start scaling things up. Quality will go up and cost per unit will go down. Often following a simple mathematical formula which describes the learning curve.
In many technologies there is something called maturity where the straight line in the diagram starts to bend and approaches a technical limit. Markets overcome this a few times by changing the technological approach of solving a problem to an approach that has a better limit. This makes the general trend continue for decades... until the point where the next technology is so expensive that noone can afford it anymore.
Thus far Silicon has won every round and chip manufacturing plants cost many billions of dollars.
Even if you're not doing chip design, anyone planning more than one chip product cycle into the future needs to take into account Moore's law. If you're building a hardware system or even a software project, ignoring Moore's law means that by the time you ship, your product might be cheaper than expected but also missing features that are now cost-effective.
Architectures MUST change radically to adapt to this or there can be no progress.
Page 37 is what you're looking for.
Is this the cost in Joules or nanoseconds?
You can get a better summary of her early career from her computer history museum oral history interview: http://www.computerhistory.org/collections/catalog/102746190 - worth your time.
These days it's not a big deal to most developers, but I think over the next few years if there aren't major advances in speed we will want to get that extra battery life and speed out of our applications and devices. Independent Developers hopefully will have a good financial reason to do that, unlike today.
Intel Lake Crest: "will enable training of neural networks at 100 times the performance on today’s GPUs, said Diane Bryant, executive vice president and general manager of Intel’s data center group"
https://venturebeat.com/2016/11/17/intel-will-test-nervanas-...
Google TPU: "The TPU...used 8-bit integer math...process 92 TOPS" (trillion operations per second)
https://www.nextplatform.com/2017/04/05/first-depth-look-goo...
Generally:
http://www.moorinsightsstrategy.com/what-to-expect-in-2017-f...
https://www.extremetech.com/computing/247199-googles-dedicat...
They promised the same sorts of things with larabee and thunderbolt and Xpoint.
All have been massive underpreforming dissapointments.
https://www.hpcwire.com/2016/07/28/transistors-wont-shrink-b...
Going from Tick-Tock to Tick-Tock-Tweak... and this year to Tick-Tock-Tweak-Tuck the fourth year of 14nm (still as compact as other companies 10nm) makes the slowdown palpable. Perhaps they will manage a 2.7x shrink at their "10nm node" with or without EUV, but it's not the straight scaling of yesteryear.
http://www.anandtech.com/show/10097/euv-lithography-makes-go...
The problem with EUV is that the source power (laser excitation of plasma) is too low, making exposure times too long for the expense of the equipment.
If they could run the EUV sources at 4 times their current power, there would be huge increases in equipment sales and conversely exposure and mask complexity would decrease substantially. Increased exposure dose could deal with some of the shot noise effects (only 3% at 10nm so less than 8% ITRS spec?).
We would also probably see equipment free up to allow 14/20/28nm litho costs go down. So it's a win-win.
This article is from 2007 so it predates the AI renaissance. Lots of AI, ML, and optimization stuff can happily eat as many cores as you want to throw at it.
Then there's the multitasking angle. On a desktop at least I often run dozens of applications, developer VMS, etc. I could definitely use 20 cores in a desktop/laptop right now. We have tests that easily max out a 24 core server that I'd love to run on my own box.
What? That's just not true. Matrix multiplication is one such embarrassingly parallel workload that can go much faster than 20 times. Ray tracing very probably too.
In essence, Amdahl's Law trumps Moore's Law.
Instead of a 19X speedup on rendering, you get an 18X speedup on both rendering, physics, and AI, and gain some extra responsivity on the mean-time.
Luxrender for example has a problem here: all threads are writing to the same output buffer, causing bottlenecks and non perfect scaling. (Might be corrected by now.) The advise then was: run 2 or more Luxrenders, and combine the output image (luxrenders flm file).
In the mean time we might see a new golden age of computer architecture where the only way to increase performance is to question assumptions about how we design computers.
I have heard the dramatic "Oh no Moore's law is coming to an end" a dozen times during computer engineering courses. Professors are usually slow to adapt new information and it is already a couple years ago that I took those courses. I think that the transistor count has been slowing down for about a decade already.