The future of computing: After Moore's law
economist.com
economist.com
As machine learning/AI/neuron emulation becomes more useful, we'll see hardware specialized for that. It's not yet clear what shape that hardware will take. Look for "works fine, but is slow on GPUs" results to lead to such hardware, rather than "build it and they will come" projects like the Human Brain Project.
There's still more improvement possible in storage devices. With the 20TB SSD drive expected in 4 years, things are looking good in that area. For compute elements, cooling limits transistor density. Storage tends not to be heat dissipation limited.
4? Samsung is shipping (a very few select partners) 16TB SSDs now! I suspect in 4 years we'll be closer to 50TB.
I would also bet on a combination of specialized hardware for certain software, something along the lines of Bitcoin miners ASICs.
At Hot Chips a couple of years back Robert Colwell, who was director of the microsystems technology office at DARPA at the time, had a very interesting presentation on where things were going. One of the things that stuck with me at the time was his contention that there are lots of ways to improve performance etc. over time but CMOS was really pretty special.
"Colwell also points out that from 1980 to 2010, clocks improved 3500X and micro architectural and other improvements contributed about another 50X performance boost. The process shrink marvel expressed by Moore’s Law (Observation) has overshadowed just about everything else."
http://bitmason.blogspot.com/2016/01/beyond-general-purpose-...
This led me to think that maybe Moore's Law is looking at the wrong metric, and is there a more fundamental law regarding increased capacity. Some proof of this is in the drop in prices of cloud computing and increases in network capacity. If chips stopped getting more powerful, in theory, we may not notice as the capability of the entire network continues to increase at the familiar rate.
However, this begs the question, did this increased capacity actually begin with Moore's Law and computer processing in general? Or is there an overarching law of progress which has always existed. To prove this, I'd need to go back and look at the growth of industrialization. I suspect as the capabilities of the factories slowed, our network speed (rail and sea, then road and then air) increased. As the network speed reached a plateau, we began to find places where we could manufacture more cheaply, thereby increasing the output per cost. Does the same growth exist in agriculture? What major industry does not fit?
There is some evidence that a law similar to Moore's law exists not only for technology. What effect does that have in our understanding of why things grow?
[edit: this is was an except from my YC application to the question regarding what have you discovered]
graph http://globedia.com/imagenes/noticias/2011/10/17/singularida...
7 TFlops for $1000 http://www.pcworld.com/article/2898093/nvidia-fully-reveals-...
Or maybe not. It's impossible to predict true breakthroughs.
I'm not sure what value that has for predicting the future. You could view the airplane as the evolution of the chariot. Or not. Does it make a difference?
http://www.johndcook.com/blog/2015/12/08/algorithms-vs-moore...
I've always believed that humans do not have the ability to program things smarter than themself, because we do not understand our own intelligence, so we have no way to reproduce it.
At the time, I said the only alternative I can think of is make random permutations and pick the best one, and go from there. But I said this as a ridiculous suggestion that no one would actually do.
But then Google actually did exactly that with their Go machine!
So, that's what I think: The future of computing will be based on randomness and the job of the programmer will be to guide it, but not program it directly. (Can you imagine programming a webpage this way? Or writing a book this way?)
The point I was making is that intelligent agents are built all the time by less intelligent agents. Humans will build more intellgent AI using exactly the same approach and these AI's can do the same.
Of couse building an AI via evolutionary training opens up a whole lot of control issues. Do we really want to be creating entities vastly more intelligent than ourselves when we have no real idea of their interests and when we occupy something that is valuble (e.g. matter and energy) to this entity?
(Well, that took more space than I wanted. I downloaded one of your circuits papers for a gander.)
http://yann.lecun.com/exdb/publis/pdf/boser-92.pdf
http://www.kip.uni-heidelberg.de/Veroeffentlichungen/downloa...
It would be interesting to see someone combine the principles of your work with analog implementations on a decent process node. Yours is kind of like a hybrid between properties of analog and digital cells. The real thing might be even more effective albeit harder to automate. There's some analog EDA but it's almost always custom work.
I would argue that this is not what they did. What they did was take something a human player does (study games to learn how to play). And then parallelized this to a level that a human can't match. Basically the equivalent of having one go player playing against an entire team of go players who are all experts.
Cost mostly, If you can replicate a human-level intelligence for 10,000 (or 100,000) and have it work 24/7 with no time off and scale out to thousands of them then you'd have something absolutely terrifying in it's capability.
Image an AWS of a 1000 von-neumann level intelligences working co-operatively 24/7 on a problem.
The stuff of sci-fi right now but maybe one day, we know it's physically possible to build human level intelligences since we prove that the rest is 'just' engineering.
Beyond simple architectural improvements, we could still move beyond basic transistor based computing. The most common example is quantum computing (which offers an asymptotic improvement in some cases), however I can imagine there beyond other classical devices that can compute certain functions more efficiently than a pure transistor based solution can.
Modern x86 and high performance ARM cores are almost unrecognizable compared to processors in the late 90s, much less the 80s. (Also, ARM was founded in 1990, not the 80s).
> While I have no doubt that the implementation of these architectures has improved to reflect modern manufacturing capabilities, this still suggests that there is room for architectural improvements in performance.
There are still performance improvements to be had, but it's not going to be anywhere near the performance scaling of Moore's law. The rate of architectural improvements is also slowing down as well (and increasingly only applicable for a smaller and smaller fraction of workloads).
> I can imagine there beyond other classical devices that can compute certain functions more efficiently than a pure transistor based solution can.
Like...? Quantum is 10+ years away right now and unlikely to get fast any time soon. CMOS has had decades and billions of dollars invested in scaling; it's going to be a long time before any of the current "CMOS killers" (virtually all of which are still transistors) reach parity.
The official Acorn RISC Machine project started in October 1983.
But everything else you say is bang on.
The number of problems for which quantum computing offers a speedup is very limited. It is absolutely not a general, all-purpose computational architecture.
Having X86 and ARM chips means you can run legacy code on them from a long time ago. Which is why they are so popular. But once we go to optical chips for quantum computing there will be no legacy apps from X86 or ARM on them. It would have to write apps from scratch.
Windows 10 is the last version of Windows for a reason. Microsoft is porting their enterprise stuff like SQL Server to Linux. https://blogs.microsoft.com/blog/2016/03/07/announcing-sql-s...
You can tell that Microsoft knows that Windows is getting long in the tooth, and has to support legacy code, and still has old code in it for compatibility reasons. They look at other operating systems like Linux to port their enterprise apps to in order to sell tech support and take a stab at Oracle, MySQL, and PostgreSQL. They know that Linux gets ported to different platforms and so can SQL Server to give them a larger marketshare.
The old X86 and ARM designs are limited due to legacy support of older programs. But they are marketed as backward compatible with older chips.
IBM's Power has been open sourced as OpenPower. http://openpowerfoundation.org/
The Dragonball CPU tries to build on the M68K family of processors. https://en.wikipedia.org/wiki/Freescale_DragonBall
So you got a lot of backward compatible CPUs out there, that are limited because they have to support legacy code. The new processors that don't have backward compatibility should run faster with fewer quirks and use new technology not from the 1980s, but programs have to be written from scratch or ported from other platforms.
Windows 10 is the last version of Windows because Microsoft knows that it will eventually have to drop compatibility in order to compete with the new systems and new processors out there that don't have legacy support. The X86-64 chips are a dead-end, and Microsoft has to look to other technologies and a different operating system. Linux is a good choice to support even if Microsoft does not officially have a Linux distro yet. If they open source their enterprise tools like Dotnet or CLR or Roslyn or Visual Studio Code SQL Server to Linux and OSX and other platforms they can sell tech support for it via their paid hotlines. Even making iOS and Android apps for Office and other things.
Microsoft is going to move away from X86-64 and Windows eventually, and focus on The Cloud instead and Azure in hosting VPS operating systems. Then when the new design of processors come out that put X86 and ARM to shame they can port their programs to that new platform.
If you remember the original Micro-Soft business model was to make programs for computers that other companies made and make them for different operating systems. They only got into DOS because IBM made them an offer they couldn't refuse. Windows was basically their attempt at making a Mac GUI for DOS, and working with IBM to bring OS/2 was yet another GUI attempt, but they quit OS/2 and focused on Windows instead. OS/2 was going to be ported to PowerPC, MIPS, Alpha, SH4 and other RISC processors because IBM and Microsoft saw the limitations of the 80X86 processors and wanted something new. But it fell apart. Then Windows NT 4.0 was ported to MIPS, Alpha, etc but abandoned. Windows RT was ported to ARM but flopped. Every attempt to move away from 80X86 processors met with disaster because people wanted to run legacy code.
But soon it will be a new day with new processors and new computers not based on 1980s designs and using Linux or some other FOSS OS and connecting with Cloud computers.
I think pg originated his "sufficiently smart compiler" startup idea in his pycon talk. You can find it online somewhere. The other take away from his talk was: just lie to customers about it being automated, manually farm out the parallelize-all-the-code tasks to works/interns/turks while saying it's "automatic," then eventually figure out how to automate it yourself later so you don't need pesky humans in the loop.
"""
Other languages skirt these issues by running on a single CPU with multiple processes to achieve concurrency; however, as Moore's Law pushes us towards increasing multicore CPUs, this approach becomes unmanageable. Elixir's immutable state, Actors, and processes produce a concurrency model that is easy to reason about and allows code to be written distributively without extra fanfare.
"""
I believe a more plausible link exists between the end of Moore's law and the rise of open hardware as explained in this article:
http://www.eetimes.com/document.asp?doc_id=1321796
TL;DR
If eight year old hardware is almost as fast as today's hardware, there is ample time to reverse engineer competitive open hardware.
Black phosphorus anyone?
For example, people did not go from buying 1 car to buying 10 cars and then 100 cars - most of us hit saturation somewhere between 1 and 2, and stayed there. Similarly, the average speed of our cars did not increase at an increasing rate except in the earliest years of automotive technology; we have a "highway maximum" between 50 and 70 MPH, and a "technical maximum" in the high 200's range for production cars which do not rely on rockets or other not street-legal tricks (a quick search brings up the Hennessey Venom GT at 270.49 MPH). Likewise applied to fast moving vehicles as a whole, including supersonic aircraft and rockets, we've already brushed up against vehicle weight and power density limits that slow the rate of improvement in acceleration.
Applied to the Moore's Law measures, that indicates we have a "endless sunset" period ahead of us where we'll still get more doublings of semiconductors, but they'll come increasingly slowly as more fundamental innovations become necessary to realize them.
What we don't know is whether there is a logistic curve on technology as a whole. Belief in the Singularity is premised on this not being the case.
That's a version of the Singularity to extreme even for Kurzweil. It's premised on technological growth not hitting the log portion of the curve before machine intelligence passes humans.
However, one interesting argument for continuation of technological advancement beyond what we might call "singularity" levels today. If technology maintains an exponential tragectory for another century or so through a few more breakthroughs then we would have some amazing tech.
So, it is not required that tech advancement is exponential -- as long as we are still early enough in the sigmoidal curve that more exponential (and linear) advancement is still to come. With biotech, quantum and the algorithmic side of AI I think we still have quite a bit of advancing to do. That said, the singularity stuff is still ridiculous woo woo.
When a technology limit is reached, but not a demand limit, interest and capital flow to other technologies for meeting that need, starting a new s-curve. eg peak oil prices lead to fracking.
Whether it will be at Moore rates we don't know. But the possibility of far superior information technology has a proof by example: biological neurons.
"And despite Silicon Valley's ostensible belief in rationality and empirical evidence, we continue to assert this despite the data strongly suggesting that it just ain't so." - Patrick Collison
The race to 7nm is expensive.
I think the next big leap will be in Chip Design software which lives under a rock inside a ditch in an anachronism that is EDA - electronic design automation. There's got to be a better way that allows a chip designer to produce a high fidelity mixed signal analog chip at speed and scale without having to pay $25K+ for the software licenses and tooling.
VLIW has been around for a while (Itanium is probably the most famous "general purpose" example) and has failed to gain traction outside of GPUs and DSP (ie not "ordinary code").
A single core processor, of course.
> What does ordinary code mean?
Say, simple random sample of all the other code being run.
> VLIW has been around for a while
Yup. I never claimed otherwise.
> Itanium is probably the most famous "general purpose" example
Yup.
> and has failed to gain traction outside of GPUs and DSP (ie not "ordinary code").
Yup.
Still, yet again, over again, one more time, once again, VLIW can get 9:1 speedup.
Why mention this? Because the OP was talking about the challenges of getting faster computing. Well, if want faster computing, one approach, that, indeed, works on general purpose code, and gives ballpark 9:1 speedup is, and may I have the envelope please [drum roll, please], VLIW. Really new? Nope. Tried with Itanium? Yup. Works? Yup.
Bottom line -- we still have 9:1 speedup available to us.
Maybe an objection is that pay a factor of 24 in transistor count and electrical power but get only a factor of 9 in performance.
Not everyone knows about VLIW.
Where did I say something wrong?
All VLIW does is move a lot of on-chip logic off-chip into the compiler. This only works for a small set of computing tasks - which is why the closest thing we have to VLIW today lives in GPUs. And why Itanium was nicknamed Itanic.
It's a non-starter for general computing because as soon as you start dealing with real-time conditions the compiler can't optimise in advance, the speed advantage turns into a speed penalty.
Yes, the compiler is involved. So what?
I think it's unfair to generalize from VLIW to everything published 20+ years ago. Plenty of the things discovered back then, or earlier, are still applicable and in production today. Your compiler picked most of its low hanging fruit ages ago. Lots of PhD theses are rehashing old ideas, often unknowingly. Results are results, what matters more is whether there's a relevant context, as you've highlighted.
"Real time"? What?
IIRC, the instruction traces didn't show 9:1 or anywhere near that.
I'm sorry about Itanium, but VLIW has to remain a possible path to faster cores. That my information is old does not mean it is wrong: The guy got 9:1 speedup on 24-way VLIW. Saying that the tricks of branch prediction, out of order execution, speculative execution, register renaming, etc. make VLIW forever obsolete is shaky without some solid references.
Moreover, if we want faster cores, then obvious, right in front of us, are two possibilities: (1) Design instruction sets that make VLIW easier to do and more productive. (2) Integrate and coordinate up and down the stack, that is, from application, e.g., collection classes, string operations, function calling, memory management, exceptional condition handling, to compilers to instructions to VLIW to the gate level logic and look for speedups. E.g., the now famous instruction sets look like they were designed for assembly language programming, and likely no compiler makes good use of all the instructions. E.g., C code, supposedly fast, forces the programmer to do the multidimensional array indexing arithmetic themselves, and to the machine language it all looks like normal work. In fact, that addressing is necessarily ubiquitous across computing, so maybe have an instruction for it. Same for string compare -- the usual C approach is just comparing one character at a time in a loop -- bummer.
In tool making, a key is to design the right tools. Else end up with a huge toolbox where most of the tools are used little or not at all. IMHO, we are still looking for the right tools in the stack from logic gates to microcode, register sets, instructions, caches, compilers, and applications. Then, VLIW in some form needs to be kept in mind.
For example, a quote from the article:
"Moore’s law was never a physical law, but a self-fulfilling prophecy—a triumph of central planning"
The physics and triumphs of engineering were all about physical law; "The end of Moore's law" was always just around the corner because of physics.
No central planning led Intel to invest their billions in R&D.
Central planning can neither force nor halt additional refinements in transistor density or alternate ways to compute.
Intel failed to keep up with every 2 years back in 2012.
CEO of Intel, announced that "our cadence today is closer to two and a half years than two.” This is scheduled to hold through the 10 nm width in late 2017.