Google supercharges machine learning tasks with TPU custom chip
cloudplatform.googleblog.com
cloudplatform.googleblog.com
I'm a bit surprised they announced this, though. When I was there, there was this pervasive attitude that if "we" had some kind of advantage over the outside world, we shouldn't talk about it lest other people get the same idea. To be clear, I think that's pretty bad for the world and I really wished that they'd change, but it was the prevailing attitude. Currently, if you look at what's being hyped up at a couple of large companies that could conceivably build a competing chip, it's all FPGAs all the time, so announcing that we built an ASIC could change what other companies do, which is exactly what Google was trying to avoid back when I was there.
If this signals that Google is going to be less secretive about infrastructure, that's great news.
When I joined Microsoft, I tried to gently bring up the possibility of doing either GPUs or ASICs and was told, very confidentially by multiple people, that it's impossible to deploy GPUs at scale, let alone ASICs. Since I couldn't point to actual work I'd done elsewhere, it seemed impossible to convince folks, and my job was in another area, I gave up on it, but I imagine someone is having that discussion again right now.
Just as an aside, I'm being fast and loose with language when I use the word impossible. It's more than my feeling is that you have a limited number of influence points and I was spending mine on things like convincing my team to use version control instead of mailing zip files around.
My guess as to why they're announcing the TPU is that they are feeling the pressure from Facebook and other AI labs, and want to reinforce their reputation as being the best place to do AI research. By revealing that AlphaGo was based on this hardware, they indicate to researchers around the world that if you want to build the most advanced ML models you need to be at Google. Same reason they talked about MapReduce/GFS back in the day.
Interesting, as the nature/science paper made no mention of this, it was exclusively trained on GPUs.
Between that time and the Lee Sedol match, the hardware running AlphaGo was switched, to these TPU.
From the paper: "The final version of AlphaGo used 40 search threads, 48 CPUs, and 8 GPUs. We also implemented a distributed version of AlphaGo that exploited multiple machines, 40 search threads, 1,202 CPUs and 176 GPUs."
It's a bit underhanded, however. IMHO, the player should be able to study recent games before the match. But this is pretty typical, there were similar late-stage improvements with Chinook (checkers) and Deep Blue (chess).
Coming out saying they can give you a service no one else can right down to a custom chip may sway a few buyers in the market.
They've probably carefully modeled the energy savings over GPU/FPGA and found it to be substantial enough, even taking into account costs of design changes.
As for testing for correctness, ASIC designs are almost always tested in software simulation instead of on FPGAs. Mainly because RTL written for FPGAs can look quite different from RTL written for ASIC synthesis. There's also a lot incidental engineering work needed to get something working on an FPGA.
It's actually very common to test ASIC RTL on FPGA. You are right that if you are targeting an FPGA you may write things differently, but this doesn't mean RTL written for an ASIC won't work it's just it might not clock as fast as it could and may not make the most efficient use of resources.
There is a lot of engineering work to take an ASIC targeted IP and run it in an FPGA, but thorough verification of an ASIC design is extremely important so it's usually worth the time to bring up an FPGA that can run it (sometimes as it's too large you have to split it across multiple FPGAs).
Especially if you try to make it go faster
You'll also break a design down into units or blocks. These can be tested separately. E.g. you might see an L1 cache as a single unit and have a testbench (piece of HDL to stimulate and check a design) to test only that. This is useful as it's far quicker than a full design test however extensive full design testing is still required as many bugs you see as the consequence of the interactions between units.
Once you've done a lot of testing you'll tape out a chip. This may well have some issues so you'll do a new 'spin'.
Obviously details and plans how it was done is where all the good stuff is, so that being hidden is understandable.
Anyone the casually follows AI knows that people have been talking about making DNN ASICS for some time. It was all a matter of time and $$$$$$
There is no doubt FB is working on them too. Which is why Google is finally publicaly saying that "we did it first ;)"
Makes sense. In that respect yeah, they probably wanted to keep it under wraps to avoid Facebook/others from getting a timeline estimate out of it and jump ahead.
That's kind of what Movidius says, too, about its Myriad 2 VPU, which is kind of a GPU (SIMD-VLIW) with larger amounts of local memory combined with hardware accelerators.
It's a "more details to follow" type of thing. Pretty standard actually.
They say 10x performance / watt, nothing about performance per unit time.
The speed at which an ASIC will run is constrained by temperature (power dissipation) and and logic timing, which itself has a dependency on temperature.
So we could call that vertical scaling, to some power ceiling which may not take us all the way to 10x, but it's not impossible.
Then there is horizontal, which I assume is applicable to these problems... running more in parallel.
In both cases, I think it's safe to assume they are getting a performance increase in the instantaneous sense.
While I agree some performance per unit increase is likely, how does a direct 10x increased based on power savings follow? Less power usage does not mean that the chip can run through more flops in the same amount of time, right?
http://electronics.stackexchange.com/questions/122050/what-l...
(see graph in the first answer)
Also, it's not known that the TPU have a way to allow to increase the clockspeed arbitrarily, nor is it known whether their architecture is capable of ensuring correctness at arbitrary clock frequencies. Some architectures make assumptions like "The time for this gate to reach saturation is very small compared to the clock frequency, so we'll pretend that it's instantaneous."
Well, that's the kind of metric you'd expect from a cloud provider. That's what's important to them.
If you're a tinkerer dabbling in TPU acceleration on your gaming/coding PC alongside with GPU acceleration, then the metric that would be interesting for you is speed increase per unit.
It's a post saying "that's how we do it" that serves a bunch of political and PR goals, but it's quite likely that this will stay an internal technology; maybe available indirectly as a cloud computing offering - in which case they'll give the specs and price of the whole solution, not of a particular model of TPU chip.
See 'low-power (inexact|approximate) computing'
A bit old but cf. DE Shaw and Anton https://en.m.wikipedia.org/wiki/Anton_(computer)
and Anton wasn't really "scale" in this sense. It was a vertically scaled single (or several) machines.
But that's just terminology/convention. One could argue a GPU is an ASIC, and a CPU is an ASIC. The only thing to argue is how specific does an application have to be to call it an ASIC instead of some other made up name.
The prime differentiators are the process, which is the term used for the many steps of fabrication of an IC (processes are often referred to as nodes, distinguished by the smallest feature size of a transistor they create), and the ability of the design engineers to create an optimal design, in terms of boolean logic and semiconductor physics.
Given the right team of engineers, and a top notch foundry, and a great deal of experience in the problem domain (machine learning in this case), a custom IC could very likely trounce a GPU.
GPUs were designed for the domain of graphics processing, which happens to have some commonality with the processing in machine learning. But, at least until recently, GPUs weren't focused on machine learning. Just graphics.
Now the GPU vendors are trying to leverage their knowledge of graphics processing and building of graphics processors to create machine learning processors, but the thing they are leveraging could also be what handcuffs them. Which gives opportunity for a company like Google to do a fresh take on the problem domain without the baggage of the knowledge of graphics processing.
Rrally interested to know how this was managed...
Competitive advantage is protected by custom hardware (and huge proprietary datasets).
Everything else can be shared. In fact it is now advantageous to share as much as you can, the bottleneck is a number of people who know how to use new tech.
"Sun's two strategies are (a) make software a commodity by promoting and developing free software (Star Office, Linux, Apache, Gnome, etc), and (b) make hardware a commodity by promoting Java, with its bytecode architecture and WORA. OK, Sun, pop quiz: when the music stops, where are you going to sit down? Without proprietary advantages in hardware or software, you're going to have to take the commodity price, which barely covers the cost of cheap factories in Guadalajara, not your cushy offices in Silicon Valley."
Developers are a complement to hardware, software, and data based businesses. Google probably spends more on employees than on hardware or data. They really want to commoditize developers that can work on their stuff.
Google's business is based on having more and better data than their competitors. In many cases they have been happy to commoditize hardware, making open their datacenter designs for instance. That helps to cheapen their costs for building datcenters.
In other cases they open source software, like tensorflow and go language. These choices are made to commoditize developers. Google wants there to be a big pool of people who know how to use the technologies that Google uses. More developers means less costs for Google to hire and train employees. Which is their biggest expense: win!
With the TPU, as long as no one else is doing that kind of hardware for machine learning, it is a proprietary advantage to keep it secret. But at some point the logic flips: when others start to do similar things Google would rather commoditize their version of the tech. Because the lesson of the last 40 years is any widely used hardware WILL become commoditized. The inertia is with software codebases and developer knowledge.
Developers developers developers developer developers!
AWSs offerings seem fairly vanilla and boring. Google are offering more and more really useful stuff:
- cloud machine learning
- custom hardware
- live migration of hosts without downtime
- Cold storage with access in seconds
- bigquery
- dataflow
- Cloud Shell
to name a couple more.
(I work on TF this year.)
(To elaborate -- it's questions like "how deep should I make this convolution? Should I use tf.relu or tf.sigmoid? How many fully-connected layers should I put here, and how big should I make them?". These are really knotty deep learning design questions, but they're often h/w independent. Not always - we certainly have some ops on TF that we only support in CPUs and not on GPUs, for example - but often.)
1. Best price/performance is tensorflow right now. So, the best software choice is platform X.
2. Then in 2 years.. Well we are using Platform X so tensorflow is clearly the best option.
In other words once you pick conv2d, you tend to also stick with whatever conv2d is optimized for. Which also means HW vendors love to help optimize popular platforms.[1] It has been climbing the charts at https://github.com/soumith/convnet-benchmarks for example.
I read "Vanilla" and "Boring" as "Horray, I don't have to spend time rewriting all this complicated code I already have!"
If I'm just dipping my toes into (say) Caffe or Theano, I don't have to rewrite it from scratch.
That is a huge advantage---not a disadvantage!---of AWS over google.
Google does boring stuff very well too.. and one can argue much better than AWS as well.. take a look at Quizlet's story: https://quizlet.com/blog/whats-the-best-cloud-probably-gcp
(shamelessly biased Googler)
I'd love to hear what other boring stuff has been a showstopper for you, in case we missed something dumb :)
If you're evaluating something today, how does it change your decision that we were late to market with Compute Engine (and in this specific case "bring-your-own-kernel")?
If it's about future boring stuff, I think the list of boring stuff isn't too long ;).
Disclosure: I work on Compute Engine.
A solid guarantee with AWS is if AWS goes down, then a multitude of Amazon's services also will go down(ex Amazonian myself), so it gives me a belief that AWS's uptime is more important to Amazon itself that it is for external customers.
(Firebase Engineer here)
Before we had custom machine types (November 2015 GA), we wouldn't have been remotely close to what they needed. I'm not even sure we've had anyone evaluate the amount of overhead KVM adds in either latency or throughput.
tl;dr: Don't let Search be your "not until they do it". We've got folks in Chrome, Android, VR, and more building on top of Cloud (as well as much of our internal tooling being on App Engine specifically).
Not implying that AWS hasn't had them. It's just that adopting GCE this early makes you a bit of a guinea pig because GCE isn't used internally at Google.
Microsoft explored something similar to accelerate search with FPGAs [2]. The results show that the Arria 10 (20nm latest from Altera) had about 1/4th the processing ability at 10% of the power usage of the Nvidia Tesla K40 (25w vs 235w). Nvidia Pascal has something like 2/3x the performance with a similar power profile. That really bridges the gap for performance/watt. All of that also doesn't take into account the ease of working with CUDA versus the complicated development, toolchains, and cost of FPGAs.
However, the ~50x+ efficiency increase of an ASIC though could be worthwhile in the long run. The only problem I see is that there might be limitations on model size because of the limited embedded memory of the ASIC.
Does anyone have more information or a whitepaper? I wonder if they are using eAsic.
[1]: http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=701142...
[2]: http://research.microsoft.com/pubs/240715/CNN%20Whitepaper.p...
If you read the academic literature, they talk about 100x-3000x energy savings with asic vs GPU. So in that light Google's 10x improvement sounds low, and could certainly fit an eAsic story.
Furthermore, eAsic has fixed wires, but logic is defined via sram configuration(AFAIK) and not fixed, so it could offer a level of programmability .
I hope we can at least see some white papers soon about the architecture--I wonder how programmable it is.
So... just use Google's machine learning cloud thingy.
The software can build the community, where the supercharging is only available when you run it on Google cloud.
(although GPU performance isn't bad either, so you don't have to, thus community)
Machine learning isn't just targeting the AI researcher market though -- it's widely used by a huge number of companies, and of course, by many of Google's most important products. I would argue that those markets combined are larger than gaming.
Google is doubling down on hosting as a source of future revenue, and they're doing that by building an ecosystem around Tensorflow.
What I think is interesting is how weak Apple looks. Amazon has the talent and money to be able to compete with Google on this playing field. Microsoft is late, but they can, too.
Where's Apple? In the corner dreaming about mythical self-driving luxury cars?
[1]: http://spectrum.ieee.org/semiconductors/design/the-death-of-...
Where Apple really looks weak is in datacenters, networking, and cloud services.
I get the feeling from today's announcements that Google sees the 2021 version of Google Now as the selling point for their 2021 Nexus line.
I don't think Apple is preparing to compete on that.
Whether that's good or not may be arguable, but it's certainly a selling point for many and I don't see Google or any other company's offerings approaching the same experience, and I suspect that's by design; they have to be more open and support all devices but that kinda dilutes everything. Apple will only get stronger in that aspect IMO.
I’d hope someone somewhere steals the blueprints and posts all of them publicly online.
The whole point of patents was that companies would publish everything, but get 20 years of protection.
But by now, especially companies like Google don’t do so anymore – and everyone loses out.
EDIT: I’ll add the standard disclaimer: If you downvote, please comment why – so an actual discussion can appear, which usually is a lot more useful to everyone.
Re EDIT: Downvotes must be comment-mandatory or not allowed otherwise.
Look at how much faster ASICs for bitcoin mining are than the GPU... orders of magnitude.
Additional die space on additional functionality might hurt the power envelope (which is where the focus on performance / watt rather than performance kicks in) but it doesn't make your chips slower per se.
Furthermore the fact that ML can be error tolerant means you also get to optimize certain floating point operations for speed or energy efficiency at the cost of accuracy. NVIDIA doesn't get to do this in their linear algebra support.
{(others, ~bottom) (google, ~top)}
Couldn't see more, but after Nvidia claiming overwhelming power with their latest GPU architecture including in the ML domain .. I was surprised.There isn't much data yet but I'm also guessing they probably have access to much more RAM than NVidia cards and can process much bigger data sets
evidence suggests they were mostly for defense customers:
http://forums.nekochan.net/viewtopic.php?t=16728751 http://manx.classiccmp.org/mirror/techpubs.sgi.com/library/m...
Sounds they also used this for AlphaGo. I wonder how badly we were off on AlphaGo's power estimates. Seems everyone assumed they were using GPU's, sounds like they were not. At least partially. I would really LOVE for them to market these for general use.
[0] If google is smart, they'd ditch +/- infinity and if they were ballsy, they'd ditch zero in their FP implementation.
For example -> https://petewarden.com/2016/05/03/how-to-quantize-neural-net...
In addition, there's a lot of literature on optimizing hardware implementations of fundamental arithmetic operations like addition and multiplication. I recall seeing a paper a while ago which talked about reducing the number of gates by allowing some bounded imprecision in the results - unfortunately, I don't remember the title right now, but it sounds like that's what they may be doing.
Link to the specialized hardware for SHA256 hashing: http://www.amazon.com/Antminer-~4-73TH-25W-Bitcoin-Miner/dp/...
http://www.catb.org/jargon/html/W/wheel-of-reincarnation.htm...
OT: Cool blog! :)
I don't think the average developer will need to understand chip design but I do think _many_ developers will need to know how to use deep learning frameworks.
Oh, and imagine a Facebook chip. :)
Then again, I've been coding for ~25 years, am an avid amateur astronomer, and have a degree in physics, so maybe the moral here is that some things just just take more time to master, beyond those first 6 weeks in a coding camp learning how to put together your first jQuery.
:-)
Maybe the next iteration will be assembly instructions specifically for Neural Nets built into CPUs...actually something like a convolution assembly instruction wouldn't even surprise me at this point.
http://www.thetalkingmachines.com/blog/2016/5/5/sparse-codin...
I wonder how general the gains from these ASIC's are and whether the performance/power efficiency wins will keep up with the pace of software/algorithm-du-jour advancements.
Per-unit manufacturing cost scales logarithmically. Even a single batch of custom silicon on yesterday's technology is only $30K. This is one of the reasons there is so much interest in RISC-V; hardware costs are not the barrier-to-entry that they used to be.
So yeah, the gaming market pushes the per-unit price of GPUs down, but even an additional 2x reduction in rackspace and power will pay for itself at the right scale.
https://2.bp.blogspot.com/-z1ynWkQlBc8/VzzPToH362I/AAAAAAAAC...
They probably didn't mean to use this version of the image for their blog - but I wonder what they were trying to indicate/measure there.
IOW, taking their claims at face value, a Nvidia card or Xeon Phi would be expected to smoke one of these, although you might be able to run N of these in the same power envelope.
But those bandwidth & throughput / card limitations would make certain classes of algorithms not really worthwhile to run on these.
Agreed. Also tells you that they don't need to communicate with the CPU much, given that it only has a PCIE. Reminds me of Knights Ferry, in this respect.
> a Nvidia card or Xeon Phi would be expected to smoke one of these
Will be very interesting to see some head-to-head benchmarks between these guys (on tensorflow and other libraries) in the next few months. Especially as Knights Landing starts to appear, and the new Nvidia card.
That seems unlikely, right? GPGPU software beating an ASIC? I guess it depends on just how abstract and adaptable the TPU IC is. Sounds like they use a custom number representation -- if they can squeeze it into less than half precision (IEEE FP16) then it would be super hard for a Phi or GPU to beat it.
But ultimately it all comes down to specific applications' bottlenecks. If they don't have to go off-card ever and their workset fits into the on-die memory, then they'd have no real advantage by using a PCIe GPU with GDDRx on it.
http://www.adapteva.com/andreas-blog/a-lean-fabless-semicond...
I'd imagine since they'd want to squeeze every performance per want out of the chip they'd want to go for smallest node possible. Virtex7 EasyPath is 16nm! It is pin to pin compatible with the FPGA version -- because they just change the mask layer and you get it in about 6 weeks. Hard to beat that.
http://www.easic.com/products/28-nm-easic-nextreme-3/
https://www.triadsemi.com/reconfigurable-full-custom-asic/
eASIC has a maskless capability where they straight-up print your silicon for prototyping/testing. Triad brought S-ASIC's to analog/mixed-signal. They're top players. eASIC's basic prototyping was $50k for 50 chips on older ones. Idk now. Triad I heard is $400k flat. Need a price quote to be sure. ;)
So, worth considering. I need to get numbers on Xilinx, though, in terms of pricing and royalties. Esp if they have something for 28nm, 45nm, or 65nm that will be significantly cheaper than other one.
Consider that they will most likely make another version, with new features in 2 years time. I'm sure the users of the chip will want that, just like any other system.
https://www.altera.com/products/general/asic.highResolutionD...
> Altera no longer offers HardCopy structured ASIC products for new design starts. Altera continues to support HardCopy for existing designs.
So what then? I certainly hope the tech sector will not just leave it at that. If you want to continue to improve performance (per-watt) there is only one way you can go then: improve the design at an ASIC level. ASIC design will probably stay relatively hard, although there will probably be some technological solutions to make it easier with time, but if fabrication stalls at a certain nm level, production costs will probably start to drop with time as well.
I've been thinking about this quite a bit recently because I hope to start my PhD in ~1 year, and I'm torn between HPC or Computer Architecture. This seems to be quite a pro for Comp. Arch ;).
I wonder if they're doing that, and to what degree.
1. It works already (IE it's already in use)
2. It works really well (or else they wouldn't be using it so broadly)
3. Considering how long this was said to be in development, it also likely means they are working on the next big improvement before these guys have even gotten the current one working.
TPUs are custom ASICs that speed up math on tensors i.e. high-dimensional matrices. Tensors feature prominently in artificial neural networks, especially the deep learning architectures. While GPUs help accelerate these operations, they are optimized first and foremost for video rendering/gaming applications -- compute-specific features are mostly tacked on. TPUs are optimized solely for doing ML-related computations.
- you can drop half the connections and it still works, in fact it works even better, during training
- you can represent the weights on as little as one bit, but still use real numbers for computing activations
- you can insert layers and extend the network
- you can "distill" a network into a smaller, almost as efficient network or an ensemble of heavy networks into a single one with higher accuracy
- you can add a fixed weights random layer and sometimes it works even better
- you can enforce sparsity of activations and then precompute a hash function to only activate those neurons that will respond to the input signal, thus making the network much faster
It seems the neural network is a malleable entity with great potential for making it faster on the algorithmic side. They got 10x speedup mainly on exploiting a few of these ideas, instead of making the hardware 10x faster. Otherwise, they wouldn't have made it the size of a HDD - because they would need much more ventilation in order to dissipate the heat. It's just a specialized hardware taking advantage of the latest algorithmic optimizations.
While reading this article, one of my first reactions was "holy shit, Google might actually build a general AI with these, and they've probably already been working on it for years".
But really, nothing about these chips is unknown or scary. They use algorithms that are carefully engineered and understood. They can be scaled up horizontally to crunch numbers, and they have a very specific purpose. They improve search results and maps.
What I'm trying to say is that general artificial intelligence is such a lofty goal, that we're going to have to understand every single piece of the puzzle before we get anywhere close. Including building custom ASICs, and writing all of the software by hand. We're not going to accidentally leave any loopholes open where AI secretly becomes conscious and decided to take over the world.
The danger of AI is more than just a random bug though. It's that an intelligent AI is inherently not good. If you give it a goal, like to make as many paperclips as possible, it will do everything in it's power to convert the world to paperclips. If you give it the goal of self preservation, it try to destroy anything that has a 0.0001% chance of hurting it, and make as many redundant copies as possible. Etc.
Very, very few goals actually result in an AI that wants to do exactly what you want it to do. And if the AI is incredibly powerful, that will be a very bad outcome for humanity.
Another way to protect against catastrophe would be to launch multiple AI agents that optimize for the goal of nurturing humanity. They can keep each other in check.
Also, humans will evolve as well. Genetics is advancing very fast. We will be able to design bigger/better brains for ourselves, perhaps also with the help of AI. Human learning could be assisted by AIs to achieve much higher levels than today.
We will also be able to link directly to computers and become part of their ecosystem, thus, creating an incentive for it to keep us around. Taking this path would enable uploading and immortality for humans as well.
My guess is that we will all become united with the AI. We already are united by the internet and we spend a lot of time querying the search engine (AI), learning its quirks and, by feedback, helping improving it. This trend will continue up to the point where humans and AI become one thing. Killing humans would be for the AI like cutting out a part of your brain. Maybe it will want a biological brain of its own and come over to the other side, of biological intelligence.
Building multiple AIs doesn't solve anything. They can just as easily cooperate to destroy humanity as to help it.
Uploading humans won't be possible until we can already simulate intelligence in computers. We can't have uploads before AI.
http://www.movidius.com/solutions/machine-vision-algorithms/...
TensorFlow on a chip....
Using it for over a year? Wow
Now these heatsinks can be deceiving for boards that are meant to be in a server rack unit with massive fans throwing a hurricane over them, but even then that is not very much power we're looking at there.
Maybe they play one move every time someone gets to go there to fix something? or could it be just a way of numbering the racks or something eccentric like that?
It would be interesting to see what the economics of this project are. I.e., what are the development costs and costs per chip. Of course it is very doubtful I will ever get to see the economics of this project, it would be interesting.
I don't think FPGAs are going to be beat out by ASICs for low volume applications anytime soon.
No wondering where you left the Torx drivers with this one.
For the curious, Optalysys has built a general purpose optics-based correlation/pattern matching machine. From some of their predecessor-company marketing material: The correlator performs pattern matching on large data sets such as high-resolution images, providing a measure of similarity and relative position between objects within the input scene. This allows large images [and general data converted to images] to be analysed far faster than electronic equivalents.
Going back to the topic of NN-based computing, I found this talk to be intriguing: https://www.youtube.com/watch?v=dkIuIIp6bl0. The main argument is that because Moore's law may no longer be in effect, it will become increasingly important to explore alternate computing solutions. (Google's TPU could be supporting evidence for this argument.) The speaker also co-authored a paper which I liked "General-Purpose Code Acceleration with Limited-Precision Analog Computation".
If not, how would one practically implement an analog computer for neural network programming (without several tables full of op-amps?)
You can implement an analog neural network yourself using a Field Programmable Analog Array. (I've never done it, but you'll see academics online writing papers about it.)
Another thing that is sort of related is Lyric Semiconductor; they built these cool application-specific probabilistic processors; they were purchased by Analog Devices a while back.
http://www.cisl.columbia.edu/grads/gcowan/vlsianalog.pdf
Brain is a bunch of components that are spread out 3D that operate like a mathematical function at slow speed. Mostly sounds analog. Results in us. So, a huge spread of analog components could get some results directly simulating something like that. Here's one of my favorites which is a wafer-scale, analog computer for neural networks.
www.kip.uni-heidelberg.de/Veroeffentlichungen/download.cgi/4713/ps/1856.pdf
The primary cost factor in anything computing is just power. You can always buy more of the things, but power is the ongoing cost and every watt in computing costs you extra in cooling power.
However, such initiatives still face the same problems as 30 years ago: custom hardware is expensive, inflexible, hard to program, and quickly becomes obsolete. It's still worth it if there's no other way to speed things up, but Moore's law is still alive and kicking, as evidenced by 15B transistor GP100 chip form Nvidia, so we can still just wait a little bit for the next gen GPUs.
Google is certainly in a good position, having developed a very popular ML framework, and having enough resources to develop good hardware (the blog post was written by Norm Jouppi - one of the best computer architects in history). It remains to be seen, however, how well these TPUs are supported in TensorFlow. What kind of models will get the advertised speed up?
Edit: to elaborate... single model training runs are possible to do quite fast now, but knowing how to tune hyper parameters remains the 'voodoo' of the field. But the best hyper params are also possible to discover through brute force: try every combination you can! Today, you can use various heuristics to improve this process, but either way, being able to train whatever X times faster just means we can search hyper parameter space that much faster. The robots are coming :)
On the energy savings and space savings front, this type of implementation coupled with the space-saving, energy-saving claims of going to unums vs. float should get it to the next order of magnitude. Come on, Google, make unums happen!
Are they saying Google Cloud customers will get access to TPUs eventually? Or that general users will see service improvements?
* Vector processing computers - not von Neumann machines [1].
* Array languages new, or like J, K, or Q in the APL family [2,3]
* The replacement of floating point units with unum processors [4]
Neural networks are inherently arrays or matrices, and would do better on a designed vector array machine, not a re-purposed GPU, or even a TPU in the article in a standard von Neumann machine. Maybe non-von Neumann architectire like the old Lisp Machines, but for arrays, not lists (and no, this is not a modern GPU. The data has to stay on the processor, not offloaded to external memory).
I started with neural networks in late 80s early 1990s, and I was mainly programming in C. matrices and FOR loops. I found J, the array language many years later, unfortunately. Businesses have been making enough money off of the advantage of the array processing language A+, then K, that the per-seat cost of KDB+/Q (database/language) is easily justifiable. Other software like RiakTS are looking to get in the game using Spark/shark and other pieces of kit, but a K4 query is 230 times faster than Spark/shark, and uses 0.2GB of memory vs. 50GB. The similar technologies just don't fit the problem space as good as a vector language. I am partial to J being a more mathematically pure array language in that it is based on arrays. K4 (soon to be K5/K6) is list-based at the lower level, and is honed for tick-data or time series data. J is a bit more general purpose or academic in my opinion.
Unums are theoretically more energy efficient and compact than floating point, and take away the error-guessing game. They are being tested with several different language implementations to validate their creator's claims, and practicality. The Mathematica notebook that John Gustafson modeled his work on is available free to download from the book publisher's site. People have already done some type of explorator investigations in Python, Julia and even J already. I believe the J one is a 4-bit implementation of enums based on unums 1.0. John Gustafson just presented unums 2.0 in February 2016.
[1] http://conceptualorigami.blogspot.co.id/2010/12/vector-proce...
[2] jsoftware.com
[3] http://kxcommunity.com/an-introduction-to-neural-networks-wi...
[4] https://www.crcpress.com/The-End-of-Error-Unum-Computing/Gus...
I agree with your overall point that we're seeing a confluence of factors. The advances in compiler technology, combined with the vectorial nature of the problems that are interesting to solve in an era of big data, mean that we can achieve a great deal of productivity by using high-level vector-capable languages.
The creator of Pandas, Wes McKinney, had a link up a few years back mentioning he was looking for people who were familiar with APL, J or K. It seems he was working on a new project/startup I think (could this have been the shuttered DataPad?). The links are dead now, but I will double check.
If the creator of Pandas is/was eyeing the older APL, and its newer brethren, I'd say it's a safe bet to keep J or K or Q on your radar because they fit. They're vector/array based; they are fast and iterative with a REPL; there is a lot of mathematical formalism in their origins and usage throughout the years, yet they are more beginner-friendly than say Haskell IMHO. I like Haskell too!
Let's see this sucker train AlexNet...
Google's always been cautious about the balance of speed and efficiency, out of concerns about programmer productivity, parallelization, and generality. See, for example, Urs's article in response to my and a few other people's crazy-academic research on using "Wimpy" nodes: http://static.googleusercontent.com/media/research.google.co... - vs http://www.cs.cmu.edu/~fawnproj/
There's a big difference between just cranking down the GHz and going for ASIC specialization. GPUs, for example, already represent a point on this spectrum -- it's true that they run at reduced GHz compared to high-end CPUs, but they're arithmetic monsters. The blog post notes, in fact, that the use of TPUs in AlphaGo let them do more searching. So why would you assume automatically that they're slow?
For if it were delivering performance on par with a $1000 Maxwell class GPU, why wouldn't you guys crow about it? That would be a really big deal wouldn't it? TitanX for 20W? That'd be awesome.
And having suffered through multiple pitches for us to buy various FPGA and boutique processors, I have yet to see someone who produced perf per watt numbers first, subsequently produce an impressive performance number. In fact, it took nearly yelling at one vendor for them to finally admit perf was abysmal.
Finally, training does not equal inference. Training requires strong scaling, but inference need only weak scale. So I suspect that Urs had to bite his tongue and buy a bunch of gpus for training networks.
Am I missing something?
That doesn't mean a TPU is faster or slower than anything in particular, it just means that quite likely that it's good for some machine learning tasks that Google cares enough about to spend the whatever dollars it cost to make the thing.
The WSJ article has a few more quotes from Norm Jouppi, btw.: http://www.wsj.com/articles/google-isnt-playing-games-with-n... (Sorry if that gets paywalled. Googling "wall street journal google tensor processing unit" got me there.)
"I'll get fired, won't have money for living and AI will take my place, but the world will be better! Yes! Progress!"
Who will benefit from this? Surely not you. Why are you so ecstatic then?
You don't need jobs as long as you have land, renewable energy sources and robots (and 3d printers). You can live in a community that is self sufficient. You will be employed by your land, as it always was up until 100 years ago. We will also have robots, maybe not the latest generation, but we don't need to go back to the 19th century agriculture.
It is you who will benefit in the end, if you can use AI to improve your life. As long as AI doesn't remain locked in the hands of one entity and we all share into the benefits, it will work out ok. In the short run we need some sort of social welfare though, and to invest in renewables and self-sufficiency technologies.
How much self-sufficient a country, city, village or small farm could be? There is a lot of potential to migrate back to small community agrarian economy with robotics and 3d printing and solar panels.
<speculation>People could trade using a different currency than that used for robotic produced goods. This currency will have to enforce differentiation of economic agents (diversity) and integration (low barriers of entry). A currency that will automatically disable the accumulation of power in a few hands and work for humans. We have to build an economy that functions more like the brain. In the brain there is no master neuron. They all share in the activity. So should be an enlightened human society.</speculation>
maybe low and middle class will have to serve people from high class who will have robots.
The jobs can can be lost - those will be lost and should be lost. Prolonging the process doesn't help anyone as well. The transition can be hard and painful, though, so speeding it up is all the better.
if jobs should be lost, I suggest you to leave you current job. and I'll want to see where you're going to get money to buy food.