Quantifying the performance of the TPU, our first machine learning chip
cloudplatform.googleblog.com
cloudplatform.googleblog.com
Makes me wonder if there are more recent generations that target training.
> if the TPU were revised to have the same memory system as the K80 GPU, it would be about 30X - 50X faster than the GPU and CPU.
Is it "hard" to interface with GDDR5/HBM? Layout challenges? Or do they need the capacity more than the speed? Why wouldn't they have used faster memory than DDR3?
a) they did not want to license a more expensive, faster design
b) while it would be faster, it would decrease efficiency to a point that did not meet their goals (for data centers, efficiency > absolute performance, within reasonable boundaries)
c) like b) just with cost of memory
d) GDDR5 and DDR3/4 have different design trade-offs. The former is optimized for sequential bandwidth (and low capacity; GDDR always was a point-to-point memory bus just to achieve the clock speeds), while DDR3/4 takes random read/write workloads into account (eh... to the amount possible with DRAM...)
--
HBM requires a silicon interposer, which is basically like another complete chip (just without the FEOL parts, "just" metallization), that has to be significantly larger than the size of all chips combined. So unless you really need that performance or have a volume product it's unlikely to be a good deal.
HBM: The interposer means you get a complex, and risky manufacturing process.
No idea on GDDR5 vs DDR3, maybe they didn't like the latency of the first.
It seems to me that Google engineers could use Tesla's or other high end GPU's for training and development, but then deploy those models on hardware optimized for forward passes...
This first generation of TPUs targeted inference (the use of an already trained model, as opposed to the training phase of a model, which has somewhat different characteristics)
This Nvidia article treats them differently, too: https://blogs.nvidia.com/blog/2016/08/22/difference-deep-lea...
But the definition of "statistical inference" on Wikipedia says "Statistical inference is the process of deducing properties of an underlying distribution by analysis of data" which seems exactly like training.
But you do find constructions using it as a synonym for training too, as in phrases like "model inference" (which means inferring models from data, aka model induction or training). Inferring things from other things is a pretty general concept, so it can be slippery without context...
Training involves finding derivatives and requires higher precision calculations which is why they don't train using TPUs.
Would love some tech details, but it seems that the paper wont be published until 5pm today
Full custom chips have the highest NRE costs and longest development cycles. Standard cell ASICs are less energy efficient, slower, and provide generally less logic density, but are cheaper and quicker to develop.
For other parts it's even less clear: Flash or hard drive controllers are pretty normal micros with some dedicated hardware bolted on - clearly ASIC, but most of it was not designed for the "AS" part, and you could just ignore the flash and SATA interfaces and use it as a regular micro.
So the distinction, if any, can't be about volume (since a lot of them are large volume parts), nor about functionality, but some fuzzy distinction by narrowness of intended use of the part (- but then again, GPUs).
And how does it apply to other domains of chips? Is a TDA7000 an ASIC?
:)
I believe that the term ASIC itself is quite overloaded. From what I've seen, many people use the term ASIC to refer to any IC that is not reconfigurable (i.e., FPGA). Going by that definition, an ASIC is any circuit that is custom designed and fabbed on a wafer. Naturally, this would include a CPU, a GPU, and whatever else you can think of.
The way I like to think of it is that an ASIC is a circuit designed to perform a specific task as efficiently as possible. Note that I used to word circuit; in other words, an ASIC could be part of some larger design.
Some examples off the top of my head:
- Digital camera CMOS sensor
- Video decryption chip (e.g., in a cable box)
- Active noise cancellation chip (if custom and not a DSP)
- Full-custom TPM
- Full-custom RSA-2048 engine
- High-performance Ethernet switch controller
- CPU cache controller (ASIC that is part of a CPU)
> I believe that the term ASIC itself is quite overloaded.
I think we're arguing the same thing, just from different angles :)