Japan to Unveil Pascal GPU-Based AI Supercomputer
nextplatform.com
nextplatform.com
If anyone wants to learn more about AI chips, I'd be happy to answer questions.
As for our training chips, without divulging too much on a public forum (I'd be happy to talk more in private or over email), I think we provide the right level of abstraction and precision which would allow a researcher to one click port a tensorflow Model (we plan to support a few others like Caffe, torch, mxnet, etc out of the gate as well) to our chips.
That 3600 W for the Model S sounds high if I understand it's just for the brainbox; that alone would consume all your 85 kWh in less than 24h.
And as always, Nvidia is twiddling the numbers, they exclude the cost of memory access when calculating 1 top/w, whereas we include it.
Dealing with a few hundred watts more isn't a deal breaker.
HVAC in automotive and industrial machines is a solved problem. I don't know why you would think otherwise.
1000 watt isn't really that much heat to deal with.
I'm aware (excruciatingly aware...) of that. Our numbers are ~10X better despite a process disadvantage. And as always, Nvidia is twiddling the numbers, they exclude the cost of memory access when calculating 1 top/w, whereas we include it.
If that estimate is on the money or high, I'd totally pay <= $20/mo for my car to drive itself. If it's low, I probably wouldn't pay $100+/mo for it.
[0] https://forums.tesla.com/forum/forums/miles-kwh
[1] X mph / (Y mi / kWh) = X/Y kW
[2] http://money.cnn.com/2011/05/05/news/economy/gas_prices_inco...
1KW is a small fraction of a typical car's heat budget.
... I'd buy it!
Hilly neighborhoods might be something else entirely.
How-so? That's 1.34 horsepower. Even assuming catastrophic losses in converting fuel to electrical power, say 90%, that's only 13.4 horsepower to generate 1000W.
AMD has announced a major AI initiative, MIopen.
My point is that all the big chip players are entering this battle. Going to be tough for any small player to make a dent.
MIopen however is basically just a rebranded Fiji/Polaris/Vega
While there may be a market for your chips, I'm curious why you think the K computer is in that market? National supercomputers, like Japan’s RIKEN "K" supercomputer, are used for many different applications (for example, physics and engineering simulations) - not just "AI." The multi-purpose use of such machines is what justifies their multi-billion dollar budgets in the first place. I can't imagine a government spending billions of dollars on a machine that only has one function (e.g. neural net training).
The history of HPC hardware is littered with special-purpose HPC microarchitectures that were eventually abandoned in favor of general-purpose processors. The one lasting exception to this has been GPUs, which have proven to be a boon to HPC applications and sparked the Deep Learning renaissance in machine learning. The difference with GPUs is that they were not strictly aimed at HPC applications. Obviously, they are used for graphics rendering in gaming, professional graphics and CAD. There are hundreds of millions of GPUs deployed for gaming and other graphics applications. The application of GPUs to HPC came later, and the specific application to deep neural networks came later still. GPUs are successful because they are a form of commodity hardware and have a wide range of applications. In a sense, hard-core gamers have become the R&D funding source for state-of-the-art HPC processors. This healthy and diversified ecosystem is what allows for the long-term sustainability of the microarchitecture.
You can always build a more efficient machine by specializing it to a narrow application. In the extreme case, you can just build a custom ASIC that has some fixed function. That would be the ultimate solution in efficiency, but things become less sustainable when you need to continuously compete with alternative solutions - the cost of competing in this space is astronomical, and there needs to be sustainable source of funding for that activity. This is why the HPC industry is completely dominated by Intel/AMD/NVIDIA processors, instead of custom ASICs that (for example) could perform some fixed matrix operations.
Having said that, there is a vague opportunity on the horizon if and when Moore's Law scaling completely fizzles out. Conceivably, after processor node scaling completely ends and the established microarchitectures have been completely optimized to death, the industry will reach a state where competition on performance/efficiency stalls because there is no major next-gen CPU or GPU because nothing more can be done to improve the product while maintaining its general-purpose applicability. At that stage, a significant opportunity could open up for special-purpose processors and it could be sustainable since the field would be far less competitive.
For the record, many of the improvements we were able to squeeze out for deep learning has also lead us to create two new designs, one for a GPU and one for a CPU. Although we're focusing on the deep learning processor for now, the ultimate goal is to develop all three and put them on a SoC. This is too ambitious in our current stage, and so we're focusing on the deep learning processor.
Phones are pretty homogeneous right now with only a couple real competitive hardware solutions implemented, but if Windows 10 proper can really make the transition to phones then we may see an explosion of increased processing demands of those devices over the next 7 years.
Best of luck!
Our primary problem is that GPUs tend to be the epitome of a bad initial market for startups to pursue, since they require large (very large) volumes and very large fixed costs. If you happen to have a better business model to get into the GPU market, drop me a line :)
We'd love to collaborate!
You can contact me at my personal email at sixsamuraisoldier[at]gmail[dot]com
There seems to be a veritable flood of various kinds of deep learning chip papers and startups at the moment. Seems the common theme is some kind of processor-in-memory (PIM) style architecture in order to reduce the energy cost of memory access, coupled with massive amounts of low-precision arithmetic units, since neural nets seem to be able to get by without the traditional FP64.
I am completely wrong in the above assessment? And what are you doing differently than all the other players? In any case, good luck!
Does your product allow to run generic frameworks like Theano or TensorFlow?
By the way, I'd love to read more hardware-startup-related news on HN.
Is a supercomputer not a cluster of multiple machines?
How far behind is AMD with GPGPU, or AI. It seems OpenCL is a dead end. CUDA won. AMD announced some CUDA to (x) code conversion which never really caught on.