Nvidia Digits DevBox
developer.nvidia.com
developer.nvidia.com
As for the business opportunity, if you are in one of these industries then one of the prime ways you differentiate from your competitors is how well your algorithm works. And how well it works depends a great deal on how much training data you are able to feed into it, which in this era of information saturation is basically a hardware limitation.
Give me a 10x more powerful machine learning system, and I'll give you a few basis points advantage over the competition, and that'll make you, not them, the dominant player.
But what I actually had in mind was QoS.
If you have an example of QoS relying on ML, I would be interested in reading more.
Beyond sensing and perception, I believe that deep learning will also be the technology that solves planning, natural language, reasoning, and creativity for machines. This is much more speculative, but you can see the beginnings of planning in the DeepMind Atari work, the beginnings of natural language processing and reasoning in machine translation and various question answering systems, and the beginnings of creativity in Deep Dream and other generative models people have done. Of course once all those pieces are solved then AI is solved. The market for a technology that solves AI is practically unlimited.
Intel is constantly searching for the next big application that will require more processing power so people need to buy faster CPUs and they can justify spending $X billion on their next fab. In recent years they've been struggling to find it. Deep learning is it. The appetite for FLOPS and memory bandwidth is unbounded for the foreseeable future. Unfortunately for Intel, CPUs are weak on both compared to GPUs. Maybe Xeon Phi can morph into a deep learning system?
What blew my mind was this lecture, given by Hinton https://drive.google.com/file/d/0B8i61jl8OE3XdHRCSkV1VFNqTWc...
The main idea is that reasoning is just a sequence of 'thought vectors' that can be encoded within a recurrent neural network.
Sounded almost outlandish to me until I watched Richard Socher's lectures and started to understand that words can be represented as vectors, and that these vectors can be then encoded into new vectors of even higher representation. 'Thought vectors' may not be so outlandish after all.
I would take this (and any such very self-assured predictions) with a grain of salt.
nVidia has already demoed hardware on that area
Formal proofs even in avionics are very limited, and planes aren't exactly dropping out of the sky. Safety critical systems need a degree of auditing and testing, to be sure, but formal proofs have never been a requirement since they aren't that practical.
https://chess.eecs.berkeley.edu/hcssas/papers/Cleaveland-pos...
> Although formal proofs of correctness are (rightly) touted as providing superior guarantees to those of testing, DO-178B makes no mention of formal methods, except in an Annex given over to techniques deemed not yet mature enough for certification purposes.
Or http://ti.arc.nasa.gov/m/pub-archive/1023h/1023%20(Denney).p...:
> In principle, formal methods offer many advantages for aerospace software development: they can help to achieve ultra-high reliability, and they can be used to provide evidence of the reliability claims which can then be subjected to external scrutiny. However, despite years of research and many advances in the underlying formalisms of specification, semantics, and logic, formal methods are not much used in practice. In our opinion this is related to three major shortcomings. First, the application of formal methods is still expensive because they are labor- and knowledge-intensive. Second, they are difficult to scale up to complex systems because they are based on deep mathematical insights about the behavior of the systems (i.e., they rely on the “heroic proof”). Third, the proofs can be difficult to interpret, and typically stand in isolation from the original code.
The situation will just get worse as we increasingly move away from hand written logic to machine learned logic. Now, there is probably a place for formal methods in this newer order, but it will take awhile for researchers to find it, and it is definitely not going to be at the level of verifying hand-encoded logic that doesn't exist.
1. Replacing testing with formal verification at Intel (2009)
http://link.springer.com/chapter/10.1007/978-3-642-02658-4_3...
2. Fifteen years of formal property verification at Intel (2008)
http://link.springer.com/chapter/10.1007/978-3-540-69850-0_8
3. seL4: formal verification of an OS kernel
http://dl.acm.org/citation.cfm?id=1629596
There is such a big need that DARPA has even come up with a contest.
http://www.darpa.mil/program/crowd-sourced-formal-verificati...
I could go on, but I would rather not spend my time combating FUD and the general incivility here.
I never said that. Where did you get that?
> Safety critical systems need a degree of auditing and testing, to be sure, but formal proofs have never been a requirement since they aren't that practical. (emphasis mine)
I was challenging this idea.
Are systems for self-driving cars the only safety critical systems? What about Intel manufactured chips used in those same cars and flights and a gazillion other devices?
You have an agenda to push and got called out on facts :)
Have a good day. (Someone just went around down voting my other posts. Not you)
Btw, I use both deep learning and knowledge based systems in my work (but not formal verification) but don't like advertising that goes in AI by academics from all camps(don't want another AI winter).
Now good luck verifying a machine learning system beyond the basics.
Then, if outliers/problems are encountered in the real world, rather than training the end user's neural network alone, the outliers are fed back to the centralised system, and later a new snapshot might be sent out.
After all, people want their self-driving cars and voice recognition phones to work out of the box. And if there's a misclassification you correct on your phone, you want it to propagate to your tablet, PC and smart fridge :)
If things are done that way, I would have thought the market would end up with only a handful of customers - there won't be a need for a deep learning chip in every phone. There might be a world market for maybe five deep learning systems :)
Deep learning--and machine learning more generally, because that's what we're really talking about here--is going to go through the same process. NVIDIA is positioning themselves to be the vertically integrated supplier in this market. The really interesting part is that no one is sure yet what kinds of applications these tools will open up -- it is likely to be an enabling platform/toolset in much the same way relational databases, app stores, or IaaSs were.
AMD and Intel are woefully behind. Intel because it's not yet a big enough market ($137B market cap compared to NVIDIA's $12B, Intel needs an opportunity to be literally 10x as interesting to be interested). Why AMD is not all over this I'm not sure... If it's interesting for NVIDIA, it should be very interesting to a $1.6B operation like AMD.
For Nvidia's market, imo we have to be careful to separate out which kind of ML specifically. Certain kinds of training of deep networks get large speedups from GPUs, which isn't true of all ML techniques. So if Nvidia wants to get a big win in this space by being a vertically integrated tool supplier leveraging GPU hardware, the strategy is tied to a bet on a specific subset of ML techniques, those where the ratio of GPU:CPU performance is especially large.
NVidia have been very good at encouraging the use of their proprietary CUDA. They also have a slight FLOPS advantage I'm told (where AMD have an integer advantage)
AMD are barely treading water in the markets they already compete in. Both Intel (CPU) and NVIDIA (GPU) have stolen back enormous mind and market share [0]. I doubt they could afford the outlay of capital and research time to build something competitive. Maybe if they were in better shape financially a braver board/CEO could pivot them away from the CPU/GPU markets they've lost to something different. Their current road-map takes AMD through to 2016, after that who knows.
[0] NVIDIA's 9XX series GPUs have been incredibly aggressively priced and marketed. AMD effectively ceded the high end consumer (gaming etc) market since September last year. Gaming parts are big PR for both companies, so it's quite a bad sign IMO.
Yes, wouldn't it be great to let the market leaders in these sectors have a total uncontested monopoly...
He's referencing this:
https://en.wikipedia.org/wiki/Thomas_J._Watson#Famous_misquo...
They're barely staying alive, and have been axing branches left and right to "focus". Like selling off their smartphone graphics branch a year after the iPhone was announced… it's now rather successful under Qualcomm (Adreno is an anagram of Radeon).
Our biggest plus factors compared to a GPU is that we are a fully standalone/independent chip that does not need to have a CPU with your main system memory attached to it. Large machine learning data sets are getting into terabytes in size, and the biggest bottleneck with GPUs is the PCIe link limiting them to 16GB/s and adds a whole lot of additional latency. In our case, we have direct connection to DRAM (We've been looking at DDR4 and HMC). In addition, we have designed the architecture to allow massive scalability with up to 384 GB/s of aggregate chip-to-chip bandwidth... NVIDIA's NVLink is aiming for 80GB/s in the 2018/2019 timeframe and will still need a connected CPU to issue jobs.
EDIT: I should also mention that our chip is fully general purpose, but we perform really well when it comes to dense matrix math (Most deep learning), with a 10 to 15x efficiency advantage over GPU. Our real killer app is FFTs, which GPUs do abysmally on, and our current benchmarks are showing a 25x efficiency advantage over the best DSPs and FPGAs built for large constellation FFTs.
A truly specialized deep learning chip probably wouldn't be useful for much else, but it would be a monster at deep learning. And the thing about deep learning is it scales really well. If you have a 10x faster machine you're almost certain to set world records on any machine learning benchmark you try.
We still have the option of including a 16 bit (half precision float) packed SIMD mode into our FPUs, which would add a bit of complexity (bringing our efficiency numbers down a bit for the double precision float, which we like to talk about as it is over 10x better than anything out there), but if there is enough customer interest we may decide to include it.
The other nice thing is that it is a superset of IEEE float, and has a "IEEE mode" where you can convert to IEEE float, which is also jokingly called the "guess" function.
As for support, that is the plan. Right now our customers have been most interested in high end signal processing, so we have been taking time to port FFTW, but supporting Theano/Torch/Caffe, etc is a relatively straightforward process.
https://en.wikipedia.org/wiki/John_Gustafson_(scientist)#Unu...
Feel free to correct/improve it anyone.
Do you have any benchmarks on *gemm or FFTs?
We will be releasing benchmark data when we have first silicon back next year. As of right now, our numbers are based on FPGA prototype implementations.
It's a very long road to a shipping chip and a lot can happen between now and then. And afterwards, it's an even longer road to the equivalent of nvcc, nvvp, cuda-gdb, cuBLAS, cuFFT, cub, cuRand, and cuDNN with solid linux/windows/mac support. And it's all free. NVIDIA achieves this through its nearly bottomless pockets from the gaming/industrial complex that has allowed them to weather many near-death experiences.
What's your burn rate BTW?
In the meantime, a $1000 GTX TitanX (which you can easily buy off Amazon) delivers 27 GFLOPS/W of FP32. So much for "a 10 to 25x increase in energy efficiency for the same performance level compared to existing GPU and CPU systems." At least get your facts straight. And while we're talking facts, the real challenge here is that they deliver 6.7 GFLOPS/$. That seems like a tough squeeze for a startup to beat. Again, what's your burn rate? For that's how it's ended so far for the other contenders. Why are you different?
I have no problem believing NVIDIA will one day be disrupted, but nothing I read on your web site made me feel in any way that it's your architecture. Suggestion: lose the router and replace it with something simpler after studying NVIDIA's memory controller and hierarchy. They considered a lot of crazy things too and there are many reasons why GPUs ended up the way they are with a very clean programming abstraction that automagically subsumes SIMD, multithreading, and multicore.
Burn rate is very low (relatively speaking)... Our full size team, which we are building up (We're hiring!), is only 7 people. Our current runway is around 18 to 20 months, and we have some unannounced (and not included in that runway) funding coming along.
Semiconductor economics are also in our favor, with our ~100mm^2 chip being a lot cheaper (per unit, and in terms of GFLOPs/$) than NVIDIA's ~650mm^2 (and up) chips.
As for "getting my facts straight" I was using actual Titan X numbers I have seen for real applications (e.g. nbody simulation)... it only gets around 4000GFLOPs single precision compared to their advertised 7000. I was being generous to NVIDIA saying 20 GFLOPs/watt at best when they are currently at ~16 (before you include the CPU). If you know about NVIDIA, you should know they have a long history of bullshitting numbers.
As for our design, there is a lot more interesting stuff that we haven't disclosed for obvious reasons, but I do stand by all of our numbers. When it comes to our network on chip, the main reason we are keeping it this way is because it is a general purpose design... of course, we would be doing things very differently if we wanted to have something application specific for machine learning.
With Nvidia throwing all their weight behind acceleration of deep learning applications, they're advancing on multiple fronts: Tegra X1 claims 1 Tflops @ 10W (half precision). That's what they're shipping today, what about 2016?
Aside from rapidly improving GPUs, you now have rapidly improving FPGAs. For example, Altera has put thousands of hard multipliers on their latest Stratix 10 chip. They claim 10 Tflops SP, and 80 Gflops/W: https://www.altera.com/products/fpga/stratix-series/stratix-... Microsoft is already using them to power Bing search (apparently they use NN based algorithms for that).
What makes you think your chip - if it's actually out in 2016, which can easily slip to 2017 - can compete with Pascal chips from Nvidia, or the next gen FPGAs from Altera/Intel or Xilinx?
Who is your market exactly?
When it comes to something like the X1 you have to remember NVIDIA loves exaggerating their benchmark results. In reality, it gets around 80% of its theoretical peak, but if you apply it to double precision, it only gets around 40 GFLOPs on Linpack (out of a theoretical 64 DP GFLOPs). Even if you take NVIDIA at face value and say they get full theoretical peak (64 GFLOPs double precision) at 10 watts, that only gets you 6.4 GFLOPs/Watt. One thing that has been completely disregarded in this and my previous posts has been that the GFLOPs numbers we have been saying have been for Linpack and matrix-matrix (Level 3 BLAS) workloads, which are a very small number of real world applications... I would say on average Level 3 BLAS benchmarks get around 90% of theoretical peak on GPU systems, but as soon as you get into other application spaces (Level 1 and Level 2 BLAS, or anything dealing with a lot of memory movement) is where GPUs really start to fall, and only get ~10-20% of theoretical peak. Our architecture is built to actually be able to reach theoretical peak in a "perfect" (which in our case, our hand written FFT kernel), in all 3 BLAS levels. In reality, I would expect us to hit at least 85-90% theoretical peak in all of those floating point domains.
As for Altera and FPGAs... they are a royal PITA to program for, and will never be all that efficient. Altera's 10 TFLOP/s number is ONLY capable of single precision float, and is based on adding up the theoretical capabilities of all the DSP slices... in reality, you would never be able to hit that with the memory limitations going to all the DSP slices. In small print, Altera even admits their 10 TFLOPs number is BS, as they list the highest FP32 number as 9.2 TFLOPs. Again, that theoretical 80 TFLOPs/Watt (which I would be surprised if it hits 50 in the real world) is only for single precision, in which we are aiming for 128. All the while we are a hell of a lot easier to program for.
As for short term competition, we know for a fact that Intel and NVIDIA will not be hitting our efficiency levels in this decade, and in addition we will have a cost advantage. Compared to FPGAs, we have the ability to actually port existing code over without huge performance sacrifices (C to gates or Altera's OpenCL work is abysmal to performance) or change your development team/learn how to write RTL (You can write code for our chip in any language that LLVM has a frontend for).
As for market, our initial target market we are actively working on is large constellation FFT type workloads... think LTE-advanced and "5G" basestation processing as that's where we have our best numbers (25x efficiency over the best DSPs in that space). Beyond that we are looking at the larger "HPC" category, and in the 5+ years out, I hope to be able to expand to more general purpose markets.
However, if you do decide to go after deep learning (seeing as it is a much faster growing, and potentially much bigger market), I have a few questions for you:
1. Will I be able to take my highly optimized Torch/Theano/custom CUDA code, and run it on your chip with minimal modifications? Especially taking into consideration that even some latest CUDA code is not compatible with older GPU architectures? 2. How much will your devbox cost, compared to the Nvidia devbox? 3. Will I get much better performance (16 bit) as a result?
Keep in mind that I'm talking about 2016/2017 version of Nvidia devbox (Pascal should have at least 30% better performance compared to current Maxwell cards, and probably more than that if they manage to move to 14nm process).
Regarding the "magic" compiler, have you been watching the development of Mill CPU? Designed by the guys with strong DSP and compiler design background. They also put a lot of emphasis on the compiler. Is that project dying? After two years of hype, it seems like it never got out of the "simulated on FPGA" phase... What can you learn from them?
https://en.wikipedia.org/wiki/Systolic_array
There are certainly applications for this sort of processor (embarrassingly parallel batches of small independent units of work), but I'd be highly skeptical that this guy has anything close to a "magical compiler(tm)" given his inaccurate understanding and significant underestimation of the competition. That's dangerously close to Intel's absurd "recompile and run" nonsense for Xeon Phi (It's anything but that)...
Nonsense. I regularly get ~5.5 TFLOPS out of them running cuBLAS SGEMMs in neural networks. You can get that down to ~1.4 if you do your best to choose stupid small values for m, n, and k, but that's a relatively minor bug and I think it's fixed in cuDNN's latest kernels.
But even if it isn't, if you're willing to download Scott Grey's maxas: https://github.com/NervanaSystems/maxas, you can hit 6.4 TFLOPS with his hand-coded SGEMM. Similarly, one can do the same for convolutional layers with Andrew Lavin's maxDNN: https://github.com/eBay/maxDNN
Your hubris is amusing, but I reiterate that until you have a shipping chip, all you have is a powerpoint processor. I'm sure you disagree. Good luck with that. But do come back when you have numbers (real numbers on real DNNs like AlexNet, VGG, or Googlenet as opposed to synthetic fantasy networks that fit your architecture well). See Nervana for a company doing this well so far.
"If you know about NVIDIA, you should know they have a long history of bullshitting numbers."
I know them very well and, well, pot.kettle.bs... IMO they're going to stay the leader in parallel computing technology right until Intel stops sniffing its own tailpipe and/or AMD hires a better driver team (they have promising HW). And both of them need to study NVIDIA's engagement with the academic community and one-up it rather than deny its efficacy.
But unfortunately, those are exactly the sort of pitches I see: unrealistic or overly simplified demos that tell me jack about the real world utility of a new architecture.
Switching to a new platform is a Mt. Doom of technical debt no matter how fantastic it is. You need to make climbing Mt. Doom at least sound like a good idea. That said, I've been on your side of the fence several times in the past. Hear me now, believe me later I guess.
You can defer and say you are targeting telecom and that's great. But you just picked a fight with NVIDIA here over deep learning and you're 7 people with <2 years of funding left in the bank. Put up or shut up.
I think we should give him some respect. He's developed a novel processor, and started a funded company before he turned 18.
Even if he fails to sell this particular design, we need more people like him.
Have we not tired of all the broken promises of Kickstarter and IndieGogo yet?
Programmed via OCL? Don't I need some kind of OS/host CPU functionality to load the inputs into DRAM and retrieve them?
And while I think one can build a deep learning ASIC, a 10x better ASIC seems like a tough bet to me. I mean on the surface it sounds good, but the devil is in the details here and by the time you've built something flexible enough to run every reasonable variant of a neural network both forwards and backwards, you start making the sort of decisions that make your processor look more and more like a GPU and that magical perf delta drops. Also if you start today, you need to target tomorrow's GPUs, not the $1000 consumer model you can buy on Amazon now.
That said, I'm looking forward to Altera's new hardcoded floating point-enhanced FPGAs. Too bad I have no idea how much they cost.
I wonder why the popular deep learning frameworks are using mainly CUDA instead of OpenCL. Is because of better Linux GPU drivers? Wondering why AMD isn't jumping on deep learning
The Altera FP capable FPGA's sound real interesting too. 10 TFLOPS, OpenCL support?
http://www.slideshare.net/embeddedvision/a04-altera-singh
Looks like they're about to be bought by Intel? http://www.electronicsweekly.com/news/business/altera-import...
Does this mean FPGA co-processors in the future from Intel?
But its performance is less than half that of a GTX 980 running CUDA. Still, AMD is silly not to try and improve on this IMO.
Probably in the same ballpark as current high end Altera chips: over $30k. If you need a thousand of them for your datacenter, think about what kind of ASIC you could design for $30 millions. Or think about how many Titan X cards you can buy for $30k :)
whoa - are you telling me that the nVidia drivers on Linux are so stable that they are building a commercial deep learning system on top of that. Is this the same thing as normal graphics drivers ?
nVidia uses CUDA or OpenGL - so its not quite the question of proprietary API.
At this point, I'm not worried about "framerate on my linux box isnt as good as windows".. its more "it works...".
You've been sold the idea that a "powerful GPU" needs to suck a lot of power all the time.
There is no real reason a "powerful GPU" shouldn't be able to scale its power usage way down when doing something simple like browsing the web. The only reason NVidia weren't able to do the "low power" thing on these systems is they weren't able to be the ones putting their GPUs on the same die as the CPU like Intel (& AMD) were. But of course they still wanted part of the action, so people ended up being sold this massive engineering bodge and told it's a good thing.
and keep in mind, it is a binary blob. proprietary to the core. in fact, CUDA the protocol itself is kinda proprietary.
if you base your solution on OpenCL, you can use all the vendors. nvidia, amd, imb, intel, altera, etc...
but NVIDIA spends billions on marketing to convince you that only CUDA matters. the same way sony spent millions to convince you that only laser disc^H^H^H^H^H^H mini disc ^H^H^H^H^H memory stick ^H^H^H^H^H blue ray, matter.
[0] https://developer.nvidia.com/deep-learning-courses
[1] http://on-demand.gputechconf.com/gtc/2015/webinar/deep-learn...
$15K is probably OK(ish) if you figure in your own time for the DIY build...probably a few days. Plus you get some vendor support, warranty on the whole package, certified working stack, future test-bed for CUDA updgrades (will work first), etc as you say.
In wild agreement. Save maybe 30% doing a custom build...so they aren't adding a huge mark-up, as they would for a gaming machine. Apparently...someone at Nvidia is looking a bit further into the future than just the short-term revenue.
Makes sense to me, since you want the best airflow possible getting to the cards in a multi-GPU setup, and unlike in conventional cases, the A540 doesn't have a drive cage between the front fans and the video cards.
Yes. You're a drywall contractor you go to Craigslist and start humping some jobs. Make some bucks and buy you a nice F150.
Conversely, there aren't a lot of "need deep learning in my dentistry" ads on CL -- and not as easy or clearcut path from point A to point B.
So the investment makes sense, if the right conditions are true. To the average developer, those conditions look very murky and uncertain to gauge. So while it's obviously some sweet tech, but it takes a lot more than tech.
Which BTW, the drywall impresario who's buying a new truck for his business isn't finding jobs on Craigslist. He's got business relationships with serious people who call him when a bid needs bidding and drywall needs hanging. It's deal flow just as it is for the sort of person who needs a 4 GPU box and CUDA code for their business. It's only a lot of money if it sits idle.
If there isn't a business case, there isn't a business case and buying it is an inefficient allocation of resources.
Of course this isn't for the average developer. Which is why they are marketing it to researchers and maybe AI startups.