Computing 10,000x more efficiently (2010) [pdf]
gwern.net
gwern.net
http://devblogs.nvidia.com/parallelforall/bidmach-machine-le...
Also on the horizon there is 3d chip manufacturing technology(3d-monolithic) ,with extremely large bandwidth between the two different layers of the chip,possibly being gpu + dram.
The energy cost of transferring a single data word to a distance of 5mm on-chip is higher than the cost of a single FLOP (20 pico-Joules/bit). 5mm =~ the distance to L2 cache or another CPU core. The cost of transferring data off-chip (3D chip and/or RAM) is orders-of-magnitude higher, see graph.
[0] http://iwcse.phys.ntu.edu.tw/plenary/HorstSimon_IWCSE2013.pd...
What you run in to are problems feeding the beast data fast enough.
But don't believe me, don't believe a word I type. Go prove the above wrong by building a killer AlexNet implementation out of this chip. See slide 9 of this presentation for why I don't think you can do this (too much data transfer):
http://www.slideshare.net/embeddedvision/tradeoffs-in-implem...
If you can reduce the size of the computation units you can have one (or several) for each layer, making it possible to hardwire the transfer between layers
Yes, I'm still waiting for memristors.
A friend of mine that worked with supercomputers back in the 90s used to say that a supercomputer was a device that converted a computational problem to a communication problem. Today this is true for mainstream computing as well. There are still problems that can take advantage of the massive amount of computing available (Some mentioned Convolutional Neural Networks), but many tasks that used to be CPU bound is now effectively IO bound.
We're seeing more like 10-100x improvements in energy efficiency and performance, not 10000x, unless the comparison point is a full blown CPU/GPU.
https://www.schneier.com/blog/archives/2015/07/friday_squid_...
Accidentally running into another group using analog selectively for acceleration is pretty neat. The coprocessor was a believable improvement showing analog power. What's your thoughts on the computing with free space and no transistors stuff? Do those other links come off as bogus to a pro or plausible enough to encourage local college students to try something with it? I think there's vast untapped potential in shifting certain functions back to analog and improving the integration of the two. Maybe in general-purpose, too. Almost certainly in INFOSEC w/ analog supporting obfuscation and tamper-detection.
What we're working on is an accelerator for the convolutional neural networks that are winning competitions like ILSVRC. Even that by itself is insufficient for a business case, though. You also have to have end application in mind too, and that end application better be power intensive or performance constrained enough that software cannot accomplish what you need it to do. Because, if software is good enough, then why take a risk on a fancy new hardware component?
"Because, if software is good enough, then why take a risk on a fancy new hardware component?"
Good point. Something that's done in many. Gotta have a clear benefit esp in price/performance/energy. This market has almost as many shut-downs as start-ups.
However, as we discover more algorithms for general intelligence, we will reach a point where the model can learn on its own - just like a human baby does. That will be the point where we will need size, speed, and power efficiency, rather than flexibility. That will be a good moment to offer a hardware solution, and that's when an analog chip will suddenly become more attractive than a digital one.
They are not able learn on chip - that is a non-starter and not particularly useful anyway. Customers dont want self-driving cars that need to learn how to drive, they want self-driving cars that already know how to drive.
That's a good point to not overlook. Plus, you mentioning this just gave me an idea for a Triad Semiconductor-style, via/metal-programmable, CNN chip tied to a specific FPGA architecture for easy prototyping and conversion. Could be some promise in there. Brain hasn't gotten further than that sentence so don't ask for details haha.
Not clear to me how you will do analog and reprogrammable at the same time unless your reprogramming is building things around the analog components that still perform pretty much the same function(s). I could see the weights, connections, location on chip, etc being configured while connected to analog, signal processing blocks scattered throughout chip kind of like FPGA's do with MAC's. My guess as a non-HW guy with a little research into these things. Am I anywhere close?
Trying to see if there's a consensus emerging in how people accelerate w/ mixed-signal chips. Might help academics figure out better place to start on next project.
Can you back up your claims with actual performance numbers? I looked at your website, and I don't see any products - do they exist? What is the flops/W for you best CNN implementation? How many ImageNet images can it process per second? What is the accuracy (assuming you can only do 8 bit precision)?
Also, how much does your chip cost?
I can say that the cost depends on what you want to do - systems can range from less than 1mm^2 to the entire reticle depending how much performance you want.
I'm sure you're aware that since Mead's retina chip there have been dozens of attempts to build NN chips, both analog and digital, and very few of them got further than the simulation stage (ETANN or ANNA chips come to mind), and no one managed to produce a commercially successful product.
Nvidia Tegra X1 claims to have 1Tops @10W for 16 bit precision, and the cost is probably under $100. They can probably double that performance if they drop precision to 8 bit. That's what they ship today, and next year they will release the Pascal version, which will undoubtedly will be bigger, faster, and more efficient. What makes you sure you can compete with them?
An ASIC is always going to be at least 10x better than a CPU/GPU for performing the same algorithm. The question isn't whether or not an ASIC can beat NVIDIA, the question is whether the target market is large enough to support an ASIC company.
At Isocline we assume that this market IS big enough to support an ASIC. Our competition is not NVIDIA, it's the future all-digital ASIC company that can do the same thing, but without all of the whiz-bang technology. If we have to, we could probably fall-back to be that all-digital company, but I'd prefer to maintain our technology advantage.
NVIDIA's advantage is flexibility, there's always going to be a lot of demand for that.
In theory, this has always been the case. Yet every single neural net ASIC built in the last 25 years has failed in the marketplace, for the same reason - "silicon steamroller". Invariably, when the ASIC was ready to ship, which was almost always much later than was hoped for, general purpose chips have caught up in performance.
I'm not attacking your startup in particular. I'm just pointing out the history behind the field of specialized neural hardware.
p.s. Your competition is Nvidia (or Intel, or Xilinx, etc), because they are well known, big players, who produce reliable products, with huge development infrastructure and expertise. Nvidia specifically has been focusing on deep learning applications, they are already targeting computer vision for cars with their mobile GPUs. If I'm Ford or Toyota, who would I consider for partnership when I need chips potentially making life or death decisions on the road? If your technology really works (big "if", because you haven't built anything yet), then your best hope is one of those big players acquires you.
History is something we need to contend with, not just for neural networks, but also for analog computing which has a similarly troubled past.
For NN history, there has not actually been a market for NN accelerators until recently. You can see this because:
1. No NN algorithm was worth accelerating until AlexNet came along in 2012
2. What commercial products even use NN now? Currently it is mostly just voice recognition which is processed server-side.
Right now we are not attempting to go after any markets that a GPU would be sufficient for the reasons you mention; we're sticking to products that can only work with our technology. By the time we went after an overlapping market our credibility would be established and that wouldn't be an issue.
I'm curious, have you considered using analog weights (e.g. floating gate transistors, or DRAM capacitors)? This could reduce multiplication from 32 transistors to just one!
1. Learning does not have to happen inside a car. But it does have to happen somewhere, and that's where the learning in hardware will be much more efficient/faster than learning on a GPU.
2. There are scenarios where local learning would be necessary/preferable to remote learning (e.g. one shot learning or continuous online learning).
What abut large scale neural networks - they aren't mentioned in your site . No plans for that ?
Part of the reason I argue for this is that there are tons of sensors which are 0.1% sensors and if you can offer the rest of the computational pipeline at 0.1% then (so long as your errors don't accumulate)_you don't lose any accuracy processing your information this way.
It also seems like this would be pretty great for graphics cards, no? I mean it'd take a lot of work to make OpenGL run on it, but once you did you could have either very inexpensive cards, very powerful cards, or both.
I think that may end up being the hardest part of all this.
https://www.linkedin.com/pub/joseph-bates/2/853/aa3
And his company's patents,
http://patents.justia.com/assignee/singular-computing-llc
Ah, I get the name now. "The Singularity Is Near" came out in 2005 https://en.wikipedia.org/wiki/The_Singularity_Is_Near