Graphics Chips Help Process Big Data Sets in Milliseconds
technologyreview.com
technologyreview.com
And now a MicroZed costs $199 and Digilent is releasing the Zybo at $149! Technology sea change, anyone? It's going to make coding REAL interesting again...
That's why I wanted to start this conversation, I think it's germane to GPU efforts and people may not be aware. I would be more excited about a project like this spreading FPGA tech, and more effort being spent there. My company is doing it and we want a bigger open-source community with us. (We're not in a position to release code yet, but we will.)
Effort is energy. I'm hoping mentioning these technologies will give people pursuing projects like this some options on where to spend their efforts. There's a finite pool of people who are working in this space and I hope it grows.
Our approach for scaling is a monster FPGA board we're developing. It has a Zynq on it. It will be extremely inexpensive in terms of value next to, say, a highly-capable production server.
What kind of projects would you be working on if you bought a Zynq board?
For deep learning, high performance in a single machine is important, as it's very difficult to distribute learning across multiple machines in an efficient way. A more powerful machine can train a bigger network in the same amount of time, and a bigger network can learn more complex things, with no known upper bound.
That said, working with the GPU requires a very specialized set of high level constructions as well; I imagine that most workloads could have significant parts of their computational path abstracted out to pretty common easy to port fpga processes.
If you ever launch a beta...let me know!
I think it's not as much a matter of making an FPGA that would be easier to program as it is a matter of coming up with an order-of-magnitude of improvement in programming techniques. What you're demanding is not far removed from saying that better programming models require improved transistors.
Also see http://www.anandtech.com/show/7334/a-look-at-alteras-opencl-... for more information
On the point itself, "accessibility" is probably understating the nature of the problem. Actually being able to program for a platform, and how easily you can do it, is a huge issue. Issues of "accessibility" apply as much to huge companies with giant server farms (note how little even GPU computing is used by them) as they do to tinkerers.
In any case, my own viewpoint is that CUDA and OpenCL are the best that has been produced by a huge effort by well financed and technically sophisticated groups. GPU computing offers advantages in flops/watt and flops/dollar, and yet is still not widely utilized because of the increased programming complexity.
Given this, I think it will be very hard for a small group to compete using a completely new architecture. On the other hand, you are starting with a blank slate and you are controlling the entire stack, so I hope you can use that to your advantage.
It depends on what you're trying to do! If you want to multiply big matrices, GPUs are great for that. A lot of scientific applications work well on GPUs. And obviously, GPUs are great for graphics.
Also, FPGAs are no longer blank templates on which you can stamp any design. The new ones all come with built-in "IP blocks" which you can't change. So you get a bunch of gates you can modify, but also perhaps dozen CPUs and a bank of memory, or so. Maybe someday GPUs and FPGAs will converge-- the former are getting more flexible, and the latter are getting less so.
The biggest problem with contemporary GPUs is that they're I/O-starved, which I don't see anywhere in your comments. PCI-e is just not enough bandwidth. That is why GPUs are a sideshow in big data. It doesn't matter how many cores you have if you're sipping your data through a straw.
The power consumption argument seems like a strawman. Replace a few incandescent lightbulbs in your home with LED ones. Congratulations. You can now run your GPU 24/7 and come out ahead in power consumption.
Umm, you took "at a ridiculously higher level of efficiency" out of my quote in order to call it silly? C'mon man.
If you've ever diligently crunched numbers on FPGA versus GPU monthly electricity costs at scale, I don't think you'd be capable of saying power efficiency is a strawman. If the idealistic efficient-resource-usage aspect is not convincing enough for you, we are talking millions and millions of dollars here (actually more). Look at Amazon GPU numbers and get back to me. We have detailed breakdowns of price comparisons that have been worked out in depth using the same algorithm. This is nothing like swapping out lightbulbs, that's an actual absurd and uninformed statement that can be verified and quantified as such.
And I do mention both IP blocks (even in quotes as you do) and how SoC's CPU-PL communication is much faster from being same-chip-integrated, hours before you did, so I'm not sure what you mean.
I'm also not getting why you use hard IP blocks as an argument that FPGAs are getting less flexible. They are helping to advance FPGA accessibility and flexibility a ton. Why do you think they're there? The development time savings are huge and the testing and optimization is by nature better.
The 7000 series has pretty heavy options, you can get a lot PL all your own on them. We've been specifically talking about an ARM+PL SoC, are you trying to say that an FPGA like a Virtex or Kintex is somehow not a "blank slate"? You want no standardization of any kind? The lack of that is a huge problem, and it's getting solutions.
I used to work at a startup company that had a plan to reduce power consumption in data centers by spinning down hard drives when they weren't in use. Too late, we learned that there isn't a lot of money in saving power. The average data center in the US might spend $500k/year on electricity. That sounds like a lot, until you consider that the fully loaded cost of a good engineer will be around $200k/year. Power becomes a factor only when you hit a wall in terms of how much can be delivered to the data center.
I was an electrical engineer in college, and I learned how to write Verilog and do register-transfer-level design. It's something that I don't think most software engineers have the background to do. Currently, I work in the area of big data, on Hadoop.
I cannot deny that Hadoop is not very power-efficient. But it offers some things that are more important to customers: infinite scalability, the ability to run on commodity hardware, and the ability to interface with the system via normal-looking software.
Actually, even writing Java code is too difficult for many Hadoop users now. They prefer to write SQL. (Even SQL is too hard for some, and they use automated tools to generate the SQL and produce charts.)
Big data is more than just querying. There are people running fancy machine learning algorithms. Those are the people most likely to use the flexibility of having a Java (or other programming language) interface.
I'm curious why you think FPGAs will win over GPUs in HPC. I struggle to imagine academics writing Verilog or VHDL. I can just barely see them sending in legions of grad students to try to write CUDA or OpenCL, but RTL design seems a bridge too far. I also haven't heard of any of the big FPGA companies trying to make a splash in HPC, whereas NVIDIA has been very active with Tesla and its other high-end offerings.
(in terms of architecture, not performance benchmarks)
One of the biggest obstacles to GPGPU exploitation is being at the mercy of each vendors proprietary software stack and the resulting fragmentation & lack of openness. It's like the pre-PC era, without Unix... Larrabee might have helped. Now Xeon Phi is a niche product.
Yeah, you have slower cores and the SIMD is twice as wide, but its main selling point is that you get to keep the regular programming model - unlike with NVidia or AMD. Along with benefits that come with it (mature software, openness, etc).
See eg. Intel's Xeon Phi Programmin Guide for an intro: http://software.intel.com/sites/default/files/article/330164...