BORPH: An Operating System for FPGA-Based Reconfigurable Computers
casper.berkeley.edu
casper.berkeley.edu
Think of it this way. At the basic level, you have logic gates. CPUs are massive ensembles of these which then run your program majorly sequentially. GPUs/GPGPUs are smaller ensembles that can be configured better for specific tasks, resulting in better performance/power ratio. At the other end of the scale is using HDLs to program the gates directly for the specific task at hand, which would offer best performance/power ratio. The development process is however more involved. In-between the last two is reconfigurable computing.
A MicroBlaze is just a regular CPU instantiated on an FPGA, most likely controlling other logic around it designed using HDLs -- HDLs portions taking over compute-intensive tasks while the CPU takes over (relatively) low-speed control logic. A reconfigurable computing environment would try to make this more symmetric.
In my opinion, there is a continuum between more complex units (CPUs) that have a large number of logic gates and offer lots of features, and between less complex units (logic gates directly) offering very small functionality. At the end of the day, it is all about finding the right architecture for this "unit", which may vary from application to application.
There is a lot of development going around these days for using OpenCL to program FPGAs directly. E.g.: http://www.altera.com/products/software/opencl/opencl-index....
Here is also a much clearer explanation taken from that link:
BORPH is an extended Linux kernel that treats FPGA resources as native computational resources on reconfigurable computers such as BEE2. As such, it is more than just a way to configure an FPGA. It also provides integral operating system supports for FPGA designs, such as the ability for an FPGA design to read/write to the standard Linux file system.
If it supports partial-reconfiguration (which it looks like it does), then it could be a very handy tool. Why? Hardware is significantly faster than software. While Linux is running, the ability to spawn hardware at will would be great for many applications.At UCSB, some of the research I did related to this very problem. What I was trying to do was have a Linux web server running on an FPGA that could dynamically reconfigure itself for different experiments. I ended up choosing a board that has an FPGA that communicates with a ARM processor. Decent size FPGAs were (and still are) expensive. For all the Linux overhead and my budget, a hybrid FPGA-processor platform turned out to be the better solution. If anyone is interested, here is the problem we were solving: http://ece.ucsb.edu/academics/undergrad/capstone/presentatio...
http://www.hpcwire.com/hpcwire/2011-07-13/jp_morgan_buys_int...
I would not at all be surprised that if that cleaned up code was ported to a GPU, you'd see a much bigger speedup.
And if they did this for the GPU code, I'm not quite sure why they needed a consultant to do the same thing before starting to work on FPGA code.
I'm not saying the numbers are wrong, but this is certainly insufficient data for any decision :)
These numbers are consistent with other computations I've seen that were well-suited for FPGAs.
(Note - it might well be that the FPGA version of that particular problem would be even faster. I merely posted the GP to point out that there was significantly more effort to tailor the model towards the FPGA than towards the GPU, which seemed to be merely a "compile for GPU" without a restructuring to make the data model fit.)
The biggest challenge for FPGAs is tools; the CASPER guys used Matlab Simulink; which works OK for small designs, but becomes challenging as the devices and designs grow. There are also a bunch of C to gates tools, Matlab to gates and Haskell inspired tools like Bluespec, and even some Python tools. In fact for the Berkeley CASPER group, as GPU through put improved (along with processing) and the size of the radio telescopes increased, it made sense to move to a hybrid FPGA+GPU backend processor.
For fun and training we built an ARM Cortex A8 (with or without GPU and vector processor) + FPGA platform called Rhino (http://www.rhinoplatform.org), with a couple of FMC IO slots for attaching ADC/DAC/Control etc then ported Borph and then linked up with the Migen project (https://github.com/milkymist/migen) to "simplify" programming and configuring the FPGA (https://github.com/brandonhamilton/rhino-gateware). We think it works great ;-)
We have used it to build a arb. waveform radar transceiver and a few other things.
Ofcource now one has the Xilinx Zynq 7K family which includes dual core ARM and tightly coupled Xilinx V7 fabric all on a single chip. Plus for a few hundred $$ you get a nice development board (www.zedboard.org) and a frontend wirless digital receiver FMC card - http://goo.gl/qCnJM
Not sure if that helps...