How many 32-bit RISC CPUs fit in an FPGA? Now vs. 1995?
forums.xilinx.com
forums.xilinx.com
At the time, it seemed obvious that, over the next decade, FPGA's would work their way into general-purpose computing, so that (for example) Photoshop filters would simply reconfigure the FPGA to run blazingly fast. Likewise with games, or video codecs, or 3D rendering, or whatever else was processor-intensive.
But that clearly hasn't happened. Instead, GPU's took off as the main computational supplement to CPU's.
Does anyone here have any insight as to why? Is there a technological reason why FPGA's never turned into general-purpose hardware standard on every desktop and laptop? Has it been a chicken-and-egg problem? A standardization problem? Or something else? Do FPGA's still have potential for general-purpose consumer computing? Or are they going to be forever relegated to special-purpose roles?
Technology sucks :(
Another likely issue with FPGAs is that for everything but raw compute they still need support circuitry (physical I/O ports, memory interfaces, etc) that are different for every app but not really "reconfigurable" in the same sense as the FPGA fabric.
Finally, I'd wonder if reconfiguration time is part of the problem - until relatively recently, reconfiguring an FPGA was all-or-nothing and could take multiple milliseconds. Not a big deal when configuring a device once on boot, but serious headwind when trying to context-switch between different jobs that need FPGA assistance.
I believe there are 4 reasons why GPUs have taken off conventionally as compared to FPGAs:
1. Last I checked, the FPGA vendors will not open their toolchains up, and not even document the bitstream formats. They will claim NDA, proprietary, etc. This has the massive side effect that you are stuck with their bloated, slow, crappy toolchains. If this were open, I guarantee hackers would be inventing all kinds of interesting ways to convert their software into FPGA bits.
2. FPGAs are VERY hard to write and debug. You have to write your design in an HDL language (either VHDL or Verilog), and you have to use a software simulator to prototype the design on first (and of course these tools are either quite pricy or if free they are usually limited or hard to use). Then you can synthesize the design and download it into the FPGA for running.
The next problem is debugging your design. The entire internal state of the FPGA is only accessible through slow scan, unless you dedicate a portion of your design to "monitors", which tap the traffic and store their values into internal RAMs. So you may have to respin the design just to get more monitors to debug where the issue is.
3. FPGA compilation is SLOW. When I used them professionally a few years ago, a Virtex5 could take multiple hours to resynthesize/place & route a medium-sized design. I believe that Virtex7 they are advertising could take over a day to respin if you change your design.
4. Most new machines already have a built-in graphics with a GPU that can be utilized as a general-purpose GPU. No one ships FPGAs in any conventional computer.
(3) is another issue, but I don't think the consumer would necessarily need to worry about compilation. The developer would just include the compiled programming files for different FPGAs in the application.
If you mean that it'll be slow on the developer's side, that's definitely a valid point. I'm sure, however, that you'll see FPGA manufacturers start to move toward remote compilations so that you're not necessarily limited by the hardware you have in-house.
[1] http://www.altera.com/products/software/opencl/opencl-index....
Altera calls it "logiclock," Xilinx has a different term, but the idea is that you don't need to re-synthesize the entire FPGA for every change. In fact, you may not want to. If you are tweaking a certain region, you're usually happier if the place & route doesn't send your stuff through a route that then kicks off a line in another block so that the timings are now off in that other block.
For an FPGA the timing is how you measure performance and getting the best timings can take quite a bit of work. Being able to lock that once you've got it right is a big plus.
It's kind of a mixed bag. It's worked okay for me if changes are truly minor, but if there are large changes to the logic it doesn't seem to be very good about "forgetting" what it learned from the previous pass. Three or four times this week I've had a design fail to make timing with SmartGuide, but work when doing P&R from scratch.
In addition, most of these tools compile to HDL, so they only add to compilation time.
My reasoning for each:
#2. More open tools would allow alternative programming models. For example, gcc already has a vhdl front end. Why not a gcc back end for an FPGA? That would open the door to more familiar languages.
#3. More open tools and specifications would allow programmers to start optimising and rethinking the FPGA compilation process, potentially leading to radical reductions in run times.
#4. People won't want FPGAs in their machines until they are easy to use. Solving #1 (and consequently #2 and #3) will make FPGAs easier to use, increasing demand, prompting manufacturers to consider including programmable logic in their machines. Granted its a chicken and egg situation between adoption and better tools, but opening the tools and specifications up could break the cycle.
Assuming the above, I'm thinking of a project, independent of gcc, that takes gcc's intermediate representation and does all the FPGA specific tasks that you mention. Yes, it would be a huge project, comparable in scope to gcc itself, and even that might be an underestimate. It could start small, to make it realistic, then incrementally expand its scope, just like linux and gcc did. Eventually, the FPGA vendors might have to choose between participation or losing customers? It might be able to exploit some of gcc's backend infrastructure in the FPGA process, but who knows?
> It just isn't a realistic goal.
Or it's a red rag to a bull, to the right person. :-)
Including using an FPGA to accelerate that process -- probably possible since from what I know there is a lot of parallelism involved in synthesising logic.
But not the kind of parallelism that's fast on an FPGA.
#3: No. The information needed for synthesis (HDL to netlist), mapping, and placing is publicly available. These topics are actively researched, yet so far no truly usable open source tool has emerged. Routing tools are not possible without the information available, though.
#4: Sure, solving #1-3 would make FPGAs easier to use, but #2 and #3 don't follow from #1 even if #1 was satisfied.
I have no idea what happened to them, but I suppose the problem was harder than they believed or at least claimed.
There are a lot of pain points in the HDLs, but it seems like Verilog has more than the others.
I saw someone working on a clojure HDL, I think it might have compiled down to or emitted Verilog. I thought it was more confusing than the HDLs to begin with, but depending on one's background it might make more sense.
I don't mean to excuse the manufacturers, but at the same time they seem to be selling into a pretty small market and it's not clear to me that opening things up will magically lead to a big expansion in chip sales that will negate the competitive risk of being the first to open up. If you have time, I'd like to learn more about this since you seem to have a lot of experience with this technology.
In my opinion there is very little for the manufacturers to gain by keeping their bitstream formats proprietary and undocumented. I don't think there is a competitive advantage, as all the manufacturers are pretty much doing the same thing. And their FPGA block diagrams are already open and documented (you can see how many flip flops, clocks, and muxes are in each logic cell, how the routing works, and where the memory cells and other units are).
I was under the impression that FPGA vendors often license functional blocks (like PCIe SERDES) to FPGA users. Might it be that part of the purpose of obscuring the bitstream format is to make it more difficult for customers to use those functional blocks without paying the toll?
Well, that's no surprise: FPGA vendors spend a lot of manpower on improving their Place&Route software. If you wanted to build something competitive, you'd need a lot of money plus access to proprietary, non-public, information.
Similarly, you can reduce the complexity of routing calculations by applying some constraints. You will potentially lose the possibility of an optimal solution, but you will gain a far faster compilation time. As always with engineering, it's a trade-off.
One of my favorite ideas on this is space-filling curves: http://www2.isye.gatech.edu/~jjb/mow/mow.pdf
It is if the original premise was to make Photoshop filters fast. A GPU can make my Photoshop filters fast now an FPGA implementation can make them fast 8 to 24 hours from now.
[0] http://en.wikipedia.org/wiki/High-level_synthesis [1] http://www.xilinx.com/products/design-tools/vivado/integrati... [2] http://www.synopsys.com/Systems/BlockDesign/HLS/Pages/defaul...
Also, GPUs might be a better match for the kinds of codes people care about. The world didn't need arbitrary bit-level computations except in rare cases, it needed insane memory bandwidth and high floating point throughput (exactly what a photoshop filter would need). The generality of most FPGAs mean they're not great for standard circuits that can be optimized. Maybe this means the FPGA market might see some success with a different trade-off of flexibility vs fixed hardware. The rise of FPGAs like the Zynq with dedicated processors or distributed RAM and DSP units is already happening.
But GPUs got shipped in volume. I think they were just cheaper for the performance level.
What???? There is a very mature set of tool for converting MATLAB to HDL.
http://www.mathworks.com/products/hdl-coder/
http://www.mathworks.com/products/hdl-verifier/
http://www.mathworks.com/products/filterhdl/
I'm sure there are similar tools for other languages. It requires you to program in a slightly different way (certain operations aren't optimal for FPGAs), but is extremely user-friendly.
At the company I work for, all FPGA programming is done in MATLAB. Unfortunately I work in a different department so I can't give you any technical details, but from what I understand, no one has written things directly to HDL in years.
Some people, when confronted with a problem, think
"I know, I'll use MATLAB." Now they have two problems.
(Apologies to jwz.)eg for something very simple you can do it for thousands or tens of thousands...but in that price range it's probably still more economical to implement in software or do it discretely with SMT unless you are very sure about the existence of a market.
However, algorithms are so rarely the stumbling block for applications; data storage, management and communications are. The major exception is graphics, hence the GPU.
The money is a huge barrier to entry for hobbyist types, but if we're talking about commercial stuff, it'll probably cost you a lot more to pay for the engineers who are competent enough to implement something that will actually work than it will to do the fabrication.
"I like FPGAs, but I doubt they will ever become widely deployed. They pay ~20x overhead, so any algorithm that is a good fit for them becomes a new instruction in the next CPU generation. The reprogrammability is only a feature in highly constrained (i.e. niche) environments."
http://www.reddit.com/r/IAmA/comments/1yj77b/as_requested_i_...
I think one of the main reason is programming the FPGA is not easy, the tools needed for FPGA development are all proprietary. Every FPGA has different way to program it.
I actualy have 2 FPGA-based devices that I use daily - polyphonic analog synthesizers, to be specific. Manufacturers are tight-lipped about how they're using the FPGAs, but it appears to be for ultra-rapid reconfiguration of analog circuit topologies without the load time delays that result from a traditional microcontroller > D/A converter arrangement. There are also field-programmable analog arrays on the market, but I have yet to see one in a commercial product and it seems like they have some way to go before being economical for audio synthesis applications.
I love FPGAs although I don't know much about how to et started with programming them. There is an ASIC coming out of patent in a year or so which I'd like to re-implement in an FPGA package, and I've thought about reaching out to the original architect who lives nearby and is a friendly fellow. I'm not sure how feasible this is, though.
It has an ARM processor on the same chip as the FPGA, which I've found to be incredibly useful.
If you don't want the ARM processor, the regular DE1 is somewhat cheaper: $150 or $125 w/ student discount.
http://www.terasic.com.tw/cgi-bin/page/archive.pl?Language=E...
Be aware that the DE1 has pretty limited resources for somethings. For exmaple, the on-chip memory can get tapped out pretty quickly when doing embedded applications, and it has less I/O and less RAM available off-chip.
Also note that the DE1-SoC actually uses a Cyclone V device, whereas the original DE-1 uses a Cyclone II, a nearly discontinued chip that is a big step down.
I found great documentation and economical boards starting at $55 available here: http://www.xess.com/store/fpga-boards/
BTW, I also found Scala based hardware construction language Chisel very interesting https://chisel.eecs.berkeley.edu/.
Thanks for the other replies on dev boards etc.
Quartus II Web Edition is Altera's free IDE, and it comes with pretty much everything you need: a big suite of libraries (most of which are also free to use), a graphical entry environment, an HDL editor, the full synthesis toolchain, and integrated Eclipse tools for writing embedded C/C++/ASM code. Other than that all you'd need to get is ModelSim-altera, the free version of the standard simulation environment.
Altera also has some pretty comprehensive (i) free online training, (ii) IP block, i.e. library, documentation, and (iii) complete reference designs. I'd recommend checking out all three, especially (i), since at the beginning it's easy to get tunnel vision just learning VHDL or whatever, and then realize that you're somewhat clueless as to how to actually get things done in any useful capacity. There's a dizzying amount of jargon and proprietary bullshit, so it's useful to just have someone tell you what everything means, and How It's Done(tm).
All that said, if that ASIC is anything complicated it might be a bit of a big project to jump into. That and it's possible that it would require more resources than an entry level board will supply.
A long time ago computers sometimes came with DSPs dedicated to Photoshop-like programs. The GPU is a logical extension of that. The FPGA is something entirely different.
Maybe FPGA's are competitive in integer math, but that's a pretty small niche ,not big enough for a coprocessor.
That is really the unique benefit of FPGAs: the implicit parallelism makes them a really great fit for the kind of high-speed bit banging that would be impossible to get right on a normal CPU (not to mention very difficult to program in the first place).
2) GPUs, being purpose-built and mass-marketed, are both much cheaper and much faster. Think of the Ford Fiesta ST vs. a Jeep Wrangler. The Jeep is simple & more reconfigurable, yet is slower and more expensive- for exactly those reasons!
3) GPUs and CPUs compliment eachother well. The only gap is heavily parallel branching code- GPUs are bad at if-statements, CPUs are bad at heavily parallel. But a branch predictor is the key to if-statements, and branch predictors are hard.
FPGAs are fundamentally a prototyping tool. Can you think of examples of prototyping tools that eventually broke into the main market, replacing the incumbent product?
Python
FPGAs can be used for prototyping ASICs, but that's by far not their major use-case.
- digital signal processing (DSP)
- network equipment (modern high-end FPGAs are able to function as 100G Ethernet switches and routers, just connect them to some PHYs and fast DRAM)
- systems on chip (SoC): processor, fast DSP, all kinds of IOs, memory controller – all on a single chip
- realtime video stream processing
For more examples, see http://www.xilinx.com/applications/index.htm
CPU gives you speed in computation. Except for a custom ASIC, nothing can match a modern CPU in single thread speed.
GPUs gives you parallelism. You can get thousands of decent CPUs in cheap boards today.
IMHO FPGAs main advantages today are bandwidth and auditability. But neither are very important in most applications.
Consequently, FPGAs don't lend themselves to an uncontrolled explosion of innovation. Compare with GPUs and CPUs where open source compilers are available and anyone can have a crack at innovating, at whatever level they choose.
As an example, if you want to dynamically generate programming for a Xilinx FPGA, you need to incorporate Xilinx's binary only PPR program somewhere into the flow. That acts as a ball and chain around the leg of reconfigurable computing, impeding its development and adoption.
[1] http://www.ifixit.com/Teardown/MacBook+Pro+15-Inch+Unibody+M...
How do you balance the need for high-performance communication between the FPGA and the rest of the computer with the inherent inability to trust whatever configuration has been loaded into the FPGA? Software security is hard enough without infinitely reconfigurable devices lurking inside our machines.
How does a possible FPGA configuration work optimally across a wide variety of FPGAs (which can have various "built-in", non-reconfigurable components) and sizes that can only be determined at "flashtime"?
What exactly do regular users need an FPGA for that isn't already handled by dedicated silicon with much greater performance and much less power usage than an FPGA would?
EDIT: I'd like an FPGA card for my PC just for running FPGA "emulators," because you can recreate a SNES or whatever in a tiny fraction of the gates that you need to make a CPU capable of emulating it with cycle accuracy at full speed. I seriously doubt there's much demand for that, though: if you're going to go to the trouble of buying special hardware for emulation, you might as well just buy the console and a flash cart.
As for regular users, personally I'd like it for video codecs. I'd like it for video transcoding. I'd like it for faster MP3 encoding. I'd like it for Photoshop filters. I'd like it for speech recognition. I'm sure I could think of things that other users would like to use it for, like 3D rendering. I don't know of dedicated silicon that does ANY of these things, except specifically for H.264 decoding. That's it. And the thing is, for all the items I've listed, it's totally fine that my computer is only ever doing one of them at a time. These are all "regular user" uses, whereas virtualization is not needed very much for regular users.
Without the annual growth of transistor density, the second best avenue to gain performance will probably be specialization, and reconfigurable specialized hardware looks more attractive than fixed one.
FPGA's are horrendously inefficient, space and power-wise, for many tasks. Much of the core is taken up by the programmable routing between different components, and generally most components will be not fully utilized (such as LUTs and RAM).
Yes, the Virtex-7 (a VERY EXPENSIVE fpga) can hold 1000 very simple 32 bit cores.... but a top end NVIDIA GPU has 2688 CUDA cores. While not entirely independent, these CUDA cores have far deeper pipelines and far superior ALU's to the one in the article. If your software fits the programming model, GPUs will handily beat a FPGA.
FPGA's are great for prototyping ASICs, and cases where timing is of critical importance - try implementing a VGA video generator on a CPU. Basically everywhere an FPGA would excel, an ASIC excels more, but FPGAs are great for low volume and or specialty hardware where a GPU is not good at accelerating the task at hand.
But the point of the FPGA is to not build cores inside it; the OP article is just an exercise. Application specific logic built with FPGA might be much more efficient than generic CUDA cores.
Also, it is a given that an ASIC will be more power efficient than an FPGA, but the FPGA will be generic and hence more money efficient.
The coding was hard and performed below expectation. The only customers that benefited were military/intelligence agencies who were using outdated techniques such as filtering in the Fourier basis.
The problem was data-transfer compounded by inflexible code generation. The data still needed to travel from RAM which meant that even without computation you couldn't get over a 50% increase in performance. It would need to be send back for more complex processing.
So, your vision of general-computation-in-silicon has occurred, just static, not dynamic.
1. They are already there in commodity HW.
2. GPU vendors make money with their hardware; they need not extort developers for a development environment like FPGA vendors think they have to do. Therefore, development for GPUs has a lower barrier of entry.
pdq, the last time I built this design, with more fully elaborated processors (control units + multiplier FUs) it took three hours and 16 GB physical RAM on a Core i7-4960HQ rMBP.
In some flows you can do a coarse floorplan of your design and route the submodules separately and then stitch them together. I imagine this is how the very largest devices are implemented in manageable design iteration times.
I don't usually worry about that, though. Since my design is just so many replicated tiles, I tend to do design iterations of 4- or 16-processor elements to test the impact on clock period / timing slack. That usually takes 2-3 minutes per design spin. Only once in a while do I place and route the whole chip to confirm some change doesn't impact timing closure.
Even with 18 years of x86 performance advances, it takes much longer to PAR the large FPGAs now than it did back in the day.
I'm currently synthesizing a 45nm ARM cortex-M0 design using Cadence Encounter flow. The complete process (RTL compiler+place+route) takes only 5 minutes!
In both tools most of my design spins take <3 minutes.
If you were building an ASIC you'd pay $$$,$$$ for such tools. The economics of FPGAs are such that the tools are either free or $,$$$. Both are reasonable and accessible to an enthusiast/practicing EE, respectively.
I am so grateful to Ross Freeman, inventor of FPGAs, and all the engineers that followed in his footsteps, for democratizing access to state of the art high performance digital logic. For $100 or so you can get a 28 nm device filled with 10,000s of LUTs and hundreds of RAM blocks and build whatever you can imagine. Amazing.
3 minutes is very fast...one of my projects takes about 30 minutes for ~30K Luts on a modern Core I7, I used too many registers.
Developing and debugging it is basically torture.
Today's quad-core Haswell is made of 1.4 billion transistors crammed into 177mm². That's almost 8 million transistors per square millimeter.
My naive best guess is that when CPUs stop getting much faster the next wave of innovation will be on bus speeds. But I'm not a chip guy so perhaps my view of the problem is overblown, but I'm very interested in learning more about this from others.
So on average we might be able to see 50x improvement in the next 30 years (according to a darpa manager).
If by faster you mean clock speed, they haven't been getting faster for a while now. Memory and IO speeds do lag behind, but we have caching to solve the former and SSDs to solve the latter. One pressing issue now is getting the power consumption down. This is especially important for laptops and mobile phones. FPGAs are pretty good at using less power, but you have to sacrifice a lot of performance and programmability.
I think that these FPGAs, when put into consumer computers, will be more like "data centers on a chip" rather than processing cores. In a simple example they could run map/reduce type operations, colocating storage and computing silicon.
Although I'm quite sure those sorts of amazing designs are already well used by the NSA. I'm sure hardware research could be the real breakthrough in cryptanalysis.
EDIT: Oh I see. Microblaze is propietary and didn't exist back then.