How FPGAs work, and why you'll buy one (2013)
embeddedrelated.com
embeddedrelated.com
I love FPGAs. They're awesome for interacting at a low level with just about any kind of hardware, at nano-second-level latency.
But although I bought one, I don't think most people will end up buying one, unfortunately; the vast majority of the consumer stuff that we considered putting on FPGAs in the past (graphics rendering, heavy math) is solved six ways to Sunday by today's GPU. The FPGA also lacks an open-source (or even shared-source) toolchain, there's no way to perform resource sharing, they have code constraints that are physical limits, and so on.
Basically, FPGAs rock when used for chip design or for powering custom prototype or small batch hardware. But they fall into a really narrow portion of the market, because there are few cost-effective consumer-level solutions that make use of FPGAs.
A shame, really. They're super cool.
They fall in a middle ground between things I can build more simply with a super cheap microcontroller (I buy Arduino nanos for $2, for example), and things I really need GPU compute for. So I'm stuck. I keep thinking I'll do a DSP project soon, but that's more because it is one of the few things I think an FPGA will be useful for, than because I need it.
So cool, yes. Practical, maybe - but not very, at least for me.
Things like the ZINC by Xilinx are working its way into things like the Red Pitaya[1] which is absolutely remarkably capable for the price range, but yeah FPGA's are pretty niche for the common end-user. Lithography on sil wafers ends up being cheaper at nearly any scale. It turns out more economical just to do that one of those cheap FPGA-to-ASIC conversion dealies (rather than using Cadence or whatever VLSI tooling you have to DFM from the start [I think UC Berkeley has a full architecture/RTL/logic/circuit/physical kit]) at the threshold of around a few thousand units.
I think this point is overstated. FPGAs are not microprocessors. I don't think an OSS toolchain would be anywhere near as beneficial as something like gcc. It would be nice to have but it's hard to think of any problems specific to programmable logic that this would address.
What would help is toolsets with less friction and less of an entry barrier.
It's worth noting that arachne-pnr (https://github.com/cseed/arachne-pnr) is written by a different person (Cotton Seed).
(109 points, 331 days ago, 107 comments)
https://news.ycombinator.com/item?id=9388751
(379 points, 999 days ago, 152 comments)
That's an understatement.
http://ark.intel.com/products/84685/Intel-Xeon-Processor-E7-... http://www.amazon.com/Nvidia-Accelerator-passive-cooling-900...
Now, I guess you can decent performance out of cheaper chips already, but paying $5000 for an Intel CPU is hardly rare.
* Fixed point.
* Resource starved. Limited # logic blocks, using DRAM from an FPGA is a massive PITA and you pay dearly for every 100kb of SRAM.
* You figure out how to pipeline so that every cell gets used on every cycle.
* HDL languages
* Terribly long compile+route times...
* ...and sometimes routing fails even if your program is well-formed and "ought to" fit on the chip!
* Proprietary toolchains
* Expensive kits (don't let the $20 kits fool you -- you'll spend thousands if you're really using an FPGA for what they're good at)
* Even more expensive test equipment (makes the thousands you spent on the FPGA kit look cheap)
The flipside is that GPUs don't do deterministic latency or talking to hardware. If you need to ship billions of samples per second from your JESD204B ADC back across your PCIe bus (to get them into your GPU, of course) then there's only one chip that fits the bill: the mighty FPGA.
What that meant was that the AVR, which was relatively slow at 8MHz, was free to run all your logic (including multithreading!) while the FPGA handled most of the timing-critical things, like motor PWM and sensor decoding.
But the coolest part was that you could dynamically modify the IO if you needed different behavior. Instead of having to find, buy, and wire up a standalone quadrature decoder IC to count axle rotations, I just had to write some code [2] and suddenly the board has a high speed quadrature decoder (or several) built in!
The biggest roadblock is the toolchain though - even though [2] was a pretty simple change, it took a long time to download and install the Xilinx ISE tools, and they're not the easiest to use.
I would love to see a general purpose AVR+FPGA board with the FPGA toolchain neatly packaged the way Arduino/Wiring has done for AVR. Verilog has a somewhat steep learning curve, so you might want to hide that in some kind of module system, with basic library modules like PWM, edge counting, or quadrature decoding that can just be mapped to the pins you want. Maybe something building on top of Icestorm?
[1] http://spacecats.mit.edu/contestants/happyboard-manual.pdf
[2] https://github.com/sixtwoseventy/joyos/commit/44975ea9bc64e5...
The cost in time and complexity of solving a problem in hardware is almost always larger.
When writing software you gain a level of abstraction. You do not have to deal with getting a special hardware (the FPGA), mess with proprietary buggy tools, and all the nitty gritty stuff of proper reset, multiple clock domains, timing closure where each rebuild of your code might take multiple hours to close. And don't get me started with writing HDL-code.
I implemented an object tracker using a FPGA and a webcam and there were multiple hardware problems that needed to be handled such as buffering data in the sdram (design a sdram memory controller), research and implement image filters and debug why the camera image sometimes got corrupted (the connection cables were too long).
Almost all of these problems would have been gone by instead connecting an USB-webcam to my PC and utilize the openCV library.
Why?
I have exactly the opposite feeling when I'm dealing with software - "if only I could design my own cache here, how much easier would it be to get around this bottleneck!"
Not until you hit a deep semantic difference between your hardware and a code you're trying to build.
For example, there are many things that are best expressed as a Network-on-Chip: a pipeline of specialised CPU cores each doing something simple. And such a notion is not that easy to represent in a pure software without a gigantic performance penalty. A pair of a CPU core and a small piece of code running on top of it can be far simpler than any piece of code written for an existing but not fitting architecture.
A simple example of such would be a dedicated Forth processor plus a Forth implementation. They two together are often much simpler than, say, a standalone Forth implementation for x86.
High-level synthesis tools on the other hand might make your claim true. They're a combo of HW and SW languages.
"A simple example of such would be a dedicated Forth processor plus a Forth implementation. They two together are often much simpler than, say, a standalone Forth implementation for x86. "
Simpler for who? The person implementing both a Forth implementation and a x86 processor? Or the person just implementing a Forth on x86? I've seen Forth implementations in SW. They're trivial. Compilers, libraries, and hardware do all the heavy lifting. I'll put it to the test, though, as I feel you might have picked a good example.
J1 Forth CPU Verilog code http://excamera.com/files/j1demo/verilog/j1.v
eForth assembler for Z80 & such in few assembly instructions http://www.figuk.plus.com/4thres/systems.htm
Forth interpreter in Ada (1985) http://www.forth.com/archive/jfar/vol3/no2/article7.pdf
Maybe it's just that I don't know Verilog. However, the Verilog code looks more like the assembler in second link. It's like a pile of gibberish. The Ada code is very straightforward. I bet the debugging was easier, too. S, I'm still not sold on hardware reducing complexity versus software implementation.
I'm 100% agreeing on improving performance esp where semantic mismatch occurs. Hell, that's what inspired me to ask you and some others to evaluate those two or three HLS tools so I could start cranking out accelerators if one was legit. ;)
Here's a great stackexchange discussion on why and how FPGAs can outperform CPUs for certain tasks: http://electronics.stackexchange.com/questions/101472/how-ca...
While the idea of programming them with VHDL or Verilog is quite straightforward, there were a lot of bumps in the road to getting a working design. Specialized tools which generate reports that go on for pages and pages (our team was completely unable to do something about the path of longest delay), complicated information about the clock, and having to work with vague documentation and intellectual property.
FPGA's are really nice when they work, but I feel like they are just too complicated for developers to get started with on their own, at the moment.