Introduction to FPGAs
smist08.wordpress.com
smist08.wordpress.com
I have an orangecrab board with a lattice FPGA, it’s super cool to be able to run a single makefile and build a riscv cpu and buildroot based linux for it.
I tried getting into FPGA development, played around a bit with simple Verilog implementations then got a cheap FPGA board an pretty much failed at Vivado. That tool is completely unusable. It makes me think of 90s Visual studio where you have to jump through 5 forms and wizards to generate a broken project that wouldn't compile or run on your board.
Are there shell tools for FPGA programming you can just set up with a Makefile or something? Having to use a GUI for stuff like this seems silly to me.
Of course, for many things, you might not need to configure vendor IP. Examples where you most definitely need to are 1) Filters, 2) FFTs, 3) DDR controllers 4) High speed transceivers. In some cases you can configure vendor IP with language templates.
If you don't need these components, then Yosys will be good for you.
But if you just want to get up and running with a dev board, use Vivado (for xilinx) and Quartus for altera. Neither are great. For Vivado, use a release that ends in a X.4 unless you want to bang your head against the wall for a few hours.
If you are not liking the graphical flow for Vivado, you can definitely just write your VHDL or Verilog and compile it without the need for "block diagrams" ,etc. Vivado calls it the "non project" flow I believe and thats what most of us in industry use.
Most people in industry use the Vendor tools. A few years ago you could get 3rd party tools but they weren't much better. The open source toolchains are coming along but still not for the beginner. The vendors will make their libraries more like foundry PDKs soon though and open source 3rd party tools will hopefully finally take off.
There's an example of using a makefile for Vivado here: https://github.com/hdlguy/make_for_vivado
Don’t worry, it’s not for everybody. It looks like other code, but it isn’t your casual code. I do it for more than decade, colleagues from other groups come me, I teach, they try and never come back.
That said, all the stuff inside the Vivado pkg is scriptable through Tcl.
When I was playing with this stuff... Between that and Fusesoc + some hooks with CMake + Verilog mode in CLion I was able to get a fairly reasonable working environment that wasn't totally awful. Still had to dump out into Vivado here and there to configure various "IP" blocks (and then take the generated Tcl and clean it up and import it into proper source control)
It was a fun hobby for awhile, but I gave up and left it behind.
It wouldn't surprise me if those 10 million LUT FPGAs end up bloating it up for the every day hobbyist.
Some form of modularity would be welcome.
But even if you remove support for all but one model of FPGA chip, it's still a very large install. And if you expand that to one entire FPGA family (maybe 100 or so chip models), that jumps it up by 50GB easily. So 500MB per FPGA model. And these models are highly similar within the same family, they just have different numbers of LUTs and layouts.
My hopeful theory is that they're doing some kind of per-device precomputation that helps accelerate placing & routing. The theory I actually believe in is that they're just duplicating lots of data unnecessarily, and don't see any value prospect in fixing it.
>To build a CPU, you need a few more elements than basic logic gates, namely memory and a way to synchronize everything to a clock, but this is a good starting point.
You can actually make all that with logic gates! Two NANDs with each's input connected to one of the other's output makes the most basic sort of static memory element. 4 NANDs and an inverter give you a D-latch that lets you synchronize your circuits to the clock signal. You can use fancier techniques that don't correspond to these so well but these are actually all you need.
It's not a great idea for a production machine but it's much more practical than the software analogy of using a Turing machine for everything.
A D latch would be enabled all the time the clock signal is high, so it is not suitable for synchronizing your circuits to the clock signal. What you need is for the device to store the input at the instant when the clock signal goes from low to high. (Using the tech terminology: the 4-NAND D-latch has level triggering, while for the clock you want edge triggering.)
SpinalHDL has been very nice to use so far instead of Verilog: https://spinalhdl.github.io/SpinalDoc-RTD/master/index.html It even has simulation and formal verification workflows built in. In the simulation you can wiggle the bits on your ports using Scala, so you can code an emulation of any peripheral or the like that you want and have your design use it. (You can also do the same thing using C++ if you use Verilator directly instead.)
You can code a bridge between a serial port in your simulated design and a TCP port, and then write a second program to bridge a real serial port to a TCP port the same way. You can then write tools that connect to the TCP port and then use those same tools against both your simulated design and your design in hardware.
Edit: By cheap I mean something like in the article or a bit more expensive, for sure < $1000. By CPU I mean something like an M1. By GPU I mean something like an Nvidia 2080Ti.
In retrocomputing for example they're useful for building accelerators that bolt much faster CPU's of a different model (like 68060 accelerators for an Amiga) or implementing a new graphics card with HD resolution, HDMI, and SDRAM controllers like the ZZ9000 https://shop.mntre.com/products/zz9000-for-amiga-preorder
Less than a microsecond of latency without even trying and 100% predictability.
If needed, you can place more processes on the chip, all running in parallel with zero interference between them.
FPGAs give you the ability to process a lot of data in parallel, with very low latency. As soon as you make use of this, a CPU has a very hard time to compete. If you have e.g. radio data from an SDR, you might want to process it on an FPGA → if you look at various SDR modules, you'll find an onboard FPGA to do exactly that. If you want to process video signals with low latency and with low power consumption, again, you might want to do it on an FPGA → look at video interface cards from BlackMagic Design and most of them will have an FPGA on board. If you have complex mathematical models to emulate e.g. some vintage analog audio hardware, which would create some serious CPU load on a PC, you might want to do it on an FPGA instead → this is what some companies like e.g. UAD do with FPGA based accelerators. If you have high-speed interfaces, like e.g. in a network router/switch, you might want to implement the packet processing in an FPGA → this is what some network equipment does, unless it uses an ASIC. There are countless applications for FPGAs.
Sometimes FPGAs are also the prototype "playground" where you can validate your design before you build an ASIC.
And finally, many FPGAs aren't really that expensive, at least if you ignore the crazy price increase from 2020+. The FPGA alone from the Basys-3 development board from the article (an XC7A35T) costs something around $20-25, at least if you don't need the exact same package. Of course that's only the FPGA and you still might want to add some external configuration flash, RAM and connectivity, but that's still very cheap. To give you an idea, a 256MB DDR3 RAM chip costs maybe $3-5. If you want to connect this to your computer, you could use PCIe, which is directly supported by these Artix-7 FPGAs. Of course there are much bigger and more expensive FPGAs available, but I think you can imagine already that you can do a lot with a $1000 hardware budget.
https://www.digikey.de/de/products/detail/lattice-semiconduc...
This price is ridiculous. I also know you can get 640 Lists for slightly more but I can't find the link.
Since they are just general digital logic machines they can be obviously implement a cpu and are great are parallel computation, but in most cases cpus and gpus are much better at general computation because they’re specifically designed for that task.
See also https://electronics.stackexchange.com/questions/69022/rtl-vs... .
Now that said, I need to pick up the RISC-V softcore introductory course on edx and finishes it.
Edit: Link to Ben's GPU project: https://eater.net/vga
[0] https://digilent.com/reference/learn/programmable-logic/tuto... specifically
[1] https://digilent.com/reference/programmable-logic/arty/refer...
Personally, I do not see point in implementing a softcore Cpu unless design absolutely requires it.
Separately, FPGA makers has shifted focus a bit from facilitating hardware design to providing coprocessors/accel cards.
Thanks a lot.
[1] https://arxiv.org/abs/1704.04760
Here's an example of such shader in D3D11: https://github.com/Const-me/Whisper/blob/master/ComputeShade...
One issue in this area is that the underlying hard logic is limited, and differs from vendor to vendor and product to product. So a nice free IP core might exist for something, say for FFT, but the complete design cannot fit on your chip's LUTs, or it may not have enough buffer blocks or clock multipliers or buses or something to run a given core, or everything "fits" but it has to be run slower to meet timing constraints.
That's not to say it's always like that.
There is probably a need though for various implementations of some algorithms, with different topologies, or for niche environments etc. Just as in embedded development we need various implementations of FFT in floating point, fixed point, using static memory, etc.
And it's not algorithmic really, but peripheral drivers contribute a ton to the community. Being able to e.g. plug a certain e-ink display into a widget without having to write the driver yourself.
opencores.org might be of interest to you (github.com/klyone/opencores-ip).