Program FPGAs with Go
reconfigure.io
reconfigure.io
Especially on the FPGA side, how does this interact with all the features of Go that seem ill-suited to an FPGA implementation? Can I write functions that generate and consume closures? Where is my garbage going on the FPGA side and how is it collected? Or is the FPGA code being written only in a subset of Go?
I understand the idea of wrapping the primitives offered by the FPGA hardware itself into channels, but I'm unclear on how one can sensibly implement a Go runtime on top of that in the FPGA without making it too difficult to understand the cost model of your Go code.
Could you elaborate on why closures are ill-suited to FPGAs? Thanks.
C++ has lots of syntactic sugar, but at the end of the day, like C and Rust it's a fundamentally pass-by-value language.
Here's a good paper which describes some of the complexities and design decisions for a full and complete closure implementation that doesn't require decoration and compiler hinting:
http://www.cs.tufts.edu/~nr/cs257/archive/roberto-ierusalimschy/closures-draft.pdfWe, a stanford lab, are pursuing similar goals but opensource and from a Scala DSL although our doc (http://spatial-lang.readthedocs.io/en/latest/tutorial/starti...) is not that up-to-date:
In Chisel, you write Scala that is compiled to Verilog(?) and you put it into your synthesis toolchain. But Chisel is mostly just that: it's a HDL, and not much more. And if you want to talk to an FPGA, especially from software, you still have to write another pile of glue that does interfacing to your device, over your peripherials, etc.
Spatial gives you more on top of Chisel. Instead, you simply write a single program and say "Accelerate this bit", and it generates both the hardware and software and glues it all together. This means you write a single program once and the compiler generates all the glue for you, so the usage is more seamless.
This is what "SDSoC" from Xilinx does, but for C/C++. You simply write a C program and annotate functions as "Accelerate" and it compiles both the hardware and software for you and generates the interconnections. Spatial is like that, for Scala.
Berkeley grad student here. The current Chisel compiler generates an intermediate representation that can be compiled to any target. Our Verilog backend is definitely the most developed, but we also have an interpreter that can directly simulate the IR.
(To be clear, I figured you compiled to a generic IR before lowering onto some chosen HDL, regardless if it's FIR or not, but wasn't sure if Chisel had multiple HDL backends strictly speaking).
There aren't official releases of this version yet, but you can get snapshots (if you don't mind some API breakage now and then). It works fairly well now and we use it in all our RTL designs.
I've always been wary of inference-based systems that target clocked logic. Most of them (especially the C-based ones) are not suited to produce efficient designs and usually require some heavy massaging of the original code to make the inference system happy.
The designers of Chisel don't speak the same "language" as your average hardware designer. Maybe this is why people have invested tons of hours to re-write risc-v in pure verilog [1] and VHDL [2].
---
The real shame of course is that Connectal is open source, although BSV is not...
https://github.com/interplanetary-robot/Verilog.jl
Wrote it in three days; although it's very young, some of the strengths are emission of human-readable verilog, and the ability to build the verilog into C (using verilator) and doing continuous testing without ever leaving julia.
You get advantages of Go's type-checker, though you're probably limited to a very small subset of the language. Note that you won't be able to use a lot of third-party languages unless their translation is really good or the third-party package code is very simple.
Docs seem to be available here (thanks to another comment in this thread): http://docs.reconfigure.io/welcome.html
I think the approach of mapping a high-level language to a CPU is not necessarily novel, but using Go for it is.
I used a different approach to build a bit-flow graph in Java for a project in the past. Rather than map the whole language to the circuit, I created some APIs that would generate the graph and export it. It looks fairly similar to what you see here.
Looking at some of the examples it seems to me that you'd still need to know hardware programming, memory etc. Now my comment seems very snarky, but I still think that it's a huge achievement to have gotten this far with this and I wish them luck! I just don't get the target user base.
The best languages to take advantage of chips that aren't compute-limited* are things like Erlang, Elixir, Go, MATLAB, R, Julia, Haskell, Scala, Clojure.. I could go on. Most of those are the assembly languages of functional programming and are also not really usable by humans for multicore programming.
I personally vote no confidence on any of this taking off until we have a Javascript-like language for concurrent programming. Go is the closest thing to that now, although Elixir or Clojure are better suited for maximum scalability because they are pure functional languages. I would give MATLAB a close second because it makes dealing with embarrassingly parallel problems embarrassingly easy. Most of the top rated articles on HN lately for AI are embarrassingly parallel or embarrassingly easy when you aren't compute-limited. We just aren't used to thinking in those terms.
* For now lets call compute-limited any chip that can't give you 1000 cores per $100
Err...I respectfully disagree. They're HDLs and more akin to hardware design than any traditional software abstraction, assembly included.
It also may depend on where you learned your craft - my decade as a logic designer seemed to show people who started life as an EE and went straight into logic design coded at a lower level than people who started as programmers (who tended to be more productive as a result) ....
But this is old data - the world changes
The exception to this are those with good hardware sympathy. They can after carry over that detail oriented thought process.
FPGAs may or may not deserve that distinction, depending on your point of view.. but even if you concede that, they're still heavily bandwidth limited. An SDRAM interface takes up quite a bit of floor space, especially if you want more than one FIFO to move data with -- and even then, you're still standing behind relatively slow memory interface. There's SRAM on newer chips, but it's still too paltry an amount to really do anything close to general purpose computing on.. especially with the languages you've mentioned.
> why Moore's law no longer works
Moore's law no longer works because we hit 4GHz in silicon, there's nowhere to go but sidways now, and that's true whether you're in dedicated or reconfigurable chips.
Not sure why they wouldn't use it instead.
Although it is technically a purely functional language, you can almost mutate variables (in reality it is creating a new immutable variable with the same name, which takes precedence)
a = 1
# a == 1
a = 2
# a == 2
Concurrency feels very natural: # concurrent
numbers = [1,2,3,4,5]
doubles =
|> numbers
|> Enum.map(fn(n) -> Task.async(fn -> n * 2 end) end)
|> Enum.map(&Task.await/1)
# doubles == [2,4,6,8,10]
# consecutive
numbers = [1,2,3,4,5]
doubles = numbers |> Enum.map(fn(n) -> n * 2 end)
# doubles == [2,4,6,8,10]This (purity) stirred my interest, but as far as I can see it's incorrect. This[1] Wikipedia page on pure languages does not list Elixir, and the Elixir Wikipedia page itself does not mention purity at all.
Can anyone clarify?
[1] https://en.wikipedia.org/wiki/List_of_programming_languages_...
People who want to use an FPGA should learn VHDL or verilog. There have been a lot of projects to make C compile to VHDL/verilog, and it's generally accepted that it does not work very well.
What is the advantage of using Go for the same purpose?
Both Xilinx and Altera have High Level Synthesis (HLS) tools. These use C or C++. If you know how FPGA work is generally done, you can separate the hype from the reality and you can understand how to use it for a real application.
The vendors have lots of libraries for IP. You don't write RTL from scratch. It would take too long to verify. You tie IP together. It can be DSP or generic maths or a video codec thing. The VHDL is done for you.
You write your algorithm in C++ in a particular format using compatible data types and calling HLS libraries. You run it all in C++ first and make sure it does exactly what you want in SW. This is where the algorithm is developed.
THEN you fire up the HLS tool and a couple of hours of synthesizing later (lol) you get to load a bitstream onto a FPGA to verify it.
Of course there can be problems in that translation. It takes good engineering to dive down into the design and find the issues.
My current work does not touch any HLS. I am doing the VHDL stuff. But I know the algorithms all started from SW first. It always does. For the bulk of the work, verification, it is somewhat irrelevant whether it is manually converted to RTL or done via tools.
Seeing how FPGAs do not operate in a linear way the way that software does on a processor, why are we trying to make them work that way? It would make more sense to me to design a high-level synthesis language with a paradigm that is also not imperative: functional programming. Like, for example, how would this kind of C code even be synthesized in hardware?:
A = 5;
B_out = A + 3;
A = 6;
C_out = A;
"A" is used as two different things, which is totally fine when the code is run sequentially, which must be what is happening when code like this is synthesized, but that's wasteful on an FPGA, because B_out and C_out don't actually have dependence on each other and could be computed concurrently, which is what would happen if we used VHDL to do something similar. We need a high-level synthesis language that describes a system which solves the algorithm we want, the same way VHDL does, except with more abstraction capabilities. In my opinion this could be a functional language.Your example is somewhat pointless. The code is written to create the HW not the other way around. I can't feed it just any crap.
You want parallelism you have to code it.
Zynq would actually be what I use! You start with SW. The ARM core is not that quick. You will use the FPGA to accelerate the tough parts. You may think you will have throughput issues but you have options via the high performance AXI ports. Your FPGA modules access the data in memory via DMAs.
KNOWING what part of the algorithm you need to accelerate actually suits FPGAs, you grab the HLS and start coding.
You have to look at some of the libraries to understand what abstraction level you are working at: https://www.xilinx.com/products/design-tools/vivado/integrat...
Video, matrices, linear algebra, encoders/decoders. Etc. I can string them together in the same way I would string HDL IP.
The advantage is I can run the algorithm in C++ first and test it all, under the assumption that the HLS library has the equivalent HW version for synthesis.
There is still a lot of HW work involved. For instance in your example with A used twice. One module would calculate B_out by reading A prior to changing its value then you would have to start the C_out module. You would need a way to coordinate the two modules to share the same memory at A. But they would be running in parallel, just not started at the same time.
It's early days for us at reconfigure.io, we're just working with a few core early users at the moment and we'll be bringing more examples, benchmarks and increased access over time.
In one of my previous lives doing embedded development, we were able able to program the FPGA using pretty plain looking C on the Nios, which just dedicated a portion of the FPGA's gates to running a simple, ARM-like processor.
It was cool for us software dudes because we could just do general-purpose computing (mostly) on the FPGA, and the verilog folks would wire it up for us to work right. It's not the cheapest way to design a product, but the stuff I worked on had crazy high profit margins, so it was a fair trade-off for better productivity.
Understanding of FPGAs is what I'd consider specialized enough that most people shouldn't be required to demonstrate it in a technical interview.
A field-programmable gate array (FPGA) is a bunch of logic gates on a chip, that can be "wired" together in almost arbitrary ways. So in a sense it's more primitive than a MCU. The "field programmable" part is that the wiring pattern is programmed on a desktop computer and fed into the IC. This allows creating combinatorial or sequential logic functions that execute extremely quickly and often in parallel. In addition to simple logic gates, modern FPGA's offer other kinds of "cells" such as memory registers.
Ironically, people have created wiring patterns that implement a complete microprocessor on an FPGA.
This is the extent of my own knowledge, based on product descriptions but no actual experience doing anything with an FPGA.
[1] http://www.terasic.com.tw/cgi-bin/page/archive.pl?Language=E...
[2] http://store.digilentinc.com/arty-artix-7-fpga-development-b...
To wit, pushing the performance of any FPGA with <insert_favorite_hdl_here> will inevitably result in a high degree of technical debt and/or vendor lock-in, e.g. instantiating device-specific hard IP.
At the end of the day, we--as developers--aren't quite at that point where we can have our cake and eat it too, making this solution yet another product lifecycle trade-off decision.
>> expensive, hard-to-source hardware engineering skills.
If only that were true.Lots of hardware engineers have moved into software.