Nyuzi: An experimental FPGA multicore GPGPU processor
github.com
github.com
They have a RISC V implementation in it so it can't be too bad
Still a lot of tooling left but my FPGA tinkering became much more enjoyable when I stumbled across MyHDL.
Every time I see imperative language X adapted to RTL/VHDL/Verilog I want to slap someone.
They aren't the same, gates are a fundamentally different primitives. You shouldn't be using a language built around serial actions for something that's inherently parallel. You bring all sorts of baggage along that you don't need.
[edit] They don't even specify if their state machine generated is Mealy vs Moore. This is the stuff that you want control over and not abstracted away by some language that you don't know how it will synthesize.
Chisel and Clash are not any different.
Verilog may have for loops but they're largely used for simulation, they're expensive in terms of synthesized hardware(mostly toolchain/process dependent on how they synthesize) and can only be fixed size.
I've got no issue with Chisel, it's a DSL which I think is a good fit. What I'm arguing against is constructs that don't have a clear mapping to hardware representations which means you have to guess at what the compiler generates(see my Mealy vs Moore comment above).
See also: a MyHDL UART: https://github.com/andrecp/myhdl_simple_uart/blob/master/ser...
You're going to spend a ton of docs explaining to the user what parts of the language you can't use rather than having a DSL spec that's clear in what's supported.
With a proper DSL you also don't have to massage language features that don't quite map(say enums for state machines) into a format that it's not meant for.
Since you don't seem interested in addressing just one issue(I'm sure I could come up with more) around state machines I don't see any reason in continuing this discussion.
Scala is not any better. It's exactly the same kind of a eDSL - no macros, nothing, just generating objects in runtime and then serialising this tree into a Verilog code. Nothing fancy.
> around state machines
What state machines?!? It does not have any more features for defining FSMs than an underlying Verilog.
114,480 logic elements (LEs)
3,888 Embedded memory (Kbits)
266 Embedded 18 x 18 multipliers
4 General-purpose PLLs
528 User I/Os
[1]: http://www.terasic.com.tw/cgi-bin/page/archive.pl?Language=E...1) OpenCL/CUDA have an OpenGL-inspired syntax with a steep learning curve and limited generalizability
2) FPGAs don't seem to be gaining the economies of scale of GPUs
I simply want to be able to emulate thousands of CPUs (millions of gates) for physics, AI, big data etc, in a way that's accessible, affordable and won't catch fire. I'm thinking MATLAB or Octave but with near-ideal speedup for embarrassingly parallel problems.
Julia fits your last sentence.
Still crazy awesome, that's a ton of work.
They are (or were) spending a huge amount of instructions on stuff that'd have dedicated hardware/instruction set support on a proper GPU. Normally rasterization and texture sampling runs in dedicated hardware, colour packing/unpacking is integrated into the memory access instructions (at least on Radeon), etc. Stuff that'd be one or two instructions on a commercial GPU instead took dozens or hundreds.
Yeah. One area I'd like to investigate is adding specialized instructions. The existing renderer is not highly optimized.