I think that one shouldn't write Verilog/VHDL code unless one clearly understands how it will be transformed into gates.
I think that one shouldn't write Verilog/VHDL code unless one clearly understands how it will be transformed into gates.
[0]: https://danielmangum.com/posts/a-three-year-bet-on-chip-desi...
Why is it necessary to understand how verilog is exactly transformed into gates? Sounds like gatekeeping to me, but please enlighten me. I have in the past written hdl and I didn’t even have a FPGA
I honestly believe that if you do not understand at least some basics, things will be way more painful and opaque than necessary, so it is a no-brainer IMO.
You also do not need to have an FPGA to learn this stuff.
Also for a design to be fast you need to use pipelining and to slice your combinational logic into thin layers. Again, you won't be able to do it without understanding how your code translates to gates.
I think you are arguing for the right cause but with wrong arguments. In C language, not understanding in details what exactly the code does, can be just as disastrous as in Verilog. Internet examples abound.
Trying hard to avoid this because it is inefficient, is (imho) a typical case of premature optimization. Who cares whether you're wasting CPU cycles (in C) or FPGA logic (Verilog & co) when you're just learning what's what?
But there's a better reason: to make sure that whoever is learning Verilog (or other HDL) actually understands what they are trying to do.
With that understanding in place, the learning becomes "what logic construct(s) would be suitable" followed by "what Verilog to write that describes that logic".
Versus: "Verilog intro course has this example, does it compile? And does it appear to do what I think the Verilog says it should do?".
In programmer's parlance: semantics vs. syntax. The algorithm + how it maps to a CPU's resources, vs. red tape required to implement it.
If you 'feel' the syntax but underlying semantics is black magic, then you could keep stumbling in the dark forever. Spitting out copypasta along the way.
But if you understand basic concepts of the underlying hw (and how you intend to use that), learning syntax to describe that is straightforward.
Chip designers intuitively visualize interconnected blocks of logic. They draw rectangles with wires between them.
To learn this intuition, writing spaghetti Verilog code that runs in a simulator is counter-productive. It's a common mistake that people coming from a software background make all the time, and that has been discussed to death. So with some appeal to authority, trust us and learn proper hardware design, hopefully it will be more rewarding and will "click" faster.
It feels condescending to read on the internet how hardware engineers clain that software engineers don't know how hardware works. Any kid that has played around with Redstone intuitively knows the difference and in fact children are more likely to have extensive experience with hardware than with software because most games have some sort of logic gate system but only a tiny minority lets you program unless the entire game is based around programming.
It gets a little more unreliable when you start accessing it in more complex ways though.
From my understanding (I'm no FPGA expert), the code in the article will infer a BRAM with two read ports and one write port. That may be fine.
I actually battled with this recently on a project. I found that the tool was not inferring a block RAM when I expected it to, so I had to modify the Verilog to gate the reads and writes so that only one could happen at a time. That wasn't an issue in my case though.
My takeaway from the exercise was that it's sort of the equivalent of relying on the optimizer of a compiler to recognize the programming pattern and do the right thing. After talking to one of the FPGA guys I work with, he seemed to feel that it's better to just instantiate a vendor IP BRAM directly. The downside is portability though.
On Xilinx, for example: a 64-bit register file doesn't map efficiently to Xilinx's RAMB36 primitives. You'd need 2 RAMB36 primitives to provide a 64-bit wide memory with 1 write port and 2 read ports, each addressed separately. Only 6% (32 of 512) entries in each RAMB36 are ever addressable. It's this inefficient because ports, not memory cells, are the contented resource and BRAMs geometries aren't that elastic.
A 64-bit register file in distributed RAM, conversely, is a something like an array of DPRAM32 primitives (see, for example, UG474). Each register would still be stored multiple times to provide additional ports, but depending on the fabric, there's less (or no) unaddressed storage cells.
The Minimax RISC-V CPU (https://github.com/gsmecher/minimax; advertisement warning: my project) is what you get if you chase efficient mapping of FPGA memory primitives (both register-file and RAM) to a logical conclusion. Whether this is actually worth hyper-optimizing really depends on the application. Usually, it's not.
1. Go to https://edaplayground.com
2. Paste the code in the right pane labeled "design.sv"; it's probably a good idea to reduce the number of registers to 2 to make the output more readable
3. On the left, select "Yosys" in tools and simulators (it seems to be the only one that works without additional configuration or fiddling) and enable "Show diagram after run"
4. Click "Run" and it should open an image with a graph rendering showing the circuit. If it doesn't work, try to disable adblockers or popup blockers (it opens the output as a pop-up apparently)
You can also presumably run yosys locally.
In this case, it implements every register as a D flip-flop; the rs outputs are implemented with muxes from the flip flops, and the D flip-flop value is set to either the current value of the register or to data depending whether write_ctrl is set and rd is equal to the register number.
E.g. in this physical layout of a Swerv RISC-V CPU, the ARF section is the register file, built out of flip-flops: https://tomverbeure.github.io/2019/03/13/SweRV.html#swerv-ph....
assign rv1 = registers[rs1];
assign rv2 = registers[rs2];