On Xilinx, for example: a 64-bit register file doesn't map efficiently to Xilinx's RAMB36 primitives. You'd need 2 RAMB36 primitives to provide a 64-bit wide memory with 1 write port and 2 read ports, each addressed separately. Only 6% (32 of 512) entries in each RAMB36 are ever addressable. It's this inefficient because ports, not memory cells, are the contented resource and BRAMs geometries aren't that elastic.
A 64-bit register file in distributed RAM, conversely, is a something like an array of DPRAM32 primitives (see, for example, UG474). Each register would still be stored multiple times to provide additional ports, but depending on the fabric, there's less (or no) unaddressed storage cells.
The Minimax RISC-V CPU (https://github.com/gsmecher/minimax; advertisement warning: my project) is what you get if you chase efficient mapping of FPGA memory primitives (both register-file and RAM) to a logical conclusion. Whether this is actually worth hyper-optimizing really depends on the application. Usually, it's not.
It gets a little more unreliable when you start accessing it in more complex ways though.
From my understanding (I'm no FPGA expert), the code in the article will infer a BRAM with two read ports and one write port. That may be fine.
I actually battled with this recently on a project. I found that the tool was not inferring a block RAM when I expected it to, so I had to modify the Verilog to gate the reads and writes so that only one could happen at a time. That wasn't an issue in my case though.
My takeaway from the exercise was that it's sort of the equivalent of relying on the optimizer of a compiler to recognize the programming pattern and do the right thing. After talking to one of the FPGA guys I work with, he seemed to feel that it's better to just instantiate a vendor IP BRAM directly. The downside is portability though.