That's actually a common optimization, e.g, Itanium doesn't have divide; just reciprocal.
I haven't spent time doing assembly optimization for a very long time, but it used to be the case that, even in x86, you were often better off using various tricks to avoid having to divide. I'd be interested in hearing if that's still the case, from someone who does that sort of thing today.
The actual design was implemented in a Xilinx Spartan-3E 1600 development board.
I love that hardware has gotten so cheap. When I was a teenager, I implemented a Sega on an FPGA, and it took Virtex II board (very high end, at the time) to handle everything. Now, an entire supercomputer fits on one of Xilinx's Spartan (i.e., budget) boards.