I wonder how the logic worked in the previous version without early start. Was it relying upon the address calculation speed to settle the outputs really quickly? Was it inserting or stretching cycles?
The memory pipeline just starts one cycle later than now. Effective address is calculated during the first cycle of the instruction. The microcode then waits for it to finish with the DLY (delay) micro-op, which releases one cycle later.
So cycle insertion. I presume that the DLY was synthetic, and was not explicitly added to the microcode ROM.
Would this be related to Next Address (NA#) pin on the 386 enabling Address Pipelining?
This is great. So proper 386 on an fpga? How cool is that.