Wasn't the kludge in some Apollos that they had two 68000s arranged so they did exactly the same thing, but with one delayed? If the leading one hit a page fault, they would generate an NMI for the trailing one before it too hit the page fault? (I'm sure I heard of someone doing that if it wasn't Apollo).
Even after the 68K family got working page fault handling there was an important different between it and that on Intel processors. The Intel processors used "instruction restart". If the page fault happened somewhere in the middle or at the end of an instruction it discarded any work it had done to that point. After the fault is resolved, the instruction would restart from the beginning.
The 68K family used "instruction continuation". When it got a page fault it would include in the exception stack frame enough internal processor state so that when it returned from the exception processing it could continue from where it left off.
Continuation is presumably more efficient than restart because you aren't discarding any work--although with continuation you have to write a bigger stack frame (and later read a bigger stack frame) because of that internal state as opposed to restart which can use the same small stack frame that ordinary interrupts and exceptions use so I'm not sure that you actually come out ahead with continuation. We aren't talking a VAX here with instructions like "evaluate polynomial" that might do a lot of work before faulting.
We had a hard t track down problem with instruction continuation. I was working at a small 68K workstation company (Callan Data Systems) which ran a swapping version of Unix. I was hacking out the process and memory handling to replace with demand paged virtual memory code, and it was going quite well--except that occasionally when I would ^C a process the damn thing would hang hard.
What was happening was that sometimes a process would page fault, the kernel would handle it, and while that was going on the kernel would let some other user process run. By the time it finished getting the faulted page, and it was time to resume the first process (the one that had page faulted) I had hit ^C on that process and so there was a SIGINT to deliver.
The way it delivered signals to user process signal handlers was by diddling the user stack frame so that it looked like the user process itself had called the signal handler just before whatever interrupt or exception had caused the process to enter kernel mode. Then when the kernel returns to user space, it returns to the signal handler.
That turned out to not be a good thing if the exception stack frame was a continuation stack frame. The processor was very much not happy to try to continue an instruction that was not the instruction whose internal state was in the stack frame.
The fix: when the kernel has a signal to deliver, first check if the stack frame is a page fault frame. If it is, instead of delivering the signal right away set the trace flag and return to user mode. That returns to the interrupted instruction, continues it, and when it finishes generates a trace interrupt, which has a normal stack frame. The trace interrupt handler can then clear the trace flag, and just go ahead and do the normal return to user mode processing which will have no trouble delivering the signal now that we are only dealing with a normal stack frame.