RV actually allows for that. But not dictating that registers get pushed to the stack, flexibility in how you manage them opens up. So for RV, and some other architectures, you have to mark the function an IRQ and the compiler will know how to figure that out.
Another gotcha is that for AXI and other burst interfaces, the hardware being able to say "I'm going to send you X words" is dramatically better for latency than each one being a single transaction. So if your stack is in a location that requires multi-cycle memory access times, this balloons in timing cost.
Sadly this is a very hard topic to condense into a few sentences. Maybe if I wrote an article on it with graphics it would help. Unsure
Yup.. If author's working with a RISC-V softcore, there are several options to reduce interrupt latency:
# If softcore saves many registers as a hardware feature: adapt it to not do so
# Save (and use) only a few registers in interrupt handler
# Reserve some registers to be nuked by the interrupt handler, and don't use those in application code
# Use another softcore with faster interrupt response
# Move to a part with cpu as hard silicon (with sane interrupt handling)
Just to name a few (or some combo thereof).