Flag checking, as opposed to trapping, incurs costs on all operations.
I know nothing about micro-architectural implementations, but as far as I can remember, all instructions that can trap (load/stores, divs, etc) take more than one clock cycle. Simple sub/adds usually have 1 cycle latency; there just might not be enough latency budget to fit the trapping machinery.
I guess that the trapping doesn't actually need happen in the critical path (i.e. the results might be available earlier) but this would mean that even ALU instructions cause speculation.
For exemple while SSE still has an fpu control word, explicit instructions were added for some stuff that historically was handled by the control register (rounding for example).