The places he suggests you can get a win seem sensible: 1) cases where the cost of dispatching instructions and handling intermediate results dominates, in which case a CISC-ish specialized instruction implemented in silicon may be a win over stringing together simpler operations; and 2) cases where you can get extra parallelization in hardware that isn't available through general-purpose instructions (e.g. doesn't map on nicely to SSE-style instructions).