The places he suggests you can get a win seem sensible: 1) cases where the cost of dispatching instructions and handling intermediate results dominates, in which case a CISC-ish specialized instruction implemented in silicon may be a win over stringing together simpler operations; and 2) cases where you can get extra parallelization in hardware that isn't available through general-purpose instructions (e.g. doesn't map on nicely to SSE-style instructions).
In fact, the concept of "doing it in hardware" (be it specialized logic or dedicated generic processors) is alive and well (and bearing fruit) on every mainframe manufactured.
The fact we all use similar x86 boxes designed to be compatible with MS-DOS is a tragedy.
The PC standard set us back at least a decade, most probably two.
Partly thanks to the MS OS/2 2.0 fiasco, also resulting it taking ten years after Intel released the 80386 before 32-bit programming became popular. Needless to say, the x64 transition went much better.
I'd prefer a clean break.
And I totally agree - I would also prefer a clean break.
But the Itanium that intel moved on to? That isn't a clean break. It's a mistake. I'm so happy that AMD was able to force them to make something useful instead.
Most of the kludginess in every modern x86 computer is dictated by the need to emulate parts of an IBM 5150.
http://news.ycombinator.com/item?id=3441885
Would Cutler or Letwin consider this acceptable?
Fwiw, Yosef K.'s own follow-up to his "HLL CPU challenge" did acknowledge several proposals he received as plausible candidates: http://www.yosefk.com/blog/high-level-cpu-follow-up.html
IMO, as is often the case, the answer lies in the middle. Look at the tremendous impact that adding AES acceleration features to x86 processors has on applications that require encryption.
What impact did you observe?