EPIC also had numerous problems. The most severe I think was that it exposed the innermost workings of the chip.
That sounds like an awesome idea, but it's actually a major issue because once you expose something you freeze it in time forever.
Pretty much all modern processors larger than in-order embedded cores are basically virtual machines implemented in hardware. The actual execution units are behind a sophisticated instruction decoder that schedules operations to achieve maximum instruction level parallelism and balance many other concerns including heat and power use in modern designs.
The presence of this translation layer frees the innermost core of the chip to evolve with almost total freedom. Even if fundamental innovations were discovered like practical trinary logic or quantum acceleration of some kind, this could safely be kept behind the instruction decoder.
EPIC on the other hand by exposing the core freezes it. I predict that if EPIC would have taken over eventually you'd have... drum roll... an instruction decoder for each EPIC "lane" or whatever that did exactly what today's instruction decoders do. It probably would have ended up evolving into a synchronous multithreaded vector processor with cores that look not unlike today's cores complete with pipelines and instruction schedulers and all the rest of that stuff.
I can imagine one scenario where EPIC and other designs that show their guts could work, but it would require the cooperation of operating system vendors. (Stop laughing!)
OSes could implement the instruction decoder layer in software, transpiling binaries from a standard bitcode like WASM, JVM bytecode, LLVM intermediate code, and/or even pre-existing instruction sets like X86 and ARM to the underlying instruction set of the processor core. Each new processor core would require what amounts to a driver that would look not unlike an LLVM code generator.
Have fun getting OS vendors to do that. Another major problem would be that CPU vendors would be incentivized to distribute these things as opaque blobs, making it very hard for open source OSes to support new chip versions. It would be a bit like the ARM / mobile phone binary blob hell, which is one of the factors making it hard to ship open source phone OSes or make open phones.
Keeping the instruction decoder on the silicon basically just avoids this whole shitshow. It lets the CPU vendors keep their innovations closed as they wish without imposing that closed-ness on the OS or apps.
The final issue with kernel compilation is that the performance probably wouldn't be much better than what we get now. We'd trade the overhead of an instruction decoder in silicon for a lot of JIT or AOT compilation and caching in the OS kernel. The performance and power use hit might be just as large or larger.