Sure, old apps might not then have that flag, but old apps aren't using avx512 either.
As mentioned elsewhere, I'm not sure this is hardware design being forced to adapt due to software implementation difficulties, so much as avx512 as a whole may not be worth the hardware area going forward. Even on the larger cores.
Note that even simple things like memcpy, which every single program out there will use, will use AVX512!
But really I think the entire ecore/pcore split for avx512 is academic, as I'm not sure tying features and flexibility to wider registers makes sense even on pcores, as I'm not sure the hardware area cost is worth the benefit. I honestly wouldn't be surprised if newer architectures don't have the wider registers options on even their larger cores.
You're chasing a pretty small market IMHO of people who have datasets large enough to benefit from large registers and wider alus, but not so large it's worth it to pass it over to an even more specialized accelerator.
The benchmarks that tend to show 512-wide simd benefits may get even bigger benefits from running on a GPU.
I'd think you could do some pretty good heuristics, like if the thread hit P-core instructions in the last time slice, don't schedule it on an E core for the next one. When a time slice starts on a P-core, leave the specialty instructions disabled, so you can monitor usage --- if you trap, enable it, and return; if that cost is still too high (which it might be), keep track of how many time slices in a row hit the trap, and maybe enable the P-core instruction preemptively for a few slices.
Or just, give the program more information and let it decide. If you've got some cores with avx and some without, maybe the process wants to schedule only on the avx cores. Or maybe it can schedule some threads on any core and others need an avx core. As long as the possible permutations at run time aren't too crazy, it's reasonable.
Why not just determine the right cpu to run on by examining the arch of the binary? Waiting for an instruction failure seems ridiculous.
If the arch is incorrect, it is a bug, and program will crash on illegal instruction. Ie like if you had an ARM binary that was incorrectly set to x64 and was run on x64.
What if the JIT generates new instructions dynamically and initially there are no AVX512 instructions but later on in the process' lifetime there are?
For this reason, most binaries built with modern Visual C++ are technically using AVX instructions even when compiled for SSE only. It doesn’t mean these binaries gonna fail without AVX support, it only means they’re capable of using AVX when available.