I don’t think it’s that simple. This design is so good because Apple designed a system, and didn’t throw together a few parts they could buy.
They should revisit all choices they made earlier. Users might want more generic ML hardware, so that they can develop the algorithms that can get burned into specific hardware on some future consumer-level hardware, a GPU that can ray trace better, more memory than you can fit on a SoC, etc.
That last issue is particularly problematic. They might have to give up the idea of having (only) unified memory.
If so, the entire design could change. In some cases, that could mean task-specific hardware isn’t (much) faster anymore than doing it in the CPU, in which case discarding it to make room for extra cores, cache, etc. might be the better choice.