That just sounds insane to me that the tools can't handle synthesis of straightforward stuff like this. My personal fantasy startup is a next generation, mostly open-source EDA simulation, synthesis/optimization and place & route toolchain. You can expose the analog cell design to the world, but generate revenue from the per-fab process-specific partnering required to work the customers through the process of making real hardware. Who's got funding for me?
[1] This has probably been thought of and dismissed a million times over by actual experts. Say, losing whatever advantages it may have due to extra micro-ops needed to move stuff back and forth to the correct register file fraction?
Have I got the paper for you! From my adviser's previous student: https://dspace.mit.edu/handle/1721.1/34012
But you're right -- the win isn't obvious. There's pain and advantages in both directions. I've always wanted to build it, but I only have so much time in the day. :(
Where what I want to see (or do, at least in my dreams) is closer to the analog end: design of parametrized cells, GPU-parallel large circuit analog simulation, circuit-dependent optimization (i.e. choose "speed" if the cell is on a critical path, "area" or "power" if not, optimally choose drive strength and buffers based on latency needs of the path, swap flip flop implementations likewise, etc...) place and route, eventually mask file generation, reverse synthesis and design rule checking (though obviously that part becomes fab-specific).
None of this stuff is "hard" in a fundamental way, but it's been hidden behind bad tooling for so long that no one seems to have tried to innovate much in the past few decades. Anyway, I got ideas aplenty.
Maybe these should be expressed as Instructions Per Second (at peak and at the point of diminishing returns?) or something like that, rather than two independent numbers. Higher clock frequency actually seems like a bad thing, all else being equal. It seems to me that throughput ought to trend toward infinity, clock frequency toward zero. ;- )
This is particularly useful when talking about a processor design, and not a specific processor in particular. As you said, there's a lot of good about slower clock frequencies, so you'll see the same ARM design being deployed at a variety of frequencies. Far easier to talk separately about a design's IPC from its achievable clock frequency (although both are important!).
Really hoping to get my hands on RISC-V-based SBCs next year.
Now, Cavium's Vulcan-based ThunderX2 seems to be beating Intel's Skylake server chips:
https://www.nextplatform.com/2017/11/27/cavium-truly-contend...
Just saying that in the context of performance/cycle/core, they're closer to a bunch of Atoms than a Xeon (which, again, is totally fine for some use cases).