> You also need to increase voltage if you want to increase clock speed.
Or use lower VT cells (they turn "on" quicker), at the expense of increased leakage power. But at these geometries, faster clock speeds is getting less feasible and you need to find increased performance in other ways.
No logic signal needs to cross the entire die in one clock cycle, there is always an alternate design. For that reason, only registers 'talking' to each other need to see a clock at the same time, and even then there is a window. Clock routing is a consideration that takes resources but it's not a problem. Realistically, a logical signal won't be going anywhere remotely near 50mm at 3GHz in 10nm, so the clock doesn't need to either.
Max die size is also limited by the vendor's tooling i.e. what their machines can literally handle. And also physical issues such as warpage. If you make a massive die and it heats up in a non-uniform manner (different bits of it get hot at different times), it expands in a non-uniform manner. This can lead to all kinds of problems.
Chips stuffed full of memory will yield better than a logic-heavy chip since large SRAMs always now include redundancy. So this too has an impact on how big you can go for a given cost. You can however get registers that are built of multiple storage elements, the output value of which is the consensus. Don't know how much these get used.