The IO and RAM for that many cores would be a bottleneck. If you manage to keep everything local to a core or a group of cores then you'll severely constrain the kind of programs you can run.
There have been many dead ends in that space, though. Thinking Machines and the Cell processor come to mind.
(My layman’s feeling always was that the GA model of a large network of small and stupid CPUs was underappreciated and hampered by GA’s merciless pricing as a possible FPGA replacement, given how painfully bad the latter are at utilizing the capabilities of modern IC technology—single layer, really?—but I’m not sure how true that is either.)
It's not clear to me actually. You can run cache coherence protocols over a network, they'll just be slow. (But not any slower than message passing would be anyway.)