We’ve been promised these magical ultra parallel machines for a decade now and either they’re much harder to build than the alarmists say, or there's simply no market.
We’ve been promised these magical ultra parallel machines for a decade now and either they’re much harder to build than the alarmists say, or there's simply no market.
What we are seeing is a hierarchy of processing cores. It's already visible even in the desktop market. Intel has their efficiency cores. Apple has the power and efficiency cores. And the GPU on top. So we already have 3 different stages here, at the top the least amount of cores but the best single threaded performance going all the way down to, let me check how many threads, almost 10000 threads on a flagship consumer GPU running simultaneously right now.
Intel tried with their Larrabee on what happens if they just toss in ton of traditional low performance CPU cores. It failed to perform. It's really hard to beat the modern SIMT style GPUs when it comes to massively parallel computation.
Now whether or not all 128 threads need to operate on the same chunk of memory or not is somewhat up for debate. It is very good for some applications (Let's Encrypt issues all of the Internet's TLS certificates with one Postgres instance), and irrelevant for others (just run 128 copies of your node app and send requests to them at random).
And on AWS there are instance types with 448 VCPUs [1] and in fact 128 VCPUs is a very common instance type. In fact with the growth of Kubernetes these larger configurations are increasingly popular because it's more cost efficient to pack the containers into fewer/larger nodes.
Also would add that Scala is in my opinion by far the best language for writing safe, highly concurrent code when you pair it with frameworks like ZIO. [2]
[1] https://instances.vantage.sh
[2] https://zio.dev
That kind of is defeating your argument though: If people pack many containers on large boxes, the individual container probably is not making use of the possible parallelization across all cores.
If you have an app that runs 24/7 at 100% bursting across all cores then Kubernetes will simply schedule a single container on that instance.