It's very absurd to break it out for each of them modules to get an independent PCI-e slot. This could raise the layout of the board significantly.
Yet the speed they currently offer (PCI-e x1 Gen2) isn't fast enough for RDMA to effectively "chain" the compute modules.
So I would see that they would have some sacrifice (e.g. fallback to the smaller and more cliche mini-PCIe like most recent fresh IOT boards out there), or even remove PCI-e expansion but offer something else (e.g. embedded SATA host/multi-NICs where only the master node could control, then the rest of the children will have to rely on RDMA, despite it will be slow and painful).
It's not quite economical for Turing Pi to implement NUMA-like architecture so I would rule this out.
From my perspective its 4*7 = 28 weak ARM cores vs 8 strong x86 cores. You can see that ARM actually had more cores, giving it an advantage plus to parallelized workload compared to x86.
Hell, maybe we can mix them in a bunch so that x86 runs powerful applications like GitLab, Prometheus and Postgres while ARM runs massively parallel workload that GPUs can't handle: Function as a Service (in AWS terms, Lambda; in CNCF's term, OpenFaaS), Linkerd handler (service mesh needs some kind of scheduling though), microservice replicas.
In the end CPU are all going to have a designated purposes, despite it should have had been "general purpose".