This implies needing to pack all of that compute as closely together physically as possible to minimize latency due to speed-of-light limits, which is substantial in massive compute clusters. It takes a long time for light to get from one side of the cluster to the other from the perspective of a computer.
This, in turn, implies needing sophisticated custom cooling since the power density is very high due to packing so much silicon into such a small footprint.
Another issue is that systems this large are at high risk of producing erroneous results due to the quantity of silicon involved creating a large attack surface for bit flips. It is very expensive to run a large supercomputing job only to find out that the output is incorrect and you'll need to start over, so aggressive mitigation, detection, and correction of sporadic data and compute errors is important.
You could run all of these codes on a stack of beige boxes. It would take so long for the computation to complete that it would be extraordinarily expensive in both time and money. These systems focus on codes that would be effectively computationally intractable if you tried to run them in, say, AWS.