Depends on the software being used but it's probably OpenMPI (or a variant). However, OpenMP is also used especially in a hybrid mode where on a node OpenMP is used for shared memory parallelism and OpenMPI is used for inter-node parallelism.
The parallel FS stuff is mainly to handle large number of nodes streaming data in and out. E.g. a few thousand nodes all reading from data sets or saving checkpoint data.
One big difference you'll see in HPC is large scale, fine grained parallelism. E.g. a lot of nodes simulating some process where you need to resync all the nodes and exchange data between them at each time step. Also checkpointing, i.e. since simulations may take weeks to run, most apps support saving application state to disk periodically so that if something crashes, you'll only lose a few hours or a day of computation. The checkpointing also causes a bunch of FS io since you need to save the application state from all the nodes to storage periodically so you'll see really high io spikes when that happens.