> why must it be so complicated to execute some simple instructions on a GPU
Because they're not simple instructions. Here's what your 2 + 2 compiles to on an AMD GPU: https://hlsl.godbolt.org/z/qfz1TjPGa
In order to fill a super wide machine, we need to give it enough work to fill. That means we need to schedule a large number of jobs, and we those jobs to have a hierarchy. The hierarchy for compute work is:
Dispatch -> Workgroup (aka Thread Group) -> Subgroup (aka Wave, Wavefront, Warp, SIMD Group) -> Thread (aka Wave Lane, Lane, SIMD Lane)
Much like any weird device like a DSP, there's a lot of power here, but you need to program the hardware a certain way, or you'll get screwed. If you want to be immersed in what this looks like at a high level, go read this post for some of the flavor of the problems we deal with: https://tangentvector.wordpress.com/2013/04/12/a-digression-...
GPU parallelism isn't magic, and a magic compiler can get us close, but not all the way.