To be more precise, GPUs are optimized for embarrassingly parallel workloads. AMD's GCN, for example, has scalar and vector instructions, where the vector instructions are 64 items wide. For graphics, this is used for running a shader on up to 64 items (vertices or pixels) simultaneously. Furthermore, the individual compute units are ridiculously hyper-threaded.
The advantage is that most of the silicon can go towards the actual computation rather than stuff like branch-prediction and out-of-order execution. The disadvantage is that branching and looping is problematic: when only one item wants to go down the other branch of an if-else-statement, the GPU has to run through both branches for all items (and execution is masked off on a per-item basis).
This works extremely well for graphics and high-dimensional numerics workloads (linear algebra, finite elements). It doesn't work at all for, say, spell-checking.