I would say that classical simd with explicit vector registers is plain parallelism, not concurrency. You can build concurrent abstractions on top of it (ISPC, any of the high level GPU languages), but when you move beyond simple simd hardware and have multiple hardware threads executing simd groups on multiple cpus, possibly with the ability to migrate threads between groups, the strong synchronization is a bit lost. But I'll admit I'm not an expert.