A "warp" is analogous to a hardware thread and you'd have up to 64 of those being scheduled on each SMX or SMM. Each of those SMX/SMMs has four warp schedulers which issue instructions to execution units. In an SMX the schedulers can issue to any of the 192 execution lanes but in an SMM each scheduler has it's own set of execution lanes. If we call a core anything that can independently issue instructions then I guess you'd call an SMX a core but on a SMM each warp scheduler looks like it's own core. But this is all further complicated by the fact that an instruction issued to one lane can be crossed over to a lane that's become idle due to predication. Which is maybe sort of like scheduling but not really.
But yes, you can't compare "CUDA cores" to actual cores and GPUs aren't equivalent to thousands of cores. The GM204 would have 64 core equivalents and most other chips would have less.
A WARP is really nothing more than a way to have work for SMXs (and computational units it controls) at as many clock cycles as possible. You need some way for masking FPU pipeline and memory latency.
> All warps are running in parallel (otherwise you won't get the performance numbers) and each has its own control path (actually each has its own code)
It's not that different from x86 hyperthreading, just with more hardware threads. Pipelined execution units are fed each clock cycle by the core. Multiple FP operations are in flight in parallel, otherwise CPUs won't get the performance numbers either.
https://www.nvidia.com/content/PDF/kepler/NVIDIA-Kepler-GK11...
https://www.nvidia.com/content/tesla/pdf/nvidia-tesla-k40-20...