1) A GPU "core" is loosely defined. Those "512 CUDA cores" are 16 streaming multiprocessors with 32-wide SIMT.
2) Big parts of GF100/110's SMs are double-clocked; most of the rest of the chip runs at 750MHz (for the clock rates in your example).
3) 5.5 Gigabits per second per data pin (faster than GF100/110). 256 data pins (fewer). 176 Gigabytes per second per chip (close).
4) OpenCL. There are and have been others, but OpenCL is probably ATI's bet for general purpose programming. You could say it's inferior to CUDA (and I'd agree), but to act like it doesn't exist cheapens the debate.
This does speak to the immense power of marketing though; it's easy to lay down such a thicket of buzzwords that you can spin a product any way you want.