1) A GPU "core" is loosely defined. Those "512 CUDA cores" are 16 streaming multiprocessors with 32-wide SIMT.
2) Big parts of GF100/110's SMs are double-clocked; most of the rest of the chip runs at 750MHz (for the clock rates in your example).
3) 5.5 Gigabits per second per data pin (faster than GF100/110). 256 data pins (fewer). 176 Gigabytes per second per chip (close).
4) OpenCL. There are and have been others, but OpenCL is probably ATI's bet for general purpose programming. You could say it's inferior to CUDA (and I'd agree), but to act like it doesn't exist cheapens the debate.
This does speak to the immense power of marketing though; it's easy to lay down such a thicket of buzzwords that you can spin a product any way you want.
The memory speed on the 6970 tops out at 5.5Gbps (1375MHz), while the 6950 hums along at 5.0Gbps (1250MHz). That puts memory bandwidth in the 175GB/s and 160GB/s ranges for the Radeon HD 6970 and 6850, respectively
The key numbers there are 175 GB/s and 160 GB/s (not quite as fast as nVidia still, but not 2 orders of magnitude away either). Even 1998's Voodoo 2 had 2.2 GB/sec of memory bandwidth :-)