One classic example of the danger of the word "gigaflop" is that of the exhaustive motion search. If we define a single mathematical operation as a "flop" (technically an iop, since this is integer math), using Sequential Elimination, an optimized exhaustive search algorithm, an 8-core Core 2 system can crank out over 2.7 teraflop-equivalents of processing.
For CPUs, the numbers using FFTW, one of the fastest FFT libraries that does take advantage of SIMD, the numbers usually do not exceed 5-6 gflops particularly for larger lengths.
OTOH, the above 55 gflops figure is also somewhat misleading since it does not include transfer time of data b/w RAM and GPU. Actual throughput is somewhere around 20gflops. On one particular project using FFT, I got around 15 gflops using GPU including transfer time while testing several FFT libraries, I never got above 3 gflops on a 2.4ghz quad-core using all four cores. The lenghts were big enough not to fit into cache thus reducing CPU performance considerably.