* This isn't parsing a CSV, this is a program written to split this exact dataset. (The code is filled with hard coded values)
* You're comparing a single-threaded run on a low-end CPU to a top-tier GPU.
* Your dataset can fit into GPU memory.
* There is a pull request for a missing semicolon, which means the posted version of the code won't even compile, so couldn't have been the version used to generate the benchmarks.
* The amount of branching in the GPU code makes it hard for me to believe that it actually ran that fast. GPU parallelism does not work well with branching since all cores in a cube must executing in lock-step, if you branch, then you now have to go back and execute all of your branches separately.