Is it the case that the CPU version could be sped up dramatically by using multiple cores and a variation on the line splitting technique?
Is it the case that the CPU version could be sped up dramatically by using multiple cores and a variation on the line splitting technique?
Seeing that it takes 14.5s with the handrolled code to parse 750GB, I however doubt the optimization is needed - it's more than 50GB every second and you need some really really exotic hardware to generate that much data. You still need to do something with the data and it's likely significantly more computational intensitive - that should be on a different thread.
That said, using the GPU would free up the CPU to do meaningful work with the data. But to be fair, assuming 1GB/s datarate, it's still 12,5m to read it all and you only gain 14.5s or less than 2% speedup if you're CPU bound.