How fast is fscanf?
github.com
github.com
- The proper conversion has enough corner cases that naive implementation practically certainly misses. Author's version produces provably different results than the library one for a lot of the same inputs.
- The fscanf must process the format string on every call, therefore we're not comparing the same functionality. There are library functions that don't need the format string to work.
- The standard library must be able correctly process Unicode and locales.
See the other comments for more details. Of course, it can be that the author doesn't need more than what he implemented, but he didn't discuss the tradeoffs or incorrect reading in the article and the other readers are advised to investigate more before using his code.
http://www.exploringbinary.com/gays-strtod-returns-zero-for-...
http://www.exploringbinary.com/glibc-strtod-incorrectly-conv...
http://www.exploringbinary.com/a-closer-look-at-the-java-2-2...
http://www.exploringbinary.com/java-hangs-when-converting-2-...
http://www.exploringbinary.com/php-hangs-on-numeric-value-2-...
Even if you have received data from an outside source in text format, perhaps the first thing you should do is convert it to a binary format on disk; this will save you a lot more overhead than any amount of scanf() optimization will.
On the other hand, if you have to crunch through huge amounts of locally stored data, try to make most of your binary file-format resemble the in-memory layout as closely as possible. That way you can, with some luck, mmap() your file and just use pointers into the bulk data. Saves at least one level of copying, from the filesystem cache to application heap. Be careful with validation of lengths and offsets, though!
If it works without problems for your data, good for you, but for general purpose it isn't nearly a replacement as fscanf() or strtod() (hopefully) work correctly and don't have weird limitations like 300 / -300 as max / min base 10 exponent. Parsing floating point numbers is not easy.
Using feof() to somehow ask if a stream has reached EOF before doing any I/O is wrong. Code like this comes up very often on Stack Overflow so it seems it's somehow intuitive for people, or suggested by some tutorials or something.
The purpose of feof() is to answer, after I/O has failed, if it failed due to the stream reaching EOF.
It's also utterly pointless in this code, it should just check the return value of fscanf() and stop when it fails to convert the desired number of values. This, on the other hand, seems deeply surprising to most people, you very rarely see code that cares about the return value when scanning. Which is strange; I/O can fail.
Further, since fscanf() basically ignores embedded whitespace, I'm not at all sure it will do as expected and read six floats per physical line of input.
Lastly, yes of course it's strange to use the super-flexible fscanf() if you're looking for performance.
Hint++: Parsing is not the bottleneck.
http://tinodidriksen.com/2011/05/28/cpp-convert-string-to-do...
As a rule of thumb the formatted I/O functions are quite slow, and floating point makes this much much worse.
It's one of those few places where rolling your own functions for your particular use case will be much more efficient than using the one size fits all standard library solutions.