Golang – encoding/csv: Reading is slow
github.com
github.com
The Java comparison here seems inapt, because it doesn't do as much as the other two. It's just a naive "split on commas" implementation that wouldn't handle quoted cells. Really, if Go's CSV reader is only 200% slower than that and 50% slower than Python's optimized C implementation, that's pretty good already.
Or some non-ascii codepage you're not told about and have to guess (e.g. Excel generates CSV in CP1250 by default, with an option to export UTF-16)
2. the Python 3 CSV library supports arbitrary codepoints as delimiter, quote character and escape character (if applicable)
Have you ever seen a "C"SV with a multibyte sequence as a delimiter? I haven't.
Even if such a thing exists, the feature is of negative utility if it slows down CSV parsing for everyone else. If you must, write two implementations, and use the slow path if your delimiter is multibyte.
Moving from runes to bytes in reading gives us a nice speedup - not quite to eliminate the gap, but it's a start. The rest is likely all the memory copies - once the data is read in a buffer, then copied byte by byte into a slice and only then converted into a string, which is another copy, because strings can't be based on pre-existing byte slices (not in the public API that is).
Opportunity cost. It slows down the parsing for support of something that nobody has ever seen in the wild (not to mention it doesn't even match the name of the format, but let's get past that since we already use ; | and others).
Plus I can't even imagine a use case that would make it a good idea to use that over a simpler delimiter, or even the special purpose ASCII delimiter character. Can you?
~~It's also unclear which version of Python is being used, the Python 2 csv module is byte-based and encoding-unaware which can lead to unexpected behaviours.~~
Go's CSV package apparently only does UTF-8, and suggestions for speeding it up in the tracker is to just remove that and work on raw bytes (FFS)
This is valid, because UTF-8 was designed to make this valid. The UTF-8 encoding of a comma, 0x2C (also the ASCII encoding of a comma), does not appear as a part of any other UTF-8 encodings. Same with the UTF-8 encoding of the double quote, 0x22. So scanning for 0x22 and 0x2C bytes, without stopping to decode other UTF-8 sequences along the way, will produce the correct result for a valid UTF-8 input string. Then you fully decode UTF-8 for the individual fields when needed (and if you're doing a string-compare for some target value that's already UTF-8, you never need to decode UTF-8 for that field at all).
A related problem is that in older UNIX terminals, pressing backspace would delete one byte, not one character. Newer UNIX kernels have code in the terminal implementation to decode UTF-8 enough to backspace an entire character.
If you want to count the number of codepoints in a string (called "rune" in Go), then you need to do so explicitly: https://golang.org/pkg/unicode/utf8/#RuneCountInString
Is Go's internal representation of the target string UTF-8?
Kinda but kinda not, a Go string is actually an arbitrary bag of bytes, but some API (such as unicode/utf8 or `range` to iterate on codepoints — runes in Go parlance) assume it's proper UTF8.
It doesn't sound right if go takes 5 hours to finish csv parsing job while Python takes 2.5 hrs.
If you're going to throw out specific numbers, you should probably get them, or at least the ratios, correct.
Numbers from the tracker are:
Go: avg 1.489 secs Python: avg 0.933 secs
If you'd like to test this on a really large dataset to come up with how long it would take for Python to perform the same operation when Go requires 5 hours, that might be a bit more useful. If we just look at the available data, then the _extrapolation_ for Python would not be 2.5 hours. There's still a gap, but there's no need to exaggerate.
Lots of modules in the stdlib are written in C.
Either the implementation or the compiler is lacking some optimization.
It's easy to be faster if you do fundamentally less. Not necessarily wrong, depending on your task, but it's not comparable.
With modern JVMs, Java can occasionally actually be faster than native compiled languages due to dynamic optimization at runtime.
In what sense is an articulated object a "conceptual drawback"? It is a richer object model and SMI and friends had the engineering chops to makes it highly performant.
Also depending on which JVM SDK is being used (Oracle Hotspot, Oracle Graal, IBM J9, HP, PTG, JET,...), the quality of escape analysis differs but it all boils down to turning those headered objects into plain structs, if possible stack allocated.
EDIT. I also did a funny thing and replaced the CPython C _csv.so extensions with pure Python version _csv.py, from PyPy. It run about 80 (eighty) times slower. It shows what wonders does JIT do (at least to some code).
I added the results of using Apache Commons CSV to the GitHub thread.
After using that, Python was actually by far the fastest. Hooray for performance sensitive code in C :)
The focus wasn't so much on performance but on initial completeness, good interface, versatility, clarity and simplicity - with faster or more specialized implementations left to the community.
There might be different opinions about that, but I personally like the approach of having a solid and ordered programming pocket knife - that also doesn't replace a Katana for cutting.
In other words....C/C++ :) I just try not to blow off my leg.
I'd be interested to see how yours works if you are willing to share it.
Edit: Plus one that is ~2x faster than Java by avoiding allocations [1].
[0] https://gist.github.com/jmikkola/6ac96ad6d6f66e772c33ec41ed2...
[1] https://gist.github.com/jmikkola/7ded8392226b7659c881f5540be...
In this case the choice to use utf-8 everywhere, including in the csv delimiters, is making it slower.
import * as csv from 'csv-parse';
import * as fs from 'fs';
type Line = [string,string,string,string,string,string];
const parser = new csv.Parser({});
parser.on('data', (line: Line) => {
if (line[0] === '42') {
console.dir(line);
}
});
fs.createReadStream('mock_data.csv').pipe(parser);
$ /usr/bin/time node parse_csv.js
43.61user 0.85system 0:45.61elapsed 97%CPU (0avgtext+0avgdata 60076maxresident)k
$ node --version
v6.4.0
Edit: Using fast-csv 24.28user 0.20system 0:24.58elapsed 99%CPU (0avgtext+0avgdata 91780maxresident)khttps://www.npmjs.com/package/csv-parser
Edit: `fast-csv` seems to be using a lot of `RegExp`s on each iteration which can't be that fast compared to csv-parser which seems to simply go over each symbol (state machine?).
4.95user 0.19system 0:05.22elapsed 98%CPU (0avgtext+0avgdata 29704maxresident)k
A lot better but that’s still 5× slower than Python.Making a version of encoding/csv that retains most of its features (custom delimiters, handling backslashes and quoting and \r) but streams like that would be a fun open source project for someone who likes Making Things Go Fast.
Depends on the speed of your storage subsystem. If you're working from RAM or from a fast PCIe SSD, you'll probably bottleneck in the encoding validation and actual parsing.
I guess "Nobody cares about your dead religion" :-)
Still, it would probably be faster, at the expense of being unreadable.