ParaText: CSV parsing at 2.5 GB per second
wise.io
wise.io
"A fast reader exploits the capabilities of the storage system"...
the graphs show that their storage system is doing 4.00 GB/sec
I wonder what processor this is running on and what their storage system this is.. multiple PCIe SSD?
I tried running a quick test but only succeeded in OOMing my 8G laptop.
Even just doing
import paratext
it = paratext.load_csv_as_iterator("/dev/shm/tmp/c.log", expand=True, forget=True)
x = it.next()
Starts eating up all my ram after about a minute of spinning the cpu... so I think they have a slightly different definition of an iterator as everyone else.Compared to
cut -d , -f 5 < c.log > /dev/null
which runs in a few seconds, or a slightly more domain specific and optimized version of 'cut'[1] that runs even faster (300-500MB/s on a single core depending on which fields you want) $ du -hs c.log ;wc -l c.log
2.1G c.log
16197412 c.log
I also wonder if that is 2.5 GB/s per core.https://github.com/BurntSushi/rust-csv does 241 MB/s in raw mode, so I find it a little hard to believe that this is 10x faster... unless that is while maxing out multiple cores.
[1] https://github.com/bro/bro-aux/blob/master/bro-cut/bro-cut.c
One high-end PCIe SSD can manage that; two easily could. A high-end NAS might, too.
> I wonder what processor this is running on and what their storage system this is.. multiple PCIe SSD?
Consumer grade NVMe disks currently achiveve 2GB/s. It is easy to use a couple of those, or just one professional SSD.
I don't think so, the difference is you're expecting it to read one row at a time and it's actually reading one column at a time. load_csv_as_iterator is iterating over the columns.
Firstly, their core parsing routine is a state machine not unlike the one found in rust-csv: https://github.com/wiseio/paratext/blob/master/src/csv/rowba...
Secondly, they do indeed appear to be achieving high parsing speed using coarsely grained parallelism. You can see the code that chunks up the CSV data here: https://github.com/wiseio/paratext/blob/master/src/generic/c...
It's pretty clever!
- all records are separated by an RS character (#x1e)
- all fields within a file are separated by a US character (#x1f)
- all instances of RS, US & ESC within a field are prefixed with an ESC (#x1b)
- there are no more rules
It's remarkable to me that ASCII defines a pretty full-featured mechanism for information interchange (start of header, file transfer &c.) and instead we continue to build mechanisms atop its alphabetic characters.It's like the original sin of computing is Not Invented Here (no doubt someone will pipe up with a story of how ASCII itself was the product of NIH!).
If we're abandoning human-readability, why even bother with ASCII? Just use a binary format. Has anyone actually used ASCII unit and record separator delimiters successfully? I'd be curious about what advantages they had over a binary format, even just a protobuf or Thrift serialized form. If we want to preserve schemalessness, there's stuff like Sereal.
You're assuming that separator characters are not human-readable. If ASCII had been used as originally intended, they'd be just as readable as line breaks. Err, carriage returns.
Computing depresses me some days …
There are all sorts of examples like this. UNIX was initially designed to work with pipes of line-oriented streams, so why did we get scripting languages where every UNIX command is reinvented as a function call or RPC frameworks where the pipe is replaced by a binary message? PHP was initially a templating language, so why did we get Smarty, Wordpress, and PEAR templates? The web was supposed to come with full support for editing & creating pages via WYSIWYG editor and have a built in mechanism (hyperlinks) for associating pages with other people, so why did we need Facebook to introduce the idea of "sharing content with other people".
In each case, there were real, pragmatic reasons that people invented new systems instead of doing what they were "supposed" to do. Line-oriented files are clumsy for representing hierarchical data or conditionals. PHP is too hard to use for most end-users, despite being built for pragmatic "just toss a webpage up" use. The editing features in the web disappeared early on, with Netscape, and they needed the critical mass of college students that Facebook provided before people felt they had an audience for anything they did.
The moral for system designers is that you can't just throw a feature out there and say "Use this." You have to adapt it to how people actually do use it, even if that usage seems brain-dead to you.
At my first job, the (proprietary) database server we used had a communication protocol that used the ASCII STX/ETX/EOT/ACK/ETB/GS/RS/US characters. When I looked up those ASCII codes I was like "Oh, that's clever", but in practice, it was just another proprietary binary format, and had all the problems of proprietary binary formats. The control codes still needed to be escaped in source code; they still weren't visible in log files or when viewing console output. They were basically just bytes we had to send to implement the protocol.
After it has been saved in that format, it is a text file consisting of one very long line (unless some of the fields have newlines, but that doesn't represent the actual structure). This makes it inconvenient to browse with a text editor or pager. Sure, you could have a special mode to make them show up as newlines, but then you can't tell the difference between real newlines and newlines inserted for convenience.
Consequently, Unix text-processing tools like cut, sed, and awk use a line feed ('\n') as a record delimiter, unless there is a special reason not to (e.g. the -print0 flag on find).
Related discussion here:
Also, the approach mentioned here is only useful if reading+parsing the data is the bottleneck (or close to it). For example, if reading+parsing takes only 10% of the total processing time, then optimizing this stage will only give at most a 11% increase in performance.
CSV, tab-delimited files and in general flat files are retarded idea to be used in banking. I know because I work for bank, these are source of all misery.
Another datapoint: data changes (price updates, number ranges) between European telecom operators are all csv files over ftp. You can recognize the security aware by their use of sftp!
10% performance increase is nothing to scoff at in a mature code base.
This comment makes me wonder how often a distributed approach is used out of some odd sense of convenience or interest, instead of taking time to create an optimized single node approach?
In other words, how often is an Hadoop-like system used when unnecessary, adding unnecessary complexity?
e.g., run a single k interpeter on each CPU, then divide and conquer
Isn't it faster to do parsing in memory and avoid I/O wherever possible?
Loading a 156mb csv file in kdb 32-bit free version, single thread:
\t trade:`sym`time`ex`cond`size`price!("STCCXH";",")0:`t.csv
1850
in paratext, 64-bit: >>> timeit.timeit('paratext.load_csv_to_dict("t.csv",num_threads=4)', setup="import paratext", number=1)
3.1176819801330566
No improvement.Loading a 1.5GB csv file in kdb:
\t quote:`sym`time`ex`bid`bsize`ask`asize`mode!("STCHXHXC";",")0:`q.csv
14135
in paratext: >>> timeit.timeit('paratext.load_csv_to_dict("q.csv",num_threads=4)', setup="import paratext", number=1)
12.962939977645874
Not too shabby! An almost 10% improvement over KDB by turning my fans on and burning my lap!However, I think they should probably make their parser faster before they waste heat trying to make slow code finish sooner.
Apple for instance use Apache 2 for Swift and they clearly aim at closed source project to use it as a library. So what's the matter?