So, maybe you did not intend it to be so, but to me the site comes off as being very sketchy and untrustworthy.
78 karma · joined August 12, 2015
So, maybe you did not intend it to be so, but to me the site comes off as being very sketchy and untrustworthy.
Re your question whether it would be useful: hard to answer because I cannot tell right now whether it solves (or intends to solve) any specific problem better than plenty of other alternatives.
zsv was built because I needed a library to integrate with my application, and other CSV parsers had one or more of a variety of limitations (couldn't handle "real-world" CSV or malformed UTF8, were too slow, degraded when used on very large files, couldn't compile to web assembly, could not handle multi-row headers (seems like basically none of the other CSV parsers do this) etc-- more details are in the repo README). The closest solution to what I wanted was xsv, but was not designed as an API and I still needed a lot of flexibility that wasn't already built into it.
My first inclination was to use flex/bison but that approach yielded surprisingly slow performance; SIMD had just been shown to be useful in unprecedented performance gains for JSON parsing, so a friend and I took a page from that approach to create what afaik (though I could be wrong) is now the fastest CSV parser (and most customizable as well) that properly handles "real-world" CSV.
When I say "real-world CSV": if you've worked with CSV in the wild, you probably know what I mean, but feel free to check out the README for a more technical explanation.
With parser built, I found that some of the use cases I needed it for were generic, so I wrapped them up in a CLI. Most of the CLI commands are run-of-the-mill stuff: echo, select, count, sql, pretty, 2tsv, stack. Some of the commands are harder to find in other utilities: compare (cell-level comparison with customizable numerical tolerance-- useful when, for example, comparing CSV vs data from a deconstructed XLSX, where the latter may look the same but technically differ by < 0.000001), serialize/flatten, 2json (multiple different JSON schema output choices). A few are not directly CSV-related, but dovetail with others, such as 2db, which converts 2json output to sqlite3 with indexing options, allowing you to run e.g. `zsv 2json my.csv --unique-index mycolumn | zsv 2db -t mytable -o my.db`.
I've been using zsv for years now in commercial software running bare metal and also in the browser (see e.g. https://liquidaty.github.io/zsv/), so I finally got around to tagging v1.0.1 as the first production-ready release.
I'd love for you to try it out and would welcome any feedback, bug reports, or questions.
OK... this isn't useful to me because I now just only use mermaid and stopped using other diagraming tools, because mermaid can be embedded now in so many places (github, in-browser editing/rendering, shareable URL, python etc), can be easily saved as text and edited (whether manually or programmatically), there are no IP issues, etc. So more useful for me would be, whatever you need to use something other than mermaid for-- do what it takes to make mermaid fill that need, so that there is no need to use anything else.
I'm interested. That is the nightmare my company obsesses about solving
Re the original analysis, my own opinion is that the outcome is only surprising when the critical detail, highlighting how the two are different, is omitted. It seems very unsurprising if it is rephrased to include that detail: "DuckDB, executed multi-threaded + parallelized, is 2.5x faster than wc, single-threaded, even though in doing so, DuckDB used 9.3x more CPU".
In fact, to me, the only thing that seems surprising about that is how poorly DuckDB does compared to WC-- 9x more CPU for only 2.5x more improvement.
But an interesting analysis regardless of the takeaways-- thank you
It would also be very helpful, imho, to indicate keys and indexes, perhaps by modifying your schema diagram, or simply (and maybe better), just dump the actual SQL schema definition (i.e. the output from sqlite3's ".schema" command)
https://github.com/liquidaty/zsv/blob/main/docs/csv_json_sql...
try xsv (https://github.com/BurntSushi/xsv) or zsv (https://github.com/liquidaty/zsv) instead (the latter of which I'm an author of)
In fact, the single-thread parser approach (with multi-thread processing) might even be better, because it is not trying to access your hard disk in 4 places at the same time. Then again, if your threads are doing some non-trivial task with each row, then IO will not be your bottleneck either way.
Obviously starts to break down if you aren't reading the whole file and you wanted to start some meaningful portion of the way in and never process what comes before it. The point is, the benefit of being able to, effectively, implicitly shard a file without saving as separate files-- might not be as impactful in practice as in theory
It is true there will be exceptions-- such as if you know you only want to read the second half the file only. In that case CSV with quoting does not give you a direct way to find that halfway point without parsing the first half.
I suppose whether this is worth the other pros/cons will be situation-dependent. For my use cases, which are daily, CSV parsing speed, when using something like xsv or zsv, has just, by itself, never been a material concern/impact on performance.
Where I think the CSV parsing downside is much greater than the fact that it must be serial (but which as described above does not prevent parallelized processing), is in type conversion not just of numbers but in particular of dates-- it can be expensive to convert the text "March 6, 2023" to a date variable. However, if you have control over the format, you could just as easily printed that as an integer such as 44991 and reduces the problem to one of integer conversion. Which is still always going to be slower than a binary format, but isn't so bad performance wise.
The benefits of CSV are:
- human readable
- does not need to be typed (sometimes, data in the raw such as date-formatted data is not amenable to typing without introducing a pre-processing layer that gets you further from the original data)
- accessible to anyone: you don't need to be a data person to dbl-click and open in Excel or similar
The main drawback is that if your data is already typed, CSV does not communicate what the type is. You can alleviate this through various approaches such as is described at https://github.com/liquidaty/zsv/blob/main/docs/csv_json_sql..., though I wouldn't disagree that if you can be assured that your starting data conforms to non-text data types, there are probably better formats than CSV.
The main benefit of Arrow, IMHO, is less as a format for transmitting / communicating but rather as a format for data at rest, that would benefit from having higher performance column-based read and compression
This was one of the reasons I wrote zsv (https://github.com/liquidaty/zsv). Maybe csvkit could incorporate the zsv engine and we could get the best of both worlds?
Examples (using majestic million csv):
---
csvcut -c 1,3 = 5.3 seconds
zsv select -n -- 1 3 = 0.19 seconds
28x faster
---
csvsql --query "select count(*) from file" file.csv = 148 seconds
zsv sql "select count(*) from data" file.csv = 0.68 seconds
216x faster
---
I will not try to recreate the arguments to the theory but here is a basic one to start with from https://www.investopedia.com/terms/p/price_discrimination.as...:
--- Wouldn’t Consumers Be Better Off If Everybody Paid the Same Price? --- In many cases, no. Different customer segments have different characteristics and different price points that they are willing to pay. If everything were priced at say the "average cost," people with lower price points could never afford it. Likewise, those with higher price points could hoard it. This is what is known as market segmentation. Economists have also identified market mechanisms whereby fixing static prices can lead to market inefficiencies from both the supply and demand sides.
Unfortunately for my emotions, there is a strong argument to support this new pricing approach-- which might be BS in specific cases but which, from a macro perspective, may nonetheless have merit-- in that fundamentally, all it is, is a mechanism for further granularity of the age-old thing called price discrimination, which is generally viewed as good for both producers and consumers (not that there cannot be exceptions to either that view, or the circumstances in which that view is valid...)
Turns out, the AI bots are really coming for us."