HNHacker News
TopNewBestAskShowJobs

mattewong

78 karma · joined August 12, 2015

submissionscomments
mattewong··on Show HN: CSVtoAny, CSV Local File Converter
The site says privacy-first and also says "we cannot lose your data if we never collect it" but it makes a WHOLE lot of POST calls passing what appear to be encrypted payloads, and refuses to work offline-- so the user has no way to verify that the limited info you claim to be collecting is in fact what is being collected. Worse, if you simply visit and use the site, you never once see any mention of terms of use, and yet those published terms-- which you will only find if you actively scroll way down to the bottom of the SPA and click on a tiny link-- and claim to be binding merely by the use of the site, which could have easily happened without the user having any knowledge or notice whatsoever that they "agreed" to something (in other words, without actually agreeing to anything). The terms also do not say anything about your data collection, though if one looks hard enough one can find it mentioned in the privacy policy, well below the contradictory opening line that says "we cannot lose your data if we never collect it". Sorry, but meta data is still data, so "we never collect [your data]" is simply false.

So, maybe you did not intend it to be so, but to me the site comes off as being very sketchy and untrustworthy.

mattewong··on The SIMD-CSV crate chose not to use simdjson tricks to parse CSV with SIMD
What is the advantage of this over the parser used by xsv? From the documentation, the only difference I can see is that xsv handles weird CSV better than this crate-- which in some situations is very important! So presumably this one must be faster? If so, how much faster? Or is there some other advantage to this?
mattewong··on Show HN: Visual, local-first data tool
Looks interesting and I gave it a whirl-- thank you. Your intro mentions filter + sort, but I couldn't find a way to do that in the web UI (maybe that's just my ineptitude).

Re your question whether it would be useful: hard to answer because I cannot tell right now whether it solves (or intends to solve) any specific problem better than plenty of other alternatives.

mattewong··on Show HN: ZSV – A fast, SIMD-based CSV parser and CLI toolkit
Hi HN, I'm the author of zsv.

zsv was built because I needed a library to integrate with my application, and other CSV parsers had one or more of a variety of limitations (couldn't handle "real-world" CSV or malformed UTF8, were too slow, degraded when used on very large files, couldn't compile to web assembly, could not handle multi-row headers (seems like basically none of the other CSV parsers do this) etc-- more details are in the repo README). The closest solution to what I wanted was xsv, but was not designed as an API and I still needed a lot of flexibility that wasn't already built into it.

My first inclination was to use flex/bison but that approach yielded surprisingly slow performance; SIMD had just been shown to be useful in unprecedented performance gains for JSON parsing, so a friend and I took a page from that approach to create what afaik (though I could be wrong) is now the fastest CSV parser (and most customizable as well) that properly handles "real-world" CSV.

When I say "real-world CSV": if you've worked with CSV in the wild, you probably know what I mean, but feel free to check out the README for a more technical explanation.

With parser built, I found that some of the use cases I needed it for were generic, so I wrapped them up in a CLI. Most of the CLI commands are run-of-the-mill stuff: echo, select, count, sql, pretty, 2tsv, stack. Some of the commands are harder to find in other utilities: compare (cell-level comparison with customizable numerical tolerance-- useful when, for example, comparing CSV vs data from a deconstructed XLSX, where the latter may look the same but technically differ by < 0.000001), serialize/flatten, 2json (multiple different JSON schema output choices). A few are not directly CSV-related, but dovetail with others, such as 2db, which converts 2json output to sqlite3 with indexing options, allowing you to run e.g. `zsv 2json my.csv --unique-index mycolumn | zsv 2db -t mytable -o my.db`.

I've been using zsv for years now in commercial software running bare metal and also in the browser (see e.g. https://liquidaty.github.io/zsv/), so I finally got around to tagging v1.0.1 as the first production-ready release.

I'd love for you to try it out and would welcome any feedback, bug reports, or questions.

mattewong··on Show HN: Convert between Mermaid, draw.io, and Excalidraw diagrams
> even if it's "this isn't useful because X."

OK... this isn't useful to me because I now just only use mermaid and stopped using other diagraming tools, because mermaid can be embedded now in so many places (github, in-browser editing/rendering, shareable URL, python etc), can be easily saved as text and edited (whether manually or programmatically), there are no IP issues, etc. So more useful for me would be, whatever you need to use something other than mermaid for-- do what it takes to make mermaid fill that need, so that there is no need to use anything else.

mattewong··on Show HN: ParquetFormatter – Convert Parquet and Ndjson to CSV (and Back)
Maybe better if the website discloses the fact that file names are getting tracked via POST to the server
mattewong··on Show HN: Healthcare Price Comparison Tool
< Technical details about the data wrangling happy to share in comments if anyone's interested in that nightmare.

I'm interested. That is the nightmare my company obsesses about solving

mattewong··on DuckDB is faster at counting the lines of a CSV file than wc
Your follow-up post is helpful and appreciated!

Re the original analysis, my own opinion is that the outcome is only surprising when the critical detail, highlighting how the two are different, is omitted. It seems very unsurprising if it is rephrased to include that detail: "DuckDB, executed multi-threaded + parallelized, is 2.5x faster than wc, single-threaded, even though in doing so, DuckDB used 9.3x more CPU".

In fact, to me, the only thing that seems surprising about that is how poorly DuckDB does compared to WC-- 9x more CPU for only 2.5x more improvement.

But an interesting analysis regardless of the takeaways-- thank you

mattewong··on Show HN: The Canada census data in a SQLite file; advice appreciated
That is great, thank you. I'd love to continue the conversation-- maybe easier in a separate forum. Can I follow-up via the email address on your profile (gaven...)?
mattewong··on Show HN: The Canada census data in a SQLite file; advice appreciated
Glad to be helpful-- I'm in the business of data process automation, so I appreciate the opportunity to learn about new use cases. If you are willing to share what your end goal was in more detail (even as simple as an SQL query that you would now run want to run against your current schema), I'd be interested to see how an optimal process could be designed to easily generate that, and possibly suggest some tooling you could find useful. You may also want to try posting questions like this in forums such as the Seattle Data Guy's discord channel, and I'm suspect you will get lots of suggestions and advice.
mattewong··on DuckDB is faster at counting the lines of a CSV file than wc
This is misleading. First, as other comments have noted, it is comparing multi-threaded/parallelized vs single-threaded, and its total CPU time is much longer than wc's. Second, it suggests there is something special going on, when there is not. Just breaking the file into parts and running wc -l on it-- or even, running a CSV parser that is much more versatile than DuckDB's-- I'm pretty confident will perform significantly faster than this showing. Bets anyone?
mattewong··on Show HN: The Canada census data in a SQLite file; advice appreciated
I am always a proponent of starting with the end goal and then working backward. What are the end results you are aiming to achieve (or aiming to allow your audience to achieve)? Is marginal precision more important than the speed impact? The optimal database design will depend on that (i.e., on what you are optimizing for...).

It would also be very helpful, imho, to indicate keys and indexes, perhaps by modifying your schema diagram, or simply (and maybe better), just dump the actual SQL schema definition (i.e. the output from sqlite3's ".schema" command)

mattewong··on How fast can you parse a CSV file in C#?
Haven't yet seen any of these beat https://github.com/liquidaty/zsv (of which I'm an author) when real-world constraints are applied (e.g. we no longer assume that line ends are always \n, or that there are no dbl-quote chars, embedded commas/newlines/dbl-quotes). And maybe under the artificial conditions as well.
mattewong··on CSVs Are Kinda Bad. DSVs Are Kinda Good
I cannot imagine any way it is worth anyone's time to follow this article's suggestion vs just using something like zsv (https://github.com/liquidaty/zsv, which I'm an author of) or xsv (https://github.com/BurntSushi/xsv/edit/master/README.md) and then spending that time saved on "real" work
mattewong··on Ask HN: Stack to Put CSV Online
so true. sometimes the best solutions are not sexy
mattewong··on Show HN: NSV, a text file format that uses newlines to separate values
While you're making a CSV variant, why not go the extra step and remove the single most problematic CSV performance problem and make NSV compatible with high-performance, parallelized processing by eliminating quoting, and instead use escapes for embedded newlines, so that a newline is always a field delimiter and two newlines is always a record delimiter?
mattewong··on Ask HN: How do you know that it's time to shutdown your startup?
Do you have revenue? If not, how close to revenue are you?
mattewong··on Analyzing multi-gigabyte JSON files locally
If it could be tabular in nature, maybe convert to sqlite3 so you can make use of indexing, or CSV to make use of high-performance tools like xsv or zsv (the latter of which I'm an author).

https://github.com/liquidaty/zsv/blob/main/docs/csv_json_sql...

https://github.com/BurntSushi/xsv

mattewong··on Ask HN: What modern tools should be standard part of a modern unixy distro?Why?
csvkit and miller are both extraordinarily slow

try xsv (https://github.com/BurntSushi/xsv) or zsv (https://github.com/liquidaty/zsv) instead (the latter of which I'm an author of)

mattewong··on Show HN: Faster FastAPI with simdjson and io_uring on Linux 5.19
Yes, that is exactly my point. You cannot start threads at 0/25/50/75 if your data is in CSV format. But what I am saying is that, if you could do that, then your performance difference will be negligible, compared to using a single thread that parses the CSV into rows and passes chunks of rows to 4 separate threads.

In fact, the single-thread parser approach (with multi-thread processing) might even be better, because it is not trying to access your hard disk in 4 places at the same time. Then again, if your threads are doing some non-trivial task with each row, then IO will not be your bottleneck either way.

Obviously starts to break down if you aren't reading the whole file and you wanted to start some meaningful portion of the way in and never process what comes before it. The point is, the benefit of being able to, effectively, implicitly shard a file without saving as separate files-- might not be as impactful in practice as in theory

mattewong··on Show HN: Faster FastAPI with simdjson and io_uring on Linux 5.19
Good point. Though, if we are talking about something coming down a network pipe, then that network connection will be serialized anyway and during the parsing process can be sharded or converted to another format or indexed or whatnot. I would still say that, a situation where anything non-trivial gets bottlenecked by the CSV parsing remains exceptionally low. If you are reading the entire file, then the difference between starting, say, 4 threads directly in positions 0/25/50/75 versus a single CSV reader that dispatches chunks of rows to 4 threads (or whatever N instead of 4) is probably nil.

It is true there will be exceptions-- such as if you know you only want to read the second half the file only. In that case CSV with quoting does not give you a direct way to find that halfway point without parsing the first half.

I suppose whether this is worth the other pros/cons will be situation-dependent. For my use cases, which are daily, CSV parsing speed, when using something like xsv or zsv, has just, by itself, never been a material concern/impact on performance.

Where I think the CSV parsing downside is much greater than the fact that it must be serial (but which as described above does not prevent parallelized processing), is in type conversion not just of numbers but in particular of dates-- it can be expensive to convert the text "March 6, 2023" to a date variable. However, if you have control over the format, you could just as easily printed that as an integer such as 44991 and reduces the problem to one of integer conversion. Which is still always going to be slower than a binary format, but isn't so bad performance wise.

mattewong··on Show HN: Faster FastAPI with simdjson and io_uring on Linux 5.19
Parsing CSV doesn't have to be slow if you use something like xsv or zsv (https://github.com/liquidaty/zsv) (disclaimer: I'm an author). The speed of CSV parsers is fast enough that unless you are doing something ultra-trivial such as "count rows", your bottleneck will be elsewhere.

The benefits of CSV are:

- human readable

- does not need to be typed (sometimes, data in the raw such as date-formatted data is not amenable to typing without introducing a pre-processing layer that gets you further from the original data)

- accessible to anyone: you don't need to be a data person to dbl-click and open in Excel or similar

The main drawback is that if your data is already typed, CSV does not communicate what the type is. You can alleviate this through various approaches such as is described at https://github.com/liquidaty/zsv/blob/main/docs/csv_json_sql..., though I wouldn't disagree that if you can be assured that your starting data conforms to non-text data types, there are probably better formats than CSV.

The main benefit of Arrow, IMHO, is less as a format for transmitting / communicating but rather as a format for data at rest, that would benefit from having higher performance column-based read and compression

mattewong··on yq: command-line YAML, JSON, XML, CSV and properties processor
Another for the mix: https://github.com/liquidaty/zsv
mattewong··on csvkit: Command-line tools for working with CSV
I wanted so much to use csvkit and all the features it had, but its horrendous performance made it unscalable and therefore the more I used it, the more technical debt I accumulated.

This was one of the reasons I wrote zsv (https://github.com/liquidaty/zsv). Maybe csvkit could incorporate the zsv engine and we could get the best of both worlds?

Examples (using majestic million csv):

---

csvcut -c 1,3 = 5.3 seconds

zsv select -n -- 1 3 = 0.19 seconds

28x faster

---

csvsql --query "select count(*) from file" file.csv = 148 seconds

zsv sql "select count(*) from data" file.csv = 0.68 seconds

216x faster

---

mattewong··on Ask HN: Programs that saved you 100 hours? (2022 edition)
FYI: https://github.com/BurntSushi/xsv is much faster than mlr (like an order of magnitude), and zsv (https://github.com/liquidaty/zsv) is even faster. But, neither support formulas. Disclaimer: I am one of the zsv authors
mattewong··on Ask HN: Programs that saved you 100 hours? (2022 edition)
https://github.com/liquidaty/zsv
mattewong··on Ask HN: Anyone tired of everything being a subscription now?
If you think it is not good for producers and consumers, you are not disagreeing with me as much as with economic theory. Personally, I think there are many distortions in practice that can lead to exceptions, but that in the bigger picture of things, it nonetheless holds water.

I will not try to recreate the arguments to the theory but here is a basic one to start with from https://www.investopedia.com/terms/p/price_discrimination.as...:

--- Wouldn’t Consumers Be Better Off If Everybody Paid the Same Price? --- In many cases, no. Different customer segments have different characteristics and different price points that they are willing to pay. If everything were priced at say the "average cost," people with lower price points could never afford it. Likewise, those with higher price points could hoard it. This is what is known as market segmentation. Economists have also identified market mechanisms whereby fixing static prices can lead to market inefficiencies from both the supply and demand sides.

mattewong··on Ask HN: Anyone tired of everything being a subscription now?
Re car manufacturers using remote "unlock" pricing, I will admit to being horrified when I first heard about it.

Unfortunately for my emotions, there is a strong argument to support this new pricing approach-- which might be BS in specific cases but which, from a macro perspective, may nonetheless have merit-- in that fundamentally, all it is, is a mechanism for further granularity of the age-old thing called price discrimination, which is generally viewed as good for both producers and consumers (not that there cannot be exceptions to either that view, or the circumstances in which that view is valid...)

mattewong··on Show HN: We scaled Git to support 1 TB repos
Thank you for posting this. Is there any way to access a Xet dataset via a URL (assuming the dataset owner has opted to share in that manner) so that, for example, one could visit a web page that contains some embedded code (JS, WASM etc) which pulls the Xet data into the page for processing?
mattewong··on Writing-Focused Startups Draw Big Bucks
"For years, tech writers have been warning about how AI will eliminate the need for all kinds of human-staffed professions from truck driving to portfolio management.

Turns out, the AI bots are really coming for us."

Page 1 of 4Next →