XMHell: Handling 38GB of UTF-16 XML with Rust
usethe.computer
usethe.computer
Steaming XML is a mature thing, libraries in most every language for decades. Oh! you had to add a "UTF-16LE", the suffering is real...
How is this hell? Did you want streaming json? it would end up being essentially the same code. YAML has to be it! (Is there a streaming parser?)
This totally feels like conjured drama.
A huge chunk of businesses are based on doing exactly this and not a lot more.
Anyone who works with texts today as a programmer who doesn't understand the basics of character encoding is likely in trouble, I find it quite ignorant to claim "Only 90s Kids Will Remember".
Also, most modern XML parser will handle/detect the correct encoding for you, so whilst you should understand character encoding, you don't need to do the heavy lifting yourself.
There are also inaccuracies in this article, the author says: "Instead, we’ll use a streaming approach (the XML folks call this “SAX”, for reasons I didn’t care enough to research)."
Erm... actually there are multiple streaming approaches available. SAX is just one of them (which was pioneered in Java).
The author also says: "I decided to go with the quick-xml library because it had good documentation, good performance, and roughly the API I was looking for."
Erm... but... quick-xml provides a StAX like API, NOT a SAX API. The former is pull and the latter is push; quite different things, with different performance characteristics depending on your needs.
I guess when everything looks like a nail to the author they reach for the rusty hammer. The amount of time and code they put into solving their problem in a limited way is rather staggering.
I would have suggested using - the correct tool for the correct job! If you have a small amount of XML (38GB isn't big these days), then work with it as XML.
Alternative options would have been to: 1. Load the XML document into a Native XML Database, which takes all of about 10 minutes whilst you make a cup of Tea. There are several Open Source ones if that's a requirement for you. Write probably less than 10 lines of query, and voila you have your CSV. 2. XSLT 3.0 has a streaming mode. You could write a small amount of XSLT 3.0, execute it using an Open Source processor from the command line (e.g. Saxon), and voila you have your CSV. 3. If you wanted to go really fast, and can spare a few dollars, they could have rented a cloud instance with more than enough memory for 1 hour, stuck it all in RAM and processed it quickly using whatever temporary scripting language they liked...
If there was an article down-vote button, I would be using it!
Are there Open source xstl 3.0 processors with streaming support?
There is also Exselt http://exselt.net (appears to be down at the moment), I don't think this is Open Source, but it was freely available for download previously.
Reminds me of the intricate machinery people build out of Excel spreadsheets and PowerBI.
P.S. I'm only half joking.
And they can create a lot of trouble!
We have to deal with a system which generates what they call XML files using simple string concats, without regard for encoding. So while the header says encoding="utf-8", the data is a mix of different encodings depending on which database the data comes from, none of them being UTF-8.
I’ve recently seen other people complain about receiving large xml and having to hack together a parser.
Can you write a how-to? Something deeper than “you can do this with XSLT” - it would be very valuable to a lot of people.
That said, pretty much every programming lang in existence has sax libraries [since this sort of thing was much more common in the 2000s], many of them will transparently handle the utf-16 encoding, so it doesn't seem like something to make that much of a fuss over.
I've used it to parse a multiple GB XML file using less than 60MB of RAM - details here: https://github.com/dogsheep/healthkit-to-sqlite/issues/7
And somewhat inconveniently you can’t hand a null builder to iterparse so for large documents you must remember to clear the elements it builds lest you explode your memory.
Eh, it was nearly enough. There were too many obscure Chinese characters but 17 bits would have fit just about everything. And for things like emoji it's not "much less", it's "a smidgen less", they're quite tiny in comparison.
And that timeline is a mess. UTF-8 came before UTF-16.
Which makes the mess that is UTF-16/UCS-2 even more unfortunate because its not like unicode had made all that many in-roads by that point. People were still mostly using various legacy code pages at that point. Imagine if we could have skip all the fuss over wide character encodings (i suppose there was already wide character encodings due to asian language encodings)
UTF-16 was an hack on top when they realized that UCS-2 wasn't enough.
Not that I'm a fan of UTF-16, but I once heard UTF-8 described as "one of the most elegant hacks in computer science", and I think I agree there ;)
Edit: Another interesting contender would be KOI8-R, as AFAIK setting its most significant bit to 0 and interpreting it as ASCII mostly results in readable text.
edit: reword
I think at this point we'd need 18 bits (19 if you still want massive amounts of private use characters) - https://www.unicode.org/versions/stats/charcountv13_0.html
Presumably that is after all where this data actually comes from. The state has it all in some ancient database, and they opened it up by using the "dump to xml" option.
Your favorite database probably has utf8/16 column and/or conversion functions too.
If you know your DB tooling and have basic XML and ETL experience, StillBored is probably right. If you have to look it all up but write software every day, a custom program is probably faster to get right.
You might still want to invest in DB/XML/ETL knowledge if more of this kind of questions come up.
If the schema changes from one month to another it would be far easier to just update the DB table schema for the next import than change parsing code.
Everything can be solved with code, but not everything should be.
Interesting as a coding exercise and perhaps that's the point of the article.
Not the way I'd do it if I wanted a stable process for consuming the data regularly.
"Very fortunately for us, we can use a library to take care of all of this for us. The encoding_rs and encoding_rs_io libraries will let us translate from UTF-16 to UTF-8 on the fly, which will let us send ordinary Rust strings into our XML parser."
Encoding is the first one: Most programs made assumptions about file encoding. You got an american program and the user wanted parts of it translated/configured in its own language? Encoding problems were almost guaranteed, and you'll have to live with losing some characters. That's in Western Europe, the rest of the world had it worse. When XML came in, you had a good chance of getting UTF-8 or else one of the other UTF variants. Finally.
Next up: every program used its own config and data file formats. Or multiple different formats for the same program. You take the manual and learn a new one every 2 weeks. Hooray, XML has only 1 set of rules to learn.
Escaping rules were different for each file format, or even incomplete or non-existent. Feeding data form 1 program to another might work most of the time, and some day it breaks because someone, for the first time in 3 years, decided to start a sentence with a space. Hooray, XML lets you escape any text you can think of, yes even PCDATA.
As a programmer, we stopped writing parsers for custom formats. Hooray, take any of the shelf XML parser and you're mostly done. Byebye, yacc and flex.
The whole schema thing was a godsend. You can validate the basic structure of your files before sending it off to a third party or putting it in a production machine. There are tools to autocomplete or even generate test cases based on schemas. Take that, custom formats from hell.
As a bonus, XML was big enough for an entire ecosystem of tooling to grow up around it, most of it cheap or free. I can open an XML file in notepad++, click on 'reformat', have syntax highlighting, validation... In the 90's this was only possible for the most common of formats, and might cost 10 000 $ or more for an extremely limited, buggy tool delivered only by the vendor.
So XML was a very big step forward. Things weren't perfect, of course. Spaces, entities, namespaces had enough pitfalls to keep you busy. The fact that he takes an 38GB file and condenses it to 92MB speaks of it verbosity (Note how the author had to sacrifice TAB/CR/LF to get there and invented a basic custom format, i.e. the 90's attitude) .
Personally, I'd liked to see JSON extended with comments and dates, a bit better numeric support, and a default usage of schemas. Then let TOML and YAML die in their respective little cradles. It would have left us with exactly 1 format for data transfer and configuration that had XMLs advantages without its verbosity and enterprisyness. But alas, too late, that oportunity has passed.
Far more importantly: you get a declaration of the encoding right at the start, so you don't have to guess.
> The fact that he takes an 38GB file and condenses it to 92MB speaks of it verbosity.
Those 92MB are not the entire payload data, he also filtered for the data of only one county.
True, but it was a lot worse than that. XML forced even the worst programmers to at least think half a second about the existence of the whole encoding concept.
Some horrors from the past: receiving so-called text files generated by concatenating multiple exports. Every chunk had a different encoding, and there is no hint when the encoding changes mid-file or what was what. As most of it is ASCII, it seems to work and people don't even understand what the problem is.
Another one: XML-support by pasting a CSV file between the beginning and end tag. <data>...CSV goes here...</data> . Of course, the CSV between the tags is not XML,violates escaping rules, and has no defined encoding.
One of the worst I ever received went like <textfile><byte value='65'><byte value='66'>.... </textfile>. Now have fun converting every element to an individual byte, then finding out the encoding of the resulting text file.
Sometimes I wish for specific configuration formats which are fit for the purpose instead of the bolted-on monstrosities like jinja-templated yaml. This has all the problems you mention, but I think we're losing something by the culture that we're not supposed to write our own parsers anymore.
The rationale for not including comments in a data exchange format is that they tend to be used for parsing and processing meta information which breaks the purpose of the data exchange format in the first place.
I don't know where the obsession to use the same file format for everything comes from, the lack of comment support in JSON makes it a bad configuration file format. That's OK, don't use it. We have .ini and TOML (and YAML if you like it rough), why not use those that are a better fit for purpose?
Hmm.. that works for data centric applications, but what about document centric applications? An advantage of XML is that it can do both, whereas JSON has no support for mixed-content.
XML indeed works for both document and data-centric applications. In practice, I mostly see data dumps from tables, and config files. Marked up documents are rare for me, and mostly in html or recently markdown. Maybe this is one of the thing that differs between jobs
All of XML, JSON, TOML,YAML only solve the first part of the job. You will need a way to put the parsed data in the data structures of your application. I will respectfully disagree about the triviality however: XML was one of the first to make this trivial. Before XML, it wasn't.
Choosing a format is forcing that preference on other ITers. If I choose, say YAML, then everybody who uses my application will be forced to learn YAML. Every format is basically easy, but has a few edge cases. Hence, it is a good practice to limit the number of formats, even if this requires compromising on a lesser format. That's why I'd like to see TOML and YAML disappear in a world that already has XML and JSON. Now clearly a world where you only have to learn XML+JSON+TOML+YAML is a lot better than the 90's world where you'd see a bigger variety of less well thought out formats every day.
There are good reasons for leaving advanced possibilities out of JSON. Part of XMLs demise was the over-complication. There is a lifecycle here: You learn X because it is easy. It isn't perfect, so you fix it up a bit to X2, X3, etc... Each fix makes the learning curve harder for beginners. One day, someone looks at X99, rightfully declares it a mess and starts a new easy Y. Now you as an X-er try to tell the Y-ers why they are wrong, and indeed they slowly grow Y2,Y3.... While the X99 and Y99 camp have a loud discussion, Z appears. One of the harder parts of an ITer's responsibility is recognizing when to invest in Y and Z even if they are still clearly inferior to X.
While the original authors dislike for XML appears apparent in a personal memory throw back read which I enjoyed, if he consults with core finance he will continue to use XML well into the foreseeable future. ISO20022 is coming to core payments everywhere in the approaching future such as with FedNow real-time payments and yes, it is XML.
(Sorry the naive but honest question. Haven't used Java since around 2005 (remember J2ME?) and Windows since 2010.)
A question about XML in rust: Does anybody know of any XML library in any language that can do tree-folds instead of stateful callbacks? I used SSAX in scheme and I now becomes sad every time I have to do any complex things using regular SAX parsers with stateful callbacks.
Oleg's site: http://okmij.org/ftp/Scheme/xml.html#XML-parser
Guile manual https://www.gnu.org/software/guile/manual/html_node/SSAX.htm... (see that section and the next one describing transformations)
> If you ever want proof of Solzhenitsyn’s assertion that “the line separating good and evil passes… right through every human heart”, consider that one man is directly responsible for both UTF-8 and the Go programming language.
It’s being served as an 1832×1465 PNG, but on screen it’s capped to 400×320, so the substantial majority of the pixels are wasted. Reducing it to a 2× image, 800×640, still PNG, drops the size by about 1.1MB. If you wanted to get fancier you could go lossy WebP and easily bring it below 50KB, but you’d probably still want the PNG as a fallback for older browsers (and even recent-but-not-quite-current Safari). At these sizes it’s probably not even worth splitting 1× and 2× into different files with srcset, though you could if you wanted to. The full markup with both of these techniques and four image files would be this:
<picture>
<source srcset="resf.webp, resf-2x.webp 2x" type="image/webp">
<img src="resf.png" srcset="resf-2x.png 2x" width="400" height="320" alt="“Rust Evangelism Strike Force” logo, in vaporwave aesthetics">
</picture>
But the first step is the most important, just shrinking the PNG from its original size to the size that you actually want to use it at. The rest is just sugar coating to make things even better, at a certain cost to you the developer if you haven’t got tooling to automate it.