Parsing JSON in 500 lines of Rust
krish.gg
krish.gg
match object(src) {
Ok(res) => return Ok(res),
Err(JSONParseError::NotFound) => {} // if not found, that ok
Err(e) => return Err(e),
}
You probably have realized that this is really tedious, and this is where macros would really shine: macro_rules! try_parse_as {
($f:expr) => (
match $f(src) {
Ok(res) => return Ok(res),
Err(JSONParseError::NotFound) => {} // if not found, that ok
Err(e) => return Err(e),
}
);
}
try_parse_as!(object);
try_parse_as!(array);
// ...
It is also possible to avoid macros by translating `Result<(&str, JSONValue), JSONParseError>` into `Result<Option<(&str, JSONValue)>, JSONParseError>` (where `Ok(None)` indicates `Err(JSONParseError::NotFound)`), which allows for shorthands like `if let Ok(res) = translate(object())? { ... }`.Also, even though you chose to represent numbers as f64, a correct parsing algorithm is surprisingly tricky. Fortunately `f64::parse` accepts a strict superset of JSON number grammar, so you can instead count the number of characters making the number up and feed it into the Rust standard library.
I wonder, given a 500 line problem, can any of the current cutting edge AIs make the code obviously, dramatically better?
How far can they go?
Just for the sake of completeness, and not to imply that you don't know this, but the JSON spec doesn't limit the size or precision of numbers, although it allows implementations set other limits.
I have encountered JSON documents that (annoyingly) required the use of a parser with bigint-support.
This actually led to a data-loss causing bug in the AWS DynamoDB Console's editor a couple of years ago. IIRC, the temporary fix was to fail if the input-number couldn't be represented by a 64-bit float. Can't remember if it ever got a proper fix.
https://gist.github.com/isaksky/6681cfad8ced1708a04b2eca92fc...
https://fsharpforfunandprofit.com/posts/understanding-parser...
% cargo run --release
[..]
Parsing speed: 103761177.00 Bytes/s
Parsing speed: 103.76 MB/s
Parsing speed: 0.10 GB/s
% doas cargo run --release
[..]
Parsing speed: 105401032.21 Bytes/s
Parsing speed: 105.40 MB/s
Parsing speed: 0.11 GB/Any random background processes might slow the system slightly during the initial run. A slightly better state of the page cache on the second run could play a role, too. Flushing the FS cache before each run might be a good idea; doing cat file.json > /dev/null could be an equally good idea.
I wouldn't think the code signing phase is run while the benchmark is running and thus affecting those numbers.
I'm pretty sure you'd already know of this, but once you've written your own version, it might help to compare and take notes from a popular, well-established, benchmarked library: https://github.com/serde-rs/json.
It uses a very similar paradigm to what the author used in the article, and provides a lot of helper utilities. I used it to parse a (very, very small) subset of Markdown recently [3] and enjoyed the experience.
[1] https://github.com/rust-bakery/nom
[2] https://github.com/rust-bakery/nom/blob/main/examples/json.r...
In general when writing a parser you should strive to minimize allocations and backtracking.
\uFEFF{"": "value", "": null}
(Yup, there's a reason those are not recommended for use)
- is this hard to do in 500 lines of rust? - is there a catch why a implementation in rust worth mentioning on HN?
don't get me wrong im realy intested in rust.
No, not that much. User posted mostly because he/she is learning the language.
> is there a catch why a implementation in rust worth mentioning on HN?
HN loves Rust and it is an fun opportunity to explore Rust semantics in use without going to the more complex part of the language. It belongs to the category of "cool first project", like a Sudoku solver, etc.
I had a goal of making it "async", so it could periodically yield for large input strings, but I'm not so sure it matters much.
Currently the API is pretty impractical to use, because of the arena structure that is set up. It could be improved
I regularly deal with situations where devs "log" by sending "json" to stdout of their container runtime, then expect the downstream infrastructure to magic structure into it perfectly. "You understand something has to parse that stream of bytes you're sending to find all the matching quotes and curly braces and stuff, right? What happens if one process in the container's emitting a huge log event and something else in the container decides to report that it's doing some memory stuff?" <blank stare> "I expect you'd log the error and open an incident?"
(the correct answer is to just collect the garbage strings in json (ha) and give them unparsed crap for them to deal with themselves, but then "we're devs deving; we don't want to waste energy on operations toil"
Later people ask "why's logging so expensive?"
Sigh.
[1] opensearch / elasticsearch, obvs
1. https://github.com/ezpkg/iter.json: last year, using go iterator, the core parsing code [1a] is around 200 lines
1a. https://github.com/ezpkg/ezpkg/blob/main/iter.json/parser.go
2. https://github.com/iOliverNguyen/ujson: 4 years ago, using callback, around 400 lines
It's probably better to set up an actual benchmark using a crate like Criterion instead [0].
Tools like that are to eliminate noise and variation, which is an entirely different issue. According to the article, "sudo" is about 70% faster. That has nothing to do with the benchmarking method.
1. Reading the entire file into RAM.
2. Providing a `const char *get_value(const char *jstring, const char *path, ...)` function with a NULL-terminated parameter list that would return the position of the value of the key at the specified path.
3. Providing a `copy_value(const char *position)` function to copy the value at the specified position.
Slow? Yup!
But, it was easy and safe[1] and used absolutely minimal RAM![2]. The recursive nature of the JSON tree also allowed the caller to use a returned value from `get_value` as the `jstring` argument in further calls to `get_value`.
I might still have a fork of it lying around somewhere.
[1] "Safe" meaning "Caller had to check for NULL return values, and ensure that NULL terminated the parameter list".
[2] GCC with `-O2` and above does proper TCO, eliminating unbounded stack growth.
So?
Many of the other JSON parsing tools build up a tree in RAM[1], so the statement "Only works if the final tree fits in RAM" is just as true, and since JSON is mostly text, the final tree is not that much smaller than the source JSON anyway.
[1] Including the one this article is presenting, and all the other solutions presented in the comments.