Show HN: Very low footprint JSON parser in portable ANSI C
github.com
github.com
http://git.qemu.org/?p=qemu.git;a=blob;f=json-lexer.c;h=3cd3...
http://git.qemu.org/?p=qemu.git;a=blob;f=json-parser.c;h=849...
Among other things, this supports streaming, is fairly fast, and has gotten a fair bit of scrutiny against malicious input.
The lexer is a hand written state machine which seems like something you should never do but turned out to be pretty reasonable.
case JSON_INTEGER:
obj = QOBJECT(qint_from_int(strtoll(token_get_value(token), NULL, 10)));
Why is it so hard to parse integer from string correctly?"Ultra fast JSON decoder and encoder written in C with Python bindings"
From the people that built the Battlefield 3 web portal.
Medium complex object:
ujson encode : 18757.01101 calls/sec
yajl encode : 6315.14030 calls/sec
simplejson encode : 5542.03928 calls/sec
cjson encode : 4651.59072 calls/sec
---------
ujson decode : 10759.69649 calls/sec
simplejson decode : 8148.35221 calls/sec
cjson decode : 7931.04387 calls/sec
yajl decode : 5887.38201 calls/sec
Link: http://codetique.com
(Same question applies to how you handle the max_memory computation).
To clarify: Any "lookup table" that maps hex values to assumed character values is a portability red flag. When using them, it's polite to add comments to explicitly call out the code page dependency and argue (from a spec or RFC, say) why that assumption is okay.
Also not sure you handle the case where the json invalidly terminates in the middle of a \u sequence.
Just from a quick glance though, may be wrong.
Could be a fun little exercise.
const json_char *cur_line_begin, *i;
...
top->u.dbl = strtod (i, (json_char **) &i);
top->u.integer = strtol (i, (json_char **) &i, 10);
IckEDIT: forgot to mention that it includes a bunch of great helper functions, too.
You should get the project listed on http://www.json.org/
I've used it and think it's pretty neat. One of these days I'll get around to releasing the helper functions we've written to make it easier to use too.
[not dissing you, just bored on a sunday afternoon...]
What would the advantage of using an enum be? (and I guess I used 4 and then removed it later.)
> is using a lookup table for decoding hex really faster than the (minimal) logic (what if it causes cache misses)?
No idea, that's just the way I did it. Feel free to try something else and profile if you're really that concerned.
> do you really think that a state machine with bit flags is the best way to express the logic here? is string_add meant to increment string_length on subsequent passes?
There's only two passes, and it increments the length on both (the first is to measure the string, the second is to know where to write in it).
> what is "[..] cur_value" supposed to do at the top of json_value_free (maybe i am missing some c trick here?)?
You're not supposed to mix code and value declarations in ANSI C, so I put it at the top of the function. It's just used to temporarily store the value while reading the parent.
[edit] on the bitfield / enum question, i've been looking around for a consistent, standard way of doing things and there doesn't seem to be any one best practice (although various people note that bit fields are normally unsigned ints, while enums are signed).
You can in C99.
https://github.com/PeterScott/json-parser/commit/db9c326f747...
Still, very nice. Comparable to jsonxx which i've been using up until now.
Maybe some kind of const json_null value to return when the key isn't found.
edit: Done that.
That's what exceptions are supposed to do. C doesn't have exceptions, so you use setjmp/longjmp.
HN may or may not work as a code review platform, but I don't think I would use myself a 3rd party software that doesn't provide tests.