Show HN: A Fast, Malloc-Free C++14 Json Parser and Encoder
github.com
github.com
I'm not sure how any JSON parser could avoid memory allocation without a difficult to use interface. Numbers, arrays, strings, and objects are unbounded by the JSON spec, so a truly malloc free library would need to provide a kind of streaming interfaces where things are returned in fixed sized chunks.
JSON wouldn't be my first choice for data storage in situations where I needed to avoid dynamic memory.
jsmn is pretty awesome, and contrary to this (std::string, std::vector) actually puts you in charge of allocations. It's also in C.
In a similar vein, and also possibly already mentioned:
- yayl (C): https://lloyd.github.io/yajl/
- microjson (C): http://esr.ibiblio.org/?p=6262
The C interface is purposely ugly but the C++ is much easier to work with.
I work a lot with very low-resource embedded systems and we like to use JSON host-side for structured data configuration, but convert to BSON (http://bsonspec.org) for target-side storage. This can be parsed and iterated with very low-overhead and no dynamic memory. Similar idea to protocol buffers but with JSON-like data.
Avoiding a build-time conversion would be nice, though, and this idea is interesting.
http://blog.llvm.org/2010/01/address-of-label-and-indirect-b...
A venerable interpreter implementation trick.
So it might be "malloc-free" for configuration files where the possible values are limited, etc. :)
I'm not sure that's entirely fair. Callback-based parsers like YAJL leave the application free to store the data in whatever data structure they want, or even to stream-process the input without storing in a data structure at all.
But regardless, the meta-programming approach described here is interesting and novel. Generating structure-specific parsing code is a well-explored area (for example, Protocol Buffers is designed entirely around this idea), but doing it as C++ metaprogramming is a novel approach (Protocol Buffers relies on compile-time code generation).
I don't actually understand how the object inspection and compile-time codegen works with this meta-programming approach; will be interesting to dig in a little deeper and learn more.
BOOST_FUSION_ADAPT_ADT (
xml_encoded<Description>,
XML_ATTR (string, "summary")
XML_TEXT (string)
)
BOOST_FUSION_ADAPT_ADT (
xml_encoded<Event>,
XML_ATTR (string, "name")
XML_SUBTREE (shared_ptr<Description const>, "description")
XML_ATTR (optional<unsigned>, "since")
)
The #defines for each macro are short, and beyond Fusion there's only a small support headerHere's a talk from CppConn describing a similar use case, but for binary data formats https://www.youtube.com/watch?v=wbZdZKpUVeg
https://code.google.com/p/chromium/codesearch#chromium/src/g...
The same goal - not parsing what's not needed - can be done with a conventional callback-based C code. You basically go through the json data, parse, say, a field and call the app asking "is this ok? shall I proceed?". If it's a yes, then you indeed proceed and parse out the value chunk and pass it to the app the same way. If it's a no, you either abort or skip over the value. The end effect is the same - parsing of an invalid json input is aborted as soon as the app flags the first bad entry; and unwanted fields are never parsed in full.
So I seriously doubt that this is a little more than a marketing spin of a proud developer -
This makes its performances impossible to match
in other languages such as C or Java that do not
provide static introspection.
I am fairly certain that vurtun's code [1] can match and most likely beat this lib's performance, with ease.A CHALLENGE!
So, er, who's up for it?
You could implement an analogue of this approach in Java. It's true that Java doesn't have language constructs that would let you do this as part of compilation, but Java has its ways. You could write an annotation processor to do this at compile time, or use a bytecode parser at runtime (this is yucky, but a fairly standard technique these days). Either way, the output would be a pair of synthetic classes which implemented the parser and encoder. A tool like this would be moderately laborious to write, but a straightforward matter of programming.
Is that true? Can classic C have default (optional) arguments?
void fun(int mandatory_arg, int optional_arg1 = 1, int optional_arg2 = 12, int optional_arg3 = 12);
No, this isn't legal ISO C. There are ways to simulate them though: https://gustedt.wordpress.com/2010/06/03/default-arguments-f...On a more general note, libraries that perform zero dynamic allocation (and instead require the library user to pass in memory) can be very convenient for systems programming where portability is a concern. For example, I use Intel XED instruction encoder/decoder, which is a library that performs no dynamic memory allocation. This allows me to use it in user space and kernel space without hijacking the malloc symbol.