JSON for Modern C++ version 3.10.0
github.com
github.com
Dear @nlohmann and contributors. Thank you for this library, I’ve used it in the past and appreciate your work and efforts.
The cJSON's C API or rapidjson's somewhat less ergonomic API has felt fine because I usually have read / write be reflective and don't write that much manual JSON code.
Specifically for the case where compile time doesn't matter and I need expressive JSON object mutation in C++ (or some tradeoff thereof), I think this library seems good.
Another big feature (which I feel like it could be louder about -- it really is a big one!) is that you can also generate binary data in a bunch of formats (BSON, CBOR, MessagePack and UJBJSON) that have a node structure similar to JSON's -- from the same in-memory instances of this library's data type. That sort of thing is something I've desired for various reasons (smaller file types with potentially better network send time for asynchronous sending / downloads (for real-time multiplayer you still basically want to encode to something that doesn't embed field string names)). I do think I may end up doing it at one layer above in the engine level and just have a different backend other than cJSON etc. too though...
Ironically, cJSON is _worse_ in my use case due to it not supporting allocators at all, as you wrote. Nlohmann fully supports C++ allocators, so it's trivial to allocate on SPIRAM instead of the limited DRAM of the ESP32. Support for allocators is why I often tend to pick C++ for projects with heterogeneous kinds of memory.
Also, Nlohmann/JSON supports strong typed serialization and deserialization, which in my experience vastly reduces the amount of errors people tend to make when reading/writing data to JSON. I've found simpler to rewrite some critical code that was parsing very complex data using cJSON to Nlohmann/JSON than fixing up all the tiny errors the previous developer made while reading "raw" JSON, parsing numbers as enums, strings as other kinds of enums, and so on.
Re: strong typed -- agreed. That's basically what I do with cJSON using a static reflection system in C++ (it recursively reads / writes structs and supports customization points). So it's kind of like using cJSON to give yourself a thin C++ layer. Agreed that a typed approach is more sensible than writing raw read / write code yourself (although that does make sense in some scenarios).
Here's what I use on ESP32 to push allocation to SPIRAM: https://gist.github.com/cibomahto/a29b6662847e13c61b47a194fa...
This is briefly how I use C++ Allocators with the ESP32 and Nlohmann/JSON (GCC8, C++17 mode):
I have a series of "extmem" headers which define aliases for STL containers which use my allocator (ext::allocator). The allocator is a simple allocator that just uses IDF's heap_caps_malloc() to allocate memory on the SPIRAM of the ESP-WROVER SoC.
I then define in <extmem/json.hpp>:
namespace ext {
using json = nlohmann::basic_json<std::map, std::vector, ext::string, bool, long long, unsigned long long, double, ext::allocator, nlohmann::adl_serializer>;
}
where `ext::string` is just `std::basic_string<char, std::char_traits<char>, ext::allocator<char>>`.
In order to be able to define generic from/to_json functions in an ergonomic way, I had to reexport the following internal macros in a separate header: #define JSON_TEMPLATE_PARAMS \
template<typename, typename, typename...> class ObjectType, \
template<typename, typename...> class ArrayType, \
class StringType, class BooleanType, class NumberIntegerType, \
class NumberUnsignedType, class NumberFloatType, \
template<typename> class AllocatorType, \
template<typename, typename = void> class JSONSerializer
#define JSON_TEMPLATE template<JSON_TEMPLATE_PARAMS>
#define GENERIC_JSON \
nlohmann::basic_json<ObjectType, ArrayType, StringType, BooleanType, \
NumberIntegerType, NumberUnsignedType, NumberFloatType, \
AllocatorType, JSONSerializer>
I am now able to just write stuff like the following: JSON_TEMPLATE
inline void from_json(const GENERIC_JSON &j, my_type &t) {
// ...
}
JSON_TEMPLATE
inline void to_json(GENERIC_JSON &j, const my_type &t) {
// ...
}
And it works fine with both nlohmann::json and ext::json.In the rest of the code, everything stays the same; I simply use ext::json (and catch const ext::json::exception&) as if it were it's default version, and it works great. FYI, I'm currently using nlohmann/json v.3.9.1.
NM. Looks like it's gcc8 which doesn't fully support C++17
Minusses: although I'd like a way to append a (C-)array into a JSON-array "at once" and not iteratively (i.e., O(n) instead of O(n log n)). Also, lack of support for JSON schema is .. slightly annoying.
I didn't realize that about the array complexity. Can you not just initialize a JSON array of N elements in O(N) (conceding a `.reserve(N)` style call beforehand if required)? rapidjson is pretty good about that sort of thing, cJSON's arrays are linked list so I basically think of its performance as at a different level and it's mostly about compile time for me.
It seems they're targeting C++11?
The string_view lookup is nearly done, but I did not want to wait for it, because it would have delayed the release even more.
I'm also working on supporting unordered_map - using it as container for objects would be easy if we would just break the existing API - the hard part is to support it with the current (probably bad designed) template API.
std::unordered_map has massive overhead in space and time. Maybe I misunderstand your idea, but outside of extremely niche use cases, std::unordered_map is basically never the right tool (unless you don't care in the slightest about performance or memory overhead).
https://stackoverflow.com/a/42588384/1593077
and the link there. Not sure that's what GP meant though.
And I don't like it when people use ridiculous hyperbole to describe what is in reality barely perceptible overhead, so it would be good if the person I originally responded to came out and defended his POV. Hopefully using actual numbers instead of agitated handwaving and screaming...
Pretty sure that makes it impossible to implement merge without pointer/iterator invalidation, which is a constraint imposed by the standard (https://en.cppreference.com/w/cpp/container/unordered_map/me...).
> it would be good if the person I originally responded to came out and defended his POV
See https://probablydance.com/2017/02/26/i-wrote-the-fastest-has... if you want numbers, or for example this talk by the same author https://www.youtube.com/watch?v=M2fKMP47slQ.
The issue is that the implementors won't change their implementation of unordered_map as it's an ABI break, to say it simply. It could be better. But also, there are other tools like flat maps/open addressing that are not done in the std library.
Not sure what your rationale is for saying it’s only useful for extremely niche use cases. I would be curious to know why you think so.
See https://www.youtube.com/watch?v=M2fKMP47slQ for a more comprehensive discussion, or the sibling threads here.
std::map is only useful if you need the data to be ordered, otherwise std::unordered_map is the right default choice.
> std::unordered_map has massive overhead in space and time.
Yes, std::unordered_map uses a bit more memory, but please show me some benchmarks where it is slower than std::map - especially with string keys!
BTW, for very small numbers of items the most efficient solution is often a plain std::vector + linear search.
This is why flat_map is a great solution. It's basically an ordered vector with a map-like interface, but without the memory allocation (and cache locality issues) for each node. Until you start getting into thousands of elements it is probably the best choice of container.
And ordered maps fall somewhere in the middle.
They might make sense if you're creating massive amounts of medium sized maps.
Or if you need range searches etc.
It's all compromises, there are no hard rules.
That's sometimes ok, particularly if you don't care about performance (beyond asymptotic behavior). But if you do, you'd probably be well-advised to try out hash map implementations that actually don't leave performance on the table by design.
You literally suggested to (almost) always prefer std::map...
> Every node has to be stored as separately allocated linked list entry.
But it's basically the same for std::map. Both containers have similar design constraints. I am aware that std::unordered_map uses more memory than std::map, but why should it be slower in the general case? After all, hashtables trade memory for speed.
> But if you do, you'd probably be well-advised to try out hash map implementations that actually don't leave performance on the table by design.
I agree. But the same is true for std::map (someone has already mentioned flat maps).
std::unordered_map is typically the go-to hash map on any given C++ day. You’re probably fine with std::map outside of large n
EDIT: being pedantic
But to be honest, I have not yet played around with PMR.
For example, C++/WinRT is a C++17 library, yet plenty of samples use C style strings and vectors.
Do you find it unreasonable to point out that what's been advertised doesn't match what's being offered?
https://github.com/simdjson/simdjson/blob/master/doc/basics....
That's a weird thing to say. Doesn't that depend on what hardware it's running on?
these projects are just commendable!
There are lots of other libraries, and surely some faster ones, but I haven't felt any need to change. I was nervous when I got an inexplicable overload ambiguity in xlclang++ for AIX, but it went away when I grabbed a newer version of the library.
away from SFO, "Modern C++" means C+11 and above.
mr. lohmann, much thanks for an easy to use, robust, complete, and performant library provided for the world to use.
every piece of software has tradeoffs, you have developed a library that is not perfect, rather it is the one people use.
nicely done.
IIRC, you would need to modify the statement to "find_package(Boost REQUIRED COMPONENTS filesystem)" if you want to use boost dependencies that need a separate .dll/.so (in this case boost::filesystem). However I rarely need to use these, so I may be wrong.
It gets far more complicated with C++11, since you also need a ton of other boost modules there.
For more Details you can read the Readme of it. https://github.com/boostorg/json
With proper CI it's not a big deal
https://github.com/cpp-pm/hunter
There is a bit of a learning curve, but it's the only dependency manager that does things right (all from within CMake)
wget boost_1_77.tar.gz
tar xaf boost*
cmake ... -DBOOST_ROOT=.../boost_1_77
Works for the huge majority of boost ; most parts that required a build were the parts that were integrated in c++ like thread, regex, chrono, filesystemThe second paragraph talks about setting a #define before including a header, as if this is still the 90's.
I'm not shitting on the library in general, I'm saying the second paragraph disproves the title.
It can be a trade-off, but becoming less and less so. I'll admit to not looking at the implementation, but when you sell it as "modern" then it's not OK to do it for performance, for example.
And almost every single use I've seen "for performance" actually has zero or negligible performance impact.
Moreover, because of `#define`s, the effect of including a file can be entirely different, so a compiler can never say "Oh, I know what's in this include file already, I don't need to include it again in other translation units".
Of course #define's have many uses, and at present cannot be avoided entirely, but C++ has been making an effort to have other mechanisms put in place so as to obviate the use of #define's
In particular, look for CppCon talks about Modules.
In a nutshell: Using compile-time facilities which are within the language itself. Compilers will expose information about the platform, about themselves, etc.
A couple of examples:
* https://en.cppreference.com/w/cpp/utility/source_location
* https://en.cppreference.com/w/cpp/types/endian
> can you build big projects across clang/g++/intel without preprocessor workarounds now?
Well, first of all, s/now/when C++20 is fully and widely adopted.
Still, a good question. Possibly not? I'm not sure. But you would need a lot less of them than previously.
> has the number of viable compilers in use dropped
Not really.
> and compatibility amongst those that live on increased?
Well, if they're compatible with the standard, that says more than it used to.
> how about stuff like debugging/instrumentation?
Maybe if constexpr(in_debug_mode) ? Not sure. I'm not on the committee...
#define definitely is not "modern". If nothing else it'll set you up for having problems modularizing it in the future.
I'm not talking about "The simplest option that comes to mind", but "modern C++".
OK, if you're getting philosophical, you're welcome to to think of the -D switch as being a type of source code generation.
But there's a big practical difference. target_compile_definitions is one line in your CMakeLists.txt (two if you add a cache variable for it) and everyone can understand it instantly. Whereas creating a myappconfig.hpp.in and using configure_file() and #including it from all your actual source files (yuck) is much greater mental overhead both to create and maintain in practice. I know which I'd prefer.
> I'm not talking about "The simplest option [available in modern editions of C++] that comes to mind", but "modern C++".
I added a little to your quote there, forgive me. But I hope it clarifies that they're the same thing. If modern editions of C++ don't add any facilities that make it easier to do conditional compilation (chosen from your build system), then #define IS "modern C++". The fact that there are other ways of doing it, which you consider cleaner but are clearly impractical, is irrelevant.
(I'm just talking about this particular usage where you #define one constant the same way everywhere. I've seen horrific tricks where you #define one way then #include, then #define something and #include again (yes there are no include guards). Clearly that isn't forgivable, whatever version of C++ you're using.)
For the record, Rust has a nicer option for conditional compilation: feature flags. In principle a future version of C++ could adopt something similar, at which point #define really would disqualify that code from being "modern C++". But clearly that's not available in the most recent standard, or even planned in future standards AFAIK.
This also means that there's no way to create a library that supports debug and not debug. And of course no ability for the user of the library to select this at runtime.
Yes, in the world of everything being compiled statically every single time (ala Google's monorepo) this can work, but code generation creates a HUGE number of gotchas, many of which are subtle.
> But I hope it clarifies that they're the same thing.
There're not, though.
void* is "the simplest option that comes to mind", but where a template solution is available, safer, and clearer, it's not "modern C++" to have code with void* all over the place.
"Easier" has nothing to do with it. There are many features in this fine library where the authors have put in effort to make it nicer, instead of just doing the easiest and quickest to implement.
You're just playing silly semantics with "simplest" here. I meant that #define is genuinely the best overall solution to this particular problem, for the reasons I already gave. Whereas in void* vs templates, templates are genuinely the best overall solution.
I have seen many many well-intentioned "but surely this time it's the best solution", where no, it's just a time bomb.
> Speed. There are certainly faster JSON libraries out there. However, if your goal is to speed up your development by adding JSON support with a single header, then this library is the way to go.
It is essentially admitting to being slow(er). We are not supposed to pay for code modernity with performance. In fact, it should be the other way around.
Does someone know the innards of this library as opposed to, say, jsoncpp or other popular libraries, and describe the differences?
You're making the false assumption that there is a compromise to be made - but there isn't. A modern JSON library should be the fastest (or about the same speed as the fastest), and vice-versa.
So what? The format is so simple that the inconsistencies doesn't even matter and hell the format can be inferred from the data and that can even be automated. Unlike JSON/XML, most software can automatically make sense of CSV; JSON needs a developer to understand it to transform it into workable data.
> CSVs often begin life as exported spreadsheets or table dumps from legacy databases, and often end life as a pile of undifferentiated files in a data lake, awaiting the restoration of their precious metadata so they can be organized and mined for insights.
CSVs are the de facto way to export spreadsheets and if you have to do this then the problem are that you are still using spreadsheets in your data flow, not CSVs. Same goes for legacy systems.
Also CSVs comes from places where having more sophisticated data serializers is near impossible and the time to code, compute, parse and analyze CSVs is trivial compared to any other format out there.
> I’ve spent many years battling CSVs in various capacities. As an engineer at Trifacta, a leading data preparation tool vendor, I saw numerous customers struggling under the burden of myriad CSVs that were the product of even more numerous Excel spreadsheet exports and legacy database dumps. As an engineering leader in biotech, my team struggled to ingest CSVs from research institutions and hospital networks who gave us sample data as CSVs more often than not.
Of course if you are working in biotech then you'll work with scientists (horrible coders most of the time) and legacy systems because it's a domain that's very adverse to change that might break stuff. And it's not the CSV format's fault if it's used as the wrong tool for the job.
> Another big drawback of CSV is its almost complete lack of metadata.
Again, emphasis on the right tool for the job. For simple tasks metada are not needed.
https://news.ycombinator.com/item?id=28221654
Your quote "The CSV format itself is notoriously inconsistent" comes from this page which the thread discusses:
If anyone out there has never clicked on the wrong button, let them cast the first stone.