Deserializing JSON Fast (2020)
blog.datalust.co
blog.datalust.co
The library is still fairly low level, and our abstraction has become complex. Still, the best performance approach we’ve been able to find.
I read once, a custom library written in JS will probably never be as fast as the JSON parser shipped in browsers.
You _can_ do it outside the engine but it'd be tricky and you'd have to either write a lot of code _or_ to trick the engine to reuse the JIT machinery for your parser.
Protocols with schemas are much much easier to parse efficiently so I suspect it'd be as fast in library written JS.
It's insanely faster to parse binary into js object without going through json at all than parsing string json, and we transformed a dumb application from slow death to "oh wow no need to optimize further, all perf problems are now fixed" simply by switching the serialization format (faster to download, faster to deserialize, faster to access).
It even helps the backend to serialize faster than writing enormous ASCII strings just to talk to the browser memory structure. Dont forget these json format are just better xml and still made for humans. Machines can work well with structures closer to the memory layout, for instance sending the header once and subsequent objects in ordered attributes with separator is closer to how an object is stored and accessed by JS.
I suppose a mistake we all make after years of JSON is thinking it represents objects for the browser but an object is always a byte array with an offset per attribute descriptor. So send the attribute descriptors and their type (which will describe a size too) first and then just fill.
Doesn’t this depend on the language? Plenty of languages implement “objects” as a hash table underneath, with properties just being keys. In fact I think most JavaScript engines even work this way, at least until some amount of JIT optimization happens to start assigning static offsets to known keys (I believe this was one of v8’s advancements when Chrome first came out.)
I was surprised I couldn't find one with a quick search.
I suppose there's other stuff that's being done that could be suboptimal, but that's the key difference.
> Why not use Protocol Buffers, or .. ?
Protocol Buffers is indeed relatively similar to FlatBuffers, with the primary difference being that FlatBuffers does not need a parsing/ unpacking step to a secondary representation before you can access data, often coupled with per-object memory allocation. The code is an order of magnitude bigger, too. Protocol Buffers has no optional text import/export.
That said, in my experience the real cost of JSON parsing turns out to be constructing the in memory representation of the sources. I'm sure JSON parsing for the purpose of deserializing specific types would be faster as you might be able to avoid some allocations (e.g. imagine a vector type `struct Vector { float elements[4]; }` a generic JSON representation would result in two allocations)
Allocations, have been a huge win. Parsing to a vector in many of the JSON benchmarks benefitted from reserving the memory up front, even if not known. In testing the cost of them, it was around 50% of the bench time. Using a bump allocator to test doubled the perf when I did it.
But actual allocation ends up being killer in the JSC parser. I even tried using structure caching to speed up object creation, but it simply did not seem to matter: either there’s too little data to amortize allocations for the necessary data structures, or you’re parsing enough that simple allocations and GC start to happen (and code in c++ can’t optimize the allocations as much as the JIT can)
With C++, there are generally just less allocations, but they probably cost more, not always and I haven't measured malloc/new for small things where they often have a pool setup before going to mmap. They still lock though. But one can often just guess when the number isn't known without much cost on most systems. Choose a size like 1kb and remove the smaller up front allocations, or if the size is known reserve that.
Woah, that's impressive.
I wonder if `str` is maybe an antipattern for a lot of JSON use cases where you can perform utf8 validation lazily or avoid it altogether for data you don't need.
Given a service architecture with mTLS, and so long as you aren't doing anything sensitive based on the data, I could see an approach like this being valuable.
That said, I also wonder if it's worth pushing JSON into use cases like this to begin with?
Are there other concerns that cause JSON to make sense?
1. The one used for ingesting data.
2. The one used for storage.
3. The one used for exporting data.
There is no reason for all three formats to be the same, and there are good reason for them not to be: the actual use cases are requiring very different properties from the data format. Since you'd expect the average log record to be ingested once, processed a gazillion times, and returned as output <<1 times, any data conversion costs would get amortized over a lot of processing. Even minor efficiency gains in the at-rest storage format should pay off quickly.
Your answer is addressing format 1: the clients are using JSON anyway, so it makes sense for the ingestion format to be JSON. That's making the bet that any data conversion costs can't be paid back over the lifetime of the data. But they are already not using the raw client input as the storage format! At a minimum they're validating / minifying / canonicalizing the JSON, so they're already paying the cost of doing a format conversion.
(But all of this is obvious, and clearly the people writing this service are smart and would have thought of it. So it seems like there's a non-obvious answer for using minified JSON for storage at rest.)
So this correlates with some of the findings the post has. One optimization that should show great promise is that if one knows that they have an array of some types like numbers/unsigned int/strings, that can greatly simplify and allow for SIMD'ification of it.
There are several factual inaccuracies in the talk. For example on-demand or in-situ parsing has been a thing especially for many high-performance C++ JSON parsers, and they are exactly designed for the constraints described in the talk (no needs for parse tree, in-place mutation or cross referencing). So everything before the specialized integer array parsing can be mostly automated without heavy hammers like SIMD---those parsers would likely use that, but I don't care. You still need to recognizes those constraints in the first place but you don't always need to write your own code for making use of them.
He did mention that dependencies are or can be liabilities, yet failed to state that more codes and contexts are also liabilities. It is correct to recognize implicit contexts that may be useful for the optimization, but it doesn't mean that every context should be utilized. For example later examples assume x86(-64), which is increasingly becoming a bad decision in general thanks to ARM servers. It is a balancing act to choose which context to utilize and which to ignore (so that the code can remain more generic for later uses or easier to understand) and the talk is mostly silent about this aspect.
He also fails to mention testing, which is the main reason I believe the OP is a better substitute for the talk. You can't exactly compare your shiny optimized code with an older version of it, because you have distilled implicit contexts into the new code so they can be functionally different. You need to somehow make those implicit contexts into explicit contracts, probably using some other codes that are not continuously run, documentations (which everyone seem to overlook!) or folk, eh, collective team knowledge. This is a hard problem by its own and again the talk is pretty much oblivious about that.
But porting SSE2 to Neon is actually pretty easy -- if you use https://github.com/DLTcollab/sse2neon, IME it's very easy to do incrementally (or avoid or postpone indefinitely, depending on your needs).
i'm wondering why the hell its not a faster format? json is awful for storing ... anything, but a good compromise given that "everyone knows what it is"