Building a high performance JSON parser
dave.cheney.net
dave.cheney.net
There are specific problems like you can't use syscalls which must be single threaded (a royal pain for container managers, since setns() is), but binding to a library like simdjson wouldn't have this problem.
1. cgo calls have much higher overhead than regular function calls, so for small documents you'd likely lose performance rather than gain, and even for large ones depending how you read from the parser it might also be terrible, callback-based libraries are even worse as calling Go from C is even slower
2. concurrency can suffer a lot as a Ccall will prevent switching out the corresponding goroutine, locking out a scheduler thread
3. cgo complicates builds, especially cross-compilation
4. it also makes deployments more complicated if you were relying on just synching a statically linked binary
5. might have improved since, but used to be most of the built-in Go development tools couldn't cross the cgo barrier, and non-go devtools generally don't support go
See issues like this: https://github.com/golang/go/issues/19574
1. Moving "length := 0" above the for loop, since it's reassigned in all the needed cases.
2. To avoid having an extra "if whitespace[c]", Including the whitespace cases in the main switch statement, even if it means duplicating or moving "s.br.release"?
Or, using a switch statement vs a lookup ("whitespace[c]"), if it must be done.
3. In the switch statement, using multiple assignment (in most cases):
length = validateToken(&s.br, "false")
s.pos = length
//
s.pos = length = validateToken(&s.br, "false")
4. In the String and default cases, inlining the length assignment within the if statement.5. Returning "s.br.window()[:length]" in each case vs breaking out of the switch statement to return. Even though it's ugly, to avoid one step.
6. I'm curious if any performance could be gained by including more cases for common characters (A-Z,a-z, 0-9), to avoid using the default case. Testing if there is a penalty for using a default case vs more cases, even if it's ugly.
7. Including additional cases for exact values to avoid extra function calls to "parseString(&s.br)" or "s.parseNumber()".
8. I'm curious in some cases, if peeking at the next character with a nested switch statement, could avoid additional iterations or function calls to validate/release.
9. In the whitespace check, peeking for common JSON formatting patterns to avoid iterations. Such as 2 or 4 spaced json, a new line, followed by tabs or spaces etc. Or possibly establishing that the JSON is "probably2Spaced/probably4Spaced" and then peeking more efficiently?
I can see how you might choose some numbers to optimize for (1..10 for example) - but strings? You could of course do a frequency analysis of the test data - but would that help in general, beyond just cheating on the benchmark?
I guess you could try for "key" and "value", and maybe "id"? Possibly adding "email" and "name"?
Also from tfa regarding numbers:
> Scanner.parseNumber is slow because it visits its input twice; once at the scanner and a second time when it is converted to a float. I did an experiment and the first parse can be faster if we just look to find the termination of the number without validation, canada.json went from 650mb/s to 820mb/sec.
I don't know how the two compare as I don't really know where the overheads happen in the go version. Assuming the analogous case is "decoding into an interface{}" the simdjson port would be considerably faster.
The ASCII table is ripe for bit twiddling (I suspect it was organized according with that in mind). You may find bit patterns in whitespace chars.
canada.json --> 31 ms, 73 MB/s
citm_catalog.json --> 13 ms, 135 MB/s
code.json --> 17 ms, 113 MB/s
example.json --> 0 ms, 73 MB/s
sample.json --> 6 ms, 124 MB/s
twitter.json --> 6 ms, 114 MB/s
[1] https://gist.github.com/youurayy/18553475c5a9f81a17345cddeeb...how does that help when you want to implement a JSON-based protocol ?
E.g., an object's "foobar" field is always an array of int, and won't suddenly become a string from one invocation to the next.
It seems incredibly strange that we're somehow not leveraging this.
But this article isn't about those protocols. It's about JSON.
Here's a pure-JS solution: https://github.com/fastify/fast-json-stringify
If only… I had to figure out how to deserialize a JSON API where a field could be: a) an object X, b) an array of X, or c) a string (mapping to X.Y)
http://amzn.github.io/ion-docs/
https://developers.google.com/protocol-buffers
https://google.github.io/flatbuffers/
ProtoBuf requires a schema for the document to be known ahead of time and involves deserialization; Ion is more like JSON in that it's self-describing, but has a rich set of data types as well as annotations on fields to further describe meaning. FlatBuffers requires a schema and provides access to the without parsing/unpacking.
On the other hand, if you're building a service and specifying its interface, rather than just specifying a data format, then there are tools like Smithy, gRPC, and Thrift.
Do you mean "JSON with a schema" (there are multiple such) or "JSON decoding to static types" (without reflection).
JSON is popular because it's schemaless; storing and versioning schemas is its own nightmare of inconsistent and inefficient hacks.
However, in practice, JSON messages almost always have an implied schema. Figuring it out on the fly shouldn't be too hard - just apply the JIT and type inference technologies that compilers have already used for decades.
(Not a problem with an immediately obvious solution, but that's why we're here, no?)
Typed json already exists in various attempts eg tjson
Really I think the problem is that JSON is great when your schema is still in flux -- once it's public/stable, it makes sense to want to strictly encode the current schema.
The solution then is that you really want a way to trivially transition from "implied schema" to "strict schema" -- though a typed json won't help you too much because the larger trouble is figuring out what your schema is in practice. I guess you really want either a code analysis tool to determine it, or more likely an automated tool to take eg a list of json responses (probably based on a test suite/code coverage) and produce a typed json schema that you can simply drop into your public API
GraphQL goes in this direction by enforcing type checks, but is designed for web frontends...
I’ve profiled many applications that were completely and thoroughly bottlenecked on JSON parsing, even when care was taken to not be profligate with JSON parsing because everyone knows it is quite slow.
It might be the slowest part of the pipeline you control, or your network might be a local network where your throughput is 10G. Modern SSDs can do 3+GB/second. Even if you're aiming at 1GB/second, your json parser has to be able to do 0.1ms/MB to be able to keep with even a normal disk. If JSON parsing is fast enough, you can use it as an interchange format high performance tasks, and if it's your input and output format, you can use existing tools for non performance critical tasks that use json already!
Of course, the original comment is saying that most “apps” are I/O bound anyway (shall we assume web apps?). I think this is a lazy argument, or at least an ignorant/self-centered one — plenty of apps are not web apps running in an embarrassingly slow context like Django or Rails. For example, I work in digital forensics/cyber security, and we have to scan through TBs of logs (sometimes in JSON).
In my opinion this is much worse than reinventing the wheel.
Especially when the result is, in the end, slower.
You say it as if reinventing the wheel was a bad thing, but it is not. Technology advances by the continued reinvention of the wheel.
Since then C has spawned many algol derivatives and the likes of Ruby and Python drew heavy inspiration from Scheme.
I'm not sure when we'll be done or what constitutes "done" but a little churn feels reasonable as a side cost to finding innovation.
Being able to allocate only a few KB to ingest any JSON is a killer feature.
There is a valid reason to reinvent the wheel : in my case I had to do something similar in C99 for a SaaS federated search engine to have the lowest memory footprint possible.