Absolute highest performance is rarely the highest priority in designing a system.
Of course we could design a hyper-optimized, application specific payload format and code the deserializer in assembly and the performance would be great, but it wouldn’t be useful outside of very specific circumstances.
In most real world projects, performance of Go and JSON is fine and allow for rapid development, easy implementation, and flexibility if anything changes.
I don’t think it’s valid to criticize someone for optimizing within their use case.
> I feel like it should deserve a mention of what the stack is, otherwise there is no reference point.
The article clearly mentions that this is a GopherCon talk in the header. It was posted on Dave Cheney’s website, a well-known Go figure.
It’s clearly in the context of Go web services, so I don’t understand your criticisms. The context is clear from the article.
The very first line of the article explains the context:
> This talk is a case study of designing an efficient Go package.
As always, try to remember that people usually aren't writing (or posting their talks) specifically for an HN audience. Cheney clearly has an audience of Go programmers; that's the space he operates in. He's not going to title his post, which he didn't make for HN, just to avoid a language war on the threads here.
It's our responsibility to avoid the unproductive language war threads, not this author's.
[{"k1": true, "k2": [2,3,4]}, {"k1": false, "k2": []}, ...]
You can amortize the overhead of the keys by turning this from an array of structs (AoS) into a struct of arrays (SoA): {"k1": [true, false, ...], "k2": [[2,3,4], [], ...]}
Then you only have to read "k1" and "k2" once, instead of once per record. Presumably there will be the odd record that contains something like {"k3": 0} but you can use mini batches of SoA and tune their size according to your desired latency/throughput tradeoff.Or if your data is 99.999% of the time just pairs of k1 and k2, turn them into tuples:
{"k1k2": [true,[2,3,4],false,[], ...]}
And then 0.001% of the time you send a lone k3 message: {"k3": 2}
Even if your endpoints can't change their schema, you can still trade latency for throughput by doing the SoA conversion, transmitting, then converting back to AoS at the receiver. Maybe worthwhile if you have to forward the message many times but only decode it once.Compare it with XML for example, which is a nightmare of complexity if you actually follow the spec and not just make something XML-like.
We have some formats which try to walk the boundary between safe/universal and fast like ASN.1 but those are obscure at best.
And whatever language you're writing in, you usually want to do what you can to maximize performance. If your JSON input is 500 bytes it probably doesn't matter, but if you're intaking a 5 MB JSON file then you can definitely be sure the performance does.
What more do you need to know about "the stack" in this case? It's whenever you need to ingest large amounts of JSON in Go. Not sure what could be clearer.
This way they only paid the parsing tax (decoding doubles, etc..) if the user used that data.
You hit the nail on the head
The first line of the article explains the context of the talk:
> This talk is a case study of designing an efficient Go package.
The target audience and context are clearly Go developers. Some of these comments are focusing too much on the headline without addressing the actual article.
Always nice to be in control over memory :)
Having done this myself, it's a massive cheat code because your bottleneck is almost always i/o and memory mapped i/o is orders of magnitude faster than sequential calls to read().
But that said it's not always appropriate. You can have gigabytes of JSON to parse, and the JSON might be available over the network, and your service might be running on a small node with limited memory. Memory mapping here adds quite a lot of latency and cost to the system. A very fast streaming JSON decoder is the move here.
That’s not something I’ve generally seen. Any source for this claim?
> You can have gigabytes of JSON to parse, and the JSON might be available over the network, and your service might be running on a small node with limited memory. Memory mapping here adds quite a lot of latency and cost to the system
Why does mmap add latency? I would think that mmap adds more latency for small documents because the cost of doing the mmap is high (cross CPU TLB shoot down to modify the page table) and there’s no chance to amortize. Relatedly, there’s minimal to no relation between SAX vs DOM style parsing and mmap - you can use either with mmap. If you’re not aware, you do have some knobs with mmap to hint to the OS how it’s going to be used although it’s very unwieldy to configure it to work well.
The latency comes from the fact you need to have the whole file. The use case I'm talking about is a JSON document you need to pull off the network because it doesn't exist on disk, might not fit there, and might not fit in memory.
I have. Many times. There's definitely not a 100x difference given that normal file I/O can easily saturate NVMe throughput. I'm sure it's possible to build a repro showing a 100x difference, but you have to be doing something intentionally to cause that (e.g. using a very small read buffer so that you're doing enough syscalls that it shows up in a profile).
> The latency comes from the fact you need to have the whole file
That's a whole other matter. But again, if you're pulling it off the network, you usually can't mmap it anyway unless you're using a remote-mounted filesystem (which will add more overhead than mmap vs buffered I/O).
I’ve seen this referred to as a pull parser in a Rust library? (https://github.com/raphlinus/pulldown-cmark)
Is it true that mindless bashing Go became some kind of a new religion in some circles?
> This package offers the same high level json.Decoder API but higher throughput and reduced allocations
Even just "Json" is problematic here as wire protocol for absolute performance no matter what will be programming language.
The only way to nudge the needle is to start exchanging direct memory dumps, which is what ProtoBuff and the like do. But this is clearly only for very specific use.
code maybe simple, but you have lots of performance penalties: resolving field keys, you need to construct some complicated data structures through memory allocations, which is expensive.
> to start exchanging direct memory dumps, which is what ProtoBuff and the like do
Protobuff actually is doing parsing, it is just binary format. What you describing is more like Flatbuffer.
> But this is clearly only for very specific use.
yes, specific use is high performance computations )
Plus no one makes people use objects in JSON. If you can send a tuple of fields as an array... then send an array.
but you need some logic which will check that key is what you want to process, and also do various type transformation (e.g. json string to long).
> Plus no one makes people use objects in JSON. If you can send a tuple of fields as an array... then send an array.
then this shouldn't be called JSON parsing anymore.
As for whether you must stuff everything in objects for it to be JSON, I mean that makes no sense. JSON arrays are also JSON. And JSON scalars are also JSON.
If this argument requires we deliberately go out of our way to be dumb, then it's a bad argument.
I'm writing a platform which has a JSON-like format for message exchange and I realized early on that in serialized data, maps are at best nominal. You process everything serially (hence the word serialization). It's a stream of tokens. The fact some tokens are marked as keys and some as values is something that can be useful to communicate intent, but it doesn't mandate how you utilize it.
Everything else is just prejudices and biases, such as "maps take resources to allocate". JSON doesn't force you to build maps when you have an object. Point in the spec where JSON mandates how you must store the parsed results of a JSON object. Should it be a hashmap? A b-tree? A linked list? A tuple? Irrelevant.
RAM on my computer works with pages loaded into CPU cache.
Rants, personal attacks, handwavings are ignored.
I didn't say "RAM pages", well vectorrized code can dump MBs of data preemptively into cache, instead of reading "next keyval sequence" each time.
> And this entire detour has no relevance to what I said in the first place,
or you just don't understand such relevance.
Hashing a string as you read it from memory and jumping to a hash bucket is not an expensive operation. This entire argument sounds like some kindergarten understanding of compute efficiency. This is not a 6052.
What are your numbers for your cool json serializer?
And you're currently telling me the bottleneck is memory.
I also said you don't need to parse a JSON object to a hashmap or a b-tree. The format suggests nothing of the sort. You can hash the key and fill it into a symbol slot in a tuple which... literally only takes the amount of RAM you need for the value, while the key is "free", because it just resolves to a pointer address.
Additionally, if you have a fixed tuple format, you can encode it as a JSON array, thus skipping the keys entirely. None of that is against JSON. You decide what you need and what you don't need. The keyvals are there when you need keyvals. No one is forcing you to use them where you don't need them at gunpoint.
I have a message format for a platform I'm working on, it has a JSON option, just for compatibility. It doesn't use objects at all, yet (but it DOES transfer object states). Nested arrays are astonishingly powerful on their own with the right mindset.
Not memory, but virtual pages implementation in linux, which is apparently single threaded and doesn't scale to high throughput. There was a patch to fix this, but it didn't make to mainline: https://lore.kernel.org/lkml/20180403133115.GA5501@dhcp22.su...