Parsing gigabytes of JSON per second
arxiv.org
arxiv.org
If you want the flexibility of JSON, I'm not sure you'll end up with something massively different from gzipped JSON
[1] https://cbor.io/
[1] this won’t work for all languages as the OP parses json into a flattish object where fields may be looked up rather than some language-specific data structures (like js objects or python dicts or whatever)
No thanks.
Use Protobufs, Parquet, Avro, etc.
It seems like a notion that started in the early internet and just refuses to die.
Everything on the internet is minified and compressed at this point, so the entire idea of "human readability" left the station years ago. Yet I'll still see a primary complaint against the likes of http2+ being "it's a binary format! On NO!".
JSON seems like a similar relic. We use it not because it's fast, but because we like the idea that it's easy to decode (Even if in practice that almost never happens).
You are almost certainly passing these through tools (even built into the browser) that are doing extra processing to make it more readable.
Let's assume, for example, CBOR ends up taking over JSON. Do you not think browsers wouldn't have a CBOR parser to make it more readable?
This isn't bolt vs welding, this is bolt vs bolt with a washer. Yes there's a small extra step, but not some sort of insurmountable hurdle.
There's no reason a binary format couldn't be as easy to read as a JSON format. All you'd really need is a "binary->json" gui and you're off.
It would take a little effort to make such a tool and you could have it integrated into every browser.
It's not like something like that is unprecedented even, After all, there's no browser tool out there that's showing you the gzipped resource on a compressed endpoint. It's always already doing the step of gunzipping. binary->Json would similarly be just an additional step the browser could do in the tools.
Those on the cloud should compute how much HTTP headers count towards their traffic egress bill.
Everyone is fine with human-readable as long as it's in English.
Binary should be the default. If writing a binary to human decoder is too much for you, you're in the wrong job.
I don't understand that point. Is there an overhead in starting the parsing, which makes regular parsing faster unless you have a large JSON file? If not, why wouldn't you want faster JSON parsing?
Personally I’m using csimdjson in a project with 100s of TB of json to burn through. This data should not be json formatted, but migrating away from json would require modifications to hundreds of different systems.
I wish we could all settle on a binary structured format thats not limited to text encoding and more like JSON (key value) that basically everyone uses with real data types (efficiently encoded numbers, dates, binary , etc).
Text files would just be something like { contentType: "text/plain" content : "text here" } while an image could be { contentType: "image/jpeg", content: <real binary data not base64> } you could also add whatever other metadata you want and all of it is retained and easily parseable. Other more structured formats obviously wouldn't just be a blob for content and every system out there would have a structured binary format viewer/editor just like there are text viewers and editors now.
Sqlite is kinda used this way and has some nice properties like indexes and transactions but is relational instead of hierarchal which can be good and bad, not as straight forward to just view or navigate. Protobuf is another used quite a bit now, honestly I don't care just something everyone agree upon that can encode more structure efficiently but still can be easily inspected everywhere.
Doubt this will any time soon but it does seem inevitable in the long run that we figure out a way to send data between system and what a string, number, date etc is and stop having text encoding and escaping issues.
Although I do like ones that have some concept of a schema to reduce repeated key size and allow validation.
The idea has been around for almost 40 years in arguably more complete format than any of the above up to and including a standardised schema representation.
It's not the lack of a suitable standard that's holding it back.
It would have to be encoded somehow, because what if the "real binary data" included the byte 0x7D which is '}'
(which I have never used but really enjoy just for its charming website design alone)
That said, probably FlatBuffer would be even an optimized json parser, but json is crazy fast if the parser takes advantage of all modern processor optimizations.
It's funny how I feel questioning JSON here on HN is like starting a discussion about politics in a family diner.
I'm forever grateful to Google for demonstrating that all "the industry" is not fully committed to HTTP, JSON and scripting languages. It's honestly a relief to have encountered some sanity somewhere.
Oh god please help me, I just did it, I questioned JSON on HN!
Some folks in the thread are suggesting using a binary format, and that's certainly a good idea for some.
My business sells software that collects moderately sized (10-100TB/day) , and this data is collected as PROTOBUF. The format is great, and coordinating the backend and client side stuff with protobuf is bliss.
However, we store the data in various relational and document databases. Neither of those use protobuf...
Most importantly, it's business critical to be able to stream this data to data lakes where it can be read by humans. None of the options are going to support protocol buffers, you're either going to have to write a parser (impossible for some) or transform to JSON before ingest (fairly expensive due to some poor choices in how to represent various fields).
It was a sound technical decision to use protoc , and a terrible choice for the business.
I think it would have been much better to use avro, still benefit from schemas but push the work of marshaling JSON back to clients...
I mean, stuff like “double click on .protobuff file and it opens in notepad and looks like json or whatever. When you click save it gets serialized back to protobuff.
Here's a short comparison of serializing formats: https://drewdevault.com/2020/06/21/BARE-message-encoding.htm...
Yes, strings, bytes, and substructures all appear in the same in the wire format. Just like strings, bytes, and dates all appear the same in JSON. If you're trying to write a generic protobuf viewer like wodenokoto suggests, you can make reasonable assumptions about the contents 99% of the time based on the data, and show the user multiple options if you're not sure. There are already lots of tools that do this, but none very well integrated into mainstream development workflows.
For example, in the wire-format, a string and a sub-message are encoded as the same type (a blob - varint + sequence of bytes), but using the schema they are clearly interpreted differently.
Sure, it is possible to make a self-describing protobuf message which includes the schema in protobuf representation, but that is a special-case.
The slightly dirty solution is https://www.schemastore.org/json/, where the IDE looks up schema from a global registry, using a fileMatch pattern.
It allowed us to have a single model for storage in the DB, for sending between services, and syncing to edge devices.
But if you're talking about data lakes, like splunk, or various other similar systems - there's no standard way for writing a parser like this across all of them and you end up implementing the same thi ng a ton of different times.
PyPI: https://pypi.org/project/pysimdjson/
There's a rust port: https://github.com/simd-lite/simd-json
... From ijson https://pypi.org/project/ijson/#id3 which supports streaming JSON:
> Ijson provides several implementations of the actual parsing in the form of backends located in ijson/backends: [yajl2_c, yajl2_cffi, yajl2, yajl, python]
I can't think of a scenario where JSON handling is our bottleneck; almost all of the massive-data-handling tasks in my entire career have usually been handled by our database of choice.
Largest individual JSONs we've had to handle like on import tasks and such were about 300mb, which proved extremely easy to manage with stuff like JSONStream -I think we used a hand-rolled analogue back then but JSONStream is actually pretty cool, check it out-
What is the use case here? data lakes?
And yeah I just streamed the data thru the processor and avoid malloc at all costs and it works in seconds
I bet this method works even faster though somehow if they felt compelled to throw research money at all at it
If it's a drop-in replacement for other JSON parsers, I think this could have a huge impact. There is a lot of value to gain by optimizing the fundamentals
The main issues see had with our codec was memory allocation and the big number of "if" in the code.
I tried to use simdjson, but there were concerns at that time about portability and long term support. It's nice to see that in retrospective we were wrong.
Now you may say we should optimize this or that but I'm not an architect and I have no say, and it's just what I gotta deal with:(
I guess if you don't have a choice then it might become a bottleneck. Parsing 300 MB of data can be very slow.
But even with small JSON messages it can become a bottleneck, e.g. check out the Xi editor's issues. They thought it wouldn't be a bottleneck and then found that sound languages don't have insanely optimised JSON parsers like this. Boom. Bottleneck.
Daniel Lemire does a ton of performance research like this. He's an unusual case of a professor who publishes fully working source code.
There's a good chance that your database of choice uses some of his stuff or was influenced by his research.