Newline Delimited JSON
ndjson.org
ndjson.org
> This page describes the JSON Lines text format, also called newline-delimited JSON.
And then I saw this at the bottom of the ndjson.org page:
> Site forked from jsonlines.org
So is this a fork of the site only, or a fork of the jsonlines standard?
There is similar issues in both projects:
extension:jsonl 62K results [1]
extension:ndjson 41K results [2]
[1]: https://github.com/search?q=extension:jsonlOnly incidentally. There's literally nothing about the format, and I hesitate to give it the implied gravitas by calling it a format, that makes the JSON part important in any way. The lines are just individual messages. JSON happens to fit, but so do many other syntaxes.
What's an array? It's an abstract sequence of items. What's a delimited stream? It's a sequence of items. Do the lines need to be JSON? Who the hell cares. If you're using JSON then they're JSON. If you aren't, then they aren't.
Honestly, the most important part of their description is "for passing messages between cooperating processes", because JSON is terrible for blindly consuming messages without cooperative explanation from the sender (http://seriot.ch/parsing_json.php).
Given that you can't actually do anything with JSON data (or most data, JSON is again not extremely special in this regard) without a concrete explanation from the sender of the expected interpretation for every field, can you tell me what you think the important difference is between "json messages delimited by newlines" as proposed here with its own custom mimetype but without any other information and "newline delimited messages that can be parsed with a json parser"?
They'd have been better off using the ASCII record separator character instead of newline. That's both already a standard and also doesn't force escaping to prevent collisions with normal text.
This is as useful/useless as JSON itself.
Besides, you do still need to deal with the situation where the character shows up in your data; using an uncommon codepoint just makes it much easier to ignore until it blows up in your face.
JSON doesn't support binary data without conversion to string first.
It is for this, but if we're already declaring a restriction to JSON then we're likewise declaring a restriction to unicode. All delimited data requires escaping if your delimiter can be part of a no-delimit sequence. At least a character explicitly intended for delimiting isn't part of any no-delimit sequences representable in unicode whereas \n is.
U+1E is a Unicode code point, and therefore requires escaping inside of a ‘no-delimit’ sequence of Unicode characters.
You know how C++ lines end with semicolons (or how Javascript lines _can_ end that way if you want)? And how that means you can put all of the instructions in a function on one hyperlong line? That's compactible. Any format where you are able to smush all the lines within a functional group together into one long line is compactible.
You can use the ASCII record separator character instead of newlines if you prefer. Separating records without colliding with newlines is literally what it's for.
See "whitespace". In order to be used in ndjson, valid json must have its newlines removed.
1. Records separated by ASCII "\n" characters, wherein each record cannot contain an "\n" character
2. Each record is UTF-8 text in JSON format
3. The schema of each record is <something>
ND-JSON takes care of the first two. You would still have to use JSONSchema for the third.
This is not very different from using JSON without the "ND" part. Except:
1. Now there's the "ND" part, and I have something standardized I can use for quick and easy understanding/agreement when collaborating, especially outside my organization.
2. Plain JSON supports exactly 1 record per stream or file. ND-JSON supports any natural number of records, or an infinite number of records. That's useful to me.
I really don't understand the resistance here. Nobody is trotting this out claiming its some hot new thing. It's an unoffensive standardization of a pretty basic data format that people have been using for 10+ years.
If you have no use for it, then just ignore it. Consider that the people using it might not be morons; they might just have different needs that are different from yours.
Unless you wrap it in an array, which is all this is doing.
A JSON array is exactly 1 "record", in that it must be parsed entirely before it can be interpreted as anything other than a string of text.
Moreover, you can write a ND-JSON file where each record is an array. The following is valid ND-JSON:
["a", "b", "c"]
["u", "v"]
["x", "y", "z"]Streaming JSON parsers do exist, but they’re rather specialized and harder to use than JSON.parse(). The advantage of ND-JSON is that you can get the same effect in a much simpler way.
{
results: [
{ ... },
{ ... },
{ ... }
]
}ndjson seems to be about how to format the content of the stream, not about how to multiplex multiple streams over a single channel. Use http2, zeromq, raw tcp or whatever else you like for that.
[ { ... }, { ... }, { ... }, ]
json_decode(“[“.implode($lines,”,\n”).”]”)
In general I use YAML for human-edited configuration files, because they generate clean diffs, and JSON for machine-to-machine communication, because of the wealth of obscenely optimized parsnips for every language.
I like YAML, and I think people hate on it more than it deserves (at least, in its current form). And yes, you can stream YAML documents by concatenating them with document separators. But ND-JSON/JSONLines/RFC-7464 is mostly a machine-to-machine format anyway.
Maybe we should be able to customize the delimiter? The media type "application/stream+json; separator=\n" would be JSONLines/ND-JSON, "application/stream+json; separator=\x1E" would be RFC 7464, "application/stream+json; separator=\n...\n" would theoretically be valid YAML, etc.
((lorem ipsum ..)
(quorat est demonstrandum ...)
...
...)
Only major line-breaks between top-level elements.In my day job we ndjson pretty extensively and it's been pretty good to us. It has far better characteristics than passing a giant json array of objects (you have to hit the closing bracket on the array to know if it's valid)
If you're really into hot having commas, maybe you can do a post-processing step with sed or in-memory string replacement to add/remove "," and \n? Then you're still getting the same JSON experience under the covers where programs process it?
Doing so would complicate adoption, since parsing is essentially just:
for line in file:
obj = json.loads(line)
> This is an 8+ year old spec and I don't think it has caught much adoptionIt's seen huge adoption, essentially becoming the standard logging format for many Node JS applications.
Usually I've heard it referred to as JSON Lines: https://jsonlines.org
There's also an RFC (7464) for more or less the same thing, except using ASCII record separators instead of newline characters: https://datatracker.ietf.org/doc/html/rfc7464
And yes, I have used this at $JOB and I intend to use it in future ${JOB}s. That's adoption.