Documentation for the JSON Lines text file format
jsonlines.org
jsonlines.org
Recently (this year) I was flicking through a repo and saw a jsonl extension and immediately guessed what it would contain. It all just fits nicely.
- ignore any line which is exactly "[" or "]" - ignore any trailing commas
This lets you consume your API content both as a JSONL stream, and as a JSON array
not sure if it's that useful though
- allowing optionally enclosing the data into an array;
- allowing comma as an optional separator between objects;
- allowing JSON objects without a newline separator;
- properly escaping Unicode newlines and / characters inside string literals for JavaScript compatibility;
- allowing binary data inside string literals;
/(\]|\})\n|(.")\n/
So in json you are simply putting in the magic character sequence: "\n" in place of the real newline character. Not unlike how many compilers will auto translate "\n" sequences you put in strings into the actual newline character.
I believe most json encoders will automatically convert newline characters into the two character sequence, so when you put it all together it goes: character sequence "\n" typed into source file -> compiler converts newline character into actual an actual newline character -> json encoder converts newline character back into "\n" sequence.
With all this in mind, since regex is basically doing the same thing as a compiler, treating "\n" sequences as a search for a newline character, you would need to escape the backslash so it is interpreted literally, so "\\n".
This gets even weirder depending on the tool/language you are doing this in though, like if you are writing out a regex within a string, the regex engine might not need to have any special logic for "\n" since the compiler is already handling it! In which case: "\\n" is an escape for the compiler not the regex engine.
TLDR: You just need to escape it: "\\n"
s = "foo\nbar\nbaz"
# 'foo\nbar\nbaz'
len(s) # 11
j = json.dumps(s)
# '"foo\\nbar\\nbaz"'
# \n is escaped as \\n
len(j) # 15[1] https://ecma-international.org/wp-content/uploads/ECMA-404_2...
Don't do this. Don't encourage people to use a value that's not registered yet. Say "x-jsonl", and you're OK.
The best argument against this format to me is that it's possible to just use JSON itself for this by wrapping all of the values in an array. I'm not totally convinced that the quality of life improvements of not needing to add commas along with the single pair of square braces necessitate an entirely new file format, but realistically I think I'll just end up following whatever happens to be supported by the tooling I use for a given project. If "JSON lines" became widely used overnight, I wouldn't be bothered by it, but if I never end up seeing it supported by anything, I doubt I'll miss it.
If you have so much data that you need to stream it, you probably shouldn't be using JSON, though...
Why not? I have a bunch of large/real-time files I'd like to read piece by piece and json is a convenient format for parsing it to something useful on the client. I understand using other, more efficient formats, but I think just wanting to stream the data isn't enough of a reason to move to them.
If you have so much data (e.g. 8GB+, yes my workstation "only" has 8GB of RAM) that it _must_ be streamed, in my personal opinion you should opt for a better format than JSON. It's likely, too, that with a more specialized/space-efficient format, you may no longer need to stream the data!
Using hidden frames for jsonp-like request and callback support, that was the fun stuff. Don't miss it.
Also don't miss XML or SOAP.
Has it occurred to you that maybe there's a reason people are using jsonl instead of an old existing standard like csv? Not just that they're such dummies they never heard of csv and would take a few years to rediscover it, but an actual reason?
Hey, you're the one who chose to recommend scanf, and claimed it would take a few years to rediscover it. I would generally assume people who read hacker news are smarter than that, and are fully aware of the security problems of scanf, and the weaknesses of csv, and are probably not crazy enough to write bulk streaming data parsers in C. And no, s.split() doesn't make me feel safer, or even do rudimentary type conversion like scanf.
Pro-tip: use a comma in your password to pollute the csv when your password gets stolen and published on the dark web.
Wobbly format? Sure. But if the data are purely internal, why not opt for convenience?
> Tactile theme by Jason Long
and the format extension is .jsonl
JSON: /ˈdʒeɪˌsɒn/ (ɒ like the o in dog, and secondary stress on the “son” part)
Jason: /ˈdʒeɪsən/ (ə, schwa, like the first a in bazaar)
For example, Datadog doesn't ingest JSONL, so if you have JSONL log files you'll have to convert it to JSON first, you can't just use sendfile to upload your log in chunks, just wrapping it with [ and ] (which can be done with ordinary writes before and after sendfile.)
Find out more at https://nestedtext.org/en/stable/
[
{"name":"Bob"},
{"name":"Alice"},
{"name":"Alison"}
]
You're already working with json anyway so you could just stay in that world rather than mixing new lines and json.https://cloud.google.com/bigquery/docs/loading-data-cloud-st...
So instead of INSERT ... VALUES, you write INSERT ... FORMAT JSONEachRow and stream your data.
https://clickhouse.com/docs/en/sql-reference/statements/inse...
Edit: ah, the streaming described in your comment is also available via --stream, but the format presented to the selection-filtering-transformation stage will be a bit different.
var data = eval(jsonString);
This is exactly why it sucks in many ways, but also why it's pretty great.``` const jsonLParser = (streamOfLines) => { for (line in streamOfLines) { if line.trim() === "[" || line.trim() === "]" -> skip if line.endswith(',') -> line = line.trimEnd(',') yield JSON.parse(line) } } ```
this allows your API to be consumed both by JSON parsers and JSONL parsers
A processor that only understands the data link format (lines) can do useful things like merge or interleave two or more such streams, delete the n-th record given n, and such.
Though, I wouldn't call it a vista of limitless possibilities or anything of the sort. :) As soon as you have to crack the JSON to do anything interesting, the line orientation becomes more or less moot.
<origin_id> <timestamp> <event_type> {payload_json} \n
This way you know what json parser to use for which line.
origin_id and timestamp is just something we always use with our EventStream
In the database it also has one iterator per origin_id, and within an origin one iterator per event_type.
This way you can listen to both an origin_id and a specific event_type while making sure you can batch fetch any missing data if the iterators skip a beat.
I THINK that the format is like a JSON file with an implicit top-level list and the top-level (only?) commas replaced by newlines? But I do not see anywhere that is spelled out.
> any help writing the RFC would be greatly appreciated (see issue).
I guess you can sign up if you want an RFC.
I have many systems that log in this format but are quite hard to visually scan during development. I typically create ad hoc filters using jq, rg or a js script.
I guess the file extension ensures that JSON is not pretty printed.
This is great because it saves you memory (don't need to store the whole thing, your can store one line at a time depending on your specific issue), latency (don't need to wait for the whole file to be read and parsed, you start immediately after the first line), etc. Huge JSON files might be also too big to load into memory. You can read one line at a time, you can stream your lines, process your input across multiple computers etc.
And, as it's just json lines, you also don't need to invent a new notation, specialized tools, etc.
Editors that remove or fail to add a final newline can go implement themselves.
> Each JSON text MUST conform to the [RFC8259] standard and MUST be written to the stream followed by the newline character \n (0x0A).
https://github.com/ndjson/ndjson-spec/
But yeah, these things should be merged and formalized. Being web-adjacent, how likely is that? What we're seeing here is the usual patterns:
- the less-useful newer thing with slightly better marketing wins eyeballs even when it's technically inferior
- an "ecosystem cornerstone" is someone's pet project from a prior employer and now neglected (https://github.com/ndjson/ndjson-spec/issues/35#issuecomment...)
- web things don't get RFCs written because web people are disconnected from old school internet people
- W3C adjacent people didn't get their RDF tuple world domination and industry standards don't matter to them
- even MDN is switching to paywalls instead of being a useful resource
What's wrong with JSON patch?
Lines-of-JSON is streaming delivery of multiple JSON documents. Most producers & consumers will only hold one item in memory at a time.
[1]: And kinda sucks at that job. Last time I benchmarked it, a gzipped lines-of-json stream that just repeated the whole document for every edit was smaller than a JSON Patch stream, and re-parsing the document was faster than applying JSON Patches.
Some discussion here: https://stackoverflow.com/questions/17224536/why-is-one-numb...
...I am by no means the harbinger of all things JSON, but it's worth pointing out that this "spec" without addressing the issues that can arise from encouraging "bare" values (not wrapped in []/{}) is highly suspect.
json is defined as one element
element is defines as ws (whitespace), value, ws
value is defined as one of: object, array, string, number, "true", "false", "null"
Wrapping in {} to make an object or in [] to make a list (which is also not an object) is not necessary to make a valid JSON document. This is also discussed in the Stack Overflow question linked.
For example, the spec explainer at https://www.json.org/ (on the righ-hand), tells you that a `json` document is an `element`, which is a `value` surrounded by any or no whitespace. This `value` can be a literal, or an object or array.
For an (IMO) easier definition, you can check how JSON values are usually defined in TypeScript e.g. https://github.com/sindresorhus/type-fest/blob/main/source/b...
> Two terms for equivalent formats of line-delimited JSON are:
> Newline delimited (NDJSON)[4] - The old name was Line delimited JSON (LDJSON).[5]
> JSON lines (JSONL)[6]
Also when you're using this with GeoJSON there's https://stevage.github.io/ndgeojson/ which has an actual RFC (https://datatracker.ietf.org/doc/html/rfc8142)
The GeoJSON one pretty much seems to exist just to hang an "application/geo+json-seq" media type registration off of. Part of me wants to say this really should have been more of a "all json subtypes are also json-seq subtypes" situation but maybe that's not really feasible with the standards/registration processes.
ndjson specifies sane newline handling, since it works with terminators.
jsonlines works with separators and thus fails to detect truncated values. As a result, it can silently produce incorrect numeric values.
On the other hand, requiring line terminators in the standard would inevitably lead to incompatibility issues. Most software would accept unterminated files, because text libraries do; and so some files will not be terminated. Some applications do not line-terminate files even on Linux (hello VSCode), and it would be even more problematic on Windows
And if you have to parse values to detect end of record anyway, there’s no point in having jsonl standard at all, since you can just try to parse until the matching brace and repeat on success.
i.e. a documemt without a newline at the end is valid jsonl, but invalid ndjson
You can detect and error out if you see an unterminated record at end of transmission. With separators, the producer might not put a separator after the last record, because there's nothing to separate it from there.
There's no justification for not using terminators, it's just a bad spec. Unsurprising story: the variant with better marketing has less technical chops.