JSON Lines
jsonlines.org
jsonlines.org
If you have a file containing this:
["Name", "Session", "Score", "Completed"]
["Gilbert", "2013", 24, true]
["Alexa", "2013", 29, true]
["May", "2012B", 14, false]
["Deloise", "2012A", 19, true]
And it gets chopped in half (a problem that is not that uncommon), you will get this: ["Name", "Session", "Score", "Completed"]
["Gilbert", "2013", 24, true]
["Alexa", "2013", 29, true]
Which is still valid. This will cause missing data rather than a clear error message. May and Deloise may end up not getting their scores and there's a good chance nobody will notice.By contrast when the file is represented as regular JSON:
[
["Name", "Session", "Score", "Completed"],
["Gilbert", "2013", 24, true],
["Alexa", "2013", 29, true],
["May", "2012B", 14, false],
["Deloise", "2012A", 19, true]
]
(not a completely ideal representation, but you get the picture)Chopping this in half will give you a clear error, forcing the developer to recover from it before continuing.
Same principle as strong vs. weak typing.
Individual JSON snippets in lines are a good replacement for plain text-only logs of indeterminate length, however.
Most JSON parsers simply parse the JSON string and represent it as an object in memory. You don't want to do that for very big files. Stream based parsers avoid this problem, but are more complicated to work with.
(disclosure: I'm the author of the Haskell ndjson-conduit library)
["Name", "Session", "Score", "Completed"]
["Gilbert", "2013", 24, true]
would be "Name", "Session", "Score", "Completed"
"Gilbert", "2013", 24, true
Which is, well, more or less just CSV.
This should work with objects too: "name": "Jane", "key": { "nested": "object" }, "foo": ["bar"]
Or mixed: "Foo", { "fnord": 23 }, trueFor the above to work you'd have to use a custom JSON-parser when reading the lines.
------------
Could also be mentioned that JSONLines-like formats are already pretty common in log-files, database exports etc. So this is more about giving it a name and standardizing it. Which I think is great!
Not necessarily. You could just re-add the square brackets before passing the lines to the JSON parser.
That would make more work for the machine, because you'd probably have to copy the whole string. But any time you can make less work for the human by making the machine work a little harder, it's usually a win.
The main unique thing here is the proposal that tools like Excel should support this format. That's definitely an interesting idea. Biggest challenge is that the JSON objects can be nested, so table tools would need a way to handle that. Still, interesting
As for CSV and TSV, they're a dead-simple format to read and write, and you can enforce a standard way to escape special characters if you aren't dealing with arbitrary user-submitted files. With numeric data escaping isn't even necessary, and if you have to deal with strings that might themselves contain special characters, the ASCII control characters 0x1E and 0x1F work well as alternatives to comma/tab and newline.
I know Python has a csv module, but I've just gotten used to writing these snippets:
def tsv_write(path, headers, data, sep='\t', end='\n'):
with open(path, 'w') as file:
file.write(sep.join(map(str, headers)) + end)
for datum in data:
file.write(sep.join(map(str, datum)) + end)
def tsv_read(path, headers, sep='\t', end='\n'):
with open(path, 'r') as file:
for line in file:
yield line.rstrip(end).split(sep)
(That tsv_read function only works if end == '\n'. Here's a somewhat-inefficient general-purpose alternative, although I've never needed it.) def read_records(file, end='\n'):
if end == '\n':
for line in file:
yield line
else:
record = []
while True:
c = file.read(1)
if c:
record.append(c)
if c == end or not c:
yield ''.join(record)
record = []
if not c:
break
def tsv_read(path, headers, sep='\t', end='\n'):
with open(path, 'r') as file:
for record in read_records(file, end):
yield record.rstrip(end).split(sep)I also wouldn't recommend you to reimplement a functionality that is already part of the standard library, especially because the standard library function is going to be much faster because it's implemented in C.
Just start by erasing the open "[", and then read one row at a time, just like you would do without the "[" and comma.
If you don't actually need a stream, I think it's a bit nicer to write a special header at the beginning and make the whole file a single JSON object:
{"header": {"title": "An example", ...},
"rows": [
["row1"],
...
["rowN"]
]}
[1] https://golang.org/pkg/encoding/json/#example_DecoderUnfortunately, it needs a root element to be valid XML, so wrap the document in <xsv></xsv>. Also, add an XML declaration.
Example:
<?xml version="1.0" encoding="UTF-8"?>
<xsv>one<c/>two<c/>three<n/>1<c/>2<c/>3<n/></xsv>
(Whitespace is passed-through, therefore it's significant.)This solves the ambiguity problems of CSV without complicating CSV by introducing things like data types (I think the virtue of CSV is its simplicity), and it should be extremely simple to write parsers based on either SAX or DOM (, etc).
Also, the use of XML provides unobtrusive places to insert application-defined metadata, for example:
<xsv columns="3">
This remains backwards-compatible with parsers that ignore attributes.XML tries to be human readable, but you will need knowledge of elements or "tags", meta data.
JSONLines also tries to be human readable, imo it does a better job, compared to XML.
But not until Excel of LibraOffice implement a JSONL, or XSV reader/writer.
Will CSV be the de-facto multi standard.
To cut back on markup in XSV, I used self-closing tags for commas and newlines. It would have been more conventional to use something like <row><column>...</column></row>. That has some advantages, maybe, but it doubles the number of tags (and it introduces things that are not strictly 1:1 with CSV, which could lead to ambiguity).
Commas are the standard column separator in CSV files, its in the name.
If you are on a French windows, it will use semi-colons as separators, encode in window cp1252 (close to latin1 but not exactly) and use CRLF (\r\n) as line returns.
But then if you use Excel for Mac 2008, Excel will generate CSVs in MacRoman (old OS9 encoding) and use CR (\r) as line returns.
Long story short Excel is not even interoperable with itself when you use CSVs. (Granted there is an advanced import tool, but it's a nightmare for the average user).
I see someone already created a GitHub linguist issue for it: https://github.com/github/linguist/issues/2217
If there was a decent visual tool (or chrome extension etc) for editing data in a format like this new one, I'd probably jump on it, but otherwise it seems like anyone currently using CSV might as well stick with it.
Each line is a version, with only the changes.
This means you can't delete though, but for blockchain-like distributed systems that's almost a feature.
I've seen so many different dialects, even from the same vendor... that's the worst part.
This is fine.