Tips on adding JSON output to your CLI app
blog.kellybrazil.com
blog.kellybrazil.com
~ % dig example.com | jc --dig | jq
[
{
"id": 61315,
"opcode": "QUERY",
"status": "NOERROR",
"flags": [
"qr",
"rd",
"ra"
],
"query_num": 1,
"answer_num": 1,
"authority_num": 0,
"additional_num": 1,
"opt_pseudosection": {
"edns": {
"version": 0,
"flags": [],
"udp": 512
}
},
"question": {
"name": "example.com.",
"class": "IN",
"type": "A"
},
"answer": [
{
"name": "example.com.",
"class": "IN",
"type": "A",
"ttl": 85586,
"data": "93.184.216.34"
}
],
"query_time": 29,
"server": "10.0.0.1#53(10.0.0.1)",
"when": "Sun Dec 05 15:12:08 PST 2021",
"rcvd": 56,
"when_epoch": 1638745928,
"when_epoch_utc": null
}
]Now, to find a query tool with a saner language than Jq...
jq can get pretty deep but for most things in this area I'm not sure how it could improve upon, but would be interested in hearing alternatives.
https://github.com/fiatjaf/jiq
Is a realtime feedback wrapper which I find useful when crafting one-off command line uses for jq and it starts getting crazy.
I actually collected a sizable list of alternatives to Jq:
https://github.com/TomConlin/json2xpath
https://github.com/antonmedv/fx
https://github.com/jmespath/jp
https://github.com/borkdude/jet
https://github.com/jzelinskie/faq
https://github.com/dflemstr/rq
And there's someone else's list of stuff around Jq: https://github.com/fiatjaf/awesome-jq
However, personally I think that next time I might instead fire up Hy, and use the regular syntax with the functional approach for any convoluted processing I come up with. Last time I mentioned this, another HNer made a Jq-like tool with Lisp-like syntax: https://github.com/cube2222/jql (from https://news.ycombinator.com/item?id=21981158).
Also, there is jellex[1], which is a TUI wrapper around jello that can help you build your queries.
[0] https://github.com/kellyjonbrazil/jello [1] https://github.com/kellyjonbrazil/jellex
Having multiple outputs is a great feature, though. I'm especially fond of tooling in Kubernetes that allows you to nicely pipe things in and out in multiple formats.
The biggest changes are typically ANSI color codes and column length. Some programs do strange things, like changing how they escape characters (ls).
JC turns off ANSI color codes when it’s output is not a terminal, but that’s it.
env NO_COLOR=1 ...(I’m aware of powershell and am ignoring it consciously)
https://github.com/lmorg/murex
You’d have to learn a new shell syntax but at least it’s compatible with existing CLI tools (which Powershell isn’t)
With murex, it does actually have a lot of intelligence built in that adapts how the output stream is displayed (eg rendering images in the terminal, “prettifying” JSON if STDOUT is a TTY, colourising STDERR red, etc.
Plus I disagree that your GP points about how JSON is handled should be a shell thing. You’re talking about data being re-encoded in different formats depending on the output. That absolutely should be a shell thing (where the logic of the pipelines are handled) with the terminal being a dumb rendering client. The last thing I want is output to be modified by the terminal leaving me scratching my head as to whether a command is running correctly or whether the terminal is displaying it weirdly.
I see now that my comment is vague in this part, but no, my idea is not about a terminal transforming the output. It’s about a presentation layer only, just like ansi escape codes (but programmable through repr-scripts). The stream remains intact, you only see it formatted. Murex does ~that, but it takes ownership of data processing by its own syntax and commands (like an IDE of CLI), while I believe that it should be done by separate tools like jq (like coreutils of CLI). It would be okay to use a dumb terminal with this system, you just would have to parse json by your mind to read output into it. See also my other comment for clarity https://news.ycombinator.com/item?id=29456775
Murex could ship it’s coreutils as external executables like GNU but that wouldn’t be a particular efficient way of doing it. Whereas Bash, zsh, Fish etc all have builtins just the same as murex (eg if/endif, for, switch, read, echo, time, jobs, etc).
The only difference between murex and bash in that regard is that murex builtins can do more intelligent data parsing than just dumb byte streams. Of course if you wanted to use jq with murex you still can. Just as you can still use sed, awq, Perl -pie and others too.
https://carreau.github.io/posts/29-JupyterCon-DisplayProtoco...
> The libxo library allows an application to generate text, XML, JSON, and HTML output using a common set of function calls. The application decides at run time which output style should be produced. The application calls a function "xo_emit" to product output that is described in a format string. A "field descriptor" tells libxo what the field is and what it means.
* https://github.com/Juniper/libxo
Then add an "--output-format" option.
As just one example, the Azure CLI defaults to human-readable output, but has an "output" parameter so you can have JSON if you want - I've never once wanted any kind of format auto-detection, and I have to say that I still don't.
However more modern shells fix this problem with having typed pipelines and builtins written to understand more than just a flat file of bytes.
Take _murex_ for example (disclaimer, I'm the author of that shell):
» jobs
PID State Background Process Parameters
2104 Executing true exec sleep 9000000
2240 Executing true exec sleep 9000000
It's readable but what if I wanted to pass it as a table? » jobs | cat
["PID","State","Background","Process","Parameters"]
[2104,"Executing",true,"exec","sleep 9000000"]
[2240,"Executing",true,"exec","sleep 9000000"]
ok, so it auto-detects it is running as a pipe and outputs it as a jsonlines table. That would be annoying in Bash. But with a type aware shell, that shell knows it's a jsonlines table, eg » jobs | debug | [[ /Data-Type/Murex ]]
jsonl
...but what can we do with a jsonlines table? Well you can select individual columns: » jobs | [ PID State ]
[
"PID",
"State"
]
[
"2104",
"Executing"
]
[
"2240",
"Executing"
]
run SQL against it » jobs | select * where PID > 2200
["PID","State","Background","Process","Parameters"]
["2240","Executing","true","exec","sleep 9000000"]
iterate through each row » jobs | foreach proc { if { =$proc[0]>2200 } then { echo $proc } }
[2240,"Executing",true,"exec","sleep 9000000"]
or even just convert it into another format, like CSV » jobs | format csv
PID,State,Background,Process,Parameters
2104,Executing,true,exec,sleep 9000000
2240,Executing,true,exec,sleep 9000000
...or YAML... » jobs | format csv
- - PID
- State
- Background
- Process
- Parameters
- - "2104"
- Executing
- "true"
- exec
- sleep 9000000
- - "2240"
- Executing
- "true"
- exec
- sleep 9000000
And it all just works without you having to think or even know what data format is traversing the pipeline.However unfortunately none of this is possible with Bash. And thus the majority of tools are forced to be dumb to compensate.
It’s a bit like having jc integrated into the shell, except that jc supports TOML, YAML, tables of various formats and all sorts - and autodetects the content type too - so you don’t have to learn a dozen different tools for managing a dozen different content types.
So it’s the best of both worlds.
I guess my claim is that the minute you need to parse the dumb bytes as anything other than strings, your problem is complex enough that you shouldn't be using shell to solve it, and you should write a proper program instead. That's not hard and fast, for example I use fish and `math` is very powerful.
`jq` is also a bad example, I think, in many ways... it explicitly doesn't follow the Unix philosophy, and instead does it's work by parsing an opaque string according to it's own DSL. Much better is a tool like gron, which allows you to process JSON using familiar tools like grep.
Ideally, I'd just prefer is libxo was common outside of FreeBSD, and I didn't have to worry about massaging stuff into structured data.
Perhaps this is something that is actually not a "fact," and is something more like an "opinion."
Advocate? No. Would I excuse it? Yes.
When I write tools, I write two kinds: ones that care intended to be consumed by something and ones that are not. If something is even marginally the former, I assume the former.
I can excuse a tool having output changing behaviour if it's the latter, and only in the case of dropping ANSI escape codes, but only because it makes sense to pipe such output into, say `pbcopy` and `tee`, without the escape codes for capture.
Automagically removing ANSI escape codes is the _most_ I'd be comfortable with, and only because is _reduces surprise_. Something sent to stdout should be the same as something sent to a pipeline, but I can forgive somebody adding ANSI codes to make things clearer to a human reader, for whom those are invisible in the stream.
Would I advocate for `ls` not to send ANSI escape codes to a pipeline? Yes. Do I think it's a good idea to pipe from `ls`? Most certainly not! `ls` is written with assumptions about the consuming terminal that defeat the principal of lease surprise.
Which is eminently reasonable. But your hard-line stance expressed above is less so.
> Bringing `ls` into this isn't a great idea.
I perceived your hardline stance as unreasonable personally, and so I wondered where the format changing functionality that most 'ls' implementations perform based on pipe detection fit in your worldview. It seems to me like your hard-line stance is not actually so hard-line. But you do a lot of mental gymnastics to get there, and you end up asserting that 'ls' shouldn't be used in pipelines. Which seems kinda nuts to me to be honest.
The better answer here IMO is that one should use "good judgment" when it comes to changing things based on automatic pipe detection. It's a matter of taste and there are more use cases for it than simply removing ANSI codes.
> I perceived your hardline stance as unreasonable personally, and so I wondered where the format changing functionality that most 'ls' implementations perform based on pipe detection fit in your worldview.
Because the output of `ls` is a trashfire for parsing already. I'd strongly discourage _anyone_ from attempting to parse its output.
> But you do a lot of mental gymnastics to get there, and you end up asserting that 'ls' shouldn't be used in pipelines. Which seems kinda nuts to me to be honest.
No, I've just had to fix a bunch of broken-ass shell scripts in the past because people didn't realise that the output of `ls` doesn't obey the principle of least surprise. No mental gymnastics here, just a lot of experience fixing other people's problems. About the only good thing that ever comes out of these is that is gives me a way to show people what they can do with safer, more predictable tools, like `find`.
I'd encourage you to pop up a level here and look at your original comment that sparked this thread:
> Yup! In fact, there's only one good reason to do automatic pipe detection, and that's if your tool normally outputs ANSI escape codes, which aren't something you want in something going to a pipeline.
There's no nuance there. No balancing of principles. No discussion about "least surprise." Just an opinion masquerading as a fact, and one that I happened to disagree with absent any other context. Your replies since then seem to think you presented a more nuanced case than what you actually did.
You've clarified a bit since then, but really haven't acknowledged that your original claim appears to be quite a bit stronger than where you've settled after a few comments into this thread.
> No, I've just had to fix a bunch of broken-ass shell scripts in the past because people didn't realise that the output of `ls` doesn't obey the principle of least surprise. No mental gymnastics here, just a lot of experience fixing other people's problems. About the only good thing that ever comes out of these is that is gives me a way to show people what they can do with safer, more predictable tools, like `find`.
So now we've moved on from discussing when and where tools should do automatic pipe detection to best practices for shell scripting. Yes, of course, I wouldn't try to parse the output of `ls` in a shell script. But that doesn't mean I'm not going to use `ls` in pipelines ever. Shell scripts are just a subset of what shell is used for. I also use it interactively. In such cases, `ls | grep foo` is a pretty common thing for me to do to see what's in the current directory. Never had a problem with it. But if I were to follow your advice, you'd want me to do what, do something like `find ./ -maxdepth 1 -name 'foo'? (Although, that includes hidden files.) That's a lot less convenient.
Output should be readable, not structured.
-c: Produce computer-readable output
So I tried it out: ~ networkquality -c | jq '{dl: (.dl_throughput / 1000000), ul: (.ul_throughput / 1000000)}'
awesome! {
"dl": 176.861488,
"ul": 6.742952
}
(Starlink in Perth, Western Australia)> This way I can easily filter the data in jq or other tools without having to traverse levels.
How is doing `jq '.cpu.speed'` any harder than doing `jq '.cpu_speed'`?
IMO as long as you aren't going insane with nesting levels, it's actually better to have a proper structure than dumping everything into an ugly flat object.
Grabbing an attribute is not necessarily any harder in a deeply nested structure, but filtering based on multiple deeply nested attributes in different branches can make a query quite complex.
It certainly seems like it does when the first example for flattening is oversimplifying an already simple structure that doesn't really need flattening. Maybe that was not meant to be a serious example but rather just for ease of understanding, but then the article should have probably said so.
> filtering based on multiple deeply nested attributes in different branches can make a query quite complex
Can you elaborate on this please? Maybe I'm just too tired to think clearly at 1 am, but I don't see how filtering is any harder. You would just do something like `jq '.foo | select(.bar.baz >= 42 and .qux.moo.asd == "abc")`.
Looking back at some other JSON output, like `ip addr`, I’m not seeing anything egregious. There are just a couple useless top-level keys in the `iostat` output that make the queries longer and don’t seem to add value.
Using the flat structure with a type field lends itself well to lazily outputting JSON lines, so maybe that’s why I tend to represent cli output that way. It’s sort of how you would format the data when sending to ELk or Splunk.
Since I initially thought of JC as a cli tool, I was trying to reduce line length. Even the name JC was selected so it wouldn’t significantly increase the line length. So not nesting data inside meta objects is something I would encourage.
For the purposes of making an example of how the concept is applied, making a verbose example is counterproductive. This doesn't require explanation because developers don't make work for themselves they don't need to, in practice.
If you are accessing a deep nested data that means you have to account for layers of existence of keys. If cpu exist then see if speed exist then access speed. Nothing wrong with deep nesting as long as you can guarantee a key and data will be generated but more often than not when the data is not being generated the JSON data and the key will not simply exist.
And people do get carried away with nesting. Also it is nice to have core information available at the surface level of JSON file.
In practice this enabled generating coverage reports orders of magnitude faster than traditional gcov wrappers like lcov
$ jc -h --dig
grep key file.json | awk -F: '{print $2}'
if you're already searching for a key, seems like you're just wanting the value.
granted, i hardly ever (have i actually ever??) interact with JSON this way, so i'm not exactly familiar with pitfalls.
(Granted grep still works, but...not nicely.)
thankfully, there are tools like jq to help maintain one's sanity
awk -F: '/key/ {print $2}' file.json "kb_read_s": 0.12
This worries me. JSON doesn’t have support for fixed point math, does it? When will some random POSIX tool spit out scientific notation at me.Also, if you just output a flat schema, is there much of a point in this vs just:
cpu: 0.2
kw_read_s: 0.12
The difference is that you can use a JSON parser vs splitting on new lines and colons?I do like the idea of JSON output as an option but before every bug and mistake gets canonized as POSIX or some other standard can we at least talk about the output format for a bit?
cpu: 0.2
kb_read_s: 0.12
mac_addr: 10:AA:FF:00:55:66 cut -d: -f2- dict((key.strip(), val.strip()) for key, _, val in line.strip().partition(‘:’) for line in text.split(‘\n’) if line and line.strip())
Also who says we can’t have a universal parser for this format just like we have for JSON? Not everyone needs to write the one liner like above, just use the libtextformat.parse(text) or whatever we would call it. {program -> human-text -> parser}+
When it could look like this: program
-> {struct-text -> program}+
[-> formatter] -> human-text . NAME="Rocky Linux"
VERSION="8.5 (Green Obsidian)"
ID="rocky"
ID_LIKE="rhel centos fedora"
VERSION_ID="8.5"
PLATFORM_ID="platform:el8"
PRETTY_NAME="Rocky Linux 8.5 (Green Obsidian)"
ANSI_COLOR="0;32"
CPE_NAME="cpe:/o:rocky:rocky:8.5:GA"
HOME_URL="https://rockylinux.org/"
BUG_REPORT_URL="https://bugs.rockylinux.org/"
ROCKY_SUPPORT_PRODUCT="Rocky Linux"
ROCKY_SUPPORT_PRODUCT_VERSION="8"
On the surface this seems great, but those quotation marks are kind of annoying. Is it possible there's an escape syntax that's used in case the name also includes quotes? eg VERSION="8.5 (Green \"Aqua\" Obsidian)"? Is it also possible you can embed newlines in between the quotes too? Who knows... Thankfully with JSON there is a simple spec. cpu: 0.2
kw_read_s: 0.12
is describing a single 'object'. How do you describe an array or list of these objects? Something like: cpu: 0.2
kw_read_s: 0.12
cpu: 1.3
kw_read_s: 0.4
Do you use a blank line to denote a new 'object'? In JSON it would be done this way, and this conforms to how the vast majority of command output maps (as they tend to be rows of columnar data): [
{
"cpu": 0.2,
"kw_read_s": 0.12
},
{
"cpu": 1.3,
"kw_read_s": 0.4
}
]
The other benefit to JSON here is that the formatting doesn't matter. This could also be expressed as a block of text with no spaces or newlines between elements: [{"cpu": 0.2,"kw_read_s": 0.12},{"cpu": 1.3,"kw_read_s": 0.4}]
Finally, this can also be streamed, using JSON Lines: {"cpu": 0.2, "kw_read_s": 0.12}
{"cpu": 1.3, "kw_read_s": 0.4} json.parse(input)Type support is lacking. Not date support for example.
Unnecessary quotes around keys make it harder for humans to read.
Different implementations allow for repeated keys.
otherformat.parse(input) // just as reasonablePlain text doesn't have support for numbers at all, which isn't much of a solution.
Let’s put it this way: if I proposed XML as the substitution for plain text, would you rather keep plain text or switch to XML?
1. You don't need to categorize every piece of data 2. You don't need to include everything in a single JSON file.
Deep nesting JSON is very annoying. The key-value pair structure of JSON is simply being abused at this point. Also, I really don't appreciate using numerical values as keys. Please use a list.
Optional header, selectable columns, one line per record, machine readable (raw) vs human readable numbers.
I’ve nothing against JSON output but I just don’t need it when you can print out two columns, select on the first, and print the second.
$ users -H -o name,hair |
> awk ‘$1 == “gorgoiler” {print $2}’
gray
Admittedly, that awk invocation is so commonly used it could probably be a lot more terse. Also, this whole house of cards collapses when you have data containing spaces.1. human review
2. scripting / automation
In case one, human readable formats are obviously preferable. But the moment you need to script a command, you want it in a machine readable format.
A perfect example of the differences between the two are how badly spaces in file names are handled. Granted POSIX deserve a lot of the blame here too.
https://devblogs.microsoft.com/scripting/learn-how-to-config...
One of the things I like most about the UNIX userland is that I can use small programs to edit vary large files, without needing lots of memory. I want programs that are designed to accomodate the possibility of line-by-line processing.
If the intent is to make output network friendly, maybe something like netstrings is useful. Easy to parse. Low memory footprint.
Seems to me this JSON idea is not designed to improve performance, agility or resource efficiency but to ignore the UNIX example in favour of a different, slower, approach that is perceived as easier for some people to use. Namely those who do not want to spend the time to learn how to use an existing, faster solution with lower resource requirements.
readable,
handy since probably every language has libs that work with it fine,
there's a lot of tools that work with jsons e.g generating code classes from json,
it's insanely popular,
really easy to learn
JC has streaming parsers that lazily output JSON lines for these types of commands. (ls, ping, vmstat, iostat, etc.)
You are misinterpreting the Unix philosophy. It's fine to use a bunch of sed, awk, grep, etc. when you're either transforming text or processing already well-structured data. But trying to write a full-fledged parser for something with only human-readable output, especially as a shell script, definitely goes against that philosophy. Congratulations, you've managed to piece together 50 commands in a pipeline and create a monstrosity that's far from the minimalist philosophy.
In fact, I would argue that by using `jc` together with `jq` you can actually create some nice pipelines for parsing the data that will be much more in line with the Unix philosophy.
Nobody ever said this was designed to improve performance, but I have a hard time believing your claims about it being significantly slower which is not backed up by any source. Most likely, eliminating the JSON conversion would be at most an unnecessary micro-optimization. But if your code was truly performance-critical, you wouldn't be piecing it together with shell pipelines that cause a bunch of unnecessary forks, you'd write it in something like C instead.
And the "JSON was designed for the web browser" argument doesn't hold much water either. You're about several decades too late for that, JSON is extremely ubuquitous and used in a lot of non-browser contexts. Sure, some people depending on their needs may use other formats like XML or protobuf, but JSON is still very common.
You are underestimating the power of unix tools. A chain of unix tools can match or exceed the performance of C programs written by average programmers. That is a true beauty of unix and partly why it is still relevant today. The author has little idea about performance and doesn't understand how unix works; otherwise he wouldn't make arrogant claims like:
> With jc, we can make the linux world a better place until the OS and GNU tools join us in the 21’st century!
> These programs do not expect large amounts of memory and many are written with the intent that they may be used to process text line-by-line
Which is only a problem if you are being very silly, don't choose NDJSON (newline-delimited JSON) and instead shove 10GB of data in a big [] array that the parser has to read in all at once. Almost every single JSON library can do NDJSON already. One of the most heavily used JSON-over-stdio applications is the Language Server Protocol, which uses JSON-RPC 2.0 and is entirely NDJSON. Same for about 15 different log-yeeting tools. Nobody has ever suggested switching LSP to plain text for performance reasons, only lower-overhead binary formats that don't throw out everything gained by having structure at all.
Large memory use by JSON is not something inherent to the encoding that plain text is somehow immune to. All sorts of CLI programs read stdin in all at once, and you don't see plain text getting slammed for exorbitant memory use.
In the context of the original post, `jc` etc, we're talking about essentially a constant sized output that's just much easier to parse, so the complaint is not relevant to those at all.
It wasn't. It was _inspired_ by JS's syntax (and that of Python), but wasn't designed for it. Crockford designed it as a lightweight data exchange format that used a familiar syntax. Quoting from the json.org website itself:
> JSON (JavaScript Object Notation) is a lightweight data-interchange format. It is easy for humans to read and write. It is easy for machines to parse and generate. It is based on a subset of the JavaScript Programming Language Standard ECMA-262 3rd Edition - December 1999. JSON is a text format that is completely language independent but uses conventions that are familiar to programmers of the C-family of languages, including C, C++, C#, Java, JavaScript, Perl, Python, and many others. These properties make JSON an ideal data-interchange language.
JSON isn't terribly difficult to parse, nor does it require "generous amounts of memory". Shy of something like s-expressions, it's about as straightforward as you can get when it comes to structured data.
Netstrings are really useful but they encode strings, not structured data.
The problem with JSON is that it does not support streaming by default. It's possible to use non-standard JSON-like formats to work around this, but then you're no longer using JSON!
This summarises the problem I have with JSON more succintly. It was not designed for streaming, thus "it does not support streaming by default".
Non-standard, line-oriented JSON formats are usable, although as a user I cannot see how they offer any significant improvement over previous approaches with fewer brackets, braces, colons, commas and quotes (BBCCQ). Consider the BSD utility mtree or the BSD-version of stat. These have options to output text in "shell-friendly" formats,1 minus all the BBCC and excessive Q. Sure, people could add options to utilities to output XML, or line-oriented JSON, but generally they don't. Why is that. Perhaps there is a reason.
You said it best: "It's possible use non-standard JSON-like formats to work around [JSON's limitations], but then you're no longer using JSON!"
Maybe JSON is just about hype or something. An attractant for today's "developers". This would explain why I am just not attracted by it.
1.
There were some upvotes then a series of downvotes as the top comment in this thread changed. Thus, opinion on this issue is mixed.
JSON of course stands for "object" notation. Once we start delimiting the "object" with a newline, or limit its size/length, it arguably starts to sound more like string and the problem of memory requirements is abated.
As the parent comment suggested, when the "object" notation is used for delimited JSON it's really not "JSON" anymore. Is it still describing "objects", or is it describing strings, with the addition of brackets, braces, commas, colons and more quotes. Let the reader decide.
It is reasonable to ask what was the purpose of JSON in the first place, and, if it was designed for sending data over a network, whether previous solutions like netstrings could accomplish the same things as delimited JSON.
ndjson is worth knowing about. We use it for things too large to stream.
While Windows isn't a whole language OS like those, .NET and COM gets pretty close to it, and that is what PowerShell knows about, instead of raw text.
This is what is missing across most traditional UNIX shells, integrate raw text, UNIX IPC (and newer ones like D-BUS/gRPC), shared libraries, structured data, into a single REPL experience.
I think it was called "recs"