Jtc – CLI tool to extract, manipulate and transform source JSON
github.com
github.com
Citation needed. The `jq` vs `jtc` section is interesting, but author seems a little full of himself with some of the explanations.
> jq is written in C, which drags all intrinsic problems the language has dated its creation
And then to follow on the C++ claims:
> Main JSON engine/library does not have a single new operator, nor it has a single naked pointer acting as a resource holder/owner, thus jtc is guaranteed to be free of memory leaks (at least one class of the problems is off the table) - STL guaranty.
That's a lot more faith than I would be willing to put into C++ or C. Sure, that claim might be correct, but there's enough edge cases and undefined behavior in the language that I take it as a fairly bold claim unless it's also been thoroughly reviewed for any places some undefined behavior may be induced.
I mean, it's probably fine, but just the willingness to make such a bold claim in that way communicates the opposite of what the author likely intended for me.
I can't even parse this sentence. What does it mean?
FWIW jq was originally written in Haskell, then ported to C. The C source for the core parts is strikingly clean. I think, it's partly from this Haskell heritage that jq gets its nice composability.
However, TBF, I do find I must refresh jq syntax whenever I use it (pr only bc I don't process json often). It will be interesting to see what jtc does - though this foreshadowing does not bode well...
Yeah, you don't know what you're talking about.
"Undefined behavior" simply means "stuff not covered by the ISO standard".
So Python and Rust are 100% UB, and people don't really care.
Perhaps the features used in this project protect against that, but my point is that undefined behavior and the way compilers deal with it is variable and problematic.
Somehow we make do in these cases and don't act like the sky is falling.
I think "there's enough edge cases and undefined behavior in the language that I take it as a fairly bold claim unless it's also been thoroughly reviewed for any places some undefined behavior may be induced" was me acting like the sky is falling. I think that's a very sane push back against a claim that I think either wasn't well founded or wasn't well explained.
Ultimately it's up to you to decide whether to accept I'm expressing what my words actually say or whether I'm expressing what you presented me as in your straw man argument you originally replied with. At this point, there's not much else I can or am willing to waste any more time saying that I haven't already.
For those, I mostly use "jq ." to get all leaves on a single line and then feed the output to standard unix tools like grep cut awk and friends.
And for more complex tasks, python, perl, even C++ if speed is needed.
There's a tool called `gron` designed for workflows like this, I've found it incredibly useful.
https://news.ycombinator.com/item?id=16727665 https://github.com/tomnomnom/gron/
There's no reason why you would need to learn the whole DSL for simple tasks. Just learn what you need. For simple stuff, the DSL is also simple.
jtc also seems to have its own DSL by the way, and it doesn't really seem more intuitive than jq's:
jtc -x'[0][:][name]<person>v [-1][children]<kids:false>f[0]<kids:true>v' -T'{"{person} has children":{kids}}' -r
That's apparently the equivalent of: jq '.Directory | map({"\(.name) has children": (.children | length == 0)}) | .[]'
There are some things that seem simpler on jtc, though, like searches of string values without regard to the json structure.IMO the ideal solution is something using pure JavaScript syntax, possibly with a library resembling jQuery for tree traversal.
Here's an explanation of the syntax in the command I posted:
jq '
# output the value at "Directory" from input object
.Directory
# pipe to map (JavaScript also has map()). The argument of map works in
# the context of each element in map's input array.
| map(
# Produce an object where the property name is a string that
# interpolates the value of the "name" property of this element.
# Instead of interpolating, we could have also used this more
# JavaScript-ish (ES5) syntax:
#
# {(.name + " has children"): ...
#
# The property value is an expression that gets the value of the
# "children" property and pipes it to the expression `length != 0`.
# `length` (which JavaScript also has) outputs the length of the
# piped input, and then we compare that with 0.
{"\(.name) has children": (.children | length != 0)}
)
# map's output is a single record which is an array. We pipe that to
# .[] to make multiple records, each an element of the array. The
# syntax here is comprised of 2 parts: `.`, which is the input object,
# and `[]` which is the subscript syntax without an index.
| .[]
'
> IMO the ideal solution is something using pure JavaScript syntaxThe greatest advantage of the current syntax is the great balance it has between legibility and terseness. I don't think making it pure JavaScript would be better.
> possibly with a library resembling jQuery for tree traversal
Using jQuery in JavaScript to traverse JavaScript objects? I don't know what to say...
jsed '
Object.fromEntries(
$("Directory").map(
x => [
x.name + " has children",
x.children && x.children.length != 0
]
)
)'
Interestingly enough there is already a tool called jsed which seems to kind of do this... https://www.npmjs.com/package/jsedEdit: Note the jQuery part of it is for advanced cases like searching for specific nodes in the tree, then navigating back up to the parent. Basically the cases like "<Work>[-1][children]" from the jtc guide, I would write as: '$("Work").parent().find("children")'
jq '.Directory | map({"\(.name) has children": (.children | length != 0)}) | .[]'HTML -> XML (via hxnormalize) -> JSON (via jtm) -> process using jtc (or even jq)
This is basically impossible to do in a way that is compatible with other tools. Things like duplicate attributes of an object can exist in XML, but not in JSON. You can still work-around these limitations if you just have a pipeline using the same toolset, but part of the point of these tools is to then convert them back to a format that some other tool can use, which is where this pattern breaks down.
Here's a list of pitfalls: https://stackoverflow.com/questions/33072812/potential-probl...
I'm not convinced that I would call this "simpler" than jq.
By contrast, this seems like one of the least intentionally designed pieces of CLI software I've seen.
$ wc -l conn.log
8505 conn.log
$ cat conn.log |time ~/src/jtc-standard.json/jtc -a -w '<uid>l' -qq |shasum
0.44 real 0.43 user 0.00 sys
167f6d638a4ccf9e0be1ff4ed74caa01a461bd9a -
$ cat conn.log | time jq -r .uid|shasum
0.10 real 0.10 user 0.00 sys
167f6d638a4ccf9e0be1ff4ed74caa01a461bd9a -
$ cat conn.log | time json-cut uid |shasum
0.04 real 0.01 user 0.01 sys
167f6d638a4ccf9e0be1ff4ed74caa01a461bd9a -
json-cut is a crappy tool I wrote that use github.com/buger/jsonparserPostgreSQL also calls the SQL/JSON path language “JSON Path”, and the data type is called `JSONPATH`.
Unfortunately, ".." just doesn't fit nicely with the "." operator of C-style languages.
I'm almost inclined to go full-XPath, and use "/", so I can use "..".
If I really wanted CLI, I would use one of those languages to write a program that reads arguments while making use of one of those JSON libraries.
It's time to stop admitting that (if the name is so terrible) and re-brand or, more likely, come up with a better backronym meaning for your name.