Comma-Separated Tree
observablehq.com
observablehq.com
name,value,color
World
Asia
China,1409517397
India,1339180127
Indonesia,263991379
becomes region,subregion,country,value,color
World,Asia,China,1409517397
World,Asia,India,1339180127
World,Asia,Indonesia,263991379
This still compresses very nicely, and allows data to be output in any order. The tree form is more of a presentation thing, which should be the last step of a data processing pipeline. As it is, the tree nodes themselves have no names or context, so it makes it harder to consume. Is the second-level entry a region? Easy to tell from this example, but if you're consuming a large data set you're forced to first give names to these columns.To get a sense of what I mean, see this editor which lets you interactively construct a treemap: https://beta.observablehq.com/@mbostock/treemap-o-matic
country,value
World/Asia/China, 1409517397
World/Asia/India, 1339180127
World/Asia/Indonesia, 263991379
An added benefit of this approach is that you can have hierarchies in multiple columns. For example: datetime, source_account, destination_account, amount
2016-07-01, Assets/ProjectFunding, Assets/Capital/Delivery, 134.2
2017-07-01, Assets/ProjectFunding, Assets/Capital/Delivery, 72.2
2016-07-01, Assets/ProjectFunding, Assets/Capital/Other, 5.0
2016-07-01, Assets/ProjectFunding, Expenses/Development, 4.7
2017-07-01, Assets/ProjectFunding, Expenses/Development, 1.6
2017-07-01, Assets/Cash, Expenses/OPEX, 0.96
2018-07-01, Assets/Cash, Expenses/OPEX, 1.62I feel like the fixpoint of this function is going back to only one separator, or something, but I am not seeing the reasoning clearly.
It can be good because each line is now an independent idea and doesn't require the indentation context, the items can be sorted or moved easily.
As for hard to work with... a lot of editors have multi-cursor support these days, which makes editing these things in bulk pretty straight forward for smaller sizes and then there's always sed and awk. ;)
Given all that, CST seems like it might be cool for building some quick trees like the given UIs you present.
I'd probably keep to to rows for longer storage though.
I've taken to using XSV to pretty-print CSV data in bug reports or feature requests for coworkers who give me CSV or spreadsheets to work with. The data doesn't need to go any further, it's already been processed, it just needs to be clearer to understand than raw CSV/JSON/XML:
https://github.com/BurntSushi/xsv
(link has examples of pretty-printed CSV)
If you do too, and you find yourself dealing with JSON a lot, take a look at gron [0], it takes JSON and denormalizes all the structure into a line-oriented equivalent.
% echo '{"root": {"left": [1, 1.4, 2, 2.88], "right": ["a", "b"]}}' | jq .
{
"root": {
"left": [
1,
1.4,
2,
2.88
],
"right": [
"a",
"b"
]
}
}
% echo '{"root": {"left": [1, 1.4, 2, 2.88], "right": ["a", "b"]}}' | gron
json = {};
json.root = {};
json.root.left = [];
json.root.left[0] = 1;
json.root.left[1] = 1.4;
json.root.left[2] = 2;
json.root.left[3] = 2.88;
json.root.right = [];
json.root.right[0] = "a";
json.root.right[1] = "b";
[0] https://github.com/tomnomnom/gron root.left[0] = 1
.left[1] = 1.4
.left[2] = 2
root.left[3] = 2.88
..right[0] = "a"
.right[1] = "b"
[0] https://github.com/jaroslov/kvinMeanwhile, the JQ tool could do your greppings and awkings but preserving and perusing the structure of the document notwithstanding its idiosyncrasies.
(A possible problem with JQ, though, is that it may be difficult to query a structure which varies, i.e. use entries on a nesting level that isn't known in advance.)
Without knowing gron, I presume the dots are escaped/quoted properly, and you would do the same in your grep/awk usage.
What if your keys contain equal signs and curly braces, how will json help you out?
> Meanwhile, the JQ tool could do your greppings and awkings but preserving and perusing the structure of the document notwithstanding its idiosyncrasies.
The structure of the data is preserved (I don't know what you mean with "notwithstanding its idiosyncrasies"). This is just a different, albeit redundant, but entirely equivalent representation of the same document. Despite their redundancy, denormalized representations do have advantages.
Oh yeah, that's certainly much easier than having a tool's parser deal with the syntax, and just working with individual keys and values instead. So much so that in fifteen years of programming I haven't seen anyone remember to include escape-sequences in their regexps. But I guess for true lovers of regexps it's just a joy to keep escaping the escape sequences that they put in, and then escaping the result once over in the coding language of choice.
> What if your keys contain equal signs and curly braces, how will json help you out?
Care to clarify why that would be a problem?
I think the tree form is great for the smaller datasets that need to be edited or reviewed manually, not bulk machine generated data. For one project I have thousands of files, each with data stored in a tree, all stored in a git repo with many editors, sort of like a data wiki.
It also works for the "presentation thing", as you say. You can use a table-backed storage system but present to the user a tree, and when they edit a node you propagate the change to the appropriate row.
Finally, some editing plugins can give you good highlighting and more context if you create a grammar file.
Source: have been tinkering with this stuff for a couple years.
I really enjoy this style of format. It makes reading and writing structured data a breeze, and can handle any format with no escaping save indentation. You can embed csvs,tsvs,psvs, et cetera, right in the tree, like in these Comma-Separated Trees. You can write a grammar file to ensure strong type checking. Finally, you can easily do conversions "fromCsv/toCsv, fromJson/toJson, fromXml/toXml, fromSql/toSql" et cetera...
Here's another example showing an embedded PSV (and also same type later expressed as a tree) and some source code embedded showing how you don't need escaping.
cobol
books
title|year|author|id|rating|ratings|reviews
Cobol Programming|1983|M.K. Roy|4944251|4.11|9|1
Structured Cobol Programming|1979|Nancy B. Stern|9030220|4.33|15|0
fileType text
year 1960
lua
fileType text
books
book
title Programming in Lua
year 2001
author Roberto Ierusalimschy
id 1321894
rating 3.97
example
function factorial(n)
local x = 1
for i = 2, n do
x = x * i
end
return x
end a
..b
..c
Would be interpreted as (a (b (c)))I still think there's something for the philosophy of developing tools that expect a strict format and fail if that format is violated. I've worked in a variety of settings and with a variety of tools, as time goes by I am finding the value of tools that are strict about expectations to provide a more maintainable product over the long haul. If I were writing a project using this style I'd prefer to receive a parse failure in the case of ambiguity, rather than carrying on and hoping it is correct.
If data correction needs to happen (i.e. one file that should be single-space indented is triple space indented) I'd prefer to explicitly pre-process the data rather than have a tool that can handle it gracefully.
Also, like S-expressions, even without a good tool, mistakes drop once you use it enough.
Similarly, with leading spaces to delimit fields--what could possibly go wrong?
Its ASCII, folks.
Edit: spelling
> to make it easier to edit the data by hand, for example to cut-and-paste some lines to move them around within the tree without having to change the leading values
I wouldn't even consider trying something like that with JSON.
I think this should be specced out early to avoid falling into the trap good old plain CSV fell into, which now now has multiple ways programs escape commas.
(and the corollary: any data can be authored or visualized in Excel).
This example is no exception. Here is how my mind interprets it...
Makes sense.
Makes sense.
Makes sense.
Makes sense.
WTF?
You might not be interested in how the parser is implemented; more likely you only care about the design and usage of the proposed data format. Which is fine! But we designed Observable to share one interface for both authors and readers under the view-source philosophy that made the early web so great: all the source is there, accessible, if you do want to dive in and understand it.
But we could do a better job of making the segue from narrative to internal implementation less jarring. I’ve edited this post to make that more explicit with an “appendix” header for the implementation. And we’ve been thinking about ways to formalize this convention, so that the code is still accessible with a click or two if you want it, but doesn’t distract from a normal read.
(It appears to be population)
(World
(Asia
(China 1409517397)
(India 1339180127)
(Indonesia 263991379)))
With an editor that supports Paredit or Parinfer, it's impossible to create an invalid tree, because you manipulate the structure of the tree (nodes and leaves) instead of manipulating error-prone text.Have you heard of org-export-json? [1]
Have you heard of json.el? [2]
[1] https://github.com/mattduck/org-toggl-py/blob/master/org-exp...
root
parent
foo: 1
bar: 1
Syntax details might be off from the top of my head, but the example in the tweet is a bit overblown.I will say it might be easier to write. do string quoting, quote escaping and newline embedded in a quoted string work?
However, the implied question here is absolutely valid: Why write your own flat-file format parser from scratch, when there's a multitude of battle-tested third-party libraries already in existence?
I mean, if you're writing something very specialized or proprietary, that's one thing. But if you're simply proposing a generalized tree shape that looks like YAML, then why not indeed use YAML? I know it's tricky to write a YAML parser... but so what, why would you? There are already a ton of great ones.
1. intermediate steps in YAML are often invalid documents (breaks the update on every key press)
2. you have to be trained to write YAML. It is quite easy to make a syntax mistake with YAML if you aren’t used to writing it. Not good for an interactive experience
<World>
<Asia>
<China>1409517397</China>
<India>1339180127</India>
<Indonesia>263991379</Indonesia>
</Asia>
</World>That said, the statement itself came off really snarky.
{World
{Asia
{China '1409517397'}
{India '1339180127'}
{Indonesia '263991379'}
}
}