> I don't think your revision of the "S-expression" grammar is in accordance with the normal use of the term; it says ()() is an S-expression, when conventionally it's considered to be two of them. If that's what you want, then adding the outer *() is fine, but you should probably drop the leading * on the recursive call *ex, because otherwise the grammar is ambiguous.
You're totally right, sorry! It's an important aspect of S-exps that I overlooked. I forgot that a typical Lisp file is not a valid S-expression, but a concatenation of a bunch of them. Plus maybe some whitespace/comments prepended.
So an S-expression syntax, as normally defined, is not closed under concatenation like Jevko.
To solve that (assuming it's a desirable feature as I do) purely syntactically, without talking about the REPL, you'd need another top-level syntax.
So why not simplify things and make one that meets this criterion by itself?
Which makes me think of another aspect where Jevko is simpler.
The restrictions Jevko puts on a unicode sequence are that the escape character must be followed by itself or a square bracket and that unescaped square brackets must be balanced.
This is pretty much an exhaustive one-sentence description that would allow you to validate or generate a Jevko.
A similarly simple description is impossible to state for S-expressions.
I tried and I came up with (for your minimal defnition of S-exp):
The restrictions S-expr puts on a unicode sequence are that it must conform to one of three alternatives:
* it must be wrapped in parens (list),
* it must be wrapped in quotes (string),
* it must not include whitespace or parens (symbol).
Also:
* If it's a list then different parts of the sequence may conform to one of these three alternatives.
* If it's a string then anything goes, provided that quote is escaped.
I don't think this is exhaustive and certainly not as clear and simple as the Jevko version. Can you come up with a simpler description?
> If you want to allow binary strings you'd probably want to define the language over bytes rather than Unicode code points, though Markus Kuhn's UTF-8B (implemented in Python as "surrogateescape") can give you the best of both worlds.
Yeah, I don't know why I thought your syntax allows binary strings. Clearly, I was confused.
Thanks for the reference, saved.
> Yes, I agree that "S-expression" is rather vague, much like "CSV".
This is an issue I want to avoid with Jevko.
The best syntax in this regard is JSON, with a relatively clear and concise specification. It does have some holes which cause problems, but I don't know anything better.
I'd like Jevko to be better.
> If you disallow having anything other than whitespace in a Jevko with >0 children, Jevko becomes a rose-tree notation, except that the root is unlabeled, so really it's more like a rose-forest notation.
Yes, depending on how you define these terms.
Important thing to note here again is that in accordance with how I intend Jevko to be used, you wouldn't be changing Jevko itself, but defining a format on top of Jevko[0].
The input of the parser/processor for this format is a Jevko syntax tree.
The output would be a rose tree. A nonblank suffix would be a (syntax) error in this format.
> If you want a rose-forest notation, you can get it in a less irregular way by declaring that Jevko infers an extra [] following a suffix containing any non-whitespace character, so foo[bar] is equivalent to foo[bar[]], foo[bar[]baz] is equivalent to foo[bar[]baz[]], and a[b]c is equivalent to a[b[]]c[], but foo[bar[] ] is equivalent to foo[bar[]] and not foo[bar[][]].
> This change would foreclose the possibility of having significant leading or trailing whitespace in some places but not others, the way I was suggesting.
Sure, you could also define a format that works like this.
Input a Jevko tree and output another tree with these extra subtrees "inferred".
The point is that you don't modify Jevko, but use it as a building block.
> Rose trees or rose forests are definitely simpler than Jevko's current data model, and they are equally powerful.
I think defining Jevko in terms of rose trees wouldn't be as straightforward as the inverse, so I disagree. But I wouldn't want to argue about this. I say go with whatever is simpler/more powerful/works for you.
> With respect to diagrams, git clone http://canonical.org/~kragen/sw/pavnotes2.git and look at {horse,johnsmith,player}.{jpg,jevko,dot}.
This is wonderful! I love it!
Thank you for putting it in the public domain.
I'd like to put this somewhere on a Jevko-related page(s), giving credit to you, linking to http://canonical.org/~kragen/.
If you prefer a different way of attribution or none at all, please let me know here or send me a gmail at darius.j.chuck
> For the diagrams I've taken the liberty of ignoring leading and trailing whitespace on arc labels (prefixes), as well as suffixes consisting only of whitespace. johnsmith and horse are your examples, but they just use Jevko as a terser syntax for JSON that suffers a whitespace problem.
Makes sense. Trimming whitespace is what is usually wanted, except in markup formats.
Perhaps it would be sensible to devise a very thin format on top of Jevko for making diagrams of this kind. It would take care of the trimming and have explicitly attached the semantics you mentioned before as proposed a tagline for Jevko:
> A minimal Unicode syntax for ordered trees with labeled nodes and labeled arcs.
It could be used as a tagline for this format!
Did you generate these diagrams from Jevko or write them by hand?
A tool to autogenerate diagrams from this format could be a good application for it.
Anyway, just throwing out ideas. Now to implement them...
> player.jevko is an example I cooked up based on Minetest's relational database schema, recast as a hierarchical schema. It also takes advantage of the ordered nature of subJevkos, the possibility of multiple identical prefixes in the same parent node, and the possibility of having both "headers" (subJevkos) and a "body" (suffix) in the same node.
That's a really nice use of suffixes!
This is really awesome work, thanks for caring and putting in the time. I'd love to see more stuff like this, people making use of Jevko. That's what it's made for!
Cheers!
[0] Another important thing to state, which I suspect may be causing some confusion in our discussion is that Jevko by design has no semantics at all. You attach the semantics on the format layer above Jevko.