How is it simpler or lower-level than s-exprs?
How is it simpler or lower-level than s-exprs?
To see how it's syntactically simpler than S-exps note that the grammar of Jevko can be condensed into one short line of ABNF:
Jevko = *("[" Jevko "]" / "`" ("`" / "[" / "]") / %x0-5a / %x5c / %x5e-5f / %x61-10ffff)
The grammar of S-exps on the other hand, I won't quote here, but I assure you it's much more complicated. How much depends on your flavor (Jevko is also simpler in this regard: there is only one flavor, clearly specified).There is no (intended) ambiguity around whitespace in Jevko: whitespace does not occur explicitly in the grammar. Whitespace characters are just characters. This is the defining feature of the syntax.
For this reason Jevko is more low-level: if you want to treat whitespace in some special way, you have to do that yourself. Although for most use-cases this is very similar and simple, e.g. https://news.ycombinator.com/item?id=33334314
But the point is that you can also leave it as-is, e.g.: https://github.com/jevko/queryjevko.js
or do something else -- it's up to your format.
Has anyone ever complained that S-expr were too complex? Adding white noise as a tradeoff doesn't seem like a win.
Ask parser writers, I guess? I had some similar devices, explicitly made so that the overall set of parsing rules would be clear, so the idea is likely not too new.
I've written many parsers, including for s-expr following rivest's RFC. S-expr take about half a dozen states to pull. A packrat parser for s-expr can fit in a single page. You don't even have to scroll to see the whole implementation. What are you talking about?
What exactly do you feel parser writers will tell you regarding s-expr?
ex = ("(" *ex ")" / *%x2a-10ffff) *%x0-20
We could write the data model corresponding to this grammar in OCaml as type ex = List of ex list | Atom of string
If we add double-quoted strings with GW-BASIC/SQL-style quoting (and resolve the ambiguity greedily): ex = ("(" *ex ")" / *%x2a-10ffff / %x22 *(%x22 %x22 / %x0-21 / %x23-10ffff) %x22 ) *%x0-20
This corresponds to the data model type ex = List of ex list | String of string | Symbol of string
This still seems both simpler and more expressive than the one-line grammar you give.
Maybe I'm missing a subtlety of S-expressions here, and I haven't tried it, but I think this correctly parses your examples like (first-name"John"last-name"Smith"is-alive true age 27
address(street-address "21 2nd Street"
city"New York"state"NY"postal-code"10021-3100")
phone-numbers((type "office"number"212 555-1234")
(type "home"number"646 555-4567"))
children()spouse())
or (
:first-name "John"
:last-name "Smith" ...
)
though of course the name-value pairing is lost there (because S-expressions lack it).(By the way, if you want to attribute your JSON example for copyright reasons, you need to attribute it to its author or authors, not to the Wikipedia, which is just the site they posted it on.)
— ⁂ —
Maybe more importantly, though, I think your abbreviated Jevko grammar is wrong in an important way, though it describes the same set of strings as the full grammar. According to the abbreviated grammar, the Jevko examples lack most of the structure of the XML, JSON, and S-expression versions. It implies that a Jevko is an ordered sequence of (Unicode!) characters and nested Jevkos (indicated by []). We could express this data model as
type jevko = atom list
and atom = Char of char | Nest of jevko
If so, then abbreviating your S-expression example (first-name "John" last-name "Smith")
the closest Jevko equivalent is not, as you claim first name [John]
last name [Smith]
but rather (supposing the hyphens were just an unfortunate concession to S-expression syntax rather than actually desired) [first name][John][last name][Smith]
We had to sacrifice the formatting white space because there's nowhere that Jevko (as specified above!) ignores it.Using this grammar, the S-expression equivalent of the Jevko
first name [John]
is rather ("f" "i" "r" "s" "t" " " "n" "a" "m" "e" " " ("J" "o" "h" "n"))
If we instead use the full Jevko grammar Jevko = *Subjevko Suffix
Subjevko = Prefix "[" Jevko "]"
Prefix = Text
Suffix = Text
Text = *Symbol
Symbol = Digraph / Character
Digraph = "`" ("`" / "[" / "]")
Character = %x0-5a / %x5c / %x5e-5f / %x61-10ffff
or, I think, equivalently in this context: Jevko = *Subjevko Text
Subjevko = Text "[" Jevko "]"
Text = *(Digraph / Character)
Digraph = "`" ("`" / "[" / "]")
Character = %x0-5a / %x5c / %x5e-5f / %x61-10ffff
then we do preserve the name-value structure you seem to be going for, which your above one-line version loses. And this allows us to write first name[John]last name[Smith]
as in your compactness examples.I think this is a more useful level of abstraction, and it's more or less the level used by, for example, queryjevko.js's jevkoToJs, although that erroneously uses () instead of []. (Also, contrary to your assertion above that this is an example of "leaving [Jevko's data model] as-is", it forgets the order of the name-value pairs as well as I guess all but one of any duplicate set of fields with the same name and also the possibility that there could be both fields and a body.)
Essentially at this level of structure a Jevko is a (possibly empty) set of name-value pairs followed by a plaintext body ("Suffix"). This is exactly like an email message, except that the values are themselves Jevkos. In OCaml we could write:
type jevko = Jevko of (string * jevko) list * string
For your example include[author]fields[articles[[title][body]]people[[name]]]
this gives the representation Jevko
([("include", Jevko ([], "author"));
("fields",
Jevko
([("articles",
Jevko ([("", Jevko ([], "title")); ("", Jevko ([], "body"))], ""));
("people", Jevko ([("", Jevko ([], "name"))], ""))],
""))],
"")
although I notice that queryjevko handles it very differently.— ⁂ —
Unlike the data model implied by your one-line grammar, I think this is an extremely useful data model. Email messages are the one and only structured data format that has remained compatible in active use and extension for over half a century; you can literally take internet email messages from 01972 and load them into a mail client today (most easily by putting them into a qmail-style maildir) and everything will just work. This is largely a result of the decentralized extensibility properties of name-value pairs: mail clients just ignore header names they don't understand, and they don't require header names that weren't originally present. This is also the basis of the extensibility of HTTP/1.0 and HTTP/1.1.
It's also very similar to the data model of popular "semistructured" or "free-form" databases of the 01990s like askSam and Filemaker. Unlike RFC-822 and HTTP/1, but like Jevko, those systems support recursively nested data. askSam even used almost the same firstname[John] syntax.
This data model does have a semantic mismatch with things like the rose-tree representation you describe at https://xtao.org/blog/rose.html, since it associates the rose-tree label with the first branch and the empty-string label with subsequent branches.
If your audience is people like me, I think it would probably be worthwhile for you to spend some time up front describing the intended semantics of a data model, as I've attempted above, rather than leaving people to infer it from the grammar. (Maybe OCaml is not a good way to explain it, though.) You might also want to specify that leading and trailing whitespace in prefixes is not significant, though it is in the suffix ("body"); this would enable people to format their name-value pairs readably without corrupting the data. As far as I can tell, this addendum wouldn't interfere with any of your existing uses for Jevko, though in some cases it would simplify their implementations.
______
Runnable OCaml code:
type jevko = Jevko of (string * jevko) list * string
(* XXX doesn't escape `[ `] `` *)
let rec dump (Jevko(hdrs, body)) = hdr(hdrs) ^ body
and hdr = function [] -> "" | (k, v) :: t -> k ^ "[" ^ dump(v) ^ "]" ^ hdr(t)
let dict kvs = Jevko(kvs, "")
let text s = Jevko ([], s)
let v = dict ["include", text "author";
"fields", dict [
"articles", dict ["", text "title"; "", text "body"];
"people", dict ["", text "name"]
]]
;;
print_endline(dump v)I'll respond to the main points and then expand on the details later.
I'm not sure what the authoritative source on S-expressions is (or even if there is one, which to me is a problem and part of the value proposition here), so I'll take R^7RS[1] as a reference.
If you look at the formal definition there (chapter 7), it's significantly more complex than both Jevko and what you beautifully constructed here.
And indeed, you have just freestyled a simplifed version of S-expressions (impressive!), but this is not the real thing.
If you would keep refactoring it with the constraints I had in mind for Jevko, you'd eventually end up with Jevko.
> though of course the name-value pairing is lost there (because S-expressions lack it).
Indeed, and that's another part of the value proposition of Jevko. The grammar is designed purposefully to take advantage of natural syntactic name-value (prefix-subjevko) pairing tendencies.
Which brings me to the next point.
**
You are absolutely correct that the abbreviated grammar matches the same strings, but doesn't have the same structure.
*The correct grammar is the one in the specification*.
I have shown the condensed version of it to illustrate the point that Jevko is indeed extremely simple. The single line captures all the essential elements and, again, matches the same strings as the full grammar.
This is unlike the similar condensed grammar for S-expressions which you sketched out here. If you would continue, it would get significantly more complex before it matches the same strings.
The OCaml type definition you wrote down should do the job of capturing the structure, although I prefer to name the elements (which may be not-so-convenient, depending on the language, so it's fine).
**
Indeed, I also think that this is an extremely useful data model. Thanks for pointing out the similarity to e-mail messages and other references which I'd love to dig into (please send if you have any links or resources about these databases).
Thanks for all the pointers, I'll take them into account.
And thanks again for your time.
[0] this seems kinda official, but there is no single standard here: https://www.s-expressions.org/standards
As for S-expressions, I agree that what I wrote above is a pretty minimal definition of S-expressions, but I still think it's valid as a definition of "S-expression", even though you'd use a modified definition in some environments. In particular, I think it successfully analyzes the structure of the S-expressions you use as examples. And I agree that it doesn't match the same strings as those modified versions; for example, it doesn't recognize $ as a symbol, or allow '.
The lexical syntax specified in R⁷RS §7.1.2 is, I think, a lot more complex than S-expressions, and it does not purport to define S-expressions; it includes Unicode, booleans, bytevectors, strings, piped symbols, vectors, circular references, quote, quasiquote, unquote, unquote-splicing, hexadecimal numbers, octal numbers, binary numbers, floating-point numbers, and so on. LISP 1.5 couldn't handle any of these, but I would be very uneasy with the assertion, "LISP 1.5 couldn't read S-expressions."
As an example of a more fleshed-out S-expression language, a few years ago when I wrote Ur-Scheme (a Scheme-subset compiler in the language it compiles), my S-expression parser supported lists (with optional dotted tails), quote, string literals, symbols, comments, signed decimal integers, booleans, and character literals. (Check out the section "Actual Parsing" in http://canonical.org/~kragen/sw/urscheme/compiler.scm.html.) This is evidently enough to conveniently write a compiler in, and I think some of it doesn't really belong to S-expressions per se — quote, comments, and booleans, for example.
One problem of my suggested data model is that, although the language is closed under concatenation (unlike, say, HTML, XML, JSON, or RFC-822), that concatenation is not semantics-preserving. Given these two Jevko documents:
x [y] z
a [b] c
concatenating them changes the label for [b] from "a" to "z\na", and perhaps more damningly, erases the whitespace before "z". But, since none of the alternative formats (except ndjson and I guess plain uninterpreted binary, ASCII, or Unicode) is closed under concatenation, maybe that's less important.I don't know if you saw the last time this topic came up I linked to https://ogdl.org/, which seems pretty close to a minimal rose-tree notation. They also put some thought into designing a query language for their rose-tree-like data model, which might be adaptable to Jevko — though they label only nodes, and Jevko labels both nodes (with suffixes) and arcs (with prefixes).
Maybe that's the subtitle for Jevko? "A minimal Unicode syntax for ordered trees with labeled nodes and labeled arcs." If that's the intended semantics it would be pretty easy to whip up a diagram in Dot to illustrate it.
Indeed we do! :)
> As for S-expressions, I agree that what I wrote above is a pretty minimal definition of S-expressions, but I still think it's valid as a definition of "S-expression", even though you'd use a modified definition in some environments. In particular, I think it successfully analyzes the structure of the S-expressions you use as examples. And I agree that it doesn't match the same strings as those modified versions; for example, it doesn't recognize $ as a symbol, or allow '.
For anybody trying to parse your definition:
ex = ("(" *ex ")" / *%x2a-10ffff / %x22 *(%x22 %x22 / %x0-21 / %x23-10ffff) %x22 ) *%x0-20
it's missing `*()` around the right-hand side: ex = *(("(" *ex ")" / *%x2a-10ffff / %x22 *(%x22 %x22 / %x0-21 / %x23-10ffff) %x22 ) *%x0-20)
;POther than that it's a indeed lovely minimal definition. And it allows binary strings!
Now if such a definition was accessible as some formal spec and implemented in various languages the way JSON is, I'd probably not be inclined to rolling my own. Or maybe I still would, because it's fun! ;)
> The lexical syntax specified in R⁷RS §7.1.2 is, I think, a lot more complex than S-expressions, and it does not purport to define S-expressions; it includes Unicode, booleans, bytevectors, strings, piped symbols, vectors, circular references, quote, quasiquote, unquote, unquote-splicing, hexadecimal numbers, octal numbers, binary numbers, floating-point numbers, and so on. LISP 1.5 couldn't handle any of these, but I would be very uneasy with the assertion, "LISP 1.5 couldn't read S-expressions."
Sure. So, would you agree that with proliferation of Lisp variants the term "S-expression" became rather vague?
> As an example of a more fleshed-out S-expression language, a few years ago when I wrote Ur-Scheme (a Scheme-subset compiler in the language it compiles), my S-expression parser supported lists (with optional dotted tails), quote, string literals, symbols, comments, signed decimal integers, booleans, and character literals. (Check out the section "Actual Parsing" in http://canonical.org/~kragen/sw/urscheme/compiler.scm.html.) This is evidently enough to conveniently write a compiler in, and I think some of it doesn't really belong to S-expressions per se — quote, comments, and booleans, for example.
Holy moly, this looks very nice! I like the style. Kudos.
> One problem of my suggested data model is that, although the language is closed under concatenation (unlike, say, HTML, XML, JSON, or RFC-822), that concatenation is not semantics-preserving. Given these two Jevko documents:
x [y] z
a [b] c
> concatenating them changes the label for [b] from "a" to "z\na", and perhaps more damningly, erases the whitespace before "z". But, since none of the alternative formats (except ndjson and I guess plain uninterpreted binary, ASCII, or Unicode) is closed under concatenation, maybe that's less important.Yes, being closed under concatenation is a feature I was aiming for and it indeed does bring with it this issue.
Just something to have in mind when devising formats. A simple solution here is to disallow having anything other than whitespace in the suffix of a Jevko with > 0 children. Then, if a format converts these labels to keys in a map, trimming leading and trailing whitespace, there is no problem. This is how I did it here:
* https://github.com/jevko/easyjevko.js
> I don't know if you saw the last time this topic came up I linked to https://ogdl.org/, which seems pretty close to a minimal rose-tree notation.
Yes, I've seen OGDL before. It's pretty nice. A similar one is https://treenotation.org/
I have experimented with indentation-based syntaxes myself, before settling on brackets.
I have found them to be problematic, at least because:
* For complex structures they become less compact.
* A grammar that correctly captures significant indentation can't really be written in pure *BNF. The way OGDL does it is this:
[12] space(n) ::= char_space*n ; where n is the equivalent number of spaces (can be 0)
[13] block(n) ::= '\' (comment|break) (space(>n) string break)+
I had the same idea. Simple enough, but still. Brackets are simpler to formalize and implement and not harder to explain.> They also put some thought into designing a query language for their rose-tree-like data model, which might be adaptable to Jevko — though they label only nodes, and Jevko labels both nodes (with suffixes) and arcs (with prefixes).
Yes, that might be interesting to look at, thanks for pointing it out. I have thought about this and came up with some ideas, but haven't decided on anything. I was thinking more along the lines of having the path DSL be simply implemented on top of Jevko, not as a completely separate grammar.
> Maybe that's the subtitle for Jevko? "A minimal Unicode syntax for ordered trees with labeled nodes and labeled arcs." If that's the intended semantics it would be pretty easy to whip up a diagram in Dot to illustrate it.
It's a nice description, but I think a little to detailed and technical to fit into a tagline. Maybe a little explanatory article with the diagram included. Would probably look something like this:
https://github.com/jevko/writing/blob/main/2022-01-10-jevko-...
Although I'd gladly see your take on it. ;)
If you want to allow binary strings you'd probably want to define the language over bytes rather than Unicode code points, though Markus Kuhn's UTF-8B (implemented in Python as "surrogateescape") can give you the best of both worlds.
Yes, I agree that "S-expression" is rather vague, much like "CSV".
If you disallow having anything other than whitespace in a Jevko with >0 children, Jevko becomes a rose-tree notation, except that the root is unlabeled, so really it's more like a rose-forest notation. If you want a rose-forest notation, you can get it in a less irregular way by declaring that Jevko infers an extra [] following a suffix containing any non-whitespace character, so foo[bar] is equivalent to foo[bar[]], foo[bar[]baz] is equivalent to foo[bar[]baz[]], and a[b]c is equivalent to a[b[]]c[], but foo[bar[] ] is equivalent to foo[bar[]] and not foo[bar[][]].
This change would foreclose the possibility of having significant leading or trailing whitespace in some places but not others, the way I was suggesting.
Rose trees or rose forests are definitely simpler than Jevko's current data model, and they are equally powerful.
Agreed about indentation.
With respect to diagrams, git clone http://canonical.org/~kragen/sw/pavnotes2.git and look at {horse,johnsmith,player}.{jpg,jevko,dot}. For the diagrams I've taken the liberty of ignoring leading and trailing whitespace on arc labels (prefixes), as well as suffixes consisting only of whitespace. johnsmith and horse are your examples, but they just use Jevko as a terser syntax for JSON that suffers a whitespace problem.
player.jevko is an example I cooked up based on Minetest's relational database schema, recast as a hierarchical schema. It also takes advantage of the ordered nature of subJevkos, the possibility of multiple identical prefixes in the same parent node, and the possibility of having both "headers" (subJevkos) and a "body" (suffix) in the same node.
(As far as I’m concerned, everyone is free to redistribute these nine files, in whole or in part, modified or unmodified, with or without credit; I waive all rights associated with them to the maximum extent possible under applicable law. Where applicable, I abandon their copyright to the public domain. To the extent that I wrote them at all, I wrote and published them in Argentina in 2022. Today, in fact. But you wrote horse.jevko and johnsmith.jevko, deriving them from things on Wikipedia, so my abandonment of copyright here shouldn't be construed as claiming that I wrote them.)
You're totally right, sorry! It's an important aspect of S-exps that I overlooked. I forgot that a typical Lisp file is not a valid S-expression, but a concatenation of a bunch of them. Plus maybe some whitespace/comments prepended.
So an S-expression syntax, as normally defined, is not closed under concatenation like Jevko.
To solve that (assuming it's a desirable feature as I do) purely syntactically, without talking about the REPL, you'd need another top-level syntax.
So why not simplify things and make one that meets this criterion by itself?
Which makes me think of another aspect where Jevko is simpler.
The restrictions Jevko puts on a unicode sequence are that the escape character must be followed by itself or a square bracket and that unescaped square brackets must be balanced.
This is pretty much an exhaustive one-sentence description that would allow you to validate or generate a Jevko.
A similarly simple description is impossible to state for S-expressions.
I tried and I came up with (for your minimal defnition of S-exp):
The restrictions S-expr puts on a unicode sequence are that it must conform to one of three alternatives:
* it must be wrapped in parens (list),
* it must be wrapped in quotes (string),
* it must not include whitespace or parens (symbol).
Also:
* If it's a list then different parts of the sequence may conform to one of these three alternatives.
* If it's a string then anything goes, provided that quote is escaped.
I don't think this is exhaustive and certainly not as clear and simple as the Jevko version. Can you come up with a simpler description?
> If you want to allow binary strings you'd probably want to define the language over bytes rather than Unicode code points, though Markus Kuhn's UTF-8B (implemented in Python as "surrogateescape") can give you the best of both worlds.
Yeah, I don't know why I thought your syntax allows binary strings. Clearly, I was confused.
Thanks for the reference, saved.
> Yes, I agree that "S-expression" is rather vague, much like "CSV".
This is an issue I want to avoid with Jevko.
The best syntax in this regard is JSON, with a relatively clear and concise specification. It does have some holes which cause problems, but I don't know anything better.
I'd like Jevko to be better.
> If you disallow having anything other than whitespace in a Jevko with >0 children, Jevko becomes a rose-tree notation, except that the root is unlabeled, so really it's more like a rose-forest notation.
Yes, depending on how you define these terms.
Important thing to note here again is that in accordance with how I intend Jevko to be used, you wouldn't be changing Jevko itself, but defining a format on top of Jevko[0].
The input of the parser/processor for this format is a Jevko syntax tree.
The output would be a rose tree. A nonblank suffix would be a (syntax) error in this format.
> If you want a rose-forest notation, you can get it in a less irregular way by declaring that Jevko infers an extra [] following a suffix containing any non-whitespace character, so foo[bar] is equivalent to foo[bar[]], foo[bar[]baz] is equivalent to foo[bar[]baz[]], and a[b]c is equivalent to a[b[]]c[], but foo[bar[] ] is equivalent to foo[bar[]] and not foo[bar[][]].
> This change would foreclose the possibility of having significant leading or trailing whitespace in some places but not others, the way I was suggesting.
Sure, you could also define a format that works like this.
Input a Jevko tree and output another tree with these extra subtrees "inferred".
The point is that you don't modify Jevko, but use it as a building block.
> Rose trees or rose forests are definitely simpler than Jevko's current data model, and they are equally powerful.
I think defining Jevko in terms of rose trees wouldn't be as straightforward as the inverse, so I disagree. But I wouldn't want to argue about this. I say go with whatever is simpler/more powerful/works for you.
> With respect to diagrams, git clone http://canonical.org/~kragen/sw/pavnotes2.git and look at {horse,johnsmith,player}.{jpg,jevko,dot}.
This is wonderful! I love it!
Thank you for putting it in the public domain.
I'd like to put this somewhere on a Jevko-related page(s), giving credit to you, linking to http://canonical.org/~kragen/.
If you prefer a different way of attribution or none at all, please let me know here or send me a gmail at darius.j.chuck
> For the diagrams I've taken the liberty of ignoring leading and trailing whitespace on arc labels (prefixes), as well as suffixes consisting only of whitespace. johnsmith and horse are your examples, but they just use Jevko as a terser syntax for JSON that suffers a whitespace problem.
Makes sense. Trimming whitespace is what is usually wanted, except in markup formats.
Perhaps it would be sensible to devise a very thin format on top of Jevko for making diagrams of this kind. It would take care of the trimming and have explicitly attached the semantics you mentioned before as proposed a tagline for Jevko:
> A minimal Unicode syntax for ordered trees with labeled nodes and labeled arcs.
It could be used as a tagline for this format!
Did you generate these diagrams from Jevko or write them by hand?
A tool to autogenerate diagrams from this format could be a good application for it.
Anyway, just throwing out ideas. Now to implement them...
> player.jevko is an example I cooked up based on Minetest's relational database schema, recast as a hierarchical schema. It also takes advantage of the ordered nature of subJevkos, the possibility of multiple identical prefixes in the same parent node, and the possibility of having both "headers" (subJevkos) and a "body" (suffix) in the same node.
That's a really nice use of suffixes!
This is really awesome work, thanks for caring and putting in the time. I'd love to see more stuff like this, people making use of Jevko. That's what it's made for!
Cheers!
[0] Another important thing to state, which I suspect may be causing some confusion in our discussion is that Jevko by design has no semantics at all. You attach the semantics on the format layer above Jevko.
FWIW the most minimal attempted standard of an s-expression would be rivest's basic s-expressions, which, as an extension of djb's netstrings can represent arbitrary binary values:
<sexpr> :: <string> | <list>
<string> :: <display>? <simple-string> ;
<simple-string> :: <raw> ;
<display> :: "[" <simple-string> "]" ;
<raw> :: <decimal> ":" <bytes> ;
<decimal> :: <decimal-digit>+ ;
-- decimal numbers should have no unnecessary leading zeros
<bytes> -- any string of bytes, of the indicated length
<list> :: "(" <sexp>* ")" ;
<decimal-digit> :: "0" | ... | "9" ;The most wonderful to me are the simplest ones, and Jevko grows out of the same spirit as them.
However, it does not attempt to be a new flavor of S-expressions and diverges in ways which to me are worth looking at. I hope it can appeal and be useful not only to minimalist syntax enthusiasts.
BTW Some time ago I've been also experimenting with binary versions of Jevko, certainly with inspiration from both netstrings and Rivest's csexps:
https://github.com/jevko/binary-experiments#asttolengthprefi...
Since then I had some more ideas which I hope to get around to implementing at some point.
I guess the main one is this:
> If your audience is people like me, I think it would probably be worthwhile for you to spend some time up front describing the intended semantics of a data model, as I've attempted above, rather than leaving people to infer it from the grammar. (Maybe OCaml is not a good way to explain it, though.) You might also want to specify that leading and trailing whitespace in prefixes is not significant, though it is in the suffix ("body"); this would enable people to format their name-value pairs readably without corrupting the data. As far as I can tell, this addendum wouldn't interfere with any of your existing uses for Jevko, though in some cases it would simplify their implementations.
You're right, things should be explained more clearly (TODO). Especially the exact role of Jevko and treatment of whitespace. I'll try to improve that.
Here is a sketch of an explanation.
Plain Jevko is meant to be a low-level syntactic layer.
It takes care of turning a unicode sequence into a tree.
On this level, all whitespace is preserved in the tree.
To represent key-value pairs and other data, you most likely want another layer above Jevko -- this would be a Jevko-based format, such as queryjevko (somewhat explained below) or, a very similar one, easyjevko, implemented and very lightly documented here: https://github.com/jevko/easyjevko.js
Or you could have a markup format, such as https://github.com/jevko/markup-experiments#asttoxml5
This format layer defines certain restrictions which may make a subset of Jevkos invalid in it.
It also specifies how to interpret the valid Jevkos. This includes the treatment of whitespace, e.g. that a leading or trailing whitespace in prefixes is insignificant, but conditionally significant in suffixes, etc.
Different formats will define different restrictions and interpretations.
For example:
# queryjevko
queryjevko is a format which uses (a variant of) Jevko as a syntax. Only a subset of Jevko is valid queryjevko.
> I think this is a more useful level of abstraction, and it's more or less the level used by, for example, queryjevko.js's jevkoToJs, although that erroneously uses () instead of [].
The `()` are used on purpose -- queryjevko is meant to be used in URL query strings and be readable. If square brackets were used, things like JS' encodeURIComponent would escape them, making the string unreadable. Using `()` solves that. "~" is used instead of "`" for the same reason. So technically we are dealing not with a spec-compliant Jevko, but a trivial variant of it. Maybe I should write a meta-spec which allows one to pick the three special characters before instantiating itself into a spec. Anyway the parser implementation is configurable in that regard, so I simply configure it to use "~()" instead of "`[]".
> (Also, contrary to your assertion above that this is an example of "leaving [Jevko's data model] as-is", it forgets the order of the name-value pairs as well as I guess all but one of any duplicate set of fields with the same name and also the possibility that there could be both fields and a body.)
I meant [whitespace] rather than [Jevko's data model].
Again, queryjevko is a format which uses Jevko as an underlying syntax. It specifies how syntax trees are converted to JS values, by restricting the range of valid Jevkos. It also specifies conversion in the opposite direction, likewise placing restrictions on JS values that can be safely converted to queryjevko.
The order of name-value pairs happens to get preserved (because of the way JS works), but that's not necessarily relevant. If I were to write a cross-language spec for queryjevko, I'd probably specify that this shouldn't be relied upon.
Duplicate fields and Jevkos with both fields and a non-whitespace body will produce an error when converting Jevko->JS.
I hope this clarifies things somewhat.
Lastly, I'll respond to this for completeness:
> (By the way, if you want to attribute your JSON example for copyright reasons, you need to attribute it to its author or authors, not to the Wikipedia, which is just the site they posted it on.)
According to this:
https://en.wikipedia.org/wiki/Wikipedia:Reusing_Wikipedia_co...
there are 3 options, one of them being what I did, which is to include a link.
I think that's all.
Have a good one!