Jevko: a minimal general-purpose syntax
djedr.github.io
djedr.github.io
1. It relies on later stages to do things like convert to native values or remove whitespace. As such an intermediary can't really "understand" the document very well. E.g., if you are using this syntax to express key/value then you can only know the keys if you know the whitespace rules that will be used to interpret the keys.
2. It reminds me more of XML than S-Expressions. If you are familiar with the ElementTree [1] representation of XML then these parsed examples become more familiar, just a prefix instead of suffix, and no attribute children.
3. Connecting whitespace and XML, I suppose this explains why so many XML formats ended up only using attributes for text values, and used tags and children only for nesting. Otherwise interpretation requires understanding how to interpret these mixed documents with unclear whitespace rules. But Jevko doesn't have attributes.
4. If an intermediary can't understand it then why define a syntax at all?
5. I don't see any escaping rules so including literal [] seems... impossible? You can essentially deserialize the parsed value to reintroduce them, but that's weird and awkward, and only works with balanced brackets anyway.
6. No way to embed binary data, or otherwise uninterpreted data. JSON has the same problem. Base64 encoding things is yet another way in which the data is not interpretable by intermediaries.
I think XML shows that "just strings" isn't fatal, and lots of things get jammed into strings in JSON too, you'll never be able to reproduce every type in a general purpose serialization format (I guess XML with namespaces attempted, but also clearly failed). So I can see a place for something that defines a tree where all leaves are strings. But this doesn't seem like a very clean way to do that.
[1] https://docs.python.org/3/library/xml.etree.elementtree.html...
From the point of view of Jevko, handling whitespace is a higher-level concern. Indeed, you typically will need additional rules to communicate effectively with it.
This is where you should specify a format (which should be standardized) that you use, e.g.:
* https://github.com/jevko/easyjevko.lua
Note: this is just a simple library I wrote recently that does the most straightforward thing imaginable. I haven't yet wrote a spec for the format.
> 2. It reminds me more of XML than S-Expressions. If you are familiar with the ElementTree [1] representation of XML then these parsed examples become more familiar, just a prefix instead of suffix, and no attribute children.
Yes, Jevko has the features of both XML and S-expressions (or neither, depending on how you look). It's supposed to be uniquely suitable for both markup and encoding of data/code in a simple way.
> 3. Connecting whitespace and XML, I suppose this explains why so many XML formats ended up only using attributes for text values, and used tags and children only for nesting. Otherwise interpretation requires understanding how to interpret these mixed documents with unclear whitespace rules. But Jevko doesn't have attributes.
A markup format built on Jevko can have the notion of attributes and rules for whitespace, e.g. see https://news.ycombinator.com/item?id=33334774 -- a nice thing about this particular format is that text nodes are explicitly specified, so you know exactly where your significant whitespace goes.
The nice thing about attributes made with Jevko over XML attributes is that you could naturally make them extensible (which is a pain in XML -- if you want to turn your unstructured string value stored in an attribute into a tree you have a problem).
> 5. I don't see any escaping rules so including literal [] seems... impossible? You can essentially deserialize the parsed value to reintroduce them, but that's weird and awkward, and only works with balanced brackets anyway.
That's not correct. There are only 3 special symbols (delimiters) and they can all be escaped. It's all in the specification.
> 6. No way to embed binary data, or otherwise uninterpreted data. JSON has the same problem. Base64 encoding things is yet another way in which the data is not interpretable by intermediaries.
Indeed, Jevko is not a binary format. Although I've been experimenting with binary equivalents, e.g.:
https://github.com/jevko/binary-experiments
At some point I might go forward with one, but that's not the focus right now.
This would be in line with how Jevko is intended to be used.
Taking the example from github page (abbreviated):
first name [John]
last name [Smith]
address [
street address [21 2nd Street]
]
why isn't it equivalent to this json? {
"first name ": "John",
"\nlast name ": "Smith",
"\naddress ": {
"\n street address ": "21 2nd Street"
}
}
somehow the whitespace is instead implicitly magiced away, but that isn't defined in the grammar anywhere as far as I can see? {
"subjevkos": [
{
"prefix": "first name ",
"jevko": {
"subjevkos": [],
"suffix": "John"
}
},
{
"prefix": "\nlast name ",
"jevko": {
"subjevkos": [],
"suffix": "Smith"
}
},
{
"prefix": "\nis alive ",
"jevko": {
"subjevkos": [],
"suffix": "true"
}
},
{
"prefix": "\nage ",
"jevko": {
"subjevkos": [],
"suffix": "27"
}
},
{
"prefix": "\naddress ",
"jevko": {
"subjevkos": [
{
"prefix": "\n street address ",
"jevko": {
"subjevkos": [],
"suffix": "21 2nd Street"
}
}
],
"suffix": "\n"
}
}
],
"suffix": "",
"opener": "[",
"closer": "]",
"escaper": "`"
}
tbh seems like a pita to clean up afterwardsWhat you are getting is the plain Jevko parse tree.
The examples explicitly talk about a format built on top of that. See this answer: https://news.ycombinator.com/item?id=33334314
Jevko itself, defined by the spec, is not the same as a format based on Jevko.
If you use a plain spec-compliant Jevko parser you for this you indeed should have all whitespace included. Although it will look more like this:
{
subjevkos: [
{ prefix: "first name ", jevko: { subjevkos: [Array], suffix: "John" } },
{ prefix: "\nlast name ", jevko: { subjevkos: [Array], suffix: "Smith" } },
{ prefix: "\nis alive ", jevko: { subjevkos: [Array], suffix: "true" } },
{ prefix: "\nage ", jevko: { subjevkos: [Array], suffix: "27" } },
{ prefix: "\naddress ", jevko: { subjevkos: [Array], suffix: "\n" } }
],
suffix: ""
}
Now on top of that you can build a format that specifies what happens to whitespace, etc. Usually you'd want it trimmed if the prefixes mean keys in a map.The simplest format/library that does just that is this:
* https://github.com/jevko/easyjevko.lua
* https://github.com/jevko/easyjevko.js
That will give you the JSON (or equivalent) that you expect (with autotrimmed keys).
Jevko leaves whitespace handling not to clients, but to formats built on Jevko which would be the thing the clients typically use (rather than plain Jevko). Such formats of course should be clearly specified, which is something I hope to get around to soon.
I already have a good candidate format for specifying first, which has 2 implementations:
* https://github.com/jevko/easyjevko.js
* https://github.com/jevko/easyjevko.lua
For more details see also:
Jevko seems "simpler" than traditional syntaxes in the same way that a heap of bricks is simpler than a wall.
This is where the rules for whitespace would be specified. These can then be taken advantage of by tools to reformat, etc.
The simplest format that I came up with and should specify next that will most likely do what you want is Easy Jevko. For details see this rabbit hole: https://news.ycombinator.com/item?id=33337098
For an alternative minimal syntax which is based around whitespace check out:
How is it simpler or lower-level than s-exprs?
To see how it's syntactically simpler than S-exps note that the grammar of Jevko can be condensed into one short line of ABNF:
Jevko = *("[" Jevko "]" / "`" ("`" / "[" / "]") / %x0-5a / %x5c / %x5e-5f / %x61-10ffff)
The grammar of S-exps on the other hand, I won't quote here, but I assure you it's much more complicated. How much depends on your flavor (Jevko is also simpler in this regard: there is only one flavor, clearly specified).There is no (intended) ambiguity around whitespace in Jevko: whitespace does not occur explicitly in the grammar. Whitespace characters are just characters. This is the defining feature of the syntax.
For this reason Jevko is more low-level: if you want to treat whitespace in some special way, you have to do that yourself. Although for most use-cases this is very similar and simple, e.g. https://news.ycombinator.com/item?id=33334314
But the point is that you can also leave it as-is, e.g.: https://github.com/jevko/queryjevko.js
or do something else -- it's up to your format.
Has anyone ever complained that S-expr were too complex? Adding white noise as a tradeoff doesn't seem like a win.
Ask parser writers, I guess? I had some similar devices, explicitly made so that the overall set of parsing rules would be clear, so the idea is likely not too new.
I've written many parsers, including for s-expr following rivest's RFC. S-expr take about half a dozen states to pull. A packrat parser for s-expr can fit in a single page. You don't even have to scroll to see the whole implementation. What are you talking about?
What exactly do you feel parser writers will tell you regarding s-expr?
ex = ("(" *ex ")" / *%x2a-10ffff) *%x0-20
We could write the data model corresponding to this grammar in OCaml as type ex = List of ex list | Atom of string
If we add double-quoted strings with GW-BASIC/SQL-style quoting (and resolve the ambiguity greedily): ex = ("(" *ex ")" / *%x2a-10ffff / %x22 *(%x22 %x22 / %x0-21 / %x23-10ffff) %x22 ) *%x0-20
This corresponds to the data model type ex = List of ex list | String of string | Symbol of string
This still seems both simpler and more expressive than the one-line grammar you give.
Maybe I'm missing a subtlety of S-expressions here, and I haven't tried it, but I think this correctly parses your examples like (first-name"John"last-name"Smith"is-alive true age 27
address(street-address "21 2nd Street"
city"New York"state"NY"postal-code"10021-3100")
phone-numbers((type "office"number"212 555-1234")
(type "home"number"646 555-4567"))
children()spouse())
or (
:first-name "John"
:last-name "Smith" ...
)
though of course the name-value pairing is lost there (because S-expressions lack it).(By the way, if you want to attribute your JSON example for copyright reasons, you need to attribute it to its author or authors, not to the Wikipedia, which is just the site they posted it on.)
— ⁂ —
Maybe more importantly, though, I think your abbreviated Jevko grammar is wrong in an important way, though it describes the same set of strings as the full grammar. According to the abbreviated grammar, the Jevko examples lack most of the structure of the XML, JSON, and S-expression versions. It implies that a Jevko is an ordered sequence of (Unicode!) characters and nested Jevkos (indicated by []). We could express this data model as
type jevko = atom list
and atom = Char of char | Nest of jevko
If so, then abbreviating your S-expression example (first-name "John" last-name "Smith")
the closest Jevko equivalent is not, as you claim first name [John]
last name [Smith]
but rather (supposing the hyphens were just an unfortunate concession to S-expression syntax rather than actually desired) [first name][John][last name][Smith]
We had to sacrifice the formatting white space because there's nowhere that Jevko (as specified above!) ignores it.Using this grammar, the S-expression equivalent of the Jevko
first name [John]
is rather ("f" "i" "r" "s" "t" " " "n" "a" "m" "e" " " ("J" "o" "h" "n"))
If we instead use the full Jevko grammar Jevko = *Subjevko Suffix
Subjevko = Prefix "[" Jevko "]"
Prefix = Text
Suffix = Text
Text = *Symbol
Symbol = Digraph / Character
Digraph = "`" ("`" / "[" / "]")
Character = %x0-5a / %x5c / %x5e-5f / %x61-10ffff
or, I think, equivalently in this context: Jevko = *Subjevko Text
Subjevko = Text "[" Jevko "]"
Text = *(Digraph / Character)
Digraph = "`" ("`" / "[" / "]")
Character = %x0-5a / %x5c / %x5e-5f / %x61-10ffff
then we do preserve the name-value structure you seem to be going for, which your above one-line version loses. And this allows us to write first name[John]last name[Smith]
as in your compactness examples.I think this is a more useful level of abstraction, and it's more or less the level used by, for example, queryjevko.js's jevkoToJs, although that erroneously uses () instead of []. (Also, contrary to your assertion above that this is an example of "leaving [Jevko's data model] as-is", it forgets the order of the name-value pairs as well as I guess all but one of any duplicate set of fields with the same name and also the possibility that there could be both fields and a body.)
Essentially at this level of structure a Jevko is a (possibly empty) set of name-value pairs followed by a plaintext body ("Suffix"). This is exactly like an email message, except that the values are themselves Jevkos. In OCaml we could write:
type jevko = Jevko of (string * jevko) list * string
For your example include[author]fields[articles[[title][body]]people[[name]]]
this gives the representation Jevko
([("include", Jevko ([], "author"));
("fields",
Jevko
([("articles",
Jevko ([("", Jevko ([], "title")); ("", Jevko ([], "body"))], ""));
("people", Jevko ([("", Jevko ([], "name"))], ""))],
""))],
"")
although I notice that queryjevko handles it very differently.— ⁂ —
Unlike the data model implied by your one-line grammar, I think this is an extremely useful data model. Email messages are the one and only structured data format that has remained compatible in active use and extension for over half a century; you can literally take internet email messages from 01972 and load them into a mail client today (most easily by putting them into a qmail-style maildir) and everything will just work. This is largely a result of the decentralized extensibility properties of name-value pairs: mail clients just ignore header names they don't understand, and they don't require header names that weren't originally present. This is also the basis of the extensibility of HTTP/1.0 and HTTP/1.1.
It's also very similar to the data model of popular "semistructured" or "free-form" databases of the 01990s like askSam and Filemaker. Unlike RFC-822 and HTTP/1, but like Jevko, those systems support recursively nested data. askSam even used almost the same firstname[John] syntax.
This data model does have a semantic mismatch with things like the rose-tree representation you describe at https://xtao.org/blog/rose.html, since it associates the rose-tree label with the first branch and the empty-string label with subsequent branches.
If your audience is people like me, I think it would probably be worthwhile for you to spend some time up front describing the intended semantics of a data model, as I've attempted above, rather than leaving people to infer it from the grammar. (Maybe OCaml is not a good way to explain it, though.) You might also want to specify that leading and trailing whitespace in prefixes is not significant, though it is in the suffix ("body"); this would enable people to format their name-value pairs readably without corrupting the data. As far as I can tell, this addendum wouldn't interfere with any of your existing uses for Jevko, though in some cases it would simplify their implementations.
______
Runnable OCaml code:
type jevko = Jevko of (string * jevko) list * string
(* XXX doesn't escape `[ `] `` *)
let rec dump (Jevko(hdrs, body)) = hdr(hdrs) ^ body
and hdr = function [] -> "" | (k, v) :: t -> k ^ "[" ^ dump(v) ^ "]" ^ hdr(t)
let dict kvs = Jevko(kvs, "")
let text s = Jevko ([], s)
let v = dict ["include", text "author";
"fields", dict [
"articles", dict ["", text "title"; "", text "body"];
"people", dict ["", text "name"]
]]
;;
print_endline(dump v)I'll respond to the main points and then expand on the details later.
I'm not sure what the authoritative source on S-expressions is (or even if there is one, which to me is a problem and part of the value proposition here), so I'll take R^7RS[1] as a reference.
If you look at the formal definition there (chapter 7), it's significantly more complex than both Jevko and what you beautifully constructed here.
And indeed, you have just freestyled a simplifed version of S-expressions (impressive!), but this is not the real thing.
If you would keep refactoring it with the constraints I had in mind for Jevko, you'd eventually end up with Jevko.
> though of course the name-value pairing is lost there (because S-expressions lack it).
Indeed, and that's another part of the value proposition of Jevko. The grammar is designed purposefully to take advantage of natural syntactic name-value (prefix-subjevko) pairing tendencies.
Which brings me to the next point.
**
You are absolutely correct that the abbreviated grammar matches the same strings, but doesn't have the same structure.
*The correct grammar is the one in the specification*.
I have shown the condensed version of it to illustrate the point that Jevko is indeed extremely simple. The single line captures all the essential elements and, again, matches the same strings as the full grammar.
This is unlike the similar condensed grammar for S-expressions which you sketched out here. If you would continue, it would get significantly more complex before it matches the same strings.
The OCaml type definition you wrote down should do the job of capturing the structure, although I prefer to name the elements (which may be not-so-convenient, depending on the language, so it's fine).
**
Indeed, I also think that this is an extremely useful data model. Thanks for pointing out the similarity to e-mail messages and other references which I'd love to dig into (please send if you have any links or resources about these databases).
Thanks for all the pointers, I'll take them into account.
And thanks again for your time.
[0] this seems kinda official, but there is no single standard here: https://www.s-expressions.org/standards
As for S-expressions, I agree that what I wrote above is a pretty minimal definition of S-expressions, but I still think it's valid as a definition of "S-expression", even though you'd use a modified definition in some environments. In particular, I think it successfully analyzes the structure of the S-expressions you use as examples. And I agree that it doesn't match the same strings as those modified versions; for example, it doesn't recognize $ as a symbol, or allow '.
The lexical syntax specified in R⁷RS §7.1.2 is, I think, a lot more complex than S-expressions, and it does not purport to define S-expressions; it includes Unicode, booleans, bytevectors, strings, piped symbols, vectors, circular references, quote, quasiquote, unquote, unquote-splicing, hexadecimal numbers, octal numbers, binary numbers, floating-point numbers, and so on. LISP 1.5 couldn't handle any of these, but I would be very uneasy with the assertion, "LISP 1.5 couldn't read S-expressions."
As an example of a more fleshed-out S-expression language, a few years ago when I wrote Ur-Scheme (a Scheme-subset compiler in the language it compiles), my S-expression parser supported lists (with optional dotted tails), quote, string literals, symbols, comments, signed decimal integers, booleans, and character literals. (Check out the section "Actual Parsing" in http://canonical.org/~kragen/sw/urscheme/compiler.scm.html.) This is evidently enough to conveniently write a compiler in, and I think some of it doesn't really belong to S-expressions per se — quote, comments, and booleans, for example.
One problem of my suggested data model is that, although the language is closed under concatenation (unlike, say, HTML, XML, JSON, or RFC-822), that concatenation is not semantics-preserving. Given these two Jevko documents:
x [y] z
a [b] c
concatenating them changes the label for [b] from "a" to "z\na", and perhaps more damningly, erases the whitespace before "z". But, since none of the alternative formats (except ndjson and I guess plain uninterpreted binary, ASCII, or Unicode) is closed under concatenation, maybe that's less important.I don't know if you saw the last time this topic came up I linked to https://ogdl.org/, which seems pretty close to a minimal rose-tree notation. They also put some thought into designing a query language for their rose-tree-like data model, which might be adaptable to Jevko — though they label only nodes, and Jevko labels both nodes (with suffixes) and arcs (with prefixes).
Maybe that's the subtitle for Jevko? "A minimal Unicode syntax for ordered trees with labeled nodes and labeled arcs." If that's the intended semantics it would be pretty easy to whip up a diagram in Dot to illustrate it.
Indeed we do! :)
> As for S-expressions, I agree that what I wrote above is a pretty minimal definition of S-expressions, but I still think it's valid as a definition of "S-expression", even though you'd use a modified definition in some environments. In particular, I think it successfully analyzes the structure of the S-expressions you use as examples. And I agree that it doesn't match the same strings as those modified versions; for example, it doesn't recognize $ as a symbol, or allow '.
For anybody trying to parse your definition:
ex = ("(" *ex ")" / *%x2a-10ffff / %x22 *(%x22 %x22 / %x0-21 / %x23-10ffff) %x22 ) *%x0-20
it's missing `*()` around the right-hand side: ex = *(("(" *ex ")" / *%x2a-10ffff / %x22 *(%x22 %x22 / %x0-21 / %x23-10ffff) %x22 ) *%x0-20)
;POther than that it's a indeed lovely minimal definition. And it allows binary strings!
Now if such a definition was accessible as some formal spec and implemented in various languages the way JSON is, I'd probably not be inclined to rolling my own. Or maybe I still would, because it's fun! ;)
> The lexical syntax specified in R⁷RS §7.1.2 is, I think, a lot more complex than S-expressions, and it does not purport to define S-expressions; it includes Unicode, booleans, bytevectors, strings, piped symbols, vectors, circular references, quote, quasiquote, unquote, unquote-splicing, hexadecimal numbers, octal numbers, binary numbers, floating-point numbers, and so on. LISP 1.5 couldn't handle any of these, but I would be very uneasy with the assertion, "LISP 1.5 couldn't read S-expressions."
Sure. So, would you agree that with proliferation of Lisp variants the term "S-expression" became rather vague?
> As an example of a more fleshed-out S-expression language, a few years ago when I wrote Ur-Scheme (a Scheme-subset compiler in the language it compiles), my S-expression parser supported lists (with optional dotted tails), quote, string literals, symbols, comments, signed decimal integers, booleans, and character literals. (Check out the section "Actual Parsing" in http://canonical.org/~kragen/sw/urscheme/compiler.scm.html.) This is evidently enough to conveniently write a compiler in, and I think some of it doesn't really belong to S-expressions per se — quote, comments, and booleans, for example.
Holy moly, this looks very nice! I like the style. Kudos.
> One problem of my suggested data model is that, although the language is closed under concatenation (unlike, say, HTML, XML, JSON, or RFC-822), that concatenation is not semantics-preserving. Given these two Jevko documents:
x [y] z
a [b] c
> concatenating them changes the label for [b] from "a" to "z\na", and perhaps more damningly, erases the whitespace before "z". But, since none of the alternative formats (except ndjson and I guess plain uninterpreted binary, ASCII, or Unicode) is closed under concatenation, maybe that's less important.Yes, being closed under concatenation is a feature I was aiming for and it indeed does bring with it this issue.
Just something to have in mind when devising formats. A simple solution here is to disallow having anything other than whitespace in the suffix of a Jevko with > 0 children. Then, if a format converts these labels to keys in a map, trimming leading and trailing whitespace, there is no problem. This is how I did it here:
* https://github.com/jevko/easyjevko.js
> I don't know if you saw the last time this topic came up I linked to https://ogdl.org/, which seems pretty close to a minimal rose-tree notation.
Yes, I've seen OGDL before. It's pretty nice. A similar one is https://treenotation.org/
I have experimented with indentation-based syntaxes myself, before settling on brackets.
I have found them to be problematic, at least because:
* For complex structures they become less compact.
* A grammar that correctly captures significant indentation can't really be written in pure *BNF. The way OGDL does it is this:
[12] space(n) ::= char_space*n ; where n is the equivalent number of spaces (can be 0)
[13] block(n) ::= '\' (comment|break) (space(>n) string break)+
I had the same idea. Simple enough, but still. Brackets are simpler to formalize and implement and not harder to explain.> They also put some thought into designing a query language for their rose-tree-like data model, which might be adaptable to Jevko — though they label only nodes, and Jevko labels both nodes (with suffixes) and arcs (with prefixes).
Yes, that might be interesting to look at, thanks for pointing it out. I have thought about this and came up with some ideas, but haven't decided on anything. I was thinking more along the lines of having the path DSL be simply implemented on top of Jevko, not as a completely separate grammar.
> Maybe that's the subtitle for Jevko? "A minimal Unicode syntax for ordered trees with labeled nodes and labeled arcs." If that's the intended semantics it would be pretty easy to whip up a diagram in Dot to illustrate it.
It's a nice description, but I think a little to detailed and technical to fit into a tagline. Maybe a little explanatory article with the diagram included. Would probably look something like this:
https://github.com/jevko/writing/blob/main/2022-01-10-jevko-...
Although I'd gladly see your take on it. ;)
If you want to allow binary strings you'd probably want to define the language over bytes rather than Unicode code points, though Markus Kuhn's UTF-8B (implemented in Python as "surrogateescape") can give you the best of both worlds.
Yes, I agree that "S-expression" is rather vague, much like "CSV".
If you disallow having anything other than whitespace in a Jevko with >0 children, Jevko becomes a rose-tree notation, except that the root is unlabeled, so really it's more like a rose-forest notation. If you want a rose-forest notation, you can get it in a less irregular way by declaring that Jevko infers an extra [] following a suffix containing any non-whitespace character, so foo[bar] is equivalent to foo[bar[]], foo[bar[]baz] is equivalent to foo[bar[]baz[]], and a[b]c is equivalent to a[b[]]c[], but foo[bar[] ] is equivalent to foo[bar[]] and not foo[bar[][]].
This change would foreclose the possibility of having significant leading or trailing whitespace in some places but not others, the way I was suggesting.
Rose trees or rose forests are definitely simpler than Jevko's current data model, and they are equally powerful.
Agreed about indentation.
With respect to diagrams, git clone http://canonical.org/~kragen/sw/pavnotes2.git and look at {horse,johnsmith,player}.{jpg,jevko,dot}. For the diagrams I've taken the liberty of ignoring leading and trailing whitespace on arc labels (prefixes), as well as suffixes consisting only of whitespace. johnsmith and horse are your examples, but they just use Jevko as a terser syntax for JSON that suffers a whitespace problem.
player.jevko is an example I cooked up based on Minetest's relational database schema, recast as a hierarchical schema. It also takes advantage of the ordered nature of subJevkos, the possibility of multiple identical prefixes in the same parent node, and the possibility of having both "headers" (subJevkos) and a "body" (suffix) in the same node.
(As far as I’m concerned, everyone is free to redistribute these nine files, in whole or in part, modified or unmodified, with or without credit; I waive all rights associated with them to the maximum extent possible under applicable law. Where applicable, I abandon their copyright to the public domain. To the extent that I wrote them at all, I wrote and published them in Argentina in 2022. Today, in fact. But you wrote horse.jevko and johnsmith.jevko, deriving them from things on Wikipedia, so my abandonment of copyright here shouldn't be construed as claiming that I wrote them.)
You're totally right, sorry! It's an important aspect of S-exps that I overlooked. I forgot that a typical Lisp file is not a valid S-expression, but a concatenation of a bunch of them. Plus maybe some whitespace/comments prepended.
So an S-expression syntax, as normally defined, is not closed under concatenation like Jevko.
To solve that (assuming it's a desirable feature as I do) purely syntactically, without talking about the REPL, you'd need another top-level syntax.
So why not simplify things and make one that meets this criterion by itself?
Which makes me think of another aspect where Jevko is simpler.
The restrictions Jevko puts on a unicode sequence are that the escape character must be followed by itself or a square bracket and that unescaped square brackets must be balanced.
This is pretty much an exhaustive one-sentence description that would allow you to validate or generate a Jevko.
A similarly simple description is impossible to state for S-expressions.
I tried and I came up with (for your minimal defnition of S-exp):
The restrictions S-expr puts on a unicode sequence are that it must conform to one of three alternatives:
* it must be wrapped in parens (list),
* it must be wrapped in quotes (string),
* it must not include whitespace or parens (symbol).
Also:
* If it's a list then different parts of the sequence may conform to one of these three alternatives.
* If it's a string then anything goes, provided that quote is escaped.
I don't think this is exhaustive and certainly not as clear and simple as the Jevko version. Can you come up with a simpler description?
> If you want to allow binary strings you'd probably want to define the language over bytes rather than Unicode code points, though Markus Kuhn's UTF-8B (implemented in Python as "surrogateescape") can give you the best of both worlds.
Yeah, I don't know why I thought your syntax allows binary strings. Clearly, I was confused.
Thanks for the reference, saved.
> Yes, I agree that "S-expression" is rather vague, much like "CSV".
This is an issue I want to avoid with Jevko.
The best syntax in this regard is JSON, with a relatively clear and concise specification. It does have some holes which cause problems, but I don't know anything better.
I'd like Jevko to be better.
> If you disallow having anything other than whitespace in a Jevko with >0 children, Jevko becomes a rose-tree notation, except that the root is unlabeled, so really it's more like a rose-forest notation.
Yes, depending on how you define these terms.
Important thing to note here again is that in accordance with how I intend Jevko to be used, you wouldn't be changing Jevko itself, but defining a format on top of Jevko[0].
The input of the parser/processor for this format is a Jevko syntax tree.
The output would be a rose tree. A nonblank suffix would be a (syntax) error in this format.
> If you want a rose-forest notation, you can get it in a less irregular way by declaring that Jevko infers an extra [] following a suffix containing any non-whitespace character, so foo[bar] is equivalent to foo[bar[]], foo[bar[]baz] is equivalent to foo[bar[]baz[]], and a[b]c is equivalent to a[b[]]c[], but foo[bar[] ] is equivalent to foo[bar[]] and not foo[bar[][]].
> This change would foreclose the possibility of having significant leading or trailing whitespace in some places but not others, the way I was suggesting.
Sure, you could also define a format that works like this.
Input a Jevko tree and output another tree with these extra subtrees "inferred".
The point is that you don't modify Jevko, but use it as a building block.
> Rose trees or rose forests are definitely simpler than Jevko's current data model, and they are equally powerful.
I think defining Jevko in terms of rose trees wouldn't be as straightforward as the inverse, so I disagree. But I wouldn't want to argue about this. I say go with whatever is simpler/more powerful/works for you.
> With respect to diagrams, git clone http://canonical.org/~kragen/sw/pavnotes2.git and look at {horse,johnsmith,player}.{jpg,jevko,dot}.
This is wonderful! I love it!
Thank you for putting it in the public domain.
I'd like to put this somewhere on a Jevko-related page(s), giving credit to you, linking to http://canonical.org/~kragen/.
If you prefer a different way of attribution or none at all, please let me know here or send me a gmail at darius.j.chuck
> For the diagrams I've taken the liberty of ignoring leading and trailing whitespace on arc labels (prefixes), as well as suffixes consisting only of whitespace. johnsmith and horse are your examples, but they just use Jevko as a terser syntax for JSON that suffers a whitespace problem.
Makes sense. Trimming whitespace is what is usually wanted, except in markup formats.
Perhaps it would be sensible to devise a very thin format on top of Jevko for making diagrams of this kind. It would take care of the trimming and have explicitly attached the semantics you mentioned before as proposed a tagline for Jevko:
> A minimal Unicode syntax for ordered trees with labeled nodes and labeled arcs.
It could be used as a tagline for this format!
Did you generate these diagrams from Jevko or write them by hand?
A tool to autogenerate diagrams from this format could be a good application for it.
Anyway, just throwing out ideas. Now to implement them...
> player.jevko is an example I cooked up based on Minetest's relational database schema, recast as a hierarchical schema. It also takes advantage of the ordered nature of subJevkos, the possibility of multiple identical prefixes in the same parent node, and the possibility of having both "headers" (subJevkos) and a "body" (suffix) in the same node.
That's a really nice use of suffixes!
This is really awesome work, thanks for caring and putting in the time. I'd love to see more stuff like this, people making use of Jevko. That's what it's made for!
Cheers!
[0] Another important thing to state, which I suspect may be causing some confusion in our discussion is that Jevko by design has no semantics at all. You attach the semantics on the format layer above Jevko.
FWIW the most minimal attempted standard of an s-expression would be rivest's basic s-expressions, which, as an extension of djb's netstrings can represent arbitrary binary values:
<sexpr> :: <string> | <list>
<string> :: <display>? <simple-string> ;
<simple-string> :: <raw> ;
<display> :: "[" <simple-string> "]" ;
<raw> :: <decimal> ":" <bytes> ;
<decimal> :: <decimal-digit>+ ;
-- decimal numbers should have no unnecessary leading zeros
<bytes> -- any string of bytes, of the indicated length
<list> :: "(" <sexp>* ")" ;
<decimal-digit> :: "0" | ... | "9" ;The most wonderful to me are the simplest ones, and Jevko grows out of the same spirit as them.
However, it does not attempt to be a new flavor of S-expressions and diverges in ways which to me are worth looking at. I hope it can appeal and be useful not only to minimalist syntax enthusiasts.
BTW Some time ago I've been also experimenting with binary versions of Jevko, certainly with inspiration from both netstrings and Rivest's csexps:
https://github.com/jevko/binary-experiments#asttolengthprefi...
Since then I had some more ideas which I hope to get around to implementing at some point.
I guess the main one is this:
> If your audience is people like me, I think it would probably be worthwhile for you to spend some time up front describing the intended semantics of a data model, as I've attempted above, rather than leaving people to infer it from the grammar. (Maybe OCaml is not a good way to explain it, though.) You might also want to specify that leading and trailing whitespace in prefixes is not significant, though it is in the suffix ("body"); this would enable people to format their name-value pairs readably without corrupting the data. As far as I can tell, this addendum wouldn't interfere with any of your existing uses for Jevko, though in some cases it would simplify their implementations.
You're right, things should be explained more clearly (TODO). Especially the exact role of Jevko and treatment of whitespace. I'll try to improve that.
Here is a sketch of an explanation.
Plain Jevko is meant to be a low-level syntactic layer.
It takes care of turning a unicode sequence into a tree.
On this level, all whitespace is preserved in the tree.
To represent key-value pairs and other data, you most likely want another layer above Jevko -- this would be a Jevko-based format, such as queryjevko (somewhat explained below) or, a very similar one, easyjevko, implemented and very lightly documented here: https://github.com/jevko/easyjevko.js
Or you could have a markup format, such as https://github.com/jevko/markup-experiments#asttoxml5
This format layer defines certain restrictions which may make a subset of Jevkos invalid in it.
It also specifies how to interpret the valid Jevkos. This includes the treatment of whitespace, e.g. that a leading or trailing whitespace in prefixes is insignificant, but conditionally significant in suffixes, etc.
Different formats will define different restrictions and interpretations.
For example:
# queryjevko
queryjevko is a format which uses (a variant of) Jevko as a syntax. Only a subset of Jevko is valid queryjevko.
> I think this is a more useful level of abstraction, and it's more or less the level used by, for example, queryjevko.js's jevkoToJs, although that erroneously uses () instead of [].
The `()` are used on purpose -- queryjevko is meant to be used in URL query strings and be readable. If square brackets were used, things like JS' encodeURIComponent would escape them, making the string unreadable. Using `()` solves that. "~" is used instead of "`" for the same reason. So technically we are dealing not with a spec-compliant Jevko, but a trivial variant of it. Maybe I should write a meta-spec which allows one to pick the three special characters before instantiating itself into a spec. Anyway the parser implementation is configurable in that regard, so I simply configure it to use "~()" instead of "`[]".
> (Also, contrary to your assertion above that this is an example of "leaving [Jevko's data model] as-is", it forgets the order of the name-value pairs as well as I guess all but one of any duplicate set of fields with the same name and also the possibility that there could be both fields and a body.)
I meant [whitespace] rather than [Jevko's data model].
Again, queryjevko is a format which uses Jevko as an underlying syntax. It specifies how syntax trees are converted to JS values, by restricting the range of valid Jevkos. It also specifies conversion in the opposite direction, likewise placing restrictions on JS values that can be safely converted to queryjevko.
The order of name-value pairs happens to get preserved (because of the way JS works), but that's not necessarily relevant. If I were to write a cross-language spec for queryjevko, I'd probably specify that this shouldn't be relied upon.
Duplicate fields and Jevkos with both fields and a non-whitespace body will produce an error when converting Jevko->JS.
I hope this clarifies things somewhat.
Lastly, I'll respond to this for completeness:
> (By the way, if you want to attribute your JSON example for copyright reasons, you need to attribute it to its author or authors, not to the Wikipedia, which is just the site they posted it on.)
According to this:
https://en.wikipedia.org/wiki/Wikipedia:Reusing_Wikipedia_co...
there are 3 options, one of them being what I did, which is to include a link.
I think that's all.
Have a good one!
I followed the link and read the page, but I’m still not sure what the point is!
Apparently for the same things as with XML/JSON/TOML.
> I followed the link and read the page, but I’m still not sure what the point is!
According to the website:
> It has no data types, no semantics, no underlying model of cons cells or anything similar. It’s as close to pure generic syntax as it gets.
> So at the lowest level Jevko is a minimal formal specification for flexible trees of text.
For its simplicity it seems quite powerful to me. I don't see it replacing JSON/markdown or even TOML. But is seems trivial to implement on very low-level, older, less bloated, or less powerful systems.
Every time I've written or otherwise dealt with JSON/XML/etc. I wished I was dealing with something simpler, so I created it. If I had Jevko as a full-fledged alternative to JSON or XML, supported by tools, etc., I'd pick Jevko in a heartbeat. I like minimalism. Some people like it too, so perhaps Jevko can serve them well. There are many good reasons.
> But is seems trivial to implement on very low-level, older, less bloated, or less powerful systems.
Indeed, this is THE feature and a realistic application. I hope!
For a naive unrealistic vision that I sketched out some time ago see:
https://github.com/jevko/writing/blob/main/2022-01-13-vision...
Say you build your own language, but don't want to write a parser for it (people who write tons of prog langs know that syntax is the most boring part (for some, I'm sure some people love it)) so you can borrow this, it gives you a syntax tree, and you just continue from there.
E.g. you introduce "Subjevko" as if it's known knowledge.
Also, the singular link to "examples" on that page, which is what I really want to see first, was broken.
My advice: for every point you make have a single, simple, illustrative example right below the explanatory text (or better yet, right above it!).
What is the "first page" that you are referring to?
Can you paste a link to it along with the broken examples link?
This Hacker News submission features the blog post under this URL:
https://djedr.github.io/posts/jevko-2022-02-22.html
Clearly, you are not talking about this page, as that contains multiple links rather than a singular link.
Perhaps you are talking about the specification which is here:
https://github.com/jevko/specifications/blob/master/spec-sta...
(linked from the blog post)
and here:
(linked from jevko.org)
All three link to Jevko examples here:
https://github.com/jevko/examples
but all these examples links seem to be correct on my end.
I agree about the importance of examples, and I try to lead with them on jevko.org and jevko.github.io (which are the front pages of Jevko -- possibly I should merge them into one).
However a formal specification is not necessarily the place to put the leading examples.
This is also where the Subjevko rule is defined. It isn't quite introduced as "known knowledge" -- the purpose of a specification is to define the unknown, more or less from the ground up. This is also why specifications tend to get a little abstract. Jevko's spec is no exception. This should be in line with expectations of authors of tools such as parsers, validators, generators, or other kinds of processors, for which the spec is the authoritative reference.
It is not necessarily the best first place to look for explanation, if you are approaching from a more casual side.
I agree that from that side a clear picture of what Jevko is and how it can be used is still lacking. I certainly should add more examples and explain the concepts with analogies.
So I appreciate the essence of your advice and hope I'll manage to improve on that.
p [Hello world!]
Changes it to: p [[Hello ] b[world!]]Every HTML tag (except for `text`) is an implementation of the base function `node`, which is:
node : String -> List Attr msg -> List Html msg -> Html msg
So in the base library, the `p` tag is implemented as p : List Attr msg -> List Html msg -> Html msg
p = node "p"
Which means that in Elm, the above content would look like this: p [] [ text "Hello ", b [] [ text "World" ] ]
i.e., every HTML tag can accept as parameters: - A list of attributes
- A list of child tags
I've found it to be very elegant to work with; and with a good auto-formatter, it's really easy to build complicated HTML without losing your place and dropping a closing tag.The advantage of a Jevko-based format over the above is that it has the simplicity and elegance parameters (IMO) cranked up to maximum! ;)
Now, for the tool support...
This is one way. I've experimented with many others. Some semi-documented here:
https://github.com/jevko/markup-experiments
Some ways don't need additional level of nesting, but it's a trade-off.
In the end something like the way you linked is my favorite, because it's very simple.
It also gives you explicit control over your text nodes.
Attributes can also be added in many ways, but I think I prefer them simply looking like child nodes, except with `=` appended to attribute name:
p [class=[sth] [content]]
Technically you can mix attributes and elements this way, but it does no harm.Jevko is the evolved version of that (I'm the author).
For a simple higher-level format, have look at Easy Jevko:
* https://github.com/jevko/easyjevko.lua – implementation in Lua
* https://github.com/jevko/easyjevko.js – implementation in JS
A way to complicate it is to add a schema like: https://github.com/jevko/interjevko.js
A lousy demo of that (with syntax coloring!) is here: https://jevko.github.io/interjevko.bundle.html
I hope to work out simple specs for these formats eventually.
Cheerio
Other examples: in XML when you write `<abc></abc>` is "abc" a string? In s-expressions (or lisp) when you write `(if a b)` are "if", "a" and "b" strings?
Likewise, in the Jevko expression `a [b]` I don't see how `b` is not a string.
Because it isn't. There is no concept of a string. `if [[a] [b]]` is similar to `(if a b)`. So in your example `a [b]` is similar to s-expression `(a b)`. Neither "a" nor "b" are strings here.
Given the Jevko syntax [b b], is that a repetition of the same object, or is that two different symbols that have the same name?
Is it like (b b) or (#:b #:b)? Why or why not; based on what documentation?
In symbol manipulation, symbols are atoms. Each symbol represents an identity. An identity cannot be fractured into sub-identities. The textual name of a symbol just provides a written notation; all semantics of symbol manipulation concerns itself only with whether two symbols are the same one or a different symbol. For instance the idea that the symbol foo is a prefix of the symbol foobar is invalid. We may only ask whether foo and foobar are the same entity or different.
A basic operation in symbol manipulation is semantics-preserving substitution. If Jevko is based on symbols then [a b b] can be replaced by [a c c] without loss of semantics. In fact, all we are doing is changing the name of the b symbol; its identity remains the same.
Suppose there is some application built on Jevko in which [a=b b] cannot be replaced by [a=b c] without destroying the semantics; suppose it has to be that in that application, if b is replaced by c, then [a=b b] must turn into [a=c c], in order to preserve a relationship that holds between a=b and b. That application has then destroyed the concept of a Jevko symbol, because a=b isn't functioning like an atom any more. It's functioning like two symbols a= b or possibly three symbols a = b.
Thus if a=b is really one symbol, and not a string of three characters, then no legitimate, Jevko-respecting application of Jevko can do anything like this, namely treat a=b as being syntax which codes for two or three atoms.
But that's a lie, because Jevko is intended for doing that sort of thing like it's going out of style.
For instance, in this code:
https://github.com/jevko/jevkodom.js/blob/master/jevkoToHtml...
can you explain what
mid.endsWith('=')? "attr": "tag"
is doing? It sure looks to me like it's asking whether a symbol (i.e. indivisible atom) ends with an equal sign, which is semantic gibberish.Jevko isn't that. It's not a programming language[0]. It's just syntax, with even less semantics than JSON and even simpler than S-expressions.
There are no symbols or references in Jevko.
It's important here to separate syntax and semantics. S-expressions are not Lisp. They are its syntax. Semantics is a separate matter.
@gnulinux and @User23 have it right.
Strings are an implementation detail. You could store the parsed text as numbers or you could Gödel-encode the entire tree into one giant number.
Saying then that Jevko is nothing but numbers is conflating things.
---
Just to clarify, in Jevko `[b b]` doesn't parse to what your description implies you think it does. `b[b]` or `[[b][b]]` would be closer to that. If you are interested in the details, please have a look at the specification or diagrams:
* https://jevko.org/diagram.xhtml
From there you should be able to easily write a parser in your favorite Lisp! ;)
---
> can you explain what
mid.endsWith('=')? "attr": "tag"
> is doing? It sure looks to me like it's asking whether a symbol (i.e. indivisible atom) ends with an equal sign, which is semantic gibberish.There are no symbols or indivisible atoms here.
What's happening here is parsing. `jevkoToHtml` is a kind of parser-transpiler which operates on a syntax tree, rather than a sequence of characters or tokens.
The syntax tree is the output of an earlier stage of parsing, done by the Jevko parser.
So you can think of this as multi-pass parsing, by analogy with multi-pass compilation.
At the same time as this second pass of parsing is happening, translation to HTML is happening as well.
Hope this clarifies things!
---
[0] To clearly see the point, here is a toy programming language which uses Jevko as its syntax: https://github.com/jevko/jevkalk
I agree; and symbol is a semantic entity, not a syntactic one.
S-expressions do not have symbols, they have tokens which become symbols when read into the machine.
Tokens are clumps of characters. i.e. strings.
> There are no symbols or references in Jevko.
> @gnulinux and @User23 have it right.
Those two are insisting that texts are symbols in Jevko, so the above two statements contradict each other.
E.g.
gnulinux> You're not understanding. These are not strings. These are symbols.
User23> Jevko language [...] has no strings. For example, a Common Lisp hosted Jevko parser implementation might choose to represent Jevko text objects as symbols.
> What's happening here is parsing. `jevkoToHtml` is a kind of parser-transpiler which operates on a syntax tree, rather than a sequence of characters or tokens.
That is obviously false. It's operating on a nested list of strings, whereby it's making a semantic decision whether a string in a certain position of a nested item in the structure ends in the = character.
In other words the syntax is character-level. The strings themselves have syntax, which is made of characters. That's in the Jevko specification. Characters are the atoms in Jevko.
The = character that the code is looking for is a %x3D element.
It's right here: https://github.com/jevko/specifications/blob/master/draft-st...
The = character is %x3D. An %x3D is Character. A Character is one of the kinds of Symbol (the other being Digraph). A Symbol is a constituent of Text. Text generates a Prefix or Suffix. A Suffix alone can be a Jevko due to this rule:
Jevko ::= *Subjevko Suffix
Thus, formally Jevko does have Symbols: a Symbol is a Character or Digraph.A Text is a sequence of Symbols, which are Characters or Digraphs. It is in no way an atom. A sequence of characters or diagraphs is a string.
User23> For example, a Common Lisp hosted Jevko parser implementation might choose to represent Jevko text objects as symbols.
This is right.
gnulinux> These are not strings. These are symbols.
This might be confusing. @gnulinux seems to be using the word `symbols` loosely here, not in the Lisp sense. I suppose that's unfortunate in this context.
What may be adding to the confusion is that Jevko spec also talks about Symbols, which are still a different thing.
So the word symbol is much too overloaded here. Let's forget about for the sake of trying to solve this thread once and for all.
The comment that started it was:
Nullabillity> I wish we could stop making up new languages in 2022 that use unenclosed strings...
And then people were trying to explain how there are no strings in Jevko.
Indeed, Jevko the syntax (Jevko is not a language) does not involve anything named string.
It does involve a rule named Text. When parsing a sequence of unicode code points, a specific parser implementation may choose to store a fragment of the sequence that conforms to this rule using a specialized object or a data structure to explicitly represent also the subfragments that conform to rules that Text is defined in terms of. But it may also choose to use a string type.
`jevkoToHtml`[0] relies on an implementation that stores the Text fragments as JS strings.
Thus:
> It's operating on a nested list of strings, whereby it's making a semantic decision whether a string in a certain position of a nested item in the structure ends in the = character.
That "nested list of strings" (not exactly an accurate description) is the syntax tree.
Those strings represent the Text fragments. This is a very convenient implementation, because most of the time we don't want to go lower than that. If we do, Symbols and Digraphs are still easily accessible and identifiable in this representation. They are just not explicitly reified.
So looking whether there is a `=` character in the string is the same as looking whether there is a `=` Character in the Text.
This happens as part of the second pass of parsing I was talking about above.
This pass does not result in any discernible intermediate representation -- it just directly guides the process of constructing an HTML string[1].
This mental model is the best for understanding this code, because it's essentially the one used to create it.
There is no intention anywhere to lie or insult the semantics or philosophy of Lisp in there, trust me.
I'm telling you this as the author of both the specification and the code.
Ultimately, you can understand this any way you like. If thinking about this as strings or numbers or whatever works for you, great!
A common mental model and agreed upon language is useful to achieve productive results in collaboration. If that's not the goal, then it's not necessary.
Anyway, thank you for looking into the specification and taking the time to understand it. Have a good night!
[0] Have in mind that these 14 lines of JS were not really written to be examined in this detail or to present the best way to do something. I just needed a quick sketch to try out an idea. It's not for use in production or lectures on parsing. :D
[1] Which is not nice, but [0]
Generally it's a simpler, more low-level thing. It could be used to build a Rebol-like language. You could add whitespace rules on that level.
A toy example: https://github.com/jevko/jevkalk
This one does not use whitespace as a separator, allowing spaces in identifiers at the expense of more brackets.
If anybody has any technical questions or comments, ask away!
It would be fun to try write that parser/ast library in a language such as haxe (haxe.org) and have it compile to the many target languages that it has - it makes it easier to try it out. When i get some free time this weekend, it'd be a nice fun afternoon hacking project.
Some tips and useful links are here: https://lobste.rs/s/xm6mm4/jevko_minimal_general_purpose_syn...
Thanks for the interest! :)
I'd love to feature it here: https://github.com/jevko/community