Mark – A simple and unified notation for both object and markup data
github.com
github.com
I don't understand this part of the readme:
> The advantage of Mark over S-expressions is that it is more modern, and can directly run in browser and node.js environments.
Does this mean I'm so out of date with JS that this syntax is actually legit JS? Or does Mark run its own parsers, at which point it's just like sexps, except it uses the "more modern" curly brace instead of "less modern" parenthesis?
http://blog.cognitect.com/blog/2014/7/22/transit http://cognitect.github.io/transit-tour/
cool project though, good to see s-expressions become more popular
No, it's close to S-expressions, but it's not a programming language like Lisp.
> just with curly braces instead of parenthesis
Well, and more syntax than S-expressions: it's got both objects and arrays as fundamental structures instead of just lists, and it has commas as noise characters.
Well, Common Lisp has both of those as fundamental atoms:
#S(foo :bar 3)
#(1 2 3)
The former creates an instance of a FOO struct with its BAR slot set to 3; the latter is a 3-item array.I agree with this: certain things need to be in the spec.
> even if you'd end up implementing them as reader macros in CL (which I don't recommend - reader macros turn your READs into EVALs, which you obviously shouldn't do on untrusted input).
This I don't agree with, because technically #\( & #\" are reader macros … they're just very well-defined reader macros. Presumably a spec which defined hash tables, regular expressions or whatever would define them as well as the Lisp spec defines lists and strings (and if not, well — it's a bad spec!).
That was my main point. In the second part of the comment I didn't mean to discourage use of reader macros - it was more of an aside that the general facility of CL-style reader macros literally makes READ "shell out" into EVAL, so you need to (diligently) disable it for untrusted input (or reimplement a limited READ by hand). So we can't say "oh, but S-exps in Common Lisp can have anything through reader macros". Presumably if hash table literals were specified as a part of basic syntax, we could depend on it being standard and part of the safe subset of READ's duties; as it is however, we can't depend on it for arbitrary inputs.
> it's got both objects and arrays as fundamental structures instead of just lists
I'm bit sad that various lisps never standardized on a format for this; had they, then maybe we would have S-expressions as a popular data interchange format.
Ah, but you mentioned objects & arrays, not hash tables grin. Agreed that hash tables would have been nice, although that does then get into issues such as canonical representations (which matter e.g. for hashing).
> Also, CL didn't clean up the Lisp space completely; right now, there are Schemes, there's Clojure, LFE, Hy, and bunch of other niche Lisps, each with their own idiosyncrasies around syntax.
It would be nice if folks who want to use Lisp would use Common Lisp rather than reïnventing various forms of more-or-less round wheel. It's a remarkably well-engineered language (not perfect, of course: argument order, upcasing, pathnames & environments all leap to mind as problematic areas), and so far as I can tell quite a bit better than any of the alternatives.
In particular, it'd be really nice to see people using Schemes for serious engineering work to use Lisp instead. It's just not well-suited to writing large systems, except by grafting on an ad-hoc, informally-specified, potentially bug-ridden subset of Lisp.
Which can easily be added to sexps (see EDN).
;; structure:
#s(type :slot1 value1 :slot2 value2 ...)
;; vector:
#(1 2 3 4)https://github.com/edn-format/edn https://learnxinyminutes.com/docs/edn/
It hits a sweet spot for me between yaml and json. Yaml is easy to type/read, but I feel it's a bit too complex on the parsing side. And json is a pain to type, so I'm reluctant to use it for human entered configuration files.
Junior talent doesn't become senior talent by just doing the right things, but by doing the wrong things and learning from them.
Hmm... I think that’s how PHP was made.
The circuit less traveled. Investigating some alternate histories of computing: https://www.youtube.com/watch?v=jlERSVSDl7Y
I'm intrigued in particular by the talk's conclusion about (I guess, again) disappearing distinction between volatile and non-volatile storage. To date, I've been a vocal advocate of hierarchical filesystems (not UNIX, but just as a unit of user-facing abstraction). The talk sent me on the way of reflecting whether I'm not just supporting another historical "wrong path". Lots of more thinking in front of me here. So thanks.
var obj = Mark.parse(`{div {span 'Hello World!' }}`);
Not sure how else executing (valid) JSON in a browser would be a recipe for disaster? `eval` was the standard way to parse JSON from trusted sources for a long time.
It has its own parsing and stringify library it looks like: https://github.com/henry-luo/mark#markjs
Secondly, to clarify what I mean by 'being more modern'. Of course, it does not mean changing from () to <> or {}, will make it more modern or something better.
Being 'more modern' means Mark takes a JS-first or web-first approach in its design. Whether we like it or not, JS has dominated the web. JSON is successful, partly because it takes a JS-first approach. Mark inherits this approach.
Being JS-first, means there'll be least adoption barrier in web.
Being JS-first, of course does not mean JS-only. Mark is designed to be generic and used by other programming languages like JSON.
Looks like s-expressions to me and it isn't legit JS.
But it's not literally lisp in the sense that the meta-syntactic stuff isn't there.
In Lisps:
- values have types (bool, symbol, number, list, array, structs, functions, ...)
- variables have by default one type: union of all the above
- in Common Lisp, you can restrict the types of values allowed in a variable
> but this uses type declarations
I must have missed this. If it's not just what struct-like entities are allowed in the markup, where did you see that?
Another is the lack of support for binary data. There's no sign of support for binary data here.
Finally, there's this claim:
> The advantage of Mark over S-expressions is that it is more modern, and can directly run in browser and node.js environments.
Is it more modern? I don't think I care.
Can it directly run in browser and node.js environments? What does that mean? It seems to need a parser. But then, S-expression parsers certainly directly run in browser and node.js environments.
---
IMO, SPKI SEXPs are much more sensible than this design and many, many other designs:
Yes, yes, ten thousand times yes! I really don't understand why, over two decades hence, the world has stuck with XPKI & ASN.1, and has invented XML & JSON, when SPKI solved the PKI problem for good & canonical S-expressions solved the flexible- and human-readable–data-exchange problems for good.
(my_dict (key value) (key value) (key value))
Un-ordered qualities for data can be useful (e.g. they allow you to reorder data to stream "important" stuff first), but I don't see it anywhere in here.In general, I'd resist specifying data as arbitrary key-value pairs, but if I decided that I indeed needed them, I'd do exactly as you suggest — and I'd mandate that the be sorted lexicographically by their keys.
Benefit of HTML: You can actually write it by hand and easily see where each element begins and ends, even when the document is longer than a screenfull. Mark has the "}}}}} problem with larger documents, so it is not as suitable for human-written markup.
It is not clear to me how mixed content like <cite>Hello <i>world</i></cite> is expressed in Mark. I expect it will be pretty convoluted.
Benefits of JSON: Maps directly to simple data structures: List, dictionaries and simple values. Similar data structures are supported in almost any language. Mark has "type names" and anonymous text content which complicates serialization and serialization a lot, and is sure to give interoperability (and perhaps security) problems.
So - worst of both worlds? Instead of tying to be an overall worse alternative to all the formats, they should rather focus on a specific niche where Mark can be a better alternative.
Take configuration files, for example. They don't have large amount of textual content like HTML, and they don't need to be transferred between disparate systems.
{size width:100 height:100}
vs <size width="100" height="100"></size>
vs {"size": {"width":100,"height":100 }}
In this case, the Mark syntax is simpler and cleaner. Mixed content is not needed, which would make the format simpler. Yeah it is basically the same as S-expressions, but that is not a bad thing.And HTML has a problem of </span></li></ul></div></div></div></div></body></html>, all spread over nine different lines, one tag per line.
Take a look at https://github.com/keithj/alexandria/blob/master/definitions... which is Lisp code styled in a standard manner. I don't see any problem there.
Paren-matching is a commodity today in all sane programming editors. It is no longer anything you could call "specialized".
When it comes to worse-case scenario, no one wins. :-(
size: {width:100, height:100 /*yay*/}
It is not clear to me how mixed content like <cite>Hello <i>world</i></cite> is expressed in Mark.
{cite "Hello" {i "world"}}Json5 is still not as "editable" as it looks though. You need to separate values with comma (except the last value), so there is more syntactic noise. So you get:
{
size: {width:100, height:100 /*yay*/},
}
This is not an issue when the text is machine-generated (as Json typically is), but is an issue when it is edited by hand as config files often is.so in practice you don't notice any brokenness
it's not like a web browser that has to work on a diversity of third party sources
I'm not even sure how to read that matrix anyway, and it does say: > The YAML Test Suite currently targets YAML Version 1.2. ... some frameworks implement 1.1 or 1.0 only
Another strength of XML is support for mixed content which seems rather awkward in {mark}. The following
<p>Some <b>bold</b> text</p>
apparently needs to written as {p 'Some' {b 'bold'} 'text'}
It would be more honest to mark support for mixed content as "verbose" in the feature table.Besides, the name {mark} seems like a bad idea. How could you find relevant results when searching for {mark} using a search engine?
XML Namespace seems to have a lot issues, thus Mark does not want to enforce something exactly following it.
Namespace in Mark, currently, is left upto the application user to define it.
We might be able to come up with a better way to define namespace.
As for the name, you can just use Mark. I use '{mark}' as an alternative name, to make it more graphical, more impressive.
XML Namespaces is syntactic vinegar.
Less is more.
Ok sure, but does it have schematron,rng, or some sort of validation? How about transformations? Xpath?
There's already a transformation library - Mark Template (https://github.com/henry-luo/mark-template) in beta release.
Mark at the moment supports CSS selector. I'm also thinking about a new Mark-specific selector.
Mark is very new. A lot to be done!
FWIW, I would suggest avoiding the (IMO) mistake of using your markup language for the schema.
E.g. like json-schemas where we need a "properties" map, "type": "string" (how many times do I have to type "type"), all sorts of syntactical overhead.
Personally, I think IDLs are much cleaner, as you can design a purpose-specific grammar. More work up front, and you don't get a parser for free, but again personally I think it's more pleasant in the long-run for developers to read and write.
Granted, not sure how that jives with your lisp/etc. way of thinking, but my two cents.
Good luck!
While it would be cool to have something that was like JSON but could deal with complex documents, I also don't see how this is a huge improvement over XML.
There seems to be a ton of s-expression parsers in npm already, that can run in browser and in node.js: https://www.npmjs.com/search?q=s-expression
Besides being able to run in js environments, what else does {mark} bring over s-expressions?
- A wise man on the Internet once said
Another comment: Coming from a Semantic Web background, and using N3 as the exchange format and N3.parse() as my client-side lib, I would advise to have a UID parameter to uniquely identify objects, and a refId syntax, so any parameter can reference some other objects of the data structure. That helps when you want to transmit a graph [1].
My humble 2 cents.
[1]: I would add that it is also useful when you retrieve some refIds that are not defined in the current data structure. You can then ask the server to dereference these refIds, and send another (portion of the) graph, that you can connect with the existing data structure.
If you don't care about parsing CPU efficiency then gzipped JSON beats protobuffers, CBOR, etc when you care about bytes sent over the wire.
If you care about CPU efficiency then protobuffers, CBOR, etc are worse than flatbuffers or capnproto.
There is not a lot of space for a new standard between these two existing categories.
Gzipped JSON does not beat gzipped Protobufs in message size. Comparing gzipped JSON to uncompressed Protobuf doesn't make sense.
> CBOR is defined in an Internet Standards Document, RFC 7049. The format has been designed to be stable for decades.
I see no reason to go with CSON over CBOR. In fact just the opposite.
Is it more modern because it is newer? There is mention of how adoption is limited, but wouldn't the adoption of a completely new syntax be even more limited :-)
There is even a canonical representation using length prefixes: https://en.wikipedia.org/wiki/Canonical_S-expressions
There might be a use case where your data is better represented in LDIF because it's hierarchical, but there's no built in LDIF support, so now you're importing a ton-o-javascript just to parse some new format.
At this point, we should realize json isn't meant to be human readable anyway. If you need to hunt through it, you put it into some type of json viewer so you can see the tree and query it. It's an interchange format, that's more compact than XML.
If you're shipping data between non-browser things like backend services, there are already binary formats like protobuff that have typing and can be optimized for small payloads.
It just that in languages like JS and Lua, where an object can be an map and a list at the same time, they'll have the convenience of mapping a Mark object into just one object, instead of many.
General JS arrays (not those TypedArrays) are actually maps indeed.
I've updated the README to be: "The advantage of Mark over S-expressions is that it takes a more modern, JS-first approach in its design, and can be more conveniently used in web and node.js environments."
Hope it's clearer now.
- Andrew Tanenbaum Computer Networks, 2nd ed., p. 254.
I think all developers go through some experience where they want to just "unify" everything because that will supposedly make it easier for them and other developers.
Overtime as you become more experienced or I guess jaded you realize that reality of a "GUT" technology platform or programming languages is a pipe dream and the effort to get people to use said new format/language/tech is more effort than what you get in return.
Anyway to be short about it I think most should just pick the best tool for the job and stop rebuilding things that don't need to. And if you do please make sure you have a plan to how you are going to replace all the old working stuff.
I think you just contradicted yourself. Sometimes the best tool for the job is something new, something improved over what already exists.
I don't think the author intends to "replace all the old working stuff". But if this tool is better for new projects, then why not? I don't get all the negativity... do people here really love XML/JSON/YAML that much? There's a whole lot to complain about in all of those!!
And yeah I don’t have a problem with XML or JSON. Those two combined with some flatbuffer other men binary protocols cover most of my use cases... like really what’s with all the XML negativity.
XSDs don't count then? https://en.wikipedia.org/wiki/XML_schema
Full formal schema definition, as in XML, is often a burden to ad-hoc scripting, which is common in JS. JSON/Mark provides sufficient type info for these adhoc usages.
XML uses the same syntax for strings, integers, and booleans, but it has mature schema/typing tools that make it easy to apply more precise typing, which you'd want to do anyway to identify timestamps, enums, and different object types.
Cons: Not seeing any advantage over JSON. If you want a type for objects just add a type field and have your code read it. Then you can use any of the existing parsers.
You could remove every '{' with 0 loss of meaning.
Some disadvantages of Mark, comparing to JSON would be:
* Mark is insecure, JSON is secure.
* Mark is slower than JSON
Passing types directly to object.constructor is of course entirely insecure. https://github.com/rurban/Cpanel-JSON-XS/blob/master/XS.pm#L... (i.e. CVE-2015-1592)
Current Mark implementation does not call arbitrary constructor during parsing. The constructors are created from scratch. But application users might want Mark to call their customer class constructor. I'm thinking passing in a callback function to Mark.parse().
However, I don't have time to do some benchmarking at the moment.
I mean JSON as a data format for api stuff is just enough as it is and you'd need some serious reason why to change from JSON and these reasons just doesn't cut it.
… with the right translator to JavaScript, which also happens to be true of S-expressions.
His table is incorrect, incidentally: S-expressions support mixed content (if I understand what he means) and are also fully generic.
He doesn't have a good example of the benefits of his proposal over S-expressions: 'more modern' just means 'undiscovered bugs.'
I respect his enthusiasm and hard work, but I believe what the world needs is hard work on existing things rather than hard work reïnventing the wheel.
Where can I read more about this feature of JavaScript?
Calling it novel to JS is a stretch. Lua does this too and I’m sure there are other languages.
And why bring YAML in the mix? Yaml isn't used for transfer I hope? Should be compared to TOML as well in that case that seems a lot better than YAML, especially for configs: https://github.com/toml-lang/toml
Or msgpack? Which also seems useful. Why not protobuf? Or just s-exps which is basically what this is.
It is very hard to make major extension to JSON and still be compatible with JS syntax. Minor changes are possible, like in JSON5.
Once it breaks JS compatibility, I don't think people will think it is JSON next any more.
I like this project.
wouldn't that make introspecting objects very annoying?
This is one of the difference between Mark object and an array. Array contents are enumerable by default.
https://github.com/tlocke/zish
It's a data serialization format with timestamp, bytes and decimal data types.
ASN.1 dates from 1984. So much for Mark being “modern”!
Not perfect, but we don't have to introduce a lot of built-in types. The syntax can be kept simple.
I think we're fine with separate languages for data and markup.
The latest trend in CMS systems, piloted by the latest content editors, like Quill, Draft.js, ProseMirror, Slate.js is to use JSON to present the content, instead of using HTML or Markdown. Using object notation, gives rise to cleaner API and data model.
So the wall between data and markup, JSON | XML may collapse one day.
In the README example, I deliberately made it resemble HTML comment, so as to make it easier for people to correlate.
begin_sl_comment ::= '//'
begin_ml_comment ::= '/*'
end_ml_comment ::= '*/'
maybe that ebnf is out of date.Then there's Mark pragma, like {!pragma}, which are preserved in the data model. HTML comment is supported as Mark pragma, not Mark comment.
Mark reserving all number-only keys is statistically likely to become a problem as a project grows larger. I'd suggest finding a different way to get out-of-band data to be fully out-of-band, rather than trying to carve out a chunk of keys.
Somewhat similarly, defining a "pragma" as "something surrounded with braces that isn't a legal object" means that if you ever want to change the definition of an object in the future, you can't, because you will turn things that used to be pragmas into objects, or less likely (because you'll try to avoid this going forward) but still possible, vice versa. You need to concretely specify what a pragma is unambiguously, in a way that you can evolve either without affecting the other. It also means errors in generation become legal pragmas instead of errors, which will cause surprises, and on the flip side, errors in parsing objects can turn them into legal pragmas rather than parse errors.
I would reserve saying "Mark is a superset of JSON" for the case when you really can feed any JSON to a Mark parser and get a (roughly) equivalent structure. Alternatively, go through the documentation with a text find option and make sure every time you say "superset" it is qualified as a "feature superset". Especially in light of "(Mark does not support hexadecimal integer. This is a feature that Mark omits from JSON5.)" The word superset should either be qualified every time or mean a strict superset; "Mark is a nearly-feature-superset of JSON" would be more accurate.
In general, a review of http://seriot.ch/parsing_json.php may be appropriate; mark addresses only one serious issue, and the other fixes are ultimately fairly superficial (the trailing comma issue, for instance, is almost never a problem for me because ninety-nine-point-I-don't-know-how-many-nines percent of the time, JSON is a thing my tools generate; the cases where that is a serious issue have generally already moved on to another format like YAML, same for comments). Also, per my comment about parse errors turning objects into pragmas, if you expect this to become a big cross-language standard it is worth reviewing a snapshot of the variability in JSON parsers, which is a simpler format. A more complicated format should expect to see even more subtle divergences in its multiple implementations and things like "misreading an object as a pragma" to become even more likely at scale.
And again let me emphasize, since you seem to be saying it again in some other replies, that "{mark} is a superset of JSON", if you mean that syntactically (as opposed to features wise), MUST mean that every valid JSON document will produce a valid {mark} parse. Nothing less than that qualifies it as a superset. Given that you reserve numeric keys I don't think that is the case; whether the grammar is a superset is harder to determine so I haven't tried. That would be something best served by taking a very complete JSON parser test suite from someone and validating that all their corner cases that are supposed to parse in JSON, parse in {mark}. Based on my own experience in the world of parsing, the odds of you passing that first try are very low; if you manage, major kudos to you as that would be a very difficult test. (Though I would imagine that since the grammar largely came from JSON a lot of the surprises would be the ways in which your parser turns out to deviate from the grammar rather than grammar errors.)
The headline reminds me of: https://xkcd.com/927/
Best of luck with {mark}.