Hope is an alternative to markup languages
poef.github.io
poef.github.io
There is wealth of research done on stand-off markup since the SGML times, with a renaissance in 2005-2010.
Historically, the main drawback (as highlighted by other comments) is that stand-off needs a way to refer to the content, i.e. a pointing mechanism. Pointers in a pointing mechanism can be either a) hard to write but easy to maintain or b) simple to write but prone to breaking after a change (and thus requiring support from authoring tools). Nobody has found a viable ergonomic balance yet.
But I'm looking at a way in which you may be able to add annotations to a document that isn't under your control. If the document is either guaranteed not to change, you won't have a problem. Documents hosted on IPFS or a versioned URL to a Hyper document fit this bill.
If you want annotations to automatically change when a remote document changes, you need to be able to 'follow' the changes. Between when you originally 'linked' to the document and now, what has changed in the document. Which parts are deleted, which parts are added. If you know that, you can update the offsets in your annotations. So I'm looking into making a versioned document format to allow just that.
IMO the best approach is inline makeup but for the text editor to have the ability to hide tags. An example is PowerGrep's docx editor. You can open an .docx file and it will be editable entirely in plain text, but when you save it, all the formatting is preserved. In some way existing text editors (e.g. Sublime) are already capable of doing this via some kind of fold by regex functionality
You _are_ limited in that you can only detect edits during interactions with the filesystem, which is strictly less information than you'd have in an editor tracking every keystroke, but a good diffing algorithm still mostly gets the job done.
I'm not sure I agree with this premise either, the hierarchical structure seems inherently semantical to me, but the argument of the OP is it isn't the case.
> what's the point of making it a text format at all?
The text part is ASCII, so it make sense to represent it as such. The "presentation" part could be binary.
As for the text part, that is simple. I like to be able to manipulate text as text. There are many many tools to work with plain text. I can use regular expressions to search through text, get ranges of the result and apply an annotation to those ranges. You lose a lot of that in any other format.
That wouldn't be much of a benefit for me because I don't use Nano. I don't think many people do either.
That's because those languages are designed to be edited by Nano. docx files aren't.
(concat (propertize "foo"
'face '(:weight bold))
" "
(propertize "bar"
'link "http://bar"))
;;=> #("foo bar" 0 3 (face (:weight bold)) 4 7 (link "http://bar"))I'm not sure how I feel about this one, but one thing I like about separation approaches in general is that it's easy to imagine how N different annotations can all exist for the same underlying document, so you could use something along these lines to share/exchange/discover/crowd-source annotations.
That said, my general yearning isn't to kick all markup out of my documents--but to kick presentational markup out. Semantic markup can add a lot of affordances to your documents, and if I'm already committing to working with something that requires a toolchain--I suspect I'd rather just generate plaintext or presentational markups from there?
Why is a plaintext unformatted file with just words and no markup any improvement? They're both editable, readable in a raw form, and probably not how you want them to actually be consumed.
So why not just keep it in one file?
100:100:svg:svg
100-100:svg:g /* group */
100-100:svg:rect
100-100:svg:rect
These elements have no text content, so their extent is 0 and offset is same. Yet they also have hierarchical relationships and you need to store extra data to express those. E.g. in this example you can guess that 'g' and 'rect's are descendants of 'svg' and the first 'rect' is a child of the 'g' but only because you have some preconceived ideas about the meaning of these elements. There's not enough data to tell if the second 'rect' is a part of 'g' or not. You need to store at least the level in the hierarchy. If you stored the data in a database (i.e. without implied order) you'd also need the relative position of the element, but here you have the order built-in.It's still a hobby research project. There's a lot I've learned building and using it, about the things that HTML is good at, and what it isn't good at.
What I'm trying these days is to see if I can get a WYSIWYG editor with a user experience that comes close to the user experience of writing markdown. This isn't easy and its interesting to explore what makes that so.
Could this project be the beginning of an answer : http://lambdaway.free.fr/lambdaspeech/
Alain Marty
I am not sure if I can see anything on the right/left. this is what I see on chromium/firefox https://i.imgur.com/quySrXT.png. No rendered veiw!
* Word break nearly always gets treated as content, denoted by the space character.
* Line break can be either: '\n' or <br>.
* Paragraph break is usually treated as mark-up: <p>.
Is there a good reason for that, or is it just historical accident?
In general the whole business of file type plus creator type felt much better than the single file type of now. E.g. in desktop publishing Illustrator EPS were very different from Photoshop EPS and it was very convenient to be able to tell them apart.
So we have a source and some associated metadata that is crafted manually. I don't see this to differ much from other cases when we need to manually sync two distinct files (e.g. C source and header files), so I don't quite understand the issue, I guess.
Nobody would put punctuation in a metadata file. But why not? It's markup in exactly the same way that a lot of HTML is. In fact formats such as Markdown can be considered extensions of punctuation just as much as they are simplifications of markup.
I remember reading a discussion of word processor document formats back in the 90s, concerning the representation of layout and embedded images and tables. One participant was saying that conceptually the correct way to treat such inclusions is as a character. It's part of the text, not some external thing imposed on the document from outside. For a page layout system such as DTP you might make a different choice because the focus is different, but for word processing I thought this was a very insightful point. A lot of markup is part of the data as well and best treated as such.
I myself don't think it's worth to actually keep content and markup in separate files, so for me the only possible use case was conflicting markup, i.e. markup that exceeds the limitations of our tools, something like two overlapping tags in XML. Yet even this case doesn't require actual physical separation; it can be solved with a different notation or maybe a different tool (like SGML, which, I believe, supported overlapping tags).
But its not meant as production ready code, instead I'm exploring the concept and what it means for editing, rendering, etc.
The editor (https://simplyedit.io/) converts a contenteditable area's content to hope, we apply some markup (annotations) and convert the result back to HTML. I wrote the Hope-HTML conversion routine once, and all the range based manipulations are now one-liners instead of treewalking horror shows.