This pattern matching of user's intentions and using DOM diffing to find out what actually happened feels right to me, especially using persistent data structures. Handling parts of the document as a tree and others as a list seems like decision that could probably make development much easier as you'd know, I imagine. I have been wondering about how to do structured inline element annotations, and lists inside block level elements certainly looks it could work.
Also, this relates to what LightTable folk's talked about in IDE as a value [1].
I'm definitely going to back this. That said, I'd like to know more about what is actually going to (hopefully) be open sourced.
Is this a stand alone app with plugins, a collection of components, or a library to build your own editor?
Is it written in purely JavaScript version 5, 6, or something else?
What about the persistent data structure implementation?
How are user intent patterns modeled, ie. what would it entail to build your own?
Have you thought about intermixing document tree with editors for other types of content, like tables and media?
[1] http://www.chris-granger.com/2013/01/24/the-ide-as-data/
It is written in ES6. The persistent data structures are not very involved, just objects that never get mutated after their initial construction.
I have been reading about the recently announced content-kit: http://madhatted.com/2015/7/31/announcing-content-kit-and-mo... and am wondering if you have considered the 'card' style concept for ProseMirror? I see ProseMirror has an interface for adding images, and says that it will support different document models in future and I'm wondering how extensible that will be.
What sort of APIs are going to be available, will it be possible to create custom 'blocks' or 'cards' of data - e.g. defining a block for adding a table / spreadsheet - similar to what you see in things like Quip or readme.io?
I've been working on a similar project based on the quill editor [1], which is in turn based on the rich-text OT library [2]. Initially, I thought to wire this all up using the sharejs project [3], but never could quite grok that codebase.
Instead, I've been working on a similar system to one you describe, with changes applied in order based on sequential version number, and concurrent updates forced to "rebase" against earlier changes (not yet open-sourced).
A question:
> Because applying changes in a different order might create a different document, rebasing isn't quite as easy as transforming all of our own changes through all of the remotely made changes.
Can you explain more about how you arrived this conclusion? From my understanding, a correct transform function should allow exactly this, according to transformation property 1 [4]. Perhaps your algorithm doesn't exactly satisfy this property; what characteristics does it have instead?
[2] https://github.com/ottypes/rich-text
[4] https://en.wikipedia.org/wiki/Operational_transformation#Con...
A correct OT transform, yes. But I'm not using OT's invariants, so this is not something my transforms do. For example, in my system, if you have "insert X at pos 5" and "insert Y at pos 5", the document will contain "XY" or "YX", depending on which arrived first.
Sorry ShareJS is such a mess right now. I did a bunch of work to make it into a kind of OT-backed database to power derbyjs. Its very much straddling two worlds at the moment and I think its doing a bad job of both.
It'll get cleaned up eventually.
How well does this approach scale to high-concurrency situations? The more active participants there are, the more often you'd be rejecting client changes and ask them to rebase, causing extra round-trips. It sounds like when a critical number of participants is reached, you'll create change requests more quickly than you can handle them. If each participant has a latency of 100ms to your server, and say you have 10 participants making 2 changes per second each, will you experience some sort of congestion?
On a similar note, what happens if you have two active clients, one with say 100ms latency and another with 500ms? If the 100ms client will be typing continously with 4 characters per second, the changes of the 500ms client won't get through until the 100ms one stops typing, right?
I realize these might not be much of a real issues in practise, but did you consider such cases? In particular, is there something in your machinery that prevents the server from doing the rebase, as opposed to sending it back to the client to do it?
As for references, I went all the way and implemented distributed OT in Gobby (https://github.com/gobby/gobby). It's somewhat doable for plain text editing, but going beyond that inflates the complexity very quickly. It gives you some nice properties like authorship tracking even across undo operations. However, the complexity goes with at least n^2 with the number of participants, so there is also an upper limit in the number of active participants that's reached rather sooner than later (even though I never actually found out where it is in practise).
I'm not related to substance.io, just someone curious about editors on the web, and so have kept track of various options. Prosemirror looks very promising.