Rich text editors and rendering engines
writer.zohopublic.com
writer.zohopublic.com
Amused by the nod to the Ladybird WASM idea at the end, I'm taking credit for that: https://news.ycombinator.com/item?id=35521878. Good ideas spread. (Joe, apologies for not getting back to your email, it's been a rather busy few months)
Are there any documented best-practices for writing your own? (aside from "don't" hah)
I get that common wisdom says it is kind of insane, but i'd like to learn why, concretely. It's possible a fraction of what i want is not the "hard part", so i feel compelled to at least explore it.
I think that was the most humbling programming experience of my life. I gave up after a few weeks of banging my head against the wall.
Like I use a custom DefaultKeyBinding.dict for my laptop using which I've defined some shortcuts for text editing. It works with Textarea but not for GDocs. Which makes editing so cumbersome for me.
I've always wondered why this was the issue until I read this article. Now I see why it's broken.
[1]: https://lexical.dev/
And even with those libraries, try implementing features like multi-columns and cross-block selections.
And the most important problem: try implementing a pure HTML editor using ProseMirror / Lexical - i.e the editor should accept HTML exactly as it is, from arbitrary sources, like how contenteditable accepts. (You can't)
Those libraries depend on tightly controlling what goes into the editor, which is an amazing trade-off if you are building a tightly controlled editor. But good luck, building an email editor that accepts any wild HTML.
Another reason our activist FTC head should focus her sights on Google after she’s done pulling Amazon apart.
[1] https://www.mozilla.org/en-US/firefox/115.0/releasenotes/
If one were to write a simple text editor from scratch - is there some place with a "spec" or list of standard modern text editor UI paradigms one would need to implement to make the editor feel natural/normal ? (so it feels like Kate/GEdit/contenteditable)
I kinda get there are a lot of little subtle quirks and corner cases that need to be covered so it doesnt feel weird or janky
For instance EMACS CUA-mode is an example of a incredibly incomplete implemention
Also Apple’s old documentation about how their text layout system works. That link is annoying to find so I don’t have it on hand.
The short is answer is no, there are very few universal truths about how text editing works. You'll experience differences between combinations of operating systems, input modes (Japanese IME for example), software keyboards on Android, languages (RTL languages) and that's just for text itself.
Then when you are thinking about more complex features simple things like "what happens when I double or triple click on this" are a complete crapshoot.
As a casual user of HTML and CSS, I'm often struck by how stupid the layout controls are. How very simple things ("centering text") are often confusing and difficult. And it's striking how these layout and content limitations make it necessary to add layer after layer of complexity (mostly in the form of JavaScript and CSS) in order for anyone to make a "modern" web page successfully. Whereas older applications that are not "web-based" allow any casual user to create rich document presentation just by pointing and clicking.
The Web today is like an Indy 500 race car with foot petals. I think we need to consider the future and how we want computers to work, and start moving towards that, rather than perpetually carrying forward the status quo.
As soon as Flash came to be, we should have thought of a standard better suited for applications.
There are a lot of problems with doing a hidden ContentEditable for input, which is what Google Docs does. Last time I looked, the input / focused DOM element was actually hidden inside the blinking caret. But that strategy is clumsy with assistive technology because the accessible bits don’t “line up” with the actual UI drawn to the screen. It also breaks some system conventions like smooth cursor movement when holding spacebar on iOS’s keyboard. You can try it in Google Docs and see what I mean.
Those kinds of issues are actually why Lexical (FB’s new text editor toolkit) doesn’t use hidden input according to Lexical’s author trueadm on Twitter (can’t find my citation). Those same issues also make us hesitate to move that direction in Notion’s editor.
If you want to use a custom “layout engine” implemented on top of the DOM today, you can use Skia CanvasKit (https://skia.org/docs/user/modules/quickstart/) or Flutter Web which is based on Skia. Although CanvasKit is kinda slow and text looks bad on iOS.
Ah!, so that's the reason why copy-paste is badly broken in google docs (and jupyter notebooks as well)? Now I understand the reason. It's because they didn't reimplement it again.
I can copy-paste between all my windows and firefox tabs, except google docs and jupyter notebooks. This is because the selection that they show it's not an actual text selection that the system understands, but just a hand-drawn color rectangle around the text. It's a fake selection, of sorts!
Basically every rich editor on the web that’s any good overrides copy/paste with custom handling.
You need to debounce/throttle the change event handler to not fire too many re renders when typing.
I used this approach for https://bigwav.app for highlighting text.
You can click the demo .bigwav file to take a look.
RTF isn’t standardized and, reading https://en.wikipedia.org/wiki/Rich_Text_Format I get the impression there are fairly wide differences between versions.
From the same page, it also isn’t simple:
- files can use at least 18 different text encodings.
- you can include Unicode code points, but if you do, they “must be followed by the nearest representation of this character in the specified code page”
- if you want to create a table in a RTF file, it seems you have to include an OLE object
- for compatibility with existing files, it seems an editor will have to support Windows meta file format.
So, I would guess any simple RTF editor would be incompatible with a significant fraction of existing RTF files.
Not sure what all you are trying to do that template literals and a simple input formatter can't do, but if you elaborate on an example or something maybe I can see where the library's failures are?
I work on non-anemic Notion and built the current @-mention menu and overhauled the text editor system to enable cross-block text selection.
The criteria I use to evaluate a rich text editor component are:
- How does collaboration work? Does the framework give me enough for OT or CRDT bindings?
- Does it use a tree-like data structure or do you need to add this later?
- Can you easily build interactions like @-mention menu or tab completion like VS Code?
- Does the editor tolerate concurrent state updates without breaking Input Method Editor composition, ie a remote user comments on text while the current user is inputting a Chinese character?
- Do demos of the above features work on Android? In Chinese?
From reading the Pell source I didn’t see much that would help with my criteria.
I'm just going to chalk this one up to a misunderstanding and hope you have a good day!