If we need to, we could split the run back out again into individual characters without losing information. And that does happen - we do that if something later gets inserted into the middle of the run of characters. Pausing doesn't help. Even the same user could later type something in the middle of a word they'd typed previously.
This limits what we can run-length encode. The code in diamond only run-length encodes items when they are:
- Typed sequentially in time
- And sequentially in insert location (position 10, 11, 12, 13, ...)
- And either all of the characters have been deleted or none of them have been deleted
Notice the other fields in that example in the post. Each subsequent character has:
- An id field = the previous id + 1
- A parent field = the id of the previous item
- The same value for isDeleted
If any of this stuff didn't match, the run would be split up.
Instead of storing those fields individually, we can store the id and parent of the first item, an isDeleted flag and a length field. Thats all the information we actually need to store. With yjs style entries (which is what diamond-types actually implements), the code is here if you can make heads or tails of it. The poorly named "truncate" method splits an item in two. Spans are only extended if can_append() returns true: https://github.com/josephg/diamond-types/blob/c4d24499b70a23...
With real data from Martin's editing trace, 182 000 inserted characters get compacted down to 12 500 list items. Which is still pretty good - thats 14x fewer entries. And it means performance doesn't stutter when large paste events happen.