Do you first query the root to see it's content blocks, then make additional queries to load the root's children block, then make additional queries to get those blocks children blocks (ie. recursively) until there are no more children?
Does that result in too many database queries? Or do you have other ways to optimize it?
[1] https://en.wikipedia.org/wiki/Hierarchical_and_recursive_que...
Also, are you adding presentational tables any time soon? :)
Most of our queries are "pointer chasing" - we follow a reference from one record in memory to fetch another record from the data store. To optimize recursive pointer-chasing queries, we cache the set of visited pointers in Memcached.
We use Elasticsearch for search features like QuickFind.
> Also, are you adding presentational tables any time soon? :)
Sorry, can't talk about future plans like that :)
From there we started looking for a narrative. We extracted out the sections you see in the final post, and removed a lot of the superfluous technical detail so we didn't end up with technology buzzword soup; for example we cut discussion of Postgres, Memcached, etc etc, how we host the web servers; the kind of details that don't actually matter to the narrative.
The illustrations were in the post from the beginning as Mermaid diagrams (https://mermaid-js.github.io/mermaid-live-editor/). As we got close to publication we polished them up in Figma.
This is really the first engineering blog post we've put out, there was a fair amount of figuring-out-how-to-do-it going on. Now that we've had the experience, we're starting to write up our playbook internally.
This seem to be a really good use case for a NoSQL database. Am I wrong ?
http://patshaughnessy.net/2017/12/13/saving-a-tree-in-postgr...
You can break out the "block" model into several tables and represent it in a relational database that way.
NoSQL = NO JOIN?
Hope that helps.
We do lean very heavily on the TypeScript type system and try to make invalid states unrepresentable.
Which is pretty much the ideal scenario for a document store. The article describes Notion as being very strictly hierarchal
The underlying persisted data doesn't necessarily have to be a bag of KV pairs.
A block is related to its parent and descendant blocks.
These relations are suitably represented in a relational database, not a document store.
EDIT: In graph theory, a tree is an undirected, connected and acyclic graph.
When comparing a document store versus a RDBMS, in terms of suitability and appropriateness, the distinction is primarily along the lines of a tree, versus an arbitrary graph (by which I mean that an RDBMS is more powerful, and more general, but not inherently as optimal in either performance, “scalability”, or UX in the places where a document store makes sense.
More specifically, the way the article describes it, you’re not interested in “give me every block of type X” — you’re only interested in “given block Y, what type is it?”.
That is, the question is one-way, and fits cleanly in a hierarchal format of a document store.
The only question posed that operates in the reverse direction is permissions, though even that’s a little odd, since it seems to me it should only go “downwards” as well — a block’s permission scope is the sum of all of its parents, and you can store it there upon iteration.
> The underlying persisted data doesn't necessarily have to be a bag of KV pairs.
It doesn’t have to be... but it can be, and appears to be.
> A block is related to its parent and descendant blocks.
Right; the singular parent, and the multiple children. A tree.
> In graph theory, a tree is an undirected, connected and acyclic graph.
When discussing trees and graphs, I think it’s obvious a distinction is being made between a graph forming a tree, and graph forming a not-tree (more complex than a tree). When I say that a square is easier to encode than a rectangle, I do not mean that a square is not a rectangle, but that a rectangle is not a square — that a square’s more specific properties give us opportunity to simplify/optimize (I only need to store one length to represent it).
A database can encode a tree just fine, but that doesn’t mean it’s the best tool to do so.
There are other properties to a document store I don’t care for, and I don’t like them in general (like the implicit schema, and total lack of data consistency validation by the data store, and the fact that you often don’t truly have a tree), but representing a tree is what’s been described, and it’s exactly what they’re specialized for.
If you want to argue against it, you need to specify why you think this isn’t a tree, because I feel it’s quite obvious it is.
https://www.notion.so/Tree-breaking-3a90e2bcd2154f4fab06a3c7...
Breaks the tree model, as 'Complete Task' would need to have both 'Subtasks' and the page itself as its direct parent.
That said… it's mostly a tree, and there may be merit to optimising for that access pattern.