Building and scaling Notion's data lake
notion.so
notion.so
Notion does not sell its users' data.
Instead, I want to expand on one of the first use-cases for the Notion data lake, which was by my team. This is an elaboration of the description in TFA under the heading "Use case support".
As is described there, Notion's block permissions are highly normalized at the source of truth. This is usually quite efficient and generally brings along all the benefits of normalization in application databases. However, we need to _denormalize_ all the permissions that relate to a specific document when we index it into our search index.
When we transactionally reindex a document "online", this is no problem. However, when we need to reindex an entire search cluster from scratch, loading every ancestor of each page in order to collect all of its permissions is far too expensive.
Thus, one of the primary needs that my team had from the new data lake is "tree traversal and permission data construction for each block". We rewrote our "offline" reindexer to read from the data lake instead of reading from RDS instances serving database snapshots. This allowed us to dramatically reduce the impact of iterating through every page when spinning up a new cluster (not to mention save a boatload in spinning up those ad-hoc RDS instances).
I hope this miniature deep dive gives a little bit more color on the uses of this data store—as it is emphatically _not_ to sell our users' data!
Full disclosure: I'm a founder of authzed (W21), the company building SpiceDB, an open source project inspired by Google's internal scalable authorization system. We offer a product that streams changes to fully denormalized permissions for search engines to consume, but I'm not trying to pitch; you just don't often hear about other solutions built in this space!
The blocks (pages are a block) in Notion are a big tree, with your workspace at the root. Some attributes of blocks affect the search index of their recursive children, like permissions: granting access to a page grants access to its recursive child blocks.
When you change permissions, we kick off an online recursive reindex job for that page and its recursive subpages. While the job is running, the index has stale entries with outdated permissions.
When you search, we query the index for pages matching your query that you have to. Because the index permissions can be stale, we also reload the result set from Postgres and apply our normal online server-side permission checks to filter out pages you lost access to but that have stale permissions in the index.
> Iceberg and Delta Lake, on the other hand, weren’t optimized for our update-heavy workload when we considered them in 2022
Curious about your thoughts here. Have you followed Icebergs progress? Do you think it'd be a tougher decision in 2024 between Hudi and Iceberg?
Have you explored a pattern like https://runtrellis.com or https://unstructured.io/ for unnesting?
Otherwise, this is the perfect app for sharding/horizontal scalability. Your notes don’t need to be queried or joined with anyone else’s notes.
For example, they mention search. But i imagine it is just searching only within your own docs. Which i presume should be fast and efficient if everything is sharded by user in Postgres.
The tech stuff is all fine and good, but if it adds no value, its just playing with technology for technology sakes
> Moving several large, crucial Postgres datasets (some of them tens of TB large) to data lake gave us a net savings of over a million dollars for 2022 and proportionally higher savings in 2023 and 2024.
What does a backing data lake afford a Notion user that can’t be done in a similar product, like Obsidian?
It can also be used to train AI models, of course.
When your data is in Postgres, running an arbitrary query might take hours or days (or longer). Postgres does very poorly for queries that read huge amounts of data when there's no preexisting index (and you're not going to be building one-off indexes for ad hoc queries—that defeats the point). A data warehouse is slower for basic queries but substantially faster for queries that run against terabytes or petabytes of data.
I can imagine some use cases at Notion:
- You want to know the most popular syntax highlighting languages
- You're searching for data corruption, where blocks form a cycle
- You're looking for users who are committing fraud or abuse (like using bots in violation of your tos)
These aren't something I would like to hear if I'm still using Notion. It's very bold to publish something like this on their own website.
This is why laws like CCPA "do not sell my personal information" exist, which I certainly hope Notion is abiding by, otherwise they'll have lawyers knocking on their door soon.
I wrote another comment about why you'd need this in the first place:
https://news.ycombinator.com/item?id=40961622
Frankly the argument "they shouldn't need to query the data in their system" is kind of silly. If you don't want your data processed for the features and services the company offers, don't use them.
Neutral party here: that's not what they said.
A) Quotes shouldn't be there.
B) Heuristic I've started applying to my comments: if I'm tempted to "quote" something that isn't a quote, it means I don't fully understand what they mean and should ask a question. This dovetails nicely with the spirit of HN's "come with curiosity"
It is disquieting because:
A) This are very much ill-defined terms (what, exactly, is data lake, vs. data warehouse, vs. database?), and as far as I've had to understand this stuff, and a quick spot check of Google shows, it's about making it so you're accumulating more data in one place.
B) This is antithetical to a consumer's desired approach to data, which will described parodically as: stored individually, on one computer, behind 3 locked doors and 20 layers of encryption.
I’ve seen 100TB+ workloads at smaller companies. Not unusual.
1. Do you identify which types of content your users use the most?
2. Do you find users who are abusing your system?
3. Do you load and process data (even on a customer by customer basis) to fine tune models for the QA service that you offer as an optional upgrade? Especially when there could be gigabytes of data for a single customer
4. Identify corrupt data caused by a bug in your code that saves data to the db? You're not doing a full table scan over hundreds of billions of records across almost 500 logical shares in your production fleet
These are just the examples I came up with off the dome. The job of the business is to operate on the data. If you can't even query it, you can't operate on it. Running a business is far more than just being a dumb CRUD API.
Also we see how at self hosting within a startup can make perfect sense. :)
Devops that abstract away things in some cases to the cloud might just add to architectural and technical debt later, without the history of learning from working through the challenges
Still, it might have been a great opportunity to figure out offline first use of notion.
I have been forced to use anytype instead of notion for the offline first reason. Time to checkout to learn how they handle storage from the source code.
Are they using this new data lake to train new AI models on?
Or has Notion signed a deal with another LLM provider to provide customer data as a source for training data?
* The team are based in the US, specifically California, and Notion Labs, Inc is a Delaware corporation.
* Their investment comes from Venture Capital and individual wealth. The investors are listed on Notion’s about page and are open about how they themselves became rich through VC funded tech companies.
There is a very open sense of panic in tech right now to climb to the top of the AI pile and not get crushed underneath. I would be amazed if there were any companies not enthralled by — and either already embracing or planning to embrace — the data-mining AI gold rush.
Notion is a great product but one would be naive to use it while also harboring concerns about data privacy.
We do not and will never sell customer data to anyone. We do not train AI models on customer data. As we state in our privacy policy for AI features (https://www.notion.so/notion/Notion-AI-Supplementary-Terms-f...):
> Notion does not use your Customer Data or permit others to use your Customer Data to train the machine learning models used to provide Notion AI Writing Suite or Notion AI Q&A [added: our AI features]. Your use of Notion AI Writing Suite or Notion AI Q&A does not grant Notion any right or license to your Customer Data to train our machine learning models.
We do use various data infrastructure, including Postgres and the data lake, to index customer content both with traditional search infrastructure like Elasticsearch, as well as AI-based embedding search like Pinecone. We do this so you can search your own content when you're using Notion.
We wrote this article to explain how Notion's AI features works with customer data: https://www.notion.so/help/notion-ai-security-practices
At ParadeDB, we're seeing more and more users want to maintain the Postgres interface while offloading data to S3 for cost and scalability reasons, which was the main reason behind the creation of pg_lakehouse.
And I'm wondering if it's possible to update the S3 files to reflect latest incoming changes on the db table?
If you know Python, here's[0] a practical example of how Iceberg works.
AWS Athena packages Trino, I’ve been using it for some queries like “find all blocks that contain @-mentions”. It’s a great tool.
Versus what Notion is doing:
> We ingest incrementally updated data from Postgres to Kafka using Debezium CDC connectors, then use Apache Hudi, an open-source data processing and storage framework, to write these updates from Kafka to S3.
Feels like it would work about the same with Bufstream, replacing both Kafka & Hudi. I've heard great things about Hudi but it does seem to have significantly less adoption so far.
I believe just not having to handle a query queue system is already.
What does that mean?
"when we considered them in 2022" is significant here because both Iceberg and Delta Lake have made rapid progress since then. I talk to a lot of companies making this decision and the consensus is swinging towards Iceberg. If they're already heavy Databricks users, then Delta is the obvious choice.
For anyone that missed it, Databricks acquired Tabular[0] (which was founded by the creators of Iceberg). The public facing story is that both projects will continue independently and I really hope that's true.
Shameless plug: this is the same infrastructure we're using at Definite[1] and we're betting a lot of companies want a setup like this, but can't afford to build it themselves. It's radically cheaper then the standard Snowflake + Fivetran + Looker stack and works day one. A lot of companies just want dashboards and it's pretty ridiculous the hoops you need to jump thru to get them running.
We use iceberg for storage, duckdb as a query engine, a few open source projects for ETL and built a frontend to manage it all and create dashboards.
0 - https://www.definite.app/blog/databricks-tabular-acquisition