Building a new database management system in academia (2017)
cs.cmu.edu
cs.cmu.edu
I have rolled out two transactional databases of my own. In both cases I had to provide very specific properties and for some reason I could not find an existing product that would meet all requirements. For example, one of them was an embedded device that was very restricted in memory, all operations needed to run with hard bounds on time and memory and the storage for the data was a flash chip without wear levelling which required the database itself to manage writes to prolong the chip's life.
The key is to notice how your database system is going to be different from others and what properties are not essential.
Also, making general purpose DBMS tends to be much more complex vs making more niche solutions where you know a bit more about what the uses are going to be and what kinds of loads you can expect.
Creating a custom engine for a given application can be very simple task because you can easily cross out requirements you don't care about and you only care that it works well for the loads that this particular application can generate.
Also, it is unlikely you are going to beat fierce competition in general purpose "and a kitchen sink" database management system market, but much easier to find a niche that is underserved and create a usable, competitive product with relatively little effort. That's how SQLite started.
Today, building a database from scratch is extremely difficult, for several reasons: 1. it anyways takes a long time; 2. there are so many successful (open-source) databases; 3. hiring top engineers are so expensive. 4. you won't get enough attention unless your system is drastically better than existing ones.
An interesting observation is that very few database was built since 2020 - almost all the newly built databases were developed on top of existing databases (PostgreSQL, ClickHouse, etc).
I started building RisingWave (https://github.com/risingwavelabs/risingwave) in early 2021. The only reason we built the system from scratch was that none of the existing systems can address the problem we are solving - distributed SQL stream processing at cloud scale. We tried Flink but gave up, as it's too heavy and it's architecture was not designed for the cloud environment.
If you want to build a database from scratch, or are simply interested in databases, we may talk.
Did the DBMS ever come into existence? (If so: link, please). If not: Why should we be interested in this announcement in 2023?
Not sure what op's intention with this was
1. OtterTune Start-up (https://ottertune.com)
2. Biological Daughter (https://twitter.com/andy_pavlo/status/1187841279260004355)
3. Pandemic
When the pandemic first started, I had a bunch of CMU students reach out to me saying that their summer internships were rescinded and that they were looking for a project to work on so that they wouldn't have a gap in their CV. I ended up taking on any student that could program C++ even if they hadn't taken my DB class before. It as my way of trying to help. But our research group grew to about 35 people. That was not sustainable and the code quality suffered greatly.
We ended killing the project and now all our self-driving work is done in the context of Postgres (https://db.cs.cmu.edu/papers/2023/p27-lim.pdf).
I also now realize that building the DBMS engine first then building the query optimizer second is the wrong order. Our future project is going to start with the optimizer first.
I was once discussing MVCC vs 2PL with an experienced Sybase and SQL Server guy, and he claimed that, when transactions are implemented properly and the database is well-designed (no surrogate keys, in particular), 2PL leads to better performance and no deadlocks, while “readers do not block writers” leads to lots of aborted transactions in a heavy OLTP workload. I verified that (I should still have the code around): lots of conflicts in PostgreSQL vs smooth concurrent execution with no retries in Sybase and SQL Server.
I have since heard similar opinions from other SQL Server practitioners: they disable MVCC and rely only on good ol’ 2PL.
https://www.vldb.org/pvldb/vol8/p209-yu.pdf
All the protocols regress to the same. This evaluation was only with stored procedures though. It would be worth doing a similar investigation with conversational DB protocols (e.g., JDBC, ODBC).
It works if you target for ~10% outcome if you have a good CI system with a decent test coverage and a ton of fuzzing.
What's your opinion of recent attempts like LingoDB, that move the query optimizer into a traditional compiler stack, in this case, MLIR?
The problem with (most) query optimizers is that they take a one shot approach at optimization. I think an optimizer should be built from the groundup to support adaptive query optimization. Something similar to Berkeley's Eddies project from 20 years ago.
I'm sure there are details I'm missing here, but I do believe the general approach could do implemented in LingoDB (or similar) as a compiler transformation, so the actually cost-to-develop this approach would remain tractable.
[0] I suspect you'd need to model the whole thing as a streaming network so that you can update the network parts as you go, effectively re-wiring the streams while not invalidating earlier results. So SAC+logic to map from one stream architecture to another. JITs that support de-optimization have to do something similar (with a lot of careful upfront design), so that's at least plausible.
A similar thing happened in physically-based rendering, with the publication of the PBTR series of books. As a result of that effort, a lot of really solid research improving various aspects of rendering occurred. Extremely influential long-term.
mutable could have a similar kind of experience if presented that way to the public.
Building a Database System in Academia - https://news.ycombinator.com/item?id=13931752 - March 2017 (15 comments)
Not a good look for people browsing at work
I hope no one had to have a hard conversation with their boss just now.
(for the curious, the link was previously *NSFW* http://macrobase dot io *NSFW*)
- Materialize: 2017
- DuckDB: 2018
- RedPanda: 2019
- TigerBeetle: 2020
Fair. I'm talking about databases with funding backing them (either by universities or otherwise).
https://use-fireproof.com/docs/architecture
It's also the easiest way to write React apps. Here are some ChatGPT expert builders that I've trained to use the CSS framework of your choice with Fireproof: https://use-fireproof.com/docs/chatgpt-quick-start/#react-ex...
[1]: https://github.com/kuzudb/kuzu
[2]: https://kuzudb.com/blog/meet-kuzu.html
[3]: https://kuzudb.com/blog/what-every-gdbms-should-do-and-visio...
[2] https://www.philipotoole.com/9-years-of-open-source-database...
https://duckdb.org/pdf/SIGMOD2019-demo-duckdb.pdf
EDIT: oh the article is old
[0] https://twitter.com/motherduck/status/1615487300523429889
Yes we have commit history from 2015 but that's from an earlier db project (noms) that we forked and built on top of
Technology-wise, writing a toy DBMS is nothing difficult. Even undergraduates can do it.