Did the DBMS ever come into existence? (If so: link, please). If not: Why should we be interested in this announcement in 2023?
Did the DBMS ever come into existence? (If so: link, please). If not: Why should we be interested in this announcement in 2023?
Not sure what op's intention with this was
1. OtterTune Start-up (https://ottertune.com)
2. Biological Daughter (https://twitter.com/andy_pavlo/status/1187841279260004355)
3. Pandemic
When the pandemic first started, I had a bunch of CMU students reach out to me saying that their summer internships were rescinded and that they were looking for a project to work on so that they wouldn't have a gap in their CV. I ended up taking on any student that could program C++ even if they hadn't taken my DB class before. It as my way of trying to help. But our research group grew to about 35 people. That was not sustainable and the code quality suffered greatly.
We ended killing the project and now all our self-driving work is done in the context of Postgres (https://db.cs.cmu.edu/papers/2023/p27-lim.pdf).
I also now realize that building the DBMS engine first then building the query optimizer second is the wrong order. Our future project is going to start with the optimizer first.
I was once discussing MVCC vs 2PL with an experienced Sybase and SQL Server guy, and he claimed that, when transactions are implemented properly and the database is well-designed (no surrogate keys, in particular), 2PL leads to better performance and no deadlocks, while “readers do not block writers” leads to lots of aborted transactions in a heavy OLTP workload. I verified that (I should still have the code around): lots of conflicts in PostgreSQL vs smooth concurrent execution with no retries in Sybase and SQL Server.
I have since heard similar opinions from other SQL Server practitioners: they disable MVCC and rely only on good ol’ 2PL.
https://www.vldb.org/pvldb/vol8/p209-yu.pdf
All the protocols regress to the same. This evaluation was only with stored procedures though. It would be worth doing a similar investigation with conversational DB protocols (e.g., JDBC, ODBC).
It works if you target for ~10% outcome if you have a good CI system with a decent test coverage and a ton of fuzzing.
What's your opinion of recent attempts like LingoDB, that move the query optimizer into a traditional compiler stack, in this case, MLIR?
The problem with (most) query optimizers is that they take a one shot approach at optimization. I think an optimizer should be built from the groundup to support adaptive query optimization. Something similar to Berkeley's Eddies project from 20 years ago.
I'm sure there are details I'm missing here, but I do believe the general approach could do implemented in LingoDB (or similar) as a compiler transformation, so the actually cost-to-develop this approach would remain tractable.
[0] I suspect you'd need to model the whole thing as a streaming network so that you can update the network parts as you go, effectively re-wiring the streams while not invalidating earlier results. So SAC+logic to map from one stream architecture to another. JITs that support de-optimization have to do something similar (with a lot of careful upfront design), so that's at least plausible.
A similar thing happened in physically-based rendering, with the publication of the PBTR series of books. As a result of that effort, a lot of really solid research improving various aspects of rendering occurred. Extremely influential long-term.
mutable could have a similar kind of experience if presented that way to the public.