SimpleDB: A Basic RDBMS Built from Scratch
awelm.com
awelm.com
And good choice to support transactions up front. Adding that in later is pretty challenging.
Mine doesn't support transactions and refactoring to support it would take a while. :(
If HN-ers are interested in more database Internals content check out /r/databasedevelopment too.
[0] https://notes.eatonphil.com/database-basics-indexes.html
[1] https://ocw.mit.edu/courses/electrical-engineering-and-compu...
CMU Intro to Databases Labs: https://15445.courses.cs.cmu.edu/fall2021/assignments.html
CMU Intro to Databases Lectures: https://www.youtube.com/playlist?list=PLSE8ODhjZXjZaHA6QcxDf...
BusTub - CMU's Version of SimpleDB: https://github.com/cmu-db/bustub
https://en.wikipedia.org/wiki/Amazon_SimpleDB
> Highly Available – It’s Amazon. Running Erlang. Whoa.
https://web.archive.org/web/20110623221347/http://www.satine...
- Why not just use mmap and let the OS handle page caching?
- Why not use a write-ahead-log for all writes, with a background thread applying transactions to the read replica asynchronously? In most cases eventual consistency is fine and you don't need to query the latest version of the data, but this can be an optional query parameter.
- Instead of embedding an unpredictable SQL-to-code compiler, why not provide direct access to physical (relational algebra) operators via function calls? eg let me write: select([fields]).where(x=2).join(table) etc ... letting me select which index to use, how to do the joins etc.
1) I think databases like to manage pages directly because the db can make more optimizations than the OS because the db has more context. For example, when aborting a transaction the db knows its dirty pages should be evicted (i'm not sure if mmap offers custom eviction). Also I believe if the db uses mmap, it loses control over when pages are flushed to disk. Flush control is necessary for guaranteeing transaction durability.
2) What you're describing here sounds similar to a LSM-tree database (e.g. RocksDB). They are used often for write-heavy workloads because writes are just appends, but they might not be great for read-heavy things.
3) This reminds me of PRQL[1] (which was trending on Hacker News last week) and Spark SQL. I'm not too familiar with this area though, so I can't really say why SQL was designed this way.
[1] https://github.com/max-sixty/prql?utm_source=hackernewslette...
2) Was thinking more of an event-sourcing model, whereby you log the SQL statements first, then update a B-Tree in the background.
Read via mmap, write by appending to a log and asynchronously applying the changes to the file.
3) Rather than yet another QL, expose a higher level API that I can target in any language
Could you please clarify the license for your code? Thanks!
The only one comment to add is
Amazon has an Amazon SimpleDB.
So different name might be before for your branding.