The “Build Your Own Database” book is finished
build-your-own.org
build-your-own.org
One small comment: the /about page only has the books listed and nothing else. Some people will probably be interested in knowing a thing or two about the author before getting the books (and there are probably a million people with the same name, so difficult to Google)
I understand that sometimes people want to keep their identity private but for an author it's tough.
https://news.ycombinator.com/item?id=34557389 (Show HN - 5 comments)
https://news.ycombinator.com/item?id=34572263 (129 comments)
https://news.ycombinator.com/item?id=35212660 (65 comments)
Why do you say the codecrafters stuff is lower quality?
Maybe the problem is the format...
I think the section on internals is a good jumping off point that leads to a lot of deeper content.
Once I had it working, I realized the KV stores I used for tags were just like columns in a relational table built within a columnar database. Querying for files based off their tags was very much just like SQL queries for table rows. So I tried using them to create relational tables.
They turned out to be incredibly fast at a variety of queries (the bread and butter of databases) without needing to create separate indexes in order to get optimal performance. I thought database experts would be intrigued when I showed how much faster my system was than other conventional RDBMS setups on the same hardware. I guess I was surprised when almost no one was even curious how it did it.
Here is a simple, short video comparing it to SQLite: https://www.youtube.com/watch?v=Va5ZqfwQXWI
The book uses Golang for sample code, but the topics are language agnostic. Readers are advised to code their own version of a database rather than just read the text.
I see the code in sparse but wanted to get the feel end to end.
Or if you see any chapter you can literally see Go code.
Other than the "if err != nil" I wouldn't have recognized the code samples. Go's error handling is a big reason I've never taken a closer look.
I don't really see any alternative. It also makes you carefully think about how you plan on managing errors in your codebase, which also seems like a very sane thing to enforce.
This is self-contradictory. In particular, the only way you can have reliable error handling is if you are forced to think about each possible failure.
I assume by "cannot easily be ignored" you mean the way exceptions blow up at runtime? I don't find that an acceptable default for any non-scripting language.
They'd prefer an easier way to not bother dealing with them with them without outright ignoring them via _
In Rust that entire check can be a single "?" symbol. How much syntactic sugar is too much is a matter of preference, but I personally think that properly handling all errors without syntactic sugar turns into an unreadable mess because there's just a lot of things which could go wrong.
I think just forwarding all low-level errors is a really bad habit, and go forces you to at least think about this.
Why exactly is that a bad habit? In almost all situations where I return an error I already have enough context, I'm just wondering what else I'd add to that.
> go forces you to at least think about this
Boilerplate code definitely doesn't incentivize thinking.
In a network environment (which is originally what go was made for) you often need to add tracing information, business-level identifiers or processing information related to your state etc.
I'm currently writing a fairly complex api in go, and to be honest this really hasn't bothered me once.
Not to say it doesn't exists, but with time i've come very suspicious of people complaints over go. Most of the time those complaints come from people that didn't realize they missed an opportunity to have written a much much more elegant solution to their problem.
I've taken a look at Go, and while it does seem pretty approachable, it's definitely not nearly as common as Python/JS, and it's always significantly harder for me to learn a new concept when the examples are also in a language I'm unfamiliar with. Maybe that's just me, though.
[1] https://survey.stackoverflow.co/2022/#most-popular-technolog...
Languages like Python or Javascript are so far removed from the system-y side of programming that the way you would implement the concepts in those languages would not translate to the way you would actually build a "real" database which is I think the purpose of the book. I think the objective isn't to teach the abstract concepts but how those concepts are expressed in real systems.
That said, it'd be an interesting read on how to make a DB in pure Python.
Regardless, you can do some crazy things in Node. See these notes about Node and MySQL:
https://github.com/tigerbeetledb/tigerbeetle/blob/main/docs/...
For sure they've got different performance profiles.
Very impressive what your group was able to do with Node, and continuing on with Zig. It's got me interested in learning more.
There's somethings the compiler will fail on like unused variable and the likes, but for the most part you need added static analysis and style checking -- some of which ships with the Go compiler.
I’ve always wanted to understand how databases work so I can build my own.
Excited about this!
https://github.com/samsquire/multiversion-concurrency-contro...
First read TransactionC.java then read MVCC.java ( or follow the methods that TransactionC calls)
11. Atomic Transactions
11.1 KV Transaction Interfaces
11.2 DB Transaction Interfaces
11.3 Implementing the KV Transaction
12. Concurrent Readers and Writers
12.1 The Readers-Writer Problem
12.2 Analysing the Implementation
12.3 Concurrent Transactions
Part 1: Modify the KV type
Part 2: Add the Read-Only Transaction Type
Part 3: Add the Read-Write Transaction Type
12.4 The Free List
12.5 Closing Remarkshttps://news.ycombinator.com/item?id=6725387
Not a book, but. :-)
It’s ... part of a book! :)
Couldn't be a better story.
2. Manage access to it through a server
There you go, you know have a database.
2. Draw the rest of the owl
It uses mmap.
All databases that are any different than these two points are just bad.
This has real "draw the rest of the owl" energy.
How to make a DBMS? 1. Get a file 2. Make a DBMS
Complex when you want to
> 2. Manage access to it through a server
A database is only "easy" IF:
- Append only
- No real "delete" or "updates" just to reiterate the above.
- Only Sequential scan
- Only need simple iterator-per-row
- No maintain secondary stuff like indexes, so not need to coordinate changes
- No concurrency
- Fit in RAM, and I mean in few MB
- No need to deal with SQL, use his own DSL (sql is so bad! so much weird stuff!, but is ok to have something sql-ish like LINQ)
- No need to deal with recursive data types, only scalars
- Is only embebed
- No need auth or security validations
Ok, after making this list, I sure forgot some other tips to make this easy!
Obviously you'd have a file per column (or index), and use directories to represent tables.
This is the correct way of doing it and yet so few databases do it.
mmap is not a panacea, it improves specific access patterns by incurring specific costs, it's definitely not true that mmap is the right choice for all databases
Of course, depending on the other things of the list this is or not a major issue. Is more about how combining several ideas leads to a easy or complex implementation.
why does generating indexes at start-up not count as having indexes?
(asking because i do this all the time)
This system was rarely restarted, so in practice it didn't matter what it did at startup, as long as it didn't take more than a few minutes, but it did place some limitations on data size. (This was a 32-bit system and everything was memory mapped.)
There you go!
it feels correct, even though it's not a complete guide
(reliably persisting changes to disk is a big part of what dbs do, but is missing here)