The “Build Your Own Redis” Book Is Completed
build-your-own.org
build-your-own.org
It all started with https://github.com/rcarmo/miniredis (which I forked to add and experiment with pub/sub), and I just found myself doing it again and again because Redis is the quintessential network service:
By implementing it, you learn about socket handling, event loop for a specific runtime, threading models, data representation, concurrency (if you want to do a multi-threaded version), etc. None of my "ports" are fully functional, but they all helped me sort out some of the above plus build tools, packaging, dependencies, etc.
It's "hello world" for core cloud native microservices, if you will (and without having to do REST or JSON stuff).
There’s a lot of interesting / subtle design choices worth studying.
https://github.com/antirez/RESP3/blob/master/spec.md
The best parts still haven’t been implemented.
There's a name I've not heard in a while!
> I actually build minimal Redis clones in every new language or runtime, or when I want to explore threading models.
100% agree with your advice; I'll definitively try to implement other parts of the Redis service in Go (eg: pub/sub, replication, clustering...) and probably repeat the same exercise when learning any new language.
“The Redis is an example of the server/client system.”
“the text reads too autogpt”
(I propose a pronunciation of “auto-jipped”).
Probably want to pick another term, since that's a homonym of a racial slur.
Or perhaps an even better sign would be the negated double turnstile, ⊭, "to denote the statement 'does not entail'" [1], making it more explicit. Hence an example would look like "the text reads too auto-jipped(⊭racism)", which can be read "the text reads too auto-jipped and this word, auto-jipped, does not entail racism in this context". Ok, done, racism solved, your move David Guetta [2].
Just like how ChatGPT fails at simple math - ChatGPT doesn't know math https://ai.stackexchange.com/questions/38220/why-is-chatgpt-...
I can now add “read” to that list. Let’s play word taboo! The rules are we can’t talk about GPT using anthropomorphic terminology.
Does GPT predict less than useful mathematical computations? Yes, and not just less than useful but basically useless.
Does GPT predict less than useful language translations, ranging from English-French, to summaries, in-the-style-of, etc? No, it’s actually quite useful as when confined to only the information contained in a prompt it doesn’t have to hallucinate an answer.
It is not useful to anthropomorphize the functionality of these tools in either a practical or legal context.
And everyone pick up a copy of Philosophical Investigations by Wittgenstein so y’all can learn about how to avoid snake-eating-tail discourse.
Personally, I'm not confident that ChatGPT wouldn't hallucinate facts when prompted to 'just' proof-read.
I would rather poor English than confident factual errors.
Our server will be able to process multiple requests from a client, to do that we need to implement some sort of “protocol”, at least to split requests apart from the TCP byte stream. The easiest way to split requests apart is by declaring how long the request is at the beginning of the request. Let’s use the following scheme.
The protocol consists of 2 parts: a 4-byte little-endian integer indicating the length of the following request, and a variable length request.
Starts from the code from the last chapter, the loop of the server is modified to handle multiple requests:
GPT suggested this instead:
Our server will process multiple requests from a client by implementing a protocol to separate requests from the TCP byte stream. The simplest method for separating requests is to include the length of each request at the start. The protocol consists of two parts: a 4-byte little-endian integer indicating the length of the request and a variable-length request. The server code from the previous chapter has been modified to handle multiple requests in the following manner:
There are no hallucinated facts because the most probable continuation of the given prompts is one that can gather all required information from the original text itself.
It's sort of like the difference between the truthfulness of analytic and synthetic claims. An analytic claim would be like "It is raining and you're outside, naked, and unsheltered so therefor water is falling on your skin from the sky." A synthetic claim would be like "It is raining outside".
Synthetic claims are said to be contingent on facts outside the text itself. These are the cases where GPT is completely useless.
The error rate for analytic claims is much lower although anyone who is writing anything should do a lot of review before publishing. Think of it like you asked your assistant to write something. You're gonna wanna read it over before you slap your name on it.
I mean, I actually don't care if you use these tools or not but your explanation of how it works will guide other readers in the wrong direction so I feel the need to correct the narrative you've presented.
GPT: Yes, I'd be happy to help you clean up language with grammatical errors. Please provide the text for me to review.
Me: Our server will be able to process multiple requests from a client, to do that we need to implement some sort of “protocol”, at least to split requests apart from the TCP byte stream. The easiest way to split requests apart is by declaring how long the request is at the beginning of the request. Let’s use the following scheme.
The protocol consists of 2 parts: a 4-byte little-endian integer indicating the length of the following request, and a variable length request.
Starts from the code from the last chapter, the loop of the server is modified to handle multiple requests:
GPT: "Our server will process multiple requests from a client by implementing a protocol to separate requests from the TCP byte stream. The simplest method for separating requests is to include the length of each request at the start. The protocol consists of two parts: a 4-byte little-endian integer indicating the length of the request and a variable-length request. The server code from the previous chapter has been modified to handle multiple requests in the following manner:"
---
That's the entirety of the interaction!
I haven't tested much but for the last day or so I've been thinking a lot about Kant, Frege, Quine and Wittgenstein!
GPT opens the door for some kind of empirical philosophy... like, what are the error rates for various kinds of tasks? Can we use a Kantian framework? How about Frege? How about Quine?
I mean, Quine is actually my favorite of the analytic philosophers because of his indeterminacy of translation argument and the notion that there really is no analytic/synthetic divide when you get down to it resonates well with me.
Death to metaphysics!
But there seems to be some use in differentiating between "All bachelors are unmarried" and "All bachelors are unhappy" if only because I'm now seeing how making a distinction can have a profound impact on the usefulness of GPT completions.
Briefly and half-assed, Quine's argument is that because you would have to be familiar with language and culture in order to understand "All bachelors are married" that the meanings of those words are fact-like and outside the scope of the proposition.
If GPT is able to do some Frege-like substitution of synonyms it is because it has this compressed language model which seems to lend credence to Quine's arguments.
I find the quality of answers you receive out of GPT drastically changes with the way you phrase questions.
I don't think I would ever have come up with asking the question in the way you did.
As someone who has english as a second language I found GPT ofthen produced incorrect and low quality answers while most of my native english speaking colleagues were getting high quality answers. Looking at their prompts compared to mine it's all down to differences in how questions are phrased.
"If you can't write code why should I believe that the prose you write is any better?"
That being said they are selling it and that's enough reason to complain.
If you want to pre-judge all technical content coming from people who have not spent huge portions of their life living in an English speaking country as being of no value, then I'm sure that will protect you from some bad content, but you're going to be missing out on an awful lot of good stuff too.
What is still missing, in my opinion, and is badly needed, is content or even an idea on how to teach taking such projects from toy prototype version to the production quality one.
By design, the service doesn't provide any documentation; it provides references to existing technical documentation (of any kind, including blog posts).
Those who expect a focused introduction to each topic will find it very tedious or hard to proceed (for example, the SQLite exercise has important details buried in a very large and confusing webpage), and likely hate it; those who like the challenge of understanding loads of raw documentation will love it.
Providing a structured path and (automated) test means can be stimulating, and can make the difference between (deciding to) learning something or not.
Some people are certainly entirely autonomous, but at the very least, there is a spectrum of need for stimulation when approaching a topic of study.
Exercism is the closer service I can think of, but it's based on simple exercises, not real-world projects.
There are a few books that have a similar target (build X), but much narrower in scope (either a single language or pseudocode, and certainly no automated testing/team features).
This model is great for someone who loses patience with all the groundwork setup.
However, I do agree that the next leg is taking this MVP/Prototype level to production and ideally sell it as a real alternative to the commercial version of Redis.
Great idea, but only accessible to the rich.
I think this is a bit much. I wouldn't buy it, but I think this is deep technical content that you might not stick on forever, and they need to pick a price for it. If it's too high or low they'll find out soon. No need for any other inputs than that.
*Saw in sibling comment this is recent addition. Without it, yeah, it may be a bit expensive for one person.
I think Redis is a great server to “build yourself” as you don’t need to start with much to get it going.
And just using a hash map in memory isn't sufficient. You'll have unbounded growth. You need a max size and need to evict. Then you want to do that efficiently so your not wasting space on useless keys...
Although people constantly bug me on attempts to build their own ClickHouse. Someone is trying to do it with Apache Arrow and DataFusion. Folks from DuckDB are trying to build their own crippled version of ClickHouse. Friends from China doing it with their Apache Doris. InfluxDB is being rewritten to be closer to ClickHouse in an attempt to make it better and so on...
I also think calling DuckDB a crippled version of Clickhouse disingenuous. That's like calling SQLite a cripped version of postgres. They have very different goals.
That's also a bit like calling Clickhouse a build-your-own vectorwise/MonetDB because they did it first.
> The end result is a mini Redis alike with only about 1200 lines of code. 1200 LoC seems low, but it illustrates many important aspects the book attempts to cover.
The techniques and approaches used in the book are not exactly the same as the real Redis. Some are intentionally simplified, and some are chosen to illustrate a general topic. Readers can learn even more by comparing different approaches.
I wouldn't emphasize the importance is understanding Redis per-se but the ideas around a system like Redis.
They mention "NGINX, SQLite, PostgreSQL, Kafka, Linux kernel, etc." - of which I'd only consider NGINX and the Linux kernel (and Redis) as "building blocks". The others might be part of their own preferred stack, but if you mention Postgres, why not MySQL? If Kafka, why not RabbitMQ?
But yeah, NGINX, Redis, and the Linux kernel are basically outside of discussion.
That confused me a bit. Otherwise this looks interesting, thanks for sharing.
There was that chicken and egg problem then, when potential employers would skip my CV because I didn't have any TDD experience and I couldn't find anywhere how to learn this.
Before I even learned that something like testing exist, I was so confident like "oh I could write something like that over the weekend, how come they needed months to do that?". Then my software would crash the first time someone other than me used it.
Anyway - what I want to say that while this book sounds like a great idea, without showing TDD and how to write code so that it can be proven it works the way intended and that it can handle unhappy paths and edge cases, it won't teach someone trying to learn programming much and it doesn't actually stand over about million other books about programming that really just scratch the surface and don't show how to write production ready code.
That's what is still missing on the market. It's almost like a well kept secret that only developers working at large corporations know.
That skill was very very difficult to acquire.
> it won't teach someone trying to learn programming much and it doesn't actually stand over about million other books about programming that really just scratch the surface and don't show how to write production ready code.
It's not a book to teach people how to program, infact the author goes out of their way to mention only C and minor c++ has been used and that it may be beneficial for learning how to build out such a POC of redis, to DIY your own.
This is not a book to hand over to an outsourcing company and expect production ready work. Nor was it described as such.
But back to the topic. I'm pretty sure there are enough books about how to write tests. And more or less all engineers understand the value of testability and coverage (not necessary tdd!). At leas I wouldn't need a book for this. But books about building something closer to complex real world systems - that's a good stuff engineers would enjoy.
> how to write code so that it can be proven it works the way intended and that it can handle unhappy paths and edge cases
That is not what TDD promises. All TDD does is ensure you have some tests at all, which is a good base on which to add more tests, so that when you find edge cases you can do more TDD, and TDD is a nice way to work. Nobody in history has come up with a way to make testing comprehensive. You can try quickcheck and fuzzing to generate lots and lots of test cases, but they are still sampling from the input space rather than covering it. You can cover the input space for a single 32-bit number input, but that’s about it. The only way to do better is to formally prove your software is correct using mathematical logic. TDD does not provide this and never will. You might have seen “100% test coverage”, but that claim is close to meaningless with respect to the range of possible inputs that your code is meant to handle. All it says is that every edge case that you did think of has a test that exercises it. A function consisting of “return 0” only will have 100% test coverage with a single test in the suite. Doesn’t mean it works. See SQLite, which still finds bugs despite 100% test coverage.
What you seem to be looking for is someone to teach you how to think of edge cases. That’s a skill that can’t really be taught. Just have a go. When you find an edge case later that you didn’t think of, great, maybe you will think of similar ones next time.
Then I found Structure and Interpretation of Computer Programs (SICP)[0] and the video lectures from MIT[1]. I had an epiphany when Sussman talked about "wishful thinking" in video for lecture 1b[0] (around 48:00 in). The lesson was something along the lines of start naming functions that would do what you needed done, and write them later. Just pretend they exist and eventually bring them into existence. SICP has so many gems. It really made a difference for me.
If applied to TDD, write a test that won't even compile because the function under test doesn't even exist. Then iterate by writing something that compiles, but will probably fail, and then improve it until the test passes.
[0]: https://mitp-content-server.mit.edu/books/content/sectbyfn/b...
[1]: https://ocw.mit.edu/courses/6-001-structure-and-interpretati...
[2]: https://www.youtube.com/watch?v=V_7mmwpgJHU&t=2s
Edit: How to Design Programs and How to Design Worlds are two other resources I enjoyed. I don't write programs in Lisp or Scheme anymore, but the experience of just running through the exercises in these and other books was enlightening.
Is Redis really that critical to modern computing?
Modern web applications? I can see it.
Redis rules and fixes a lot of problems, especially when you probably already have a server/cluster for something else. A building block of modern computing is a bit overzealous to me.
KV persistence is covered in Chapter 3, right from the start. Redis is also mentioned as an example for an in-memory store with „weak durability by writing to disk asynchronously“.
It's not like there are decades of paying the price of such decisions...
This is exactly the sort of playground effort where those who do want to learn can get their hands dirty. If this is their first time doing this kind of project, they shouldn't use any of this code in production. There will be all sorts of cheated corners and vulnerabilities, and not just in obvious high-risk places like this. That's how hands-on learning goes.
https://nvd.nist.gov/vuln/detail/CVE-2021-32675 https://nvd.nist.gov/vuln/detail/CVE-2021-41099 https://nvd.nist.gov/vuln/detail/CVE-2021-32761
I agree it's instructive, but on the other hand, the time might be right to start teaching why not to do this stuff. Modeling a protocol parser in a high-level language that can spit up correct and highly-optimized C code would be just as instructive and perhaps even more fun.
Yes, "read in a buffer, cast that buffer to the protocol-specific struct, read the various fields from memory, etc" are all operations one should generally NOT do. This was instructive maybe three decades ago, when the internet was relatively safe.
Knowing how to do this RIGHT is important.
The place this is still helpful is low-level programming, and embedded is a far better place to learn. If one part of your microwave is talking to a sensor, security isn't really an issue, since you control both ends.
(And rereading, I don't mean "growing out" in a derogatory way -- I grow into and out of a lot of things -- and low-level programming is something anyone can enjoy for a few years)
Kafka?
We almost configured it, but instead implemented our own Cache web service and used the built in memory/cache management of that. Yes, it's only accessible via http but it's given us a lot of flexibility. We are primarily using it for caching of large datasets (hundreds of MBs). When service has to be restarted, it makes a call to get all the items it needs.