Scaling One Million Checkboxes to 650M checks
eieio.games
eieio.games
I think you hit every type of interruption and point of failure, except storage space, and it is great to see your resolutions.
I wasn't aware Redis could do the Lua stuff which makes me very interested in using it as an alternative state.
As for the bandwidth - one of my biggest gripes with cloud services as there is no hard limit to avoid billing overages.
FWIW I certainly hit storage space in some boring ways - I didn't have a good logrotate setup so I almost ran out of disk, and I sent my box-check logs to Redis and had to set something up to offload old logs to disk to not break Redis. But neither of those were very big deals - pretty interesting to have a problem like this where storage just wasn't a meaningful problem! That's a new one for me.
And yeah, thinking about bandwidth was such a headache. I was on edge for like 2 days, constantly checking outbound bytes on my nic and redoing the math - not having a hard cap is just really scary. And that's with Digital Ocean, which has pretty sane pricing! I haven't used the popular serverless stuff at all, but my understanding is that you get really gouged on bandwidth there.
(also yes, lua-in-redis is really incredible and lets you skip sooo many hard/racey problems as long as you're ok with a little performance hit, it was a joy to work with)
Probably the key takeaway that many early-career engineers need to learn. Scaling's not a problem until it's a problem. At that point, it's a good problem t have, and it's not as hard to fix as you might think anyway.
Scaling those systems is a total bitch.
One Million Checkboxes - https://news.ycombinator.com/item?id=40800869 - June 2024 (305 comments)
[0] 22mb: https://blog.winricklabs.com/images/pixmap-rewind-demo.gif
regarding storage of the log, compression is a thing :)
Think the total cost was about $850, which was (almost) matched by donations.
I made a mistake and never really spun down any infrastructure after moving to go, and also could have retired the second redis Replica I spun up; think I could have cut those costs in half if I had been focused on it. But given that donations were matching costs and there was so much else going on I wasn't super focused on that.
I kept the infra up for a while after I shut down the site (to prepare graphs etc) which burned a little more money, so I'm slightly in the hole at this point but not in a major way.
This should, especially after your Go rewrite, be enough to host everything?
If I was planning to keep the site up for long I would have moved, but in this case I knew it was temporary and so toughing it out on DO seemed like a better choice.
I recently had to rebuild my box because I left a postgres instance open on 5432 with admin:admin credentials, and without default firewalls in place it got owned immediately.
That would have been less painful on DO for sure.
Painful to use managed redis when I was debugging (let me log in! let me change stuff! give me the IP!! ahhhhh!!!!) but really nice otherwise. A little painful to think about giving that up, although I coulda run a really hefty redis setup on hertzner for very little money!
Kudos to the author - your projects are great.
I don't think it's unscalable at all - it's just that if you do this there's not a great story for adding a second machine if the first one can't handle the load (but ofc a beefy machine with a fast implementation could handle a lot of load).
When we did the go rewrite we considered just getting one beefier box and doing everything there (and maybe moving Redis state in memory), but it felt easier and safer to do a drop-in rewrite.
I mean this is basically your redis situation right? Just with a very specialized "redis".
You could scale this out, even after the pretty massive ability to scale up is exceeded. Have some front-end servers that act as a connection pooler to your datastore. Or shard the datastore, and have clients only request from the shards that they are currently looking at.
The entire timeframe of this project was 2 weeks, and the critical period (most activity / new eyes) was a couple of days.
> "is my specialized datastore gonna be faster than Redis"
Absolutely! With how efficient this code would be, you'd likely never need to scale horizontally, and in that case it is extremely easy to compete with a network hop (at least 1ms latency) versus L1 cache (<50ns)
The comparison with redis otherwise only applies once you do need to scale horizontally.
There's also the fact that redis must be designed to support any query pattern efficiently; a much harder problem then supporting just 1-2 queries efficiently.
Oh sure, yeah, I think we just agree then!
But that's based on my skillset and the timeframe of the project; I'm not disputing that the project is bounded by Redis's performance.
However... You'd need to persist those booleans somewhere eventually, of course, if you want the state to survive a process restart. And if you want multiple concurrent connections from the same box, you have to somehow allow multiple writers to the same object. And if you want multiple boxes (for redundancy, load spreading, geo distribution...), you need a way to do the writing and reading from several different boxes...
By this time you're basically building redis.
Sorry that some of the stuff went over your head! I wanted to include longer descriptions of the tech I was using but the post was already suuuper long and I felt like I couldn't add more.
Very happy to answer any questions that you've got here!
I'm not sure how you'd simplify the architecture in a major way to be honest. There are certainly services you could use for stuff like this, but I think there you're offloading the complexity to someone else? Ultimately you need:
* A database that tracks which boxes are checked (that's Redis)
* A choice about how to put your data in your database (I chose to just store 1 million bits for simplicity)
* A way to tell your clients what the current state is (I chose to send them all 1 million bits - it's nice that this is not that much data)
* A way for clients to tell you when they check a box + update your state (that's Flask + the websocket)
* A way to tell your clients when a box is checked/unchecked (that's also Flask + websockets. I chose to send both updates about individual boxes and also updates about all 1 million boxes)
* A way to avoid rendering 1 million dom elements all the time (react-window)
The other stuff (nginx for static content + a reverse proxy) is mostly just to make things easier to scale; you could implement this solution without those details and the site would work fine, it just wouldn't be able to handle the same load.Probably this is not the simplest thing to do if you want a certain degree of reliability. Should be definitely easier than writing the entire storage engine, but likely an overkill for this kind of overnight hobby projects.
you could also implement a bitset over AtomicLongArray.
more complicated: partition into x*x chunks and rw-locking those. this could be backed by an mmap'ed a million bytes for persistence, but no idea if that'd make the app disk io bound or something.
It may have been possible to do it all in-memory on a single large host, but then if it is unable to meet the demand or fails for whatever reason then you are completely out of luck.
Always wonder why clearly smart people put themselves through these heroic gymnastics just to avoid learning a bit of <sys/socket.h>
Still a cool write-up B).
My takeaway from this experience is that I have no idea which of my sites is going to be popular - most of my projects are not nearly this successful - and that I should be optimizing for speed of creation, not speed of an individual project. My comparative advantage is "writing ok code and making ok decisions very quickly" and I'm going to continue to lean into that.
Will your next post be a statistical analysis of which checkboxes' were the less/most checked ?
I remember scrolling way down and being kind of sad that the one I choose was almost instantly unchecked.
When I go to https://onemillioncheckboxes.com/ nothing is checked and in the JS console I just see
{"total":0,"totalGold":0,"totalRed":0,"totalGreen":0,"totalPurple":0,"totalOrange":0,"recentlyChecked":false}
> We passed 650 million before I sunset the site 2 weeks later.
https://gist.github.com/jeff-hykin/4cdebafd8698298d021f103e2...