A distributed key value store in under 1000 lines
github.com
github.com
- The API supports GET/PUT/DELETE (no range queries).
- It uses 1 master server and N volume servers.
- The master server keeps track of all keys and all requests need to go through this server. The values are stored in one or more volume servers using rendezvous hashing.
- The volume servers are just Nginx + DAV which stores data on disk.
- There's no live support for rebalancing. If you change the backend volumes, then you need to take the master server down, run a re-balancing job and then bring it up again.
- Objects are locked in the master server during modification. This means that while a server is processing PUT/DELETE all GET requests for that key will be blocked.
- There's no handling of errors from the volume servers. E.g. if one volume server fails to store data, then it won't rollback the changes done to the other volume servers nor do anything in the master server.
- The benchmarking tool (thrasher.go) does "10000 write/read/delete" in 16 threads/goroutines. This means that it generates 10k random keys (i.e. no contention at all), and for each key it does a PUT, then a GET and then a DELETE. Their result show that it's capable of handling 3815.40 keys/sec (i.e. multiply this by 3 to get the request count).
> - The API supports GET/PUT/DELETE (no range queries).
There are range query on keys:
# list keys starting with "we" curl -v -L localhost:3000/we?list
> - There's no handling of errors from the volume servers
If a GET/PUT/DELETE fails it is communicated to the master server, who doesn't write anything in its own database
All in all I believe it does quite a lot for less than 1000 lines already. Knowing its limitations I'm curious to see how well it handles production usage
I took it to mean the webserver doesn't support "HTTP range requests"[0], as opposed to ranges of keys.
[0] https://developer.mozilla.org/en-US/docs/Web/HTTP/Range_requ...
Ah, I missed this!
> If a GET/PUT/DELETE fails it is communicated to the master server, who doesn't write anything in its own database
Yes, but there's no protocol for cleaning it up. Example: If the second replica fails to store it, then the first replica will still have the data and there doesn't seem to be a way that it will be removed. This might end up "leaking" storage. Also, if you then attempt to do a rebuild it seems to me that the file will be revived.
> All in all I believe it does quite a lot for less than 1000 lines already.
Oh sure. I wasn't trying to be negative. I think it's great that people play around with these ideas! Distributed systems are complicated beast, and there's only so much you can do in 1000 lines.
Another thing I realized: It buffers all received values in-memory in the master server. You might need quite a lot of memory if you intend handle uploads of big files in parallel.
Like you try running this in production and 3 months later find out it instantly dies if there's more than 4% packet loss between you and the DNS server or something
I'm not an expert in those products but from what I saw people are usually aware that geohot's system is a work in progress, though so far he pulled it off because it's cheap(relatively), not worse than the competition and constantly updated.
https://comma.ai/ says it’s been driven 30 million miles (over years), has anything gone wrong?
> openpilot ALC and openpilot LDW do not automatically drive the vehicle or reduce the amount of attention that must be paid to operate your vehicle. The driver must always keep control of the steering wheel and be ready to correct the openpilot ALC action at all times.
While changing lanes, openpilot is not capable of looking next to you or checking your blind spot. Only nudge the wheel to initiate a lane change after you have confirmed it's safe to do so.
Many factors can impact the performance of openpilot ALC and openpilot LDW, causing them to be unable to function as intended. These include, but are not limited to:
Poor visibility (heavy rain, snow, fog, etc.) or weather conditions that may interfere with sensor operation. The road facing camera is obstructed, covered or damaged by mud, ice, snow, etc. Obstruction caused by applying excessive paint or adhesive products (such as wraps, stickers, rubber coating, etc.) onto the vehicle. The device is mounted incorrectly. When in sharp curves, like on-off ramps, intersections etc...; openpilot is designed to be limited in the amount of steering torque it can produce. In the presence of restricted lanes or construction zones. When driving on highly banked roads or in presence of strong cross-wind. Extremely hot or cold temperatures. Bright light (due to oncoming headlights, direct sunlight, etc.). Driving on hills, narrow, or winding roads. The list above does not represent an exhaustive list of situations that may interfere with proper operation of openpilot components. It is the driver's responsibility to be in control of the vehicle at all times.
I wouldn't call it jaded, I'd call it experienced.
1. It solves a more limited set of use cases. It's perfectly reasonable to want to solve a wide range of use cases, but there's value in more limited software which solves a more limited set of use cases, too.
2. It caters to a more limited set of users. Suckless is perhaps the epitome and most extreme flavour of this, where they just excluded things like config files in favour of a config.h file. This doesn't make the software less stable, but it does make it less usable for some users.
For example, I am working on a CI system right now, as I ran out of Travis "open source credits" pretty fast and didn't care much for the alternatives (or Travis, for that matter, but it did work) and I figured there's some space for a somewhat different kind of CI. It's currently functional but unfinished at abut 1500 lines of code, and I'd be surprised if the finished version will be more than ~3000 lines.
It's smaller because it excludes features that are useful, but not for my (and I suspect many people's) use cases, and because it makes some assumptions that you roughly know what you're doing instead of abstracting everything. Combined, this drastically reduces the code size. This also means it's not useful for everyone, but I'm okay with that and it's a deliberate trade-off: not every project needs to solve all use cases or cater to all users.
So, it may be more reliable than you think by virtue of that selective counting of codebase lines.
I watched a video of a "cabin anyone can afford and build" recently. It's 120 sq ft. He isn't saying "Why isn't everyone building their house for $2000?". It's a project to show that things people consider too complex can actually be achievable from the ground-ish up.
It was a bit of an exaggeration yes, but it made my point about neglecting rarer failure cases very succinctly.
> Also, isn't splitting things into small services instead of large beasts full of hidden behavior the basis of the argument in favor of microservices?
With microservices your goal should always be simplicity rather than size. Simple code is often small, but small code is not always simple.
In fact, 3 well built microservices will usually have more lines of code between them than a single monolith that does the same job would as they have to have all the boiler plate for talking to each other and accepting RPC etc.
(Also converting a large system that's hard to reason about into microservices generally results in a large distributed system that's hard to reason about - system design is hard!)
A truly distributed KV-store would use a consensus algorithm to prevent the master server being a single point of failure.
For instance, CockroachDB uses Raft [0].
[0] https://www.infoq.com/presentations/cockroachdb-distributed-...
GFS/NFS predates many modern data stores, had SPOF, and are considered important contributions in distributed computing.
This looks really interesting, but I think I'd sell it as "using nginx as a key-value store" rather than a key value store in less than 1000 lines.
I realise you're being snarky but systems and verification schemas like that do exist. The terms I most often associate with them is "high-assurance" or "critical supply chain" software. There's bound to be many others.
These requirements crop up in stuff like the firmware for ATM keypad. Or Google's source code (they vendor everything in their own trees). Or Vegas slot machines.
In an amusing twist, software supply chain assurance is such a massive problem that large security consultancies offer code escrow services. If you, as a software seller, can't guarantee that you'll be around 20 years down the line, the buyer can require you to submit your code to such a third-party service. Should you go out of business, the source code will remain accessible to the buyer for their future development needs.
All that other code needs maintainers. Is it really that interesting to say "look how much I can punt to others?"
It's not as if anyone except the authors will use these micro libs, we might as well encourage metrics that produce interesting ones.
Excluding ansible/chef recipe to manage the machine, I could deploy Redis and have no lines of code to maintain.
FWIW, the app does import and run Go's HTTP server for "server" commands: https://github.com/geohot/minikeyvalue/blob/master/src/main....
SeaweedFS has filePath -> fileIds -> locateVolumeServers -> lookupOnVolumeServerByFileIds. These levels of indrection give the system more control of the data management and placement. For example, one file can be mapped to multiple file ids.
minikeyvalue has key -> locateVolumeServer(by nginx) -> lookupOnVolumeServerByKey. This removed the file id indirection level. This means the data is placed on statically organized volume servers. And the whole value needs to be on one volume server. So the value can not be too large. Obviously this large value is not a requirement.
There are no need to over-engineer anything. If this is what is needed, no need to add more levels of indirection.
And no need of being critical about this project. It is efficient and it works. It is a piece of code that the author can easily fix if any problem happens. The code and the language are just tools to get the job done. His goal is autonomous driving, not a general-purpose key value store.
If you have the time to go through the codebase, it was an interesting read. Nothing too surprising, but cool to see how it all comes together.
Okay.
This is basically a front-end with key balancing.