Ways to shoot yourself in the foot with Redis
philbooth.me
philbooth.me
- Connection pooling / pipelining and circuit breaking is a must at scale. The clients are a lot better than they used to be but it's important developers understand the behavior of the client library they are using. Someone suggested using Envoy as sidecar proxy, I personally wouldn't after our experience with it with redis but it's an easy option. - Avoid changing the cluster topology if the CPU load is over 40%. This is primarily in case of unplanned failures during a change. - If something goes wrong shed load application side as quick as possible because Redis won't recover if it's being hammered. You'll need to either have feature flags of be able to scale down your application. - Having replicas won't protect you from data loss so don't treat it as a source of truth. Also, don't rely on consistency in clustered mode. - Remember Redis is single threaded so an 8xl isn't going to be super useful with all those unused cores.
Things we have alarms on by default: - Engine utilization - Anomalies in replication lag - Network throughput (relative to throughput of the underlying EC2 instance) - Bytes used for cache - Swap usage (this is the oh shit alarm)
Someone on the team decided to use Redlock to guard a section of code which accessed a third-party API. The code was racy when accessed from several concurrently running app instances, so access to it had to be serialized. A property of distributed locking is that it has timeouts (based on Redis' TTL if I remember correctly) - other instances will assume the lock is released after N seconds, to make sure an app instance which died does not leave the lock in the acquired state forever. So one day responses from the third party API started taking more time than Redlock's timeout. Other app instances were assuming the lock was released and basically started accessing the API simultaneously without any synchronization. Data corruption ensued.
The classic solution is leases: assume bounded clock drift, and make lock holders promise to stop work some time after taking the lock. This is only correct if all clients play by the rules, and your clock drift hypothesis is right.
The other solution is to validate that the lock holder hasn't changed on every call. For example, with a lock generation epoch number. This needs to be enforced by the callee, or by a middle layer, which might seem like you've just pushed the fault tolerance problem to somebody else. In practice, pushing it to somebody else, like a DB is super useful!
Finally, you can change call semantics to offer idempotency (or other race-safe semantics). Nice if you can get it.
i disagree here. instance b did the right thing given the information it has. instance a should realize it no longer owns the lock and stop proceeding. but in reality it also signifies concurrency based limitations in the api itself (no ability to perform a do and then commit call). https://microservices.io/patterns/data/saga.html
You're right that A could try to stop, but I think it's more complicated than that. A is calling a third party API, which may not have a way to cancel an in-flight request. If A can't cancel, then A should refresh its claim on the lock. A must have done neither in the example.
The second is to use a redis client that has its own thread - your application blocking on a third party API response shouldn't prevent you from updating/reacquiring the lock. You want a short timeout on the lock for liveness but a longer maximum lock acquire time so that if it takes several periods to complete a task you still can.
The third is to not use APIs without idempotency. :)
But now you have a released lock and a client that thinks they have the lock.
This has nothing to do with the redis server. This is bad application code monopolizing a single connection waiting for an unrelated operation. A stateless request / response to interact with redis for the individual operations does not hold any such locks.
Well, yes. That is why the preceding sentence, which you didn't quote, said "poorly-implemented application logic". So thanks for agreeing with my post, I guess.
The point, in case you missed it, was to advertise ways I'd fucked up and hopefully help others not to fuck up the same way in future. It was never my intention to say Redis was the problem and I'm sorry if it made you think that.
https://redis.io/docs/reference/clients/#maximum-concurrent-...
With MTLS, a good security posture is to log every connection establishment, with basic metadata about the certificate involved - it's SAN and public key hash are the best bet. For troubleshooting, do that logging before the authentication decision. But anyone can make their own certificate, so keeping network controls keeps that list free of clutter.
I would love to see some numbers on this. My intuition says there are probably some workloads where JSON strings are better and some where one key per property is better.
thankfully it was on a staging env, I think he's at google now.
If a junior dev can cause catastrophic harm from one wrong command, it's the org's fault for not having safeguards in place, not the dev's fault for an (understandable) error.
incompetent is maybe a bit harsh, but i did say he was junior, and junior devs make mistakes, and this guy was well meaning and messed up. you don't get from junior to senior or principal or staff without some mistakes, and it's the responsibility of the more senior devs to not have them in a position where their mistakes are catastrophic.
# To disable:
rename-command FLUSHALL ""
# To rename:
rename-command FLUSHALL DANGER_WILL_ROBINSON_FLUSH_ALLSince then I treat any prod server terminal like I’m entering launch codes for a middle system.
Anything outside of ls or cd I’m very careful, read the command a couple times before executing, etc.
We tend to go ahead and either use a runbook, or whatever experience we might have, to setup a pretty detailed plan of what to run on which systems with which purpose. You can then throw these plans at someone else to review. Sure, it takes an hour or two more to setup a solid plan and waiting for a review takes time as well. But this has turned into a great tool to build up experience in weird parts of the infrastructure.
one of my frustrations of bigco software is people taking basically maintenance roles where the computer tells them what to do, because the lava flow legacy code base is too scary to touch.
however, you can automate your daily clean up tasks. it's certainly shellacking more mud on the ball, but if you're not going to even try scripting your repetitive tasks, then i don't know why you're a programmer.
Takes a lot of effort to stop people opening up a shell in prod or grabbing a prod DB dump or even just connecting to the prod datastore directly from their local env.
That's the kind of mistake you only make once!
I've noticed this particularly with Ruby where the official gem has cluster and sentinel support, but many other gems that depend on Redis expose their own abstraction for configuring it and it isn't compatible with the official package.
Of course, I think that running Redis in clustered mode is actually just another way to shoot yourself in the foot, especially if a standalone instance isn't causing you any trouble, as you can easily run into problems with resharding or poorly distributing the keyspace. Maybe just try out Sentinal for HA and failover support if you want some resilience.
Better than nothing though.
We do expire but we don’t think we have a thundering herd problem with them all happening at the same time.
What? Was this inside a MULTI (transaction) or something? This isn't a flaw of Redis being single-threaded. Honestly all of these "footguns" sound like amateur programmer mistakes and have zero to do with Redis.
> If you're particularly naive, like I was on one occasion, you'll exacerbate these failures with some poorly-implemented application logic.
Then a few paragraphs above that is this sentence:
> The gotchas that follow were all occasions when I didn't use it correctly.
I'm not sure how to make it more clear that I'm criticising myself, not Redis, in the post, but that's the intention. If you have suggestions how I could make it more obvious, please let me know.
> I'm not sure how to make it more clear that I'm criticising myself
"Mistakes I made while building applications on Redis"
PS Love the aesthetic of your blog!
(I also wonder if "shoot yourself in the foot" is an idiom that doesn't translate well; at least in the UK I think it's fairly well understood to put blame squarely on the person doing the shooting, rather than the firearm they happen to be holding at the time)
Unless you're using AOF mode with fsync always, you can lose writes. If you're doing that, you should be using a real database instead.
> The open source, in-memory data store
Finally, check out Redis-compatible alternatives that don't require the data set to fit in RAM. [0]