I'm not sure what reality this design comes from, but it's not one I care to inhabit.
I'm not sure what reality this design comes from, but it's not one I care to inhabit.
The normal flags you're looking for are: noout, noscrub, nodeep-scrub. Do your maintenance. Unset those flags.
The noout is what tells Ceph to essentially not shuffle data when OSDs go down. This is also referenced in the article.
Think of the case where you have more than the three servers describe in this incident. You really don't want to disable cluster-wide automated repairs because of some other kind of maintenance. Something else WILL happen during your scheduled work and... ooops.
Why 600 or 900? Unavailability events in larger clusters mostly tend to last only 10-15 minutes, for routine things such as regular node reboots, per page 2 of https://static.googleusercontent.com/media/research.google.c...
The reason for "noout" is you generally want to minimize IO operations while you're performing your maintenance. You should be closely monitoring your cluster while performing maintenance, and if something else goes wrong, you abort and unset noout, wait for it to finish rebuilding, and reassess.
EDIT: This used to be 300 in Jewel, and seems to have changed to 600 in a later version.
I'd still argue that you want to provision N+2 capacity, so you can withstand a planned event and an unplanned one at the same time, without having to go through manual tweaking. I'd leave that for truly exceptional cases, such as when the whole cluster is hit by some nasty bug or has entered a spiral that calls for drastic measures.
In an ideal world, we would have had multiple people to dedicate to building the kind of cluster they wanted to build. In the real world, after a dozen people left, I'm afraid they may well have ended up in a situation not much better than the University in question.
http://docs.ceph.com/docs/master/rados/troubleshooting/troub...
My memory is generally quite poor, but I vaguely recall this feature being present for as long as I've been familiar with Ceph. I obviously have no way of knowing what happened in that meeting you were in, and maybe the other people who were proposing using Ceph were not very familiar with performing maintenance on it, or maybe there was other constraint or use case they had in mind, but I'm quite confident in saying that the normal case for Ceph maintenance involves only a very marginal amount of data movement (to bring the temporarily-down OSD back up-to-date with changes that occurred while it was down).
Assuming that you run with 3 failure domains, and only maintain one failure domain at a time. Noout mostly gets the job done. What it doesn't do for you is save you from an actual failure in a different failure domain during maintenance. EC pools & k+(m>=2) or replication > 3 - would cover this as well.
We've had mostly great success with noout + maintain failure domain at a time, wait for recovery, proceed to next failure domain, repeat until done. To the point where we've been comfortable leaving a lot of the babysitting & work to machines.