Done correctly, Ceph is extremely reliable, resilient, and fast. Once you get over the initial learning curve, dare I say, even a joy to work with.
1,355 karma · joined December 15, 2015
Done correctly, Ceph is extremely reliable, resilient, and fast. Once you get over the initial learning curve, dare I say, even a joy to work with.
ZRAM is a compressed block device that is stored in RAM. It's great!
Previously, if I ever had high memory pressure situations, I really dreaded the slowdowns. Now, with swap sitting on top of /dev/zram0 it's a completely different experience.
I have ZRAM enabled on all of my personal machines, both laptops with limited memory, and desktops with 64 or 128GB of RAM. It's rarely used, but it is nice to have that extra room sometimes.
The performance of a zram device is so much faster than even the latest NVMe drives.
I run Ceph at work. We have some clusters spanning 20 racks in a network fabric that has over 100 racks.
In a typical Leaf-Spine network architecture, you can easily have sub 100 microsecond network latency which would translate to sub millisecond Ceph latencies.
We have one site that is Leaf-Spine-SuperSpine, and the difference in network latency is barely measurable between machines in the same network pod and between different network pods.
They've been active in the Ceph community for a long time.
I don't know any specifics, but I'm pretty sure their Ceph installation is pretty big and used to support critical data.
I'm genuinely curious.
A modern furnace works via a heat exchanger, where the combustion produced pollutants never mix with the indoor air being pushed through. All pollutants are expelled outside via a property functioning chimney. This is one reason why you should have the furnace (and chimney function) inspected annually. Aging heat exchangers will show hotspots before there is a possibility of air being mixed, giving plenty of time to plan for a replacement. Of course there is a possibility of failure, which is why you should have a carbon monoxide detector.
I recently got a sewing machine for an unrelated project and around the same time I ordered it I had one of these cloth reusable bags rip, because I put too many heavy things in it. When I got the sewing machine, for practice I decided to see if I could fix the bag. It turned out to be surprisingly quick and easy. I didn't use any extra material besides the thread, and I believe the bag is much stronger now.
With upmap and balancer it is very easy to run a Ceph cluster where every single node/disk is within 1-1.5% of the average raw utilization of the cluster. Yes, you need room for failures, but on a large cluster it doesn't require much.
80% is definitely achievable, 85% should be as well on larger clusters.
Also re scale, depending on how small we're talking of course, but I'd rather have a small Ceph cluster with 5-10 tiny nodes than a single Linux server with LVM if I care about uptime. It makes scheduled maintenances much easier, also a disk failure on a regular server means RAID group (or ZFS/btrfs?) rebuild. With Ceph, even at fairly modest scale you can have very fast recovery times.
Source, I've been running production workloads on Ceph at fortune-50 companies for more than a decade, and yes I'm biased towards Ceph.
I only want to add a small suggestion. I get that large distributed production systems will occasionally go down, but it would be great if you could look into reducing the latency of your status page.
By my count there was at least a 35 minute delay between when things broke and before the status page (https://fastmailstatus.com) was updated.
Also, I think it would have been nice to have a bit more explanation on this event than simply "database issues" [1]. Being able to know that this was related to an upgrade would have made me feel a bit better during the time the status page was updated and until the issue was resolved.
Thank you for your hard work and an excellent email service!
-A long time customer.
It started out as not being able to search, but the situation is quickly deteriorating and now I'm unable to open pretty much any email message.
Some content seems to briefly show up and then it quickly disappears and after that, it's as if cache has been invalidated and you can't get back into it.
Laptop: LG Gram 16"
There are several variations of the laptop with spec differences, the main thing you want is Intel CPU with integrated graphics.
Great screen (16:10 ratio), great battery life (80Wh), dual NVMe slots (if you care about bit rot).
Last, but not least, very, very light.
You can also tell Ceph to use a single disk as your failure domain. No one does that either. Homelabbers maybe, but then why are you comparing such setups with Google?
We run Ceph with a failure domain of an entire rack. We can literally take down (scheduled or unscheduled) an entire rack of 40 servers, and continue to serve critical, latency sensitive applications, with no noticeable performance loss.
We have a Ceph footprint 5x larger than CERN run by a team of 4-5 people.
I ended up returning all of the expensive stuff and only keeping a bunch of really cheap Amcrest PoE cameras.
Here is one such example: https://www.microcenter.com/product/634071/amcrest-5mp-ultra...
I've been very impressed with Amcrest cameras.
They support being configured without an outside internet connection.
They support dual streams.
They have all kinds of tuning settings, and come with sane defaults.
They support H.265 encoding, so you get good quality at small file (and bandwidth) sizes.
They also have MicroSD slots, and some cloud stuff that I don't use.
They work great with Fridate and Google Coral with is what I use them with.
Highly, highly satisfied and recommend.
It sometimes helps with defragmenting memory, memory pressure in general, and depending on your workload and other things running on the box, you could have better performance by being able to have more things in page cache.
My home cluster is running on some 7 Raspberry Pis currently, so it's not very performant, but the uptime over the past 2-3 years has been unbeatable.
With cephadm it's very easy to stand up, and upkeep has been basically zero.
If you have your own hardware, Ceph is solid.
I wouldn't put more than 200 million objects in a single bucket, but other than that it can be very reliable.
I have ~ half an exabyte on-prem, and sleep very well at night.
The 3x20TB btrfs setup I mentioned is configured with RAID1C3 for metadata, and RAID1 for data, and works just fine with even or odd number of drives.
It's funny how people assumed RAID5 when they saw 3 drives.
I switched to this after years of running on ZFS, and for my workloads btrfs is faster on Linux (not to mention the licensing/packaging mess).
Hearing the drives has been really nice actually, and got me noticing all kinds of interesting and sometimes unexpected behavior going on with my system, and actually helped find a bug with my terminal multiplexer.
With 64GB of RAM my entire home directory fits, so only writes go to the drives, and it's been surprisingly performant for my workloads.
However, I feel like your criticism re fzf is not really fair, because I run into this with other tools quite often. So often in fact that it took me only a second to convert your command in my head to this:
git checkout @($(git branch |fzf).strip()) ls -U1 |wc -l
Can be very fast with lots of files on a modern Linux box.My wife, a licensed teacher in our state, was switching jobs from one school district to another, and watching her go through that process was interesting...
The interview process for any kind of school administrative position consisted of multiple rounds of interviews.
The interview process for a teaching position consisted of them confirming that my wife had a pulse.
https://forums.ivanti.com/s/article/Recovery-Steps-Related-t...
(Linked from TFA)
"The reason we haven't been hacked is because we still run 16‐bit POS systems, and today's hackers don't know how to fix anything is such a small address space."