Rare things become common at scale (2014)
longform.asmartbear.com
longform.asmartbear.com
Not just the hard crash software/hardware edge cases either. Regulatory edge cases, abuse vector edge cases, localization and accessibility edge cases will all need to be dealt with.
I remember people exclaiming 'musk is right, Twitter doesn't need 1000s of developers' only for Twitter to start failing a lot of regulatory and abuse vector edge cases shortly after their layoffs. It's no good judging business needs based on a basic garage implementation of Twitter. The real world and a few billion users makes for an entirely different set of problems.
If you have 10 servers, a developer can't spend a couple of months to get a 0.1% performance boost--the benefit will never cover the cost. If you have 100,000 servers, it just might. If you have 1,000,000 servers, it almost certainly is worth an entire team looking for that much performance, year in and year out.
This is true of performance, avoiding downtime, or whatever else. You could never justify such an expense at a startup, but with sufficient scale they pay for themselves many times over.
And how would developers mitigate that? Isn't that the kind of thing you need human content reviewers to handle (assuming any automated tool Twitter had was still there after the layoffs)?
It's strange this is the lesson you took from Twitter. Twitter fired 90% of their developers, and almost nothing changed. Twitter is a site that exists in the real world, at scale. People don't even talk about the fact that Twitter fired 90% of their devs anymore, they just began firing devs themselves.
Twitter even provides mute, block, and whatnot functionality to prevent specified things from even showing up in your line of sight to begin with. And if the app is really bothering you, you can always set it down and go outside, take a walk, meet somebody new, do something that will put a smile on your face on your deathbed.
Lumping in mean comments online, with actual abuse is approaching risible. Words have meanings, we shouldn’t dilute or distort them.
By Twitter “not taking action,” sounds like your friend is upset that he or she can no longer co-opt the proprietors of the site into enacting punitive measures on people who draw his or her ire.
Maybe some mean things were said or whatever, but at the end of the day it’s just text on a screen isn’t it? And there’s a lot more to life than text on a screen, isn’t there?
It’s also weird how you mention the technical functioning of the site, then bring up the “Trust & Safety Org” when the legacy of “Trust & Safety” is a small cabal with extremist views arbitrarily deciding what information to censor and suppress based on their own viewpoints, whims, and influence from government agencies.
That has nothing to do with the technical functioning of the site which is a matter of reproducible, specifiable, determinate functions implemented in computer code to produce a useful product. The kind of thing that really turns the mind of an autist on.
P.S. Not to be too blasé about your friend, mean words can be an issue, especially an ongoing pattern, but anonymous strangers online seems like less of an issue than irl, and was this really an issue where block or mute wasn’t sufficient? How so?
X's systems for block and mute require the abuse to occur before you have an avenue to respond. Considering that all you need to get an X account is an email account, it's a pretty low bar for brigading. And that's to say nothing about organized campaigns to falsely report an account for abuse.
For individuals, I suppose you can make some kind of argument that those tools are sufficient, but if you're the poor social media manager for some township or minor government agency that draws the ire of the internet hate machine, you have to deal with all the abuse that goes with it. You are barred by the constitution from blocking people (and rightly so), and you have no real power to prevent them from creating sock puppet accounts to continue the abuse. PTSD is pretty common amongst (former, since they fired them all) twitter content moderators, because being consistently exposed to that stuff can eventually be pretty traumatizing.
CSAM, beheadings, videos of the worst things imaginable are what trust and safety deal with on a daily basis.
Lots of reliable/production software was built by teams much smaller than 1000's. Things like operating systems (e.g. the Linux kernel), compilers, tools or libraries, come to mind. Some even by single people.
That's not to say that all weekend garage projects are better than what 1000+ developer teams produce but don't knock the ability of small teams to do amazing things. Twitter isn't exactly the pinnacle of engineering accomplishment either. I'm pretty sure Twitter suffered from bloat similar to many other large tech companies. Elon taking a sledgehammer to that is probably not the best approach though. That said some people were saying the whole thing will fall apart in days, and it didn't.
[EDITed for typo]
Obviously manually interceding whenever you have a server failure is unsustainable when you have 1000 (or more servers). Why do people somehow believe that manually interceding to prevent bad action with surveillance is somehow sustainable, when previously the rate limiter was that the surveillance itself was manual?
There are a lot of things that are unintuitive at larger scales. AI and intellectual property is another.
And I have sympathy for people who said that algorithmically amplified speech on large platforms is qualitatively different from regular town square speech. It is. I don’t know how the laws should change in response but to pretend that Twitter is just like a bar is pretty naive imo.
A CS rep who “truly cares” is gonna set you back around 50K[0] in salary, call it 75K total cost to employ. I don’t know what their average customer value is, but seems they start doing phone support at 40USD, and 24x7 chat support for all customers. Let’s be generous and assume 50USD/customer average.
That means it takes at least 125 average customers to pay a CS rep’s salary.
Now bear in mind it’s 24x7, so you need at least 8 CS reps. That means you need to be retaining 1000 customers per year just to break even on your CS team. That’s around 20 a week.
The best customer support is the one your users never have to contact. If you can automate fixes (bonus points for automated pro-rated credits on bills based on downtime impact), or improve reliability, the customers with the highest value to you will stick with you.
That’s not to say “give a shitty support experience because it doesn’t matter”, it’s just that it’s a solution to a different problem than the one presented here.
You can probably get away with 1-2 good CS reps, provided the lower tiers have the right tools to triage things effectively. Put another way, if you have 1 CS rep that really cares, that does the followup calls, then you can employ 7-8 front-line CS reps at half the cost who just take diligent notes, and handle the common errors that customers make. I sincerely doubt that the person I have to call to help me connect my cable modem to my ISP makes even 40K a year. But all they have to do is follow a script, write down my answers to their questions, and escalate if the script calls for it.
I agree, completely, however, that ideally, every real issue happens precisely once and results in a change in the automated failure response.
That’s sort of my point… do you feel as though that rep really cares? Is that an outstanding customer support experience for you? Many folks would say ISPs and telcos have the worst customer service.
You’re not wrong though that a good script can help a lot of users with simple issues, and as it turns out, that’s the kind of thing that’s actually pretty easy and effective to automate.
The main reason bad customer service happens is because the customer has a problem that the CS rep is not empowered to fix or escalate the problem. Its not a question of "caring", its a question of results. Its really hard to care about your job when faced with unrealistic expectations from the customer, and insufficient resources from the employer. A lot of folks with customer service departments could save a bundle on labor and earn a lot of goodwill if they would recognize that.
… the actual reality is that companies employ 7-8 front-line CS reps that can't even read the ticket, and end up asking questions already asked & answered by the form I was required to fill out during the ticket!
I'm not buying upthreads' math either, though. Either the support plan is paid, or not. In the case of paid support plans. AFAICT given the support vs. the pricing for plans I've been a part of, a customer is just profitable. I don't get a full time, 40h/wk support agent, I only get them for the duration of time they're attending my tickets, and that proportional yearly salary cost is < the support plan's price in the cases I've seen; support plans are just ludicrously expensive, but so many companies feel they're obligated to have them, whether for DR, compliance, or whatever, that they don't question it. There are some companies that just do support as part of the included purchase (this is the right way, IMO), in which case, yeah, it takes a few customers. But in this case, how many customers/support rep is dependent on how much you can or are spending on operating costs.
> I agree, completely, however, that ideally, every real issue happens precisely once and results in a change in the automated failure response.
Today's customer support zeitgeist is utterly opposed to this idea. I agree that it's the correct response, but support's goal — that is, the metric which they are apparently measured by, in nearly every case I've seen is, is to close the ticket as fast as possible. This means not waiting around for permanent fixes: if a problem can be kludged around & the ticket closed faster, that wins. But then the extra due diligence of leaning on the engineering side becomes optional, extra … and just doesn't happen.
I've literally had tickets closed because "there hasn't been a response on this ticket for some time" — with the ticket plainly in the provider's court — and "we don't want to waste time your time with long running tickets" — nor a fix to my problem, clearly.
Customer Support is Goodhart's Law.
The latest change is that now I have to fight an LLM that cannot address my problem to get to that front-line rep who won't read my problem. Yay, progress! /s
First, I can't believe its taken me this long to start italicizing when I quote somebody. I'm going to do that for all future quotations.
Second, I can't easily imagine a better illustration of Goodhart's law than this. I can certainly see how they got to "average time to close", even though obviously the important metric is "how many customers eventually reached a satisfactory conclusion". Its just that its hard to answer that without bugging a bunch of people, reminding them of that time when the product shit the bed and they had to call support.
And "time to close" is not a bad metric, but it can't be the only metric. This reminds me of something that seems like a corollary of Goodhart's law: the easier it is to measure something, the less likely it is to be useful as a metric.
I don't even see any upsell to 'premium support' or anything like that for retail users. However, probably enterprise customers do indeed get separate pricing for support.
All told though, I'm not really saying that my maths was supposed to be on the money, just that the costs of running a CS team like the article suggests are high, and having an expensive CS team is only justified if they drive enough revenue, either through attracting new customers or retaining existing ones.
And to be clear, having outstanding CS is indeed a differentiator. But it's not a differentiator compared to reliability/automation, it's a differentiator compared to other companies' CS teams.
I think good customer support/relations has to be committed to for reasons that don't immediately show up on spreadsheets, with the trust that they will pay long-term dividends. I realize that is antithetical to The Way Things Work Now, but in both my personal and professional lives I try only to do business with companies that agree.
There are two ways to think about automating response to technical problems: a reactive way and a proactive way. Reactive automation looks to diagnose, repair, or work around system faults as they happen, within the constraints of the design of the system. The proactive approach happens early, at design or architecture time, in designing and building the system in such a way that it is resilient to rare failures and designed to be automatically fixed.
For example, think about a standard primary-failover database system with asynchronous replication. Reactive automation would monitor the primary, ensure it stayed down once it was deemed unhealthy, promote the secondary, provision a new secondary, and handle the small window of data loss. This works great for "fail stop" failures of the primary, and can often recover systems within seconds. Where it becomes much more difficult is when the primary is just slower than usual, or if there's packet loss, or if it's not clear whether it's going to come back up. Just the step of "make sure this primary doesn't come up thinking its still the primary" can be very tricky.
A more proactive approach would look at those hard problems, and try and prevent them by design. Could we design the replication protocol in such a way that the old primary is automatically fenced out? Could we balance load in a way that slower servers automatically get less traffic? Should we prefer synchronous replication or even active-active to avoid the hard question of what to do with lost data or slow primaries?
Looking through a certain lens, the answers to all those questions just add complexity to something simple. But, from another perspective, they take the hardest thing (weird at-scale failures) and turn them into something much simpler to handle. It's an explicit trade-off between system simplicity and operational simplicity. Folks without experience running actual systems tend to have poor intuition about this trade-off. The only way I've seen to help people build good intuition, and make the right decisions, is to get the same folks who are building large-scale systems deeply involved in the operations of those systems.
Race conditions just become conditions, because they will happen all the time. It has never been more important to know which syscalls are actually atomic and under what conditions that holds true.
From what I can tell of the last 20 years of corporate strategy, they just decide to not do these things instead of automating them.
To the extent where I am unaware of a customer service process that is not a Kafkaesque call-tree at this point
I will say though that I am a curious what they mean by fatal server error. I once worked at a well known company with about 400 production servers that were all running near capacity. I cannot remember a single serious hardware problem that truly killed a server (we did failover often to a backup but that was usually for software or upstream infra reasons rather than a hardware failure). I understand the scale is lower than in the article, but a server failure every day with a fleet of 2000 server feels like a lot to me.
At any rate, assuming they were just spitballing on a number, the point stands that you need to design and plan for failure even if it is rare. You really don't want the one time the server fails to be the time the CEO is demoing your product to your highest value customer.
OTOH it often works against you in the forward direction, e.g. tricking people into thinking we should screen everyone for everything all the time, because they don't appreciate how many incremental adverse events this will result in at scale (vs. the improvements in morbidity/mortality stats at scale)
Medicine is reactive. It is individually-focused. Something goes wrong, so we visit the doctor.
Public health is proactive but possibly based on bad predictions. It deals with populations instead of individuals. Concepts like "herd immunity" are from public health.
These approaches are oppositional and perhaps complementary.
https://www.ajpmonline.org/article/S0749-3797(11)00514-9/ful...
Put millions of cars on the road and instances of unlikely bitflips in memory will occur. The engine software apparently wasn't resilient enough to handle this.
E.g. a certain car part fails so many times out of a million. Okay, change something in the design. Hey look, now it's failing 1/3rd as many times, and it's a bit cheaper. Rinse, lather, ...