Rapid release at massive scale
code.facebook.com
code.facebook.com
I like that Facebook separated out deploy from release. It means that you can roll the release out relatively slowly, checking metrics as you go. Bad metrics mean blocking the release, which means turning off the feature via feature flag. I think for the rest of the world, that would mean halting the release and notifying the developer.
Disclosure: I work with smart people who spend lots of time thinking about this and writing blog posts like "Deploy != Release": https://blog.turbinelabs.io/deploy-not-equal-release-part-on...
Works like a charm, our product owner and engineering team are happier than ever, since it takes away a big part of the risks and makes rollbacks a lot faster.
It's simple. Never make a breaking release.
How hard can it be to not do something? Not very difficult.
It pays off in the end, but not worth making it a "criteria for success" when breaking out from branch-based to trunk-based continuous delivery, otherwise the trunking will most likely end up never happening.
shameless plug: at goeuro.com we shifted from branch-based to trunk-based CD in a short time (<3 months) with a diverse set of services and workloads, by applying a holistic socio-cultural, technical and process approach. Could be of interest if anyone is trying to make a switch: https://youtu.be/kLTqcM_FTCw
This is critical regardless in SoA / anything other than strict blue-green deployment.
In my experience of using Facebook (primarily groups), it is a highly buggy platform and it's very hard to say that it behaves consistently or even that features are really ready before release - this, IMO, implies that both development and testing as well as rolling out the releases are messed up. This post talks about building a better "conveyor belt", so to speak, to release changes, but if the basic product is buggy and didn't get good attention in design/dev/test, no improvements in the "conveyor belt" can help make it awesome.
Standard features that have existed for long may or may not work (how good are the regression tests then?). Posts and comments in groups may sometimes just disappear (thank goodness an admin activity log was added sometime in the last several months so we can stop wondering if an admin deleted anything). There's a feature in groups to mandate people wanting to join a group to answer some questions setup by the admins. Most people submit the answers but that never gets saved, and it's unknown what the trick is to get the answers to stay (this has been around for several months now?). New features aren't always announced.
I see Facebook as a platform that's used to share ephemeral things. So this level of quality is probably just ok (though I don't believe it justifies the company's revenues and valuation).
Since I do not conform to Facebook's ridiculous policy on using "real/authentic names", I don't even venture into contacting support if I see any issue, lest my presence be obliterated (yes, I try to keep away from Facebook, but do need it for some important awareness building because there's a large audience there).
As a platform used by billions, Facebook still has a very long way to go in being reliable.
And with so many people checking in to the master branch, how is it not permanently broken? With 1000 devs pushing code, you're bound to have severe bugs daily.
A lot of orgs don't have continuous deployment because of reasons such as:
- they don't have a good enough automated testing suite (or at least don't trust it fully), and thus rely on "sign offs" to have people commit to saying it's quality
- they don't measure in production properly (no real error alerts, no way to measure release success), and often deal with things in a "go or no-go" type way
- they don't canary test. To me this one is critical - the only way to get real production use is to have real production users actually using the site/platform/app, just a sample of them, to see what could go wrong, especially with new features
A lot of managers I've worked with are shocked whenever I pull out the "continuous deployment is easy. doing it well is hard" line.
I had the advantage of starting with fresh codebases and a small team. Obviously, adding this to an existing organization is much more difficult.
When you setup your system correctly from the start, it also becomes a great hiring tool. Once you show developers the environment you work in, mouths drop and they almost beg to work for you.
Doing it is hard, but not impossible. At this point there really is no excuse to not start things off like this. It is about 1-2 weeks of effort to setup a new project with all the right tooling on top of GCP thanks to the features they give you as part of their platform.
Completely agree that from a tech side there is no excuse, however a lot of QA culture has persisted through orgs and they want to keep that feeling of control (even though automation does it way better than them)
Facebook is not important. It has no impact when it's broken.
Started with 2 engineers (myself and another guy) and grew it to about 15. Zero QA, Zero DevOps.
We had CI/CD and a full test suite. We deployed from master as many times a day as we needed / wanted.
It can work if you open your mind to it and you hire the right people who know what they are doing.
Just because it releases quickly and has no QA doesn't mean it's a good thing.
The only metrics that matters is calls from your users. Facebook doesn't even have a number to call when it's broken.
Also: People tend to call the police: http://time.com/3071049/facebook-down-police/
Facebook is "7th most valuable company in the world" important.
FTFY
Releasing less often is a way to guarantee that bigger bugs will get through at some point, requiring hotfixes etc. The more you release, the higher quality releases you have, and the smaller production incidents.
The point isn't to remove QA, it's to trust that the automation in place is high quality and will catch the majority of issues before they are issues (i.e. if something makes it through the automation, then it should be caught in the internal release, or at least the 3% canary group), and then on the back of issues, make the automation more robust.
The more people who are introduced into a process, the more likely it is to fail at some point - the fact is that Facebook has a pretty low rate of huge production issues compared to most software companies, they must be doing something right.
Both of those metrics are incredibly shit ways to measure productivity.
I guess thats one explanation as to why the facebook app is 200mb+. They've been superproductive with all those lines of code.
A better metric would rely on actual features or bugs, imo.
Also, there is a difference between using them as metrics that you want to raise vs metrics you just don't want to drop or fluctuate wildly over time.
Why even bother measuring lines of code/number of commits.
Presumably your consulting business has really taken off now that you have such unique insight into a historically unempirical area of our field?
See also: https://users.cs.duke.edu/~ola/courses/cps108/code/java/harp...
Throughout my career imve found that number of changes comitted correlates pretty well with features delivered and bugs fixed. Also with feature and bug size.
Yes there are edge cases when it takes a day of debugging to make a 1-line fix, but those are rare. Just like it’s very rare to deliver a useful new feature by changing a single line.
Yes there are also features that are tracked as a real ticket and require a 1-line copy change and nothing else. Nobody thinks doing those makes you hella productive, it just needs to be done.
As for padding lines and changes. That’s what code review is for.
One should, but companies often care more about effort than results. They can manage based on effort, they can't manage based on results.
If you spend 2 days getting the same results as somebody else does in 5 days, guess what, they don't want you milling around those extra 3 days and bringing morale down. Gotta give you more work!
BUT, if a team of developers is not padding their code, nor has shitty programmers who write page after page of nonsense, then in the end the proportion of functionality to LOC should be fairly stable, and thus LOC is a useful and meaningful proxy metric for the amount of new or changed functionality in a codebase.
Its not difficult to track "velocity" via story points,features or bugs and is at least a step in the right direction.
LOC going down is often adding the most value.
People also take offense because a management weenie can misread the metric and penalize people who delete code, or write concise code, or solved a super tricky critical bug after a month of debugging with a one-line change.
But I still claim, these all normalize in the end. If a developer on my team has committed only 30 lines in three months, sorry but I'm gonna be skeptical of their actual productivity.
No, but you can implement a "story", i.e. the removal of code is presumably a ticket somewhere in the system. (Or could become one.)
> I've never been in a codebase that did more and more by getting smaller.
Run on a smaller system? Run faster?
(I'm not sure whether or not I agree with the parent's point, nor yours, just offering suggestions.)
If LOC is what you rely on, I'd be sceptical of many things, sorry not sorry.
I've just committed 3 lines that fix a bad intermittent race condition that took over a month to track down (partially because I can't test it and have to rely on other people who aren't 100% available). Is that unproductive?
On the other end, I'd fully expect someone doing that producing : - 3 lines of fix - A comprehensive documentation of what happened here, how to diagnose it and track it down, and how to avoid it in the future, (ideally if the company supports it with an oral explanation to concerned teammates).
In the end, I'd be very skeptical of one month of work letting 3 LOC as the only tangible output too.
The problem and diagnostic conversations are documented in the JIRA ticket; the solution is documented in another ticket which "fixes" the first one (for abstruse internal reasons); and more commentary is provided in both the comments around my fix and my verbose commit message.
(Also there's a whole bunch of angry Slack ranting about this issue but that probably doesn't count as "documentation".)
I suppose it depends what people mean by LOC. I guess my 3 LOC isn't actually just those 3 LOC. Definitely an interesting thing to think about at least. Ta!
LOC or commits are certainly bad ways of comparing two different people, teams, etc. But they are fine ways of measuring whether a specific event at a specific time changed productivity, assuming a decently large team etc.
If the build system is super slow and just breaks all the time, your going to see a drop in the number of commits landed per day. If git is completely down and commits per day drop per 0, then you have an outage ;)
This is 'productivity' in the sense of the release eng team attempting to measure the volume of code changes they're pushing out and whether changes they made to the push/release process impacted other teams' ability to iterate. Unless you have some insight as to why adopting a more continuous deployment process would change the ratio between LOC pushed or code-changing-diffs pushed to features-or-bugs specifically at Facebook such that they would not function as acceptable proxies, questioning the metric chosen seems hollow. It's certainly not 'productivity' in the sense of whether people on product teams are considered to have done a good job or are deserving of bonuses / promotions.
In theory yes. In practice, when you know your team and their average output and quality, no.
>I guess thats one explanation as to why the facebook app is 200mb+. They've been superproductive with all those lines of code
Or you know, most of it is assets, debugging symbols, stray frameworks and libs not stripped, etc. As is the case with almost any app beyond a few MB.
Measuring programming progress by lines of code is like measuring aircraft building progress by weight.
That said, one of my favorite metrics is number of deleted LOC.
Call me an old fart, but if you are in a situation where you need to make that many changes per day, you are utterly fucked from almost every angle.
Every aspect of this article sounds to me like people have no idea what they are trying to do, so they write code and push it, and it goes live. And everyone is very happy about this, for some insane reason.
No offense to anyone, but this is not a reality I want to live in. And the article doesn't do much to defend the notion.
- They are doing a lot of fixes. The number of bugs found is proportional to the number of users you have and number of changes you are making.
Across their whole product line (client apps, Hip Hop VMs, Flow, whatever), there might be millions of lines of code.
But what exactly seems to actually change year over year on FB itself (server/client of the actual social website) that warrants so many commits?
FB does this stuff all the time.
Maybe you need to slow down.