How to Deploy Software
zachholman.com
zachholman.com
To start with get something to monitor errors/exceptions and email you. To name a few services:
https://github.com/errbit/errbit (can be hosted on Heroku for free)
Also make sure that you have accessible logs that log useful information (timestamps, the user making the request, unique request ID). Then use syslog or a SaaS service to aggregate logs from all servers in one place, and keep them for as long as you can.
I use a combinations of Zabbix and a very agressive Smokeping mesh deployed with docker (10 pings per minute from all separate DCs all other DCs) to monitor our worldwide nodes and be able to account for routing problems on the backbone rather than our nodes themselves. Very handy tool, particularly if you work with latency sensitive applications. I've been surprised some of the datacenters we use appear not to have access to similar...
We do use ES as a core part of our infrastructure, so our only real barrier was setting up Heka.
I haven't had much of a need to look at logs since I started using Airbrake three years ago. For the price, it's really hard to beat the value when you compare the price to various log aggregation systems.
It has been rock solid for us with no event limits and cheap pricing.
(not an employee, just a fan)
Logs are only one part of monitoring though, and it's easy to miss the wood for the trees if you're drinking from the logs firehose. Some metrics monitoring is also vital, naturally I recommend http://prometheus.io as I'm one of the core developers :) https://blog.raintank.io/logs-and-metrics-and-graphs-oh-my/ discusses this a bit more.
I have been using https://www.pingdom.com for uptime monitoring and some basic perf/latency tracking. It works well and doesn't cost a lot, plus it can hit my services from multiple regions.
That's a sure way to go to jail.
Don't store which IP requested what for more than necessary — in many countries, there are limits if a few months on how long you may store metadata.
If it's true, you'd think it would be widely-known and publicized in those countries (whichever ones they are) but I can't find any reference to such a policy existing anywhere.
That's the reason Google Analytics, Piwik and others have an option to blank out parts of the IP address.
Jail is a bit dramatic though (at least here), unless you do something bad with them fines + maybe damages would be the maximum, and it is quite unlikely you'll actually get hit. But if you are ever in a situation where this would come to light (= already in legal trouble), it could make things a lot more complicated, especially if you process data for other companies.
There is currently a court case as to whether dynamic IP addresses from ISPs counts as personal information. But static IP address probably count. And all the other information (user X logged in at time Y and send message to other user Z) would be PI too.
Not in (any of) the United States.
Personally, I don't see much of a problem when my IP is stored somewhere since you'd need a reference, but others have apparently thought more about it and want it to be anonymized.
(I've now come up with scenario: If you have access to multiple such sets of data, you could connect accounts from different systems. I. e. when the youporn logs leak, someone could tie the data to your HN account etc.)
Combine that with leaks of a social network where you have real names and IP addresses, and you could actually identify the name, job, address, etc of everyone who watched a specific clip of porn.
That’s very bad.
Metadata. Not even once.
Don't make any plans to have a physical EU footprint and you can care less what the EU thinks of your internal policies. EU customers can take it or leave it. Their choice. It's not a warm stance to take, but you've probably got a dozen other balls in the air that you're juggling. EU data compliance probably isn't one your worried about as a small startup.
(especially european users, and businesses)
We'd love to get some feedback on our product when we launch to beta, so if you're interested, please sign up for our email list at http://logdebug.com
Though the pedantic butt-head in me wants to point out that out of all of our paying customers, they're all using that level of service. :) So technically...
One thing about the plans that made me cringe was that "SSL" encryption was not available to the Startup plan. I feel like we're entering an era where HTTPS isn't really something that should be offered as a perk, ya know?
Careful, you can run into Data Protection legal issues with that approach.
It's a great way to think of it.,
Three areas that I think would have been worth including:
1. Pre-production.
You deploy to test & other pre-prod environments more often than prod. They should use the same scripts/tools/processes as production deployments, only with different permissions.
2. Configuration.
Test and production environments will always have different config settings, so no team will ever be able to deploy to more than one environment without encountering this problem. I think there's still an open question around whether those configuration settings should live in the same source control as the code, in a different source control repository, or a dedicated system. Source control systems and sensitive values (passwords, API keys, etc.) don't always mix.
3. Build your binaries once.
The article is more focussed on dynamic languages, but for compiled languages, I think this is important. If you branch, compile, deploy to test, test it, get the all clear, then compile again and deploy to production, there's a lot of opportunity for differences between what you tested and what goes to production to sneak in.
In fact even for dynamic languages, this might be a valuable practice. What if the JS minifier on one build server is different to another, and the deployed script ends up being different in production to what was tested.
Disclaimer: I'm the founder of Octopus Deploy, and these practices might be biased towards enterprisey .NET/on-premises deployments rather than cloud hosted, dynamic language projects.
Assuming your builds are solidly reproducible like they should be, how do such differences "sneak in"?
Granted, truly reproducible builds in the first place are Really Hard(tm).
https://octopus.com/blog/build-your-binaries-once
A real world example of this is that when .NET 4.5 shipped, the compilers in .NET 4.5 for 4.0 code produced different output than the .NET 4.0 compilers would have. So installing a system-level update on Windows on a build server would mean Test and Production got different results:
http://blog.marcgravell.com/2012/09/iterator-blocks-missing-...
Also, releases move through environments at different times:
Monday: 1.0 goes to Test
Tuesday: 1.1 goes to Test
Wednesday: 1.0 goes to Prod
Between 1.0 and 1.1, perhaps you updated Node, or went to Python3 - so your build server had to be updated. Reproducing Monday's build is going to be more difficult the more time that goes by.
As you said, truly reproducible builds are hard. Why not just zip the artifacts and use the same files, instead of rebuilding?
Pre-production should test production in a non-production environment. Specifically, if a system connects to, say, SalesForce, and you've been using a local sandboxed instance of SalesForce in dev and UAT, the pre-prod should connect to the real, live SalesForce. I have a pre-prod environment for my current project, and it's useless because org policy states that only prod environments can connect to live instances of anything. Completely invalidates the reason for having pre-prod.
2. Configuration ...will always have different config settings...
Yes, aside from pre-prod. Pre-prod and prod MUST be mirror images, INCLUDING config settings. Identical. The same in every way except for host names.
It interests me in how many combinations of pre-prod environments exist, as I (naively) expected this to be standardized, but it's not (e.g. https://en.wikipedia.org/wiki/Deployment_environment#Environ... ). I can see the need for two vs four environments for handling different stages of the life-cycle depending on business needs (and budget). Addressing how to choose the right set of environments for a given company/product would be useful.
I also want to say that your product is phenomenal and it has improved the quality of our shop's build and deploy pipeline by an order of magnitude. Thank you.
As you might expect, it wasn't really the fault of feature flags but rather indicative of a process problem. We'd either have to prod the team or person in charge of it, verify that work was actively pushing things forward to ship, and so on.
At various times we'd have, for lack of a better phrase, the "No (wo)man", who would come in and say "no" to a lot of things. One of the best ways to achieve this is to see that nothing has happened on a feature flag for awhile and then send a pull request to remove the feature flag (and the feature entirely). This got people out of the woodwork who said waiiittttt a minute let me just finish that up, or if no one vehemently disagreed with the pull then you could actually just remove it entirely. But it did take some explicit reflection on whether the flags were defensible to remain flags.
Why do you have so many feature flags at a single time? Do you have a team of 20-40 people or more? Are your features taking months to create? Are you trying to create a feature flag for every new line of code before deploying it?
A lot of times code can be deployed without using a feature flag, because it's small enough, maybe you are overlooking those cases?
>Also please re-read my post carefully. The details are there
Clearly there is something wrong with your project that you aren't mentioning. Plenty of other projects do flags without problems. Your team has issues. * Always be developing against the current running features. No brainer.
* Design things so that they integrate feature flags, not work around them. This usually means pushing feature flag determination to more generic/common code.
* Separate backend/frontend changes into separate feature flags when possible. Turn on backend changes early and often to better measure your feature's impact.
* Give individual features their own flag, but also have a global flag that manages the entire experience. This makes it easier to manage your gradual dial up as well as shut off problematic features that would otherwise mess up the launch.
* Be diligent about removing feature flags once they're turned on. Schedule it into sprint time, reward teams that remove them, make it a management mandate, whatever. Just get rid of them once they're no longer needed.
* Invest in monitoring around your services that (ideally) can correlate failures with features. you should turn on features over the course of a few hours/days to mitigate customer impact in the event of failures and gain data about performance at 50/50.
I think the answer to your specific question of "testing every combination" is that you can't, easily. But by keeping the number of feature flags that are inactive low (< 150 is very liberal) for a given service, having everyone develop against the current running features + dev overrides, and using gradual dial up with integrated monitoring to catch poor interactions when the impact is small, you'll have mitigated a lot of your concerns.Flags were there because of two main reasons:
* Our app was deployed more frequently than a service we depended on (e.g. every commit for app, nightly for service in staging, every 2 weeks for service in prod, etc)
* We wanted to have it rolled out to a subset of users for testing
We did not run into exponential trees of combinations because we rarely had two different flags interacting. Maybe it was a happy accident of trying to make work parallelizable or maybe because our feature flags never lasted more than a month or so.
The code was intentionally dumb and the flags were stored directly in the source code (not in a database table or another config file or something). Simple, stupid calls to `FeatureFlipper.isFooReportEnabled()`. We did not test this class because each method was a simple boolean check that the current user appeared in a list or ENV != prod.
We stubbed out `FeatureFlipper` when using it throughout the app. Stub the feature to be enabled, check behavior. Repeat for disabled.
Most of the features were simply hidden at the view level. For users, if there is no button, it doesnt exist. We didn't particularly care if the user would "guess" the url of a feature flagged page -- not worth the effort.
Doing a "full rollout" of the feature was a non-event. Just delete the method on FeatureFlipper and go fix the compile errors :)
The only thing you have to remember is: Its either going to production, or its getting dropped.
There is an expiry date built into the feature flags framework (90 days), and either they need to be explicitly cleaned up and deprecated, or extended with a valid reason attached.
Also, we code review every commit. If any diff is coding around/through a feature flag, the first question is always: "Can you clean up that experiment first?"
Great writing. Spaces all the way.
Spaces are fixed. Tabs are adjustable via most user interfaces.
So, in theory, you'd think tabs are superior because everyone can have their own amount of spacing and that's that.
Therefore tabs > spaces
(I should probably leave now)
A simple solution is to use tabs, because then every developer can set the tab distance however they like, and they are happy. It breaks down when you have code that is aligned beyond the indent, like this:
if(a && b &&
c && d &&
e && f ) {
In that case, increasing (or decreasing) the tab distance will ruin the intention.The best solution is to use tabs to the point of indentation, and then spaces thereafter, but a lot of code editors don't support that, so in practice it's hard to implement, so people use spaces to preserve their formatting when it gets uploaded to github.
Python and spaces for me.
> Python and spaces for me.
Sounds like you are a very open-minded individual. Not.Open a file where it is all spaces, and it looks the same on every machine. Open a file with a mix of spaces and tabs, and it often turns out an absolute mess.
Longer version: First understand that there is no tabs vs. space. There is only tabs + spaces vs. only spaces. (Because not all indentation line up with tab spaces, and someone may wish to line up assignments, lists, etc.) Only one developer in the history of your project that uses another tab stop standard is then enough to mess up your indentation. And that's just one way to mess it up, in a sufficiently large project some creative developer will find another.
Spaces just work and ensures your guidelines are followed. The only downside is a few wasted bytes. It might have been a religious debate in a long distant past when someone actually counted bytes, but today it's mostly young developers who don't know better who engage in it (with a few exceptions). Linus uses spaces. OpenSSL changed to spaces as part of cleaning up their codebase. It's the default behaviour of GNU indent. It is a good idea.
Tabs width can be set to whatever you want, so everyone can use the spacing they prefer.
And thank goodness Microsoft hard-coded in a 10 hour delay to ensure the KDS root key of a domain is propagated, even if I create it before I create additional domain controllers...
Such fun!
Add-KdsRootKey –EffectiveTime ((get-date).addhours(-10))
See https://technet.microsoft.com/en-us/library/jj128430.aspxIf we are trying to deploy migration #1, first deploy a version B of the code that supports (but provides the same set of features as before) the db both after migration #1 and before migration #1 (maybe with the help of a feature flag that is set in the migration). Do the migration. The deploy a version C of the code which removed the feature check above. But all this requires 2 different versions of the code and a lot of process just to ship out one migration. It gets combinatorically worse if you have more than one migration to deploy.
Zero (no crash) downtime deployments seems like too much effort for too little gain.
That and some other methods are described in a post on the Yelp engineering blog[1].
1: http://engineeringblog.yelp.com/2015/04/true-zero-downtime-h...
If you can, always make migrations backwards compatible with the previous version of the code, so they don't need to be rolled back if the code needs to be rolled back. Having a good migration rollback procedure is nice too, but usually unnecessary if testing has gone well.
If you need to add a new field to the model then always create it as nullable, whether or not you use a default value. That'll allow this particular database migration to run on previous versions of the code. You can test this by generating the migration SQL and running it on the database for your test/dev environment which is running your current production code.
Once the new code and migration is running in production and you're satisfied with it, immediately create a new change which makes the previous migration more strict (remove nullable, use a data migration or default value for existing nulls). Test, stage, deploy that.
Having backwards compatible migrations leads into zero downtime deploys also. Once the data migration has run, your app servers can be running version X or X-1. Before you push new code to a node, do a graceful shutdown (allowing queued requests to complete) or remove the node from the cluster (haproxy socket api for example), update the code, bring the node back into the cluster.
Caveats:
- not all database migrations can be backward compatible, but most can be made so by breaking your change into two or more changes. First - make the change without strict integrity checks. Second, enforce integrity checks and provide defaults.
- zero downtime deploys requires multiple nodes (counter example would be welcome, but I can't think of one).
I would have agreed with you before working at GitHub, but the end result is that deploying is so easy that 3 deploys does not feel like "a lot of process." On most of our applications I can do this in 15 minutes or less.
As far as database migrations are concerned, GitHub (and others, of course) takes the perspective of migrating before the code that uses those migrations go out. In other words, as @herge says in a sibling comment, the code that gets pushed needs to support two branches of code and two branches of data simultaneously. It's certainly some extra work (and can be pretty gnarly depending on the scope of the migration), but once you get to a certain point it's kind of the only way to do no-downtime migrations.
There's many possibilities to help with the actual migration process, depending on what database you're using. With MySQL, for example, you can do something like the process in lhm: https://github.com/soundcloud/lhm
Zero-downtime deploys aren't super difficult in Rubyland anymore (many have written how they achieve it in Unicorn, for example), although I'm not as familiar with how other folk do it across other languages and platforms these days.
...well, once someone says it out loud. :)
After that, of course it makes sense that they would have a need for an API compatibility window across at least two versions. It's exactly the same issues as supporting a backend and client side where you can't instantaneously force an update to all clients. With your DB, you're in control of the version, but you're certainly not in control of making it "instantaneously", so the same rules apply as when you're waiting for some curmudgeon user to update: API versioning and a support window.
Now, if only we had an automagic way to make it less painful for a project to support two different versions of an API from the same codebase....
[1] http://www.se-radio.net/2012/06/episode-186-martin-fowler-an...
It's a controversial article though.
Sofar I'm the only one regularly doing it without feature branches. Running with the idea that just because you can branch cheap doesn't mean you should. Of course toggles are technical debt to be managed, but so are branches.
I've found good practice has been mentioning which toggles are available on the README with their defaults (could be generated)... they should be tracked and removed ASAP. I read a newer article that breaks them down into categories http://www.infoq.com/news/2016/02/featuretoggles. Toggles over branches are showing value as we run different variants on staging without having to redeploy different builds, but instead changing a launch variable. It's especially clear with 2+ WIP features. We're using environ with Clojure which doesn't have any fancy runtime toggling, but that'd be another thing to look at.
Personally, I use both, based on what is best for the feature. Why not, after all?
But day-to-day, stakeholders want to test feature X that they've heard is going well but it's not stable enough for develop? Okay, let's [engineering time and $$$]. vs. adding FEATURE_NEWDB=true to the upstart script. We've already got automation engineers helping us out, but until the deployment problem is solved toggles are more practical in our case.
It also rendered equally nicely on my android phone and my desktop browser.
I saw that the author has a github repo for an older blog style for jekyl, but I'd like to see a similar thing for this one.
Thanks
The software world > web servers.
(Sorry to be picky, but some people seem to assume that all software is developed for the web these days whereas the web world is just a significant & vocal minority).
As he discussed, feature flags also work for downloadable software (desktop/mobile) - the multiple deployments obviously don't make that much sense in that case though.
Unfortunately (or fortunately), we in the embedded software world tend to be less vocal than the others.
If its one piece of software I would think of it as an install. If it's coordinating multiple pieces its a deployment.
That said, the article does come off a bit as trying to be authoritative, but at the same time it doesn't leave enough room for possibilities where alternative approaches may have merit as well (i.e. "this is how to do it" vs. "what worked well for us, ymmv"). Newbies that read this article will think that the principles described are the canonical way and even try and apply them in scenarios where alternatives may prove superior.
Other than that, a lot of good advice, well done!
He argues for pushing about as often as possible. With our small team thats very do-able, every push gets tested and linted by the 'blue' or 'green'. You're supossed to only push passing code which you easily can by running the tests and lint locally. So instead of all the pain points mentioned in the post you write passing code, pull and rebase on other passing code, and then push. Little code review, no worries about hasty reverting, few / early conflicts keeping us from trilling each other up or writing incompatible features.
The reason I argue for Git Flow? Our tree is an absolute mess. Most often a single chain, of often linearly scrambled features. In other words removing one feature would be hard and require a bunch of legwork, not a couple Git commands.
If anyone strongly feels there's a better way for a small team than lightning fast CI let me know!
I ended up going with Distelli, it's a SaaS but it's fantastic. These days deploys often involve more than just one app or language, and I really prefer a tool that can ship anything. Also, having a GUI to see deployment statuses is invaluable. With those requirements none of the Nodejs tools can stand up to the other, more mature utilities. And rather than have to write all my deploy logic in another language, I just purchase the service.
Disclaimer: I'm the founder at @distelli
[1] http://nvie.com/posts/a-successful-git-branching-model/#crea...
Behind the scenes, it just does an octopus merge of all selected branches into master. Since the codebase was reasonably large, we almost never encountered problems with merge conflicts.
I would respectfully argue that this is not a desirable feature.
Although you might save some time by deploying multiple branches at once, you cloud the waters of what and how to roll back.
I think a better idea is to make deploys easy, quick, and revertible so that you can deploy early and deploy often, and in the event of a rollback, you can rollback just the broken feature.
For example, when do you do testing? If we test as soon as the pull request is opened, we know that master is going to change a lot between now and when we finally deploy this code so the tests might not be valid.
If, on the other hand, we wait to test just before we deploy to live we risk locking up the queue for too long. This might not be an issue if you're test only take a couple of minutes to run but if you have lots of integration tests (like facebook for example [1]) then it could become a big issue.
Is the solution to this, you just accept that the codebase you're testing won't exactly be the same as what's deployed to production and the risk that comes with that?
[1] https://developers.facebooklive.com/videos/561/big-code-deve...
That is why you care what is on master: because you need to rebuild if your runtime changes.
If you need to rebuild a specific version, it's as easy as checking out the tag.
To get the disclaimer out of the way I'm a co-founder of Vamos Deploy. Our product addresses many of the deployment problems that have been discussed here so I thought I'd mention it. We are looking for feedback on the product and an early adopter or two - https://vamosdeploy.com
I'd like to cover some techie details here. Vamos Deploy encapsulates an application with it's dependancies and runtime config so it can be deployed as one to any number of machines irrespective of the OS (well, Linux and Windows at the moment). This encapsulation is achieved by configuring a 'grid' with all the application package versions, library/runtime dependancies, runtime property values and local repository names (hostnames usually). When the grid is deployed (all via CLI) the respective local repos get updated. You can have multiple grids on a host (in a local repo) thus enabling multiple, differently configured, encapsulated applications that don't conflict. It avoids duplication by the grid sharing the underlying application packages and libraries in the local repo. There is a audit log of all actions for traceability and transparency. A simple ownership model prevents non-prod code getting into production and restricts who can deploy to production. It can be combined quite easily with any config management tool for release orchestration. You don't need RPM/Deb packages or deal with Yum repos. We have concentrated on making it easy to learn and use so max benefit can be attained quickly.
I'll stop there. Be interested in anyones feedback here or https://vamosdeploy.com#contact for a chat.
I'd be much more interested to learn about how people develop mobile and web apps, where feature flags are far less useful as you need to push the entire app to the AppStore, so your iteration time is much slower.
Instead of branching we tag every deploy and use dependency management heavily (ie maven, npm, etc). That is the project that gets deployed never really has any branches but is composed of lots of smaller projects each in their own repository which may have branches but they have to be released.
This approach cuts build time, improves coupling/cohesion as well as facilitate a possible transition of OSS useful components (that do not provide a competitive advantage nor or proprietary).
I have seen way too many projects that have this giant monolithic source tree (particularly PHP projects) and thus have to rely on branching much more heavily. I firmly believe this is the wrong approach.
Does anyone actually do this? This seems counterproductive - what if there are multiple branches?
But I've never deployed a feature branch into production like the blog post suggests. I had my own questions about this lower down in the comments.
That should work fine with multiple branches in most cases, so long as you have a system to stop anyone else deploying their branch while yours is running.
The bold version of the font is also available as a separate font family (AvenirNextLTW01-Bold). It looks much more like a "normal" font weight and is incredibly readable [3].
[1]: http://i.imgur.com/uVHXptR.png [2]: http://i.imgur.com/IXwR6EU.png [3]: http://i.imgur.com/xgz7EtR.png