LambCI – A continuous integration system built on AWS Lambda
github.com
github.com
There are an increasing number of build systems that encourage squashing these bugs. The resulting build outputs are simpler and are often more portable. They're also easier to reason about. That translates to simpler deployment, simpler operations, and fewer edge cases to debug.
IMHO, the most promising answer to the 5 minute limit is finer granularity and better caching of dependency inputs.
I have aspirations for QA (https://github.com/ajbouh/qa) to learn this trick, but it needs some lambda-specific smarts before it gets there.
> No root access > 5 min max build time > Bring-your-own-binaries – Lambda has a limited selection of installed software > 1.5GB max memory > Linux only
There's a reason Jenkins is still used so widely. It's not because of utilization or all the other things pointed out. When your project gets big enough, managing the CI pipeline turns into a distributed systems problem with distributed queues, locks, error/failure recovery, and all the other headaches that such systems bring. Heck, reporting alone on a test suite with 12k tests is a problem in and of itself.
When a build system can only be effectively invoked by CI/CD, it starts to pervert developer incentives. People need to check things in before they can be sure they work. They don't bother with tiny fixes because of the inertia. Flaky jobs get a quick rebuild, because reproducing a build failure locally is complex enough that they'd prefer to avoid it if they can.
Over time, these add up to a system that grows through accretion, which is the enemy of both agility and understandability.
Better is a build process that uses simple, reusable components that work equally well on developers' machines. These tools can be tested, refined, and replaced incrementally, using the same build processes that the rest of your code base does. You can do this without needing coupling your build processes to the specific way(s) that Jenkins (and company) model builds or their configuration.
I know how we've solved the problem to provide as much validation as possible before shipping something to production and at pretty high rates of code churn. Whereas what you're suggesting is untenable on a large enough project. That's like saying drink your milk and have a hearty breakfast. Nice platitudes but not actual engineering. Our solution is not unique in fact. Shoppify and other big shops follow exact same practices (https://www.youtube.com/watch?v=zWR477ypEsc). Not because they don't know any better and haven't heard of setting up proper build pipelines using principles from immutable infrastructure but because at large enough scale you need mutability.
Jenkins was just an example. We don't use Jenkins but you do need something that manages workers and their lifecycle. Saying reduce your test runtime to 5 minutes and have better engineers and tools doesn't cut it.
Isn't the architecture of your build directly related to both the architecture of your system and your deployment?
If so, why would somebody think that a monolithic app, even one with threading and workers built in, be better than simply engineering your own as you go along? After all, this is supposed to be engineering, right? Not "How to use Jenkins"
I agree that platitudes aren't solutions, but code smells are the kind of thing that lead one to actually take ownership instead of perhaps using the same paradigm only larger, yes?
Apologies if I missed the point, dkarapetyan.
If any of our experiences or insights can help others in their own environments, all the better!
In terms of reporting on tests that run in parallel, we built a tool that specializes in exactly that. It collates output from parallel tests, it times out on tests that are hung, it makes sure the build system doesn't kill it if tests are too silent. It also tracks which tests have run against which versions of the codebase in the past and what their outcomes are. We use supporting tools to analyze test flakiness and understand when they are introduced. We have had a lot of success with this approach, as developers debugging weirdness across many tests is less miserable when they can use the same tools that CI does.
Critically, when bugs in those tools are discovered, developers can pinpoint and fix those bugs locally with reasonable ease. Deploying fixes to the test runner (or the logic that allocates workers for the test runner) is like any other change. No need to tinker with Jenkins (or buildbot, etc) config. No need to take the build system down to test that the change is correct. No need to bring up a test version of the build system and experiment with your change there.
We've gone to great lengths to make our system something that's a joy to work with and helped us be very productive across the many different environments we need to operate in.
It's tough to know how much detail is appropriate in comment threads like these. You're absolutely right that there's a lot that needs to come together to make something like what I've described work. I know because we pulled enough of it together to support our own large and heterogeneous projects.
It sounds like you have also thought about this problem a lot. Can you share more about the sorts of tests (language, test library, etc) you have? Perhaps we can break new ground where each of our respective experiences and intuition intersect.
My contention was the emphasis on local reproducibility. In the past I would have said yes, local reproducibility should be a feature of any well designed CI pipeline but nowadays I'm not sure anymore.
Local development environments are optimized for iteration speed at the cost of reproducibility and stability. Whether this is the right decision or not can be debated. CI environment on the other hand is designed for reproducibility and stability. Those sets of requirements are somewhat at odds and you can't optimize for all at the same time. Tools should be shared across local and CI environments as much as possible but not when it comes at the cost of compromising the requirements for each environment.
Your whole project should build straight from very fresh boxes, and doing builds on developers machines will never be fresh boxes. It's hard to remove the cult knowledge from a development team. My project just discovered a new unlisted dependency when doing a deploy, because every developer knew about it and installed it on their machine beforehand. It was explicitly listed in the build/devtools dependencies, but not supposed to be a runtime dep. Had the developers run tests on a fresh machine, they'd have run into it. (Of course, the CI team had also installed it on the CI boxes, because they used it for debugging.)
For a large class of tests, having the developers run them in the build environment is fine. But you also need to run them in the deploy environment, and to that end developers should hit the CI system. Every test that the CI system runs on each commit should be runnable before the commit. You should have tooling and spare capacity such that the CI system is used to run tests immediately before commit, not right after - That's too late. You should run them whenever a dev sends off for code review. You should run whenever a dev feels like it; If they're in a good spot, run the unit tests, run the CI tests, see what's broken.
CI runs the same exact build system (though with a few different options so the outputs are easier to during and after the build).
Passing CI is compulsory, as humans aren't allowed to release changes on our team. Humans may only do code review. If and when a change passes code review, it will be deployed automatically once it passes CI.
We use some of the same compute capacity that our CI system uses to scale test runners across many physical machines (though tests run against a pool of freshly cloned VMs using delta disks so we get a pretty big speedup and lots of control over the environment that tests run in).
There's a fascinating correlation between developer machines and build slaves. It's been my experience that needing to install system software of any kind on one usually leads to a headache later. We've gotten it down to just Xcode on OS X and almost just build-essential on Ubuntu.
So in spirit we do exactly what you're saying, we've just found a way to do it while using the same tooling on both CI and developer machines. We also demand that the build slave images are generated straight from install media and a fixed set of files (like those that install Xcode), so the only simple way to add dependencies (i.e. build tools or libraries) is via our build system. Use of apt, homebrew, etc is completely separate for our developers. And if they mess with the build system in a way that allows those files to leak in, the fact that build slaves are pristine means that their change will fail CI and never be deployed.
Does my explanation make sense? Happy to answer follow up questions. Also happy to be shown where our rigor is lacking :)
It was this sharing of artifacts that provided some of the impetus to use a sandbox, since a polluted output could poison the cache in hard to detect ways.
Jenkins is a huge pain the behind, fragile, the machines always inevitably turns into special snowflakes, configs in git are poorly supported. It's also very expensive in any non-trivial setting and requires full-time babysitters. Cloud CIs still charge quite a lot for any meaningful parallel builds, and there's the security aspect of uploading your code to third-parties, I trust my S3 bucket permissions more than a random CI SaaS. This seems like a great start for a sweet spot between on-prem and SaaS CI.
- No (semi-)stable releases, requiring mixing and matching git revisions of Hydra and Nix until we find a combination which builds. nixpkgs master (i.e. unstable) contains a Hydra package and module, but I think that's been added after the latest stable nixpkgs release (16.03).
- As a consequence of the above, we need to build a lot of stuff locally since they're not in the binary caches.
- While nixpkgs master contains a NixOS module for Hydra, these modules don't yet work outside NixOS. Hydra can be run on, say, Ubuntu, but it requires manual, imperative configuration of Hydra, Postgres, etc.
- Configuration relies on interacting with the Web UI. "Declarative jobsets", which allow jobset specification via Nix, are in master, but I couldn't get any git revision containing that to build.
- Build slaves must be running NixOS. As I'm stuck with an Ubuntu host, this forces me to use VMs, which has overhead and makes some architecture choices more difficult. NixOS containers may fix this, but they currently only work on NixOS.
- Since I'm forced into VMs, I might as well deploy them with NixOps. The standard NixOS image used by NixOps is only 10GB, and NixOps doesn't support resizing it (there are open pull requests for this).
- The Hydra Web UI allows logging in, configuring, etc. I would prefer if there were a Web UI without write access to the DB, i.e. once the jobsets are set up (via a declarative approach or a write-enabled Web UI), there should be a read-only view of the progress, outputs, logs, etc. (new builds would be triggered by git pushes, or by SSHing to the server and running a command).
- It would be nice if the DB requirements were a little lighter; e.g. if there were an SQLite option. Using the Postgres NixOS module is easy enough, but it's a bit overly-complicated (e.g. setting up credentials and sharing them between hydra and the DB, one-shot systemd jobs to initialise the DB, etc.); plus, if anything goes wrong I'd have to learn how to use Postgres (I already know far more about MySQL, MariaDB, SQLServer and Oracle than I'd like!).
> distributed queues, locks, error/failure recovery, […] [test] reporting
Jenkins doesn't help with any of those things.Also, a lot of those issues can be worked around by using ecs for the builds.
This looks very close to the ideal CI infrastructure. I'm used to waiting on queues and long VM or container boots and configuration on other services.
We can almost certainly count on Lambda getting longer execution times and higher memory limits. We can also count on containerization solving the root problem.
We should also be building software with the goal of tests that run within reasonable limits like this.
`time make test` takes 39 seconds on my businesses Go projects. I'd consider a 5m test suite serious tech debt. The time that developers wait for feedback on tests and deployment is becoming a business bottleneck in the continuous delivery age.
Note that doesn't mean you should switch to Go, as the ridiculously fast compile times could be outweighed by other factors (retraining, lack of specific features, etc.)
I would however investigate it as an option as it seems to have found a place in the hearts of many former rails shops!
This might help:
http://martinfowler.com/bliki/Serverless.html
Google has a seperate cloud functions offering:
Every time I see something nice, there is this increasing chance that I'm gonna end up sad because it requires some sort of external provider like AWS, DO, Heroku, GCE... I don't have any of those and I don't want any of them.
The design intention for serverless applications uses pub-sub events rather than client-server calls. That's a major enabler of async processing and containers-on-demand, and it looks like LambCI has followed the pattern to a tee.
If you have to accept requests from a non-event-driven world, AWS offer their API Gateway to provide a listening server endpoint, but I think it's telling that this was not available when Lambda was released, and that LambCI does not need it.
How you can tell OP has very little experience with AWS services being broken for hours before the little "i" gets added to the service on their status board.
> AWS services being broken for hours before the little "i"
Are these failures being systematically tracked and documented somewhere?Down detector seems to show when they've had historical issues: http://downdetector.com/status/aws-amazon-web-services
Other HN comments about AWS' Status Board not being "faithful":
https://news.ycombinator.com/item?id=9809315
https://news.ycombinator.com/item?id=11839846
"It will take a direct hit with nuclear weapon on the datacenter for Amazon to change icon to red on service status page."
"Still down 5 hours later. ELB won't register instances. Ugh"
https://www.google.com/search?q=aws+status+hackernews
"Serverless" is not magic. Its only hiding the pain until its excruciating (service down, you wait until someone else fixes it).
Should read:
"It will take a direct hit with multiple nuclear weapons on all the data centers in an AWS region for Amazon to change icon to red on service status page."
Not sure if I'm pointing out how good Amazons AZ/region model is or how bad they are at updating the status page. :)
As a huge fan of FaaS its sad that half the discussions become arguing over what to call it.
There is no "cloud", there is only "someone else's computers".