How do you deploy in 10 seconds?
paravoce.bearblog.dev
paravoce.bearblog.dev
Then the BORG came and assimilated us. Our deploys take easily 45+ minutes to really start shifting traffic
See https://stackify.com/kestrel-web-server-asp-net-core-kestrel... - comparison table 3/4 of the way down.
- On commit/push, your build runs once, and stores in an artifact. If nothing has changed, don't rebuild, reuse.
- Your build gets packed into a Docker container once and pushed to a remote registry. If nothing has changed, don't rebuild, reuse.
- Every test and subsequent stage uses the same build artifact and/or container. Again, this is as simple as pulling a binary or image. Within the same pipeline workspace, it's a file on disk shared between jobs.
- Using a self-hosted CI/CD runner, on the same network and provider as your artifact/container registry, means extremely low-latency, high-bandwidth file transfers. And because it's self-hosted, you don't have to wait for a runner to be allocated, it's waiting for you; it's just connected to remotely and immediately used. K8s runners on autoscaling clusters make it easy to scale jobs in parallel.
- Having each pipeline step use a prebuilt Docker container, and not having each step do a bunch of repetitive stuff (like installing tools, downloading deps...) when they don't need to, is essential. If every single job is doing the same network transfer and same tool install every time, optimize it.
- A kubernetes deploy to production should absolutely take its time to cycle an old and new pod. Half of the point of K8s is to prevent interrupting production traffic and rely on it to automatically resolve issues and prevent larger problems. This means leaning on health checks and ramping traffic for safety. But actually running the deploy part should be nearly instantaneous, a `kubectl apply` or `helm upgrade` should take seconds.
The only exception to all this is if you (rightly) have a very large test suite that takes a while to go through. You can still optimize the hell out of tests for speed and parallelize them a lot, though.
I prefer to run tests locally whenever possible, for instance as a git hook, rather than in a CI instance. If you need auditability for something like PCI, that approach probably won't work, but I think the small web (i.e. most of the web) can do just fine with it.
How is this a good excuse? Is it really that difficult for a developer to spend an afternoon understanding GitHub Actions and Docker, at least at a superficial level so they can understand what they're looking at?
In the words of Jake the Dog, "Sucking at something is the first step to getting good at something."
Many ways to skin the cat. This is just one of them.
On my VM, keep this running:
while true; do ssh pipe.pico.sh sub deploy-app; docker compose pull && docker compose up -d; done
On my local machine: docker buildx build --push -t ghcr.io/abc/app .; ssh pipe.pico.sh pub deploy-app -e
https://pipe.pico.sh/ docker buildx build --push -t ghcr.io/abc/app . && ssh myvm 'cd /my/app/path && docker compose pull && docker compose up -d'
?That might work if you have a single VM but it's a little more complicated when you have an app on multiple instances.
pipe is a multicast pubsub which means you can have many subscribers.
I can see a count of readers per day on each post. It also shows counts of devices, browsers, countries, and referrers. Here's what it looks like: https://herman.bearblog.dev/public-analytics/
You also just lost all your guardrails and collaborative controls, as well as created a dependency on all engineers being equally capable.
In other words, unless you are DHH and don't have to scale (both in terms of workload and terms of company), this scenario doesn't apply in the real world.
If you're hitting this, you need to take a look into the service as the problem, not blame the infra layer.
k8s can absolute roll out a deployment in <60s, if not <10s. The bottleneck I see, when it is slow, is slow app termination. If your service takes 5 minutes to terminate, it isn't going to matter what the infrastructure layer does. Sometimes this is failing to handle SIGTERM (resulting in k8s having to fall back on timing out & SIGKILL'ing) … but sometimes it's just the app is slow to terminate, 'cause bugs. But it's those bugs that should get fixed.
You can somewhat workaround it by setting the surge to 100%. (And … even if you have a fast app, 100% surge might be a good idea, too. As always, it depends. If surging is going to eat up all the available RAM or CPU … maybe not.)
And most importantly: the underlying principles guiding k8s's behavior are going to apply just as equally to a shell script. app.service doesn't respond to SIGTERM[1]? You're going to have to decide what to do. Surge or not? Same thing. Potential for surge to result in resource depletion…?
> bash script
A service's program/code should generally be owned by root:. A service (generally) does not need the ability to re-write its own code.
> Only a few at my company understand Docker's layer caching + mounts and Github Actions' yaml-based workflow configuration in any detail.
… the working knowledge of either of those two things is not rocket science. The docker caching is probably the worst of the two; but you only need to understand it to speed up builds.
While GHA's YAML isn't pretty … it's also hardly complex. And for the most part, if your action simply defers to a script in the repository (e.g., I keep these in ci/), then it's mostly reproducible locally, too. (And there are some tools out there to run a GHA workflow locally, too, if you need more completeness than "just run the same script as the workflow".)
> Show you how I provision a Debian server using Ansible
I have spent enough years with Ansible to know all the problems with it, and I'd rather not go back to it.
([1] although a vanilla systemd service is going to have an "advantage" in that the default SIGTERM handling is different from a container. So it might look faster, in the case of buggy-app-with-no-SIGTERM-handler will die instantly … but it's probably still a bug, as ax'ing the service is probably also just dropping requests on the floor.)
> How do you deploy in 10 seconds?
By editing in production, of course.
Thank you for your attention.
We were a PoS app, and had no SaaS web offering.
The head office was keen to get one, so someone there prototyped one, then that prototype impressed the board so much, they got a couple of contractors in (with no knowledge of the actual product development team) to flesh out the prototype.
Eventually the product team admitted what they were doing and brought in the developers to take a look at how they were working.
They were using git, but not really using it. Because their actual mode of working was to SSH into the production machine, and use vim to edit the code.
Multiple users. SSHing into the same space, editing the same files. Sure, they had a git log, but it wasn't exactly best practices.
It was also all PHP, but we were a VB6/.NET shop, so there was also some friction around whether we should all learn PHP and embrace that going forward.
I was half impressed they managed to get so far with their prototype, half horrified by what I found.
There was not just no security consideration, they didn't know what an IDOR was.
I was incredibly impressed by the professionalism of the security auditor they got to review the code. I learned a heck of a lot from him in the few days he was with them. I regret to this day not writing down his name.
But yeah, those deployments were probably the fastest they would ever deploy.
Despite being impressed by the forward thinking of the product team to force the issue and get some kind of SaaS web presence, I sharpened up my CV and got a job which didn't involve either VB6 or PHP.
Now I have to file an exceptions for a found buffer overflow vulnerability in libfdisk1 identified in my miminal container image running in a locked down, read only container context. Because ITSac has processes for it.
Because this story being flagged in WILD ya'll.
We deploy in <1s. New sites, mods, small sites, large.
If your site is 100MB, write the new site to disk and read it into memory, kill the old server and bind to port on new server, in 1 second.