Experiment: The costs of slow build times
github.blog
github.blog
Without this beta github offer: 2-core CPU (x86_64), 7 GB of RAM, 14 GB of SSD space
Your last development machine that was thrown away is so much faster than this (and we do have great tools for administer stuff like this nowadays). Hell, the computer I bought as a student in 2008 is comparable(!) (though it didn't have an SSD). And it will have so much better network connectivity with everything else on your network.
Whenever you hear cloud, realize that the dedicated time you will get (unless you specifically pay for it (in which case it will be more expensive than self-host)) is in the same order the phone you had two generations ago. That is why gmail search sucks. Because they can't afford to really search your messages and can not even do exact matching properly.
So yes, apparently github is fixing this now and if paying for this problem makes sense for you do it. But this is a problem that was partly invented by the cloud in the first place.
>the computer I bought as a student in 2008 is comparable
>the phone you had two generations ago
Good for Americans (or first-world citizens)
The laptop cost about 110% of my monthly entry-level salary.
GitHub CI also has much better network connectivity (both speed and especially ping) compared to most parts of my country.
By your logic poor people actually have more money to spend on renting cloud services.
In my country I've seen people rent cars, rent a smartphone, and literally rent someone else's Uber account just to work as a driver.
If Mighty had developed their product for ancient Windows laptops instead of Macs, they would have had millions of customers. It still wouldn't be profitable, because they'd be very poor customers.
Sounds silly, but suppose you are a starving artist whose ancient laptop can't even run Photoshop. It might make sense to remote into a powerful computer that can run Google Chrome, then from that computer use a web-based Photoshop replacement like Figma.
Unfortunately the people at Mighty chose to focus on Mac users, because they had to impress wealthy investors to get them to invest, and those investors are Mac users.
Self-hosting is an option for sure, but an in-between one would be to run your own gitlab and set up your own runners. I actually found out it's pretty easy to configure a runner on your own device, but with the same process you can rent a server or VPS at higher specs than the default offering and make your builds faster.
There is a significant IPC increase with every generation.
~Comparable. Or maybe I got it wrong, both are still dog slow compared to anything people actually use though.
My single threaded code runs about half as fast in Actions compared to my 3900x. They're not fast instances.
In order for the quad core, eight thread 960 to be slower than an Actions instance, there'd need to have been an 8x uplift in single core performance since '08. It's been more like 2x.
This is comparable to what several of my friends are using right now, two of whom are trying to get into IT (so mostly doing frontend development), because the economy fucked them over for the third time in their lives. They're also in their 30s. This is nothing unusual at all.
It's pretty sad how out of touch this place can be. Not everyone in this world makes $20k a month and can buy macbooks as they come out.
2) if your friends are doing frontend and cant afford a computer that is less than 10 years old; they need to negotiate better
3) most people aren't developing on their own personal computers
4) $20k is about half of what I take home a year in the EU; yet somehow I am running a quad core laptop and 16G of ram... for the last 8 years.
> $20k is about half of what I take home a year in the EU
So… roughly $3300 a month? One of my friends makes 10× less than that. It's right around the mean salary.
Edit: I originally stated (incorrectly) that $300 is close to the median salary. It's even worse — it's a mean. Our bureau of statistics doesn't publish the median, and I mixed them up.
I also make significantly less than that, and it's considered a good salary here.
> computer that is less than 10 years old
I checked several online stores just to be sure. They have lots of new laptop models with 2-core Celeron CPUs (or low-tier Intel i3s), and 4 GBs RAM, for $350 or a bit more. They've been selling like hot potatoes since the start of the pandemic because many can't afford anything else (and didn't have a computer before that — too expensive).
But yeah, we should just stop being poor. I get it.
I'm not sure where this sentiment is coming from. We're talking about business equipment with the aim of making a profit for your company here. If your company can only afford to provide $350 laptops to the software development team then they're going to have a hard time. That's not a value judgement, that's just a statement of fact.
It's like getting annoyed at logistics enterprises because they point out you need to lay out a few million dollars for trucks and warehouses. I can't afford to do that so it looks like, even if I wanted to, I'm not going to start a heavy goods transport business.
Though of course if your area is ever discovered (including remote work) the high value companies will come after you for the cheap good developers, give them better machines and repeat the rewards of good developers.
That holds true for other industries. Taking logistics: maybe a company uses motorbikes to move smaller packages around, rather than lorries, storing outdoors, etc. It still doesn't answer my question of where the "just don't be poor" sentiment came from. Maybe the OP is just having a bad day :-)
However, and please don't read this as charity, I would be really happy to send you some of my outdated equipment.
I have a Precision 5520 (Xeon E3-150M v6, 512G NVMe, 16G DDR4, ANSI Keyboard) which I bought with my own money that is collecting dust, it will fry any of those machines you mentioned.
I will send it for you for no money, just shipping (from EU).
Uh, how should it be read then?
$300 is a high, but not unheard of hourly rate here...
Something for $1000 bought a decade ago would be much better.
Getting something used you could get way way lower than that.
I'm aware that I am privileged (barely a tiny fraction of $20k a month though, and wouldn't buy a macbook even if I thought it was worth the money)
I find that using regular cloud instances (eg EC2), with a custom runner for some CI platform (Gitlab, teamcity, whatever you prefer) is a really sweet spot.
At QuasarDB, our C++ builds only take about 20 minutes this way, as long as we’re using a 128 vCPU instance. It’s a decent sweet spot for us.
Keep in mind we have 168 hours a week
1. company with <=20 engineers (or slow paced open-source) in one time zone. Except night time, you need constant running server for ~10 hours a day. (except weekends) = 50 hrs/week
2. company with 20 engineers in different time zones. CI running everyday except weekend = 120 hrs/week
3. if you have more engineers, then probably you need running CI every day.
If you are a small company, in one region, it might make sense to fully rely on cloud, for other scenarios dedicated servers could be cheaper option. Highest cost is initial setup if automation is difficult.
Some CI/CD services like GitHub Actions or TravisCI will take a configuration file which declares the CI/CD steps to run. These services will run the build on some machine; but the point of "self hosting is good" is that you'll be able to run the CI/CD build with a more powerful host, and complete the CI/CD build quicker compared to the 'managed' cloud builders.
Doesn’t address local tools dependencies, but those tend to be less (tho golangci-lint is like ten minutes of downloads for the over 100 deps (and due to or terrifically slow MITMed network proxies) I finally just made all the build tooling a published base image for the actual builds).
I just spent three days making it dockerized and the builds are often 5 minutes, due to go build getting all the deps at the go mod tidy stage which is after the copy source tree into docker step and hence not cached at all. Deployment will be on docker so whatever.
After a day of writing go, I feel happy and smart and content. After a day of getting docker files to works I feel pleased but a little like I have eaten too much candy.
One frustrating part is that a single GitHub runner can handle one job at a time, and our repos have 4-8 highly parallelized jobs. So we need multiple runners. To do this in an economic way, I made a Docker image and run 10 instances of it a single dedicated host from Hetzner. For ~$50/mo, we have unlimited CI and job run as fast as theoretically possible.
The CI box has no inbound ports open to the internet, and the OS is set to auto-update, so the maintenance burden is low.
You might not want to overload one instance.
And running it like you do, is a no brainer anyway
I understand why GitHub has this constraint - to avoid clashes between jobs. If multiple jobs (from the same repository) are executing simultaneously within the same stateful environment, they are more likely to clash over shared resources (/tmp, cleanup tasks, database names, etc.). However, even if my jobs are clean and idempotent, GitHub is nudging me to think about runners as "VMs" rather than "containers" (because often CI jobs involve their own containers, and docker-in-docker is a pain), and "self-hosting a bunch of CI VMs" becomes expensive rather quick.
but, if it does work for someone, doing what you've done will give a much better experience to their developers.
my employer uses single-user VMs for it's runners; it works well, but sometimes actions invocation is high and it can take a few minutes for a runner to come around to taking my job. that would be much less of a problem with dockerized runners.
One alternative is to use Firecracker on the CI host and run entire VMs for jobs (still quick). However, I am not comfortable with that stack to do "set & forget" implementation with no ongoing burden.
a) Developer builds locally
b) Developer tests locally
c) Developer pushes to a repository
d) CI starts
e) CI ends
f) Wait for human code review and approval
g) Merge and deploy
h) Observe that nothing broke / no need to revert.
Is it even possible for (e) - (d) to be short enough, let alone (f) - (d), to keep the developer's attention instead of context switching? Most devs I know just context switch immediately after (c). If you care about developer productivity, you're much more likely to get results from focusing on (a) and (b), by hooking development laptops into caching and restricting system/test scope, than you are by reducing CI times, unless your CI takes some ungodly amount of time to run (several hours).For point of comparison: article examines how using monster machines can reduce the build time of Fedora to 27 minutes (not really a comparable example to most companies, but OK). My devs complain about an (admittedly unoptimized) CI time of 20 minutes (on a much simpler project than Fedora) that introduces context switching. Is the article really trying to get me to believe that Fedora developers wouldn't context switch on a 27 minute build, twiddling their thumbs for 27 minutes, but that a 35 or 55 minute build would? Something about getting under the 30 minute bar gets developers to keep their focus? I call bullshit.
I'm currently working on a computer vision pipeline. Some pipeline components test and build in 3-5 minutes, which is the absolute upper limit on my attention. Other components take 10-20 minutes, and I always switch away to something else.
I will say blocking review on CI is unneeded and, depending on your system, dramatically slows down down your whole process. PRs being posted is the main notification people get. If people instead poll for CI being done, then you'll get significant lag. This shouldn't negatively affect things. The biggest risk is CI pointing out a change big enough to invalidate the review but that has been rare in my experience. CI should just be treated as another reviewer.
Say you find a typo in a comment, or a small linting error/extra whitespace/etc, and you want to fix it. You know there's no behavioral change. Should be easy right?
If your build and test suite takes hours, and the infrastructure is flaky (sometimes there are not enough job runners, sometimes the test setup fails, etc), you will often decide it's not worth it and move on. Or you'll decide to roll that change into the next PR, at which point PR's start to have a lot of unrelated work because "well, I gotta run CI for <new feature> anyway, may as well fix some linting errors while I'm at it", which is often not great practice.
Or, somebody leaves a comment in your PR that is a nitpick about something which isn't that important, but you agree you should probably change it, but tests finally finished/passed. If you make a change to address the nit, tests will take hours to run again. Do you bother, or do you "leave it for a followup"?
Or worse, do you force-merge the amended PR because there's no way that little nitpick will break anything? Leading to a culture of merging over incomplete tests?
Tests taking a long time can gradually degrade the quality of a project in this manner. Little things never get done, or the culture starts being around skipping tests for changes that "shouldn't break anything", etc. I've seen it happen in $dayjob and it really kills the culture of quality for a project.
I agree that there's a hazard here that should be defended against. My question is, if it's a small non-functional issue like that, why not mark the PR for auto-merge after you open it? Then it doesn't really matter how long the (wasteful) tests take to run, right?
> tests will take hours to run again. Do you bother...?
I think it's easy to blame CI in that case, but really, that sounds to me like a classic case of management prioritizing deadlines over quality. If management gives you the breathing room to and imposes the expectation not to merge unless the quality bar is reached, then you sit back and relax while you let this thing that's outside of your control take its time.
Look, clearly there's a number for which the CI test suite is "too long". IME that number is several hours. What I fail to see is how the higher opex of larger machines justifies small decreases in build time, particularly when there are much better improvements to be realized by caching etc. that don't necessarily carry the additional opex.
Anecdotally, at work I usually contribute to 2 codebases. The first, we've invested into optimizing CI heavily, CI is generally not flakey and almost always takes less than 10 mins to run. It's also engineered to generally fail early most of the time, so if it's going to break, you often find out in < 2mins.
The other codebase takes about 20mins, often takes >15mins to fail and is generally quite flakey.
Subjectively speaking, the code quality on the first codebase is much higher, since making those small changes is generally quite painless.
However, I do agree with you. For me it's so much more pleasant working on the faster repo and it feels like it's because of the faster build. But the above criticism again applies to my feelings.
Am I by-definition a worse developer because I find myself in this scenario? If I were a better developer, what would I do? Change teams to the CI team so I can fix their shit? Just quit and join a different company?
Not everyone is in the position to fix the things that are causing them pain.
Everyone is in a position to fix it because everyone can quit. Life is too short.
Take a look at the kubernetes/kubernetes repo and look for kind/flake to get an idea of what a project looks like that has a tremendously complex integration testing, and how often the integration tests aren’t 100% reliable. I’d wager any organization that’s “all-up”/“end-to-end” testing software past a certain level of complexity (especially anything that touches cloud service integration) is going to suffer this problem. If you’re a developer on such a project, flakes are just sorta the way of things.
There are things like ci.skip option in git for this.
I'm currently working on a low priority story in a completely different area from the important feature my boss needs done because there is too much WIP from other stories. (I need to poke some people to look at my reviews right now - but I know they are working on critical must be done yesterday fixes)
> It’s cheaper—and less frustrating for your developers—to pay more for better hardware to keep your team on track.
I agree with this point, 100%. And not just in the context of build hardware. It's often cheaper to pay for specialized services rather than paying in-house developers to do those things. I've lost count of the number of times in my career that I've spent days of effort (at $200/hr) so that the company I was working for could avoid paying for some $15/mo SaaS.
But you know, when you're scared of the unknown unknowns, the risk, managed services are attractive.
But with automation already in place cost to both maintain that and that once-a-month visit to replace a hard disk in datacenter, or once-a year half-day to install few servers is far cheaper than anything cloud, especially if you need bandwidth
It is the lost productivity of the developer not making the things you’re paying them for.
The reason the business is willing to spend $x a year on the developer is because the developer produces >$x a year of value, in the form of revenue generating software.
If you divert 5% of your time into doing something else, you are taking money the business wanted to invest into capital (new software features) and spending it on… something else. Lunch orders or whatever. Not making capital, anyhow.
Buy a rack and a dozen cheap-ish rack chassis with some middle range CPUs with lots of cores but not necessarily the highest specs otherwise (Last gen we used 9700K and for a recent refresh it's AMD 7700X). You can get quite far at $1k per node and $12k for a rack. Then you can run 3-4 build agents per node so you have a 48 agent cluster in your rack. Electricity and management will be on top of the purchase price of course, but it's still a bargain.
Even if you do use cloud compute for builds, it's worth having some self-hosted nodes for the base need and just using the cloud for peak use scaling.
Gitlab runners on an autoscaling group handles load for me and scales up and down with need. Took 2-3 hours to set up, and has lasted 2 years without incident.
The value for me is cost (especially in time) to implement, mental burden, and risk of delaying projects.
Notably absent from the comparison is the cost for buying a build host yourself.
Build servers are fine for the main branch, but why does a dev have to build the whole thing to quickly check the last 50 lines that they've written?
Also, more hardware does not make builds faster magically all the times. one needs to spend time to actually use more/ better hardware.
Finally, build load comes in spikes. end of the working day, you have lots of builds all of a sudden. Few runners get overloaded in peak or many runners waste a lot of money. May be we need builds being done in some other timezones where its off-peak time, and even better, if people could pool their resources (and thats what SaaS runners do for us.)
I die a little inside every time my build is waiting for a runner while my M1 Mac is just sitting there.
GitLab behaves similarly
What could be more convincing is the effect on product value that build times could have, and they certainly could (and I, for one, believe that they do, especially if they can be made to be less than 15 seconds or so). For example, it's certainly conceivable that very short build times make it easier to write more tests, which could result in more correct software with better features and/or a shorter time-to-market.
Some products targeted at developers also make the mistake of focusing their marketing message on the cost, perhaps because that's the part of the business that developers personally feel, but it's not the message that would convince their employers.
What matters is development velocity and quality - these are the things you will be competing on - not really cost of production.
Slow builds impact both of the above of course.
Salaries are going to be considered largely immovable and they won't "go down" or show as a lower number on the P&L as a result of your devs getting some time back. Whereas your infrastructure costs jumping up (even by a small amount, esp if you have any amount of scale) will set of alarms, rightly IMO, in the Finance department.
It's rarely only a question of just allocating money to more/better hardware, it's also a question of policy and willingness of your organisation to keep CI time short/feedback fast.
It's great for keeping things tidy and not having to break the whole thing up in modules like "UI 21.0.12.11 needs Backend 12.2.1.11 or greater". Instead everything builds more or less into a monolith. But it does have one drawback: making a commit invalidates every bit of the build output immediately.
While largely fixed now, older C# projects suffered from separating dll references from the reference to the package that provided them.
This lead to the situation where you may be referencing a package that hasn't been downloaded. Due to the package cache, the solution would typically build on a "dirty" environment like the developer's machine or anything that had previously built the project but would fail on any fresh environment.
I don't think is has to end like that. You can have separate queues and separate levels of assurance. For example, does every commit have to be tested in each of 20 possible configs? You can run one common config by default and allow unblocking the rest on demand. Then enforce all of them only as a merge gate.
If you can also split them into separate queues that don't block each other, you get both larger matrix and faster dev turnaround.
So you do need policies to actually be allowed to optimise CI.
Also my local machine has resources entirely dedicated to me and isn't held back because someone else decided to rebuild the world.
Scale up ahead of time, so there’s always a machine ready. Prefetch the repository / packages when CI workers start, so they’re already installed before a build is triggered. Use your imagination - CI doesn’t have to suck.
If you're paying for an engineering team, that's a rounding error.
Effort and money will only be expended to make things faster when the build times are perceived as intolerable.
While more cores can certainly help with certain types of projects, such as those that can be easily parallelized, this is not always the case. For example, web app projects won't benefit as much from additional cores.
Another important factor to consider is the single-core performance of each vCPU. Many server-class CPUs, such as those used by GitHub, are built with a very high-core count but with a very low single-core speed. In contrast, BuildJet uses consumer CPUs, such as the 5950x, which offer slightly less core count but an excellent single-core speed.
It's quite astonishing how slow "the cloud"/server-class CPUs can be, we compared my old MacBook Pro 2015 vs. a 2vCPU GitHub actions runner and the MBP 2015 won most of the time.
BuildJet's bet is that single-core performance is critical for a fast CI, and it appears that the self-hosting comments here on HN also agree.
(We're working our own CI, DM me if you're interested in the fastest CI on the market)
You'll be able to customize the runner image however you want, and it will be running on very fast NVMe SSDs.
We use Rust, so I think a sane caching story is more important than anything else. Not sure about single vs multi core tbh, I can add both helping.
Do you have any experience with customers building rust and docker rust images?
All our runners are self hosted. There are fewer large runners, so developers are inclined to go small. Hypothetically, you can run all your jobs in large but then you may have to wait for availability but in practice we haven't seen any real contention.
We run over 15,000 jobs a week on 6 physical servers with this pattern.
Not to mention the tendency to craft convoluted and heavy solutions to rather standard problems. Because "this is how it's done now".
Standing here, holding my hammer. Lo, on the horizon: a nail!
Some might just not care that much, or nobody might speak up about the problem, not wanting to make it seem like they're just complaining, in the case of nobody being able to provision more resources, or optimize the build. For example, in some environments, getting a better CI node might be associated with red tape, lots of back and forth with a different department or something like that, especially if it's all on-prem instead of using managed services or cloud resources.
> At the same time you probably shouldn't be throwing more hardware at the problem without trying to understand it first. I've witnessed numerous times throughout my career what can be saved by tuning a few parameters.
This is an excellent point!
In a context of containers (just an example, though illustrates some pitfalls nicely) there's a world of difference between using a base image that has the runtime and tools you need, vs installing them every time. There's a lot of difference between pulling in your dependencies every time, vs using a cache for it or an instance of Nexus that acts as a caching proxy (though that's not very nice to configure in most languages). There's even difference between making sure that packages install correctly from the cache, vs having them already be installed in some intermediate layer that you can just reuse when building your container.
Even without containers, there can be lots of things to take into consideration, like whether you could build your back end and front end (in the context of webdev) or other components in parallel, whether you could run any tests you might have in parallel, or whether there's not duplicated work going on in there somewhere.
I gotta wait the 5min before I can even debug it as we aren't allowed to run it locally
I'd kinda love to have transparent "workplace" that builds local by default but just have "build in cloud" option for those few cases where I do need to do a ton of compute at once (say getting to bisect something) but moving completely to the cloud seems just worse experience overall compared to just relatively beefy machine
My only complaint is I haven't been able to quite figure out how to make the developer machine cache independent from the CI/CD one, which has caused a couple of small issues we were able to work around.
One thing I have to say though, is I don't use NX plugins except for `@nrwl/jest`. I just use their Task runner capabilities. I found the plugins were a hassle if you weren't adopting the integrated[1] approach. For whatever reason, Jest worked well in the package based monorepo (for whatever reason, our IDE of choice, JetBrains, doesn't support TSConfig Paths properly in TSX files so it killed DX. So we had to go package based)
That said, it was easy enough to get up and moving with the task runner and we've seen major benefits. Did require us to think a little bit differently in how to run tasks in a parallel friendly way though.
[0]: https://nx.dev/
I definitely agree on investing to make builds faster, but when they're a few minutes and the build will go down to slightly fewer minutes, those 'marginal' gains don't seem worth the increased build cost.
Just throwing more compute at the problem isn't always the answer either. The build tools I've been using over the past few years (Webpack and now Vite) have gotten so much faster on their own, which shows that there's loads of slack in our code and lots of room for improvement there before we need to throw more compute at it.
Also, I'm certainly not just twiddling my thumbs while it's running...
Am I the odd one out? Feels like a bit of a straw man argument from MS/GH trying to justify a blanket switch to more powerful runners 'because the data says so'
Do the math _for your business_ and _your use-case_
Is a remote dev environment an economically viable alternative? I suspect not but haven't run the numbers. It's not clear to me what tier of virtual server would beat an M1 macbook pro with 32gb ram and sufficient number of cores.
I always see it as a hierarchy; editor, local runtime, pre-commit / pre-push hooks, CI. At the moment, if this setup works, my pre-push hook should only run various tools (linters, unit tests) on the files that were affected by the commit. I'll leave it to CI to run the full suite of verifications and end-to-end tests.
My CI is basically 5% git checkout (varies), 2% build (a few minutes), 80% run tests, 5% static analysis, the rest is misc crap like uploads of artifacts, code signing, etc.
CI software in general is very slow. But even building locally, there are many projects that have just ridiculous build times because it's just not a priority at all. I've seen small-medium sized Java projects with a 25 minute test suite.
This has been in part due to extremely inefficient tests that keep assembling the entire system over and over each test, often with extreme redundancy in coverage, ensuring you need to keep running the entire suite whenever a test breaks. Nobody is going to run tests for half an hour unless you hold a gun to their head. So you make your change by 10 AM, usually something trivial, and then you spend the afternoon fixing the 30 tests that broke, running the suite over and over to find more broken tests to fix. It's almost comically inefficient.
It's rarely obvious what these types of tests actually test. It's like... just a bit of everything for the sake of good coverage. That and the inertia from having to wait half an hour for the test results makes it very compelling to start just "fixing" the tests by changing the assertions to expect the new value.
If you have a more sensible testing strategy, it's obvious what is being tested (and thus also relatively clear what broke), and the turnaround is seconds possibly low single-digit minutes, so you don't have to rebuild the mental context every time you run the tests.
Oh my... I spent hours trying to figure out why I broke a test, only to figure out that my change actually fixed incorrect behavior! The previous dev who worked on the code had "fixed" the test failure by changing the assertion to expect the incorrect behavior!
Automatic test suites with thousands of tests sound great in theory, but in practice it's not all rosy...
Our integration test suite is absolutely horrid. Like, the thing takes literally 25 hours of real time to run, so it’s split into 50 pieces running in parallel, ultimately taking just under 30 minutes. It has like 3000 unique scenarios and pretty nearly 95% of every scenario is exactly the same. It’s a _huge_ waste.
If it breaks, I pretty much only know that it’s broken (often because of some flaky behavior, or timing related thing, not an actual issue). Sometimes pinpointing the issue goes meta and is a matter of looking at which tests are broken, instead of how they fail.
But when all those tests succeed, I know that my system is just as stable as it was before. All the major flows are covered, and if it runs successfully, I can be fairly certain my system works.
Now imagine dealing with the codebase that leads to such a test suite without it…
As you say, 95% of every scenario is exactly the same, which I think is relatively common in this type of code base. The ideal test time is thus 90 seconds instead of 30 minutes. Maybe that unattainable, but I wager there exists a test suite that has the same coverage but with significantly less redundancy.
Test engineering is actually hard. If testing is just some afterthought primarily aimed at coverage with not much consideration put toward performance, what you typically get is redundant tests with abysmal run-times.
I think it's a mode of development more than a framework that's at fault, even though the two tend to correlate.
1/ Get an actual, real machine. We run all of our CI for 10 devs out of a single 13900K / 64GB RAM machine, from a Jenkins setup that I've done over a weekend that has a single runner. 5-10 minute runs for PR validations (500kloc app, shitloads of tests, etc), 3 minutes for app builds. Buy it and host it yourself, or rent from Hetzner. You don't need a sysadmin for that. Machine is fucked ? Just restart it lol. Dependencies out of date ? Just upgrade your dockerfile, or your ansible setup and reinstall it all. Used to use Gitlab runners, Bitrise, Github Actions, etc. They're all too slow, because all these services make money on allocating as little as possible to you to stack as many clients as possible on the same server. Pretty UIs are not worth 45 minute builds.
2/ Cache. Cache. Cache. Cache. Cache. It isn't hard if you're not looking for perfect. Nobody cares if sometimes, randomly you have one small module that rebuilds even though it shouldn't have: the beefy machine is here to make sure it takes 10 seconds anyways. All that matters is that you never have a cache that says even though it shouldn't. And to solve this: don't make the cache yourself. Your build tool either has it (Gradle, Maven, Webpack, etc), or there's a widely used tool that works (ccache, etc.). How do you store your cache ? Not on the build machine. Because that fucking sucks. Make it a remote cache (If using gradle, you can pay for enterprise if you're feeling optimistic, host a gradle remote cache yourself, or use https://github.com/androidx/gcp-gradle-build-cache and forget about it. The running costs of it have been well under a hundred dollars over a year).
3/ Less important, but know your build tool. I've spent a lot of time rewriting Gradle projects to avoid cross configuration, allow configuration caching, allow parallel configuration, etc. The same applies to everything else. Your build tool should be part of your knowledge.
So far no complaints to be honest.
This should be at the very top of the article. It basically negates entirely the whole experiment because no developer will just sit and wait for hours for a build (unless they planned on taking that break anyway).
The cost is not 0 but it is absolutely negligible.
An interesting experiment would be to find the switch-time threshold for a wide range of developers.
Typically I don’t rely on a full build for getting things accomplished, and not CI if I can help it. Local, incremental, and minimal path to feedback is what I aim for.
Saves more money that way, and you don't have to care much about build duration.
It’s also easy to contract on your terms. You just say “I want that for very cheap” and get flooded with interested contractors.
/s
What you will achieve with this policy is minimizing the number of build processes, from a few times a day to once or twice a week (or less), which is exactly what you do not want from a CI/CD / DevOps environment.
Why would I be paying for bathroom breaks and idle time? I wish there was software to count how much time is lost when users minimize all windows or switch to desktop, so I could deduct it from their pay. I know some of this idle time is spent thinking about problems to solve related to the job, but it's something that can be done in your free time, or while commuting.
You want your workers to compile the code before committing it and moving on to the next task, right? But don't feel like you should pay for that?
Or are you assuming that people are like fleshy machines that can just store the data they just used on disk and work on something else while the compilation is running and then they just put that data back into RAM once the compilation is done? People don't actually multitask that well.
> [...] some of this idle time is spent thinking about problems to solve related to the job, but it's something that can be done in your free time, or while commuting.
Now that just sounds like you don't feel like you should pay workers that do intellectual work for their work at all – they can do their work in their "free" time? By that logic I don't see why you don't just pay for the 1 minute it takes to commit the code once it's finished – in the end that's all you care about. Why would it be your problem how they get there?