In large part, industry could largely use a significantly simpler model than git for VCS since many features are largely moot on a daily basis. Many could probably get away with SVN in large part, for example (although branching isn't nearly as good).
I think you raise a great point in that we need to look at how processes have evolved ontop of version control and look at adapting those to a similar model. In some cases it's just not practical because the way infrastructure testing an deployment works but it's the direction I think we should be going. In fact, the first step would be to make issue systems and project management interfaces provided by efforts like GitLab and GitHub available in a local capable, distributed fashion. Clearly communications often require some degree of centralization or at least peer message propagation but there's no reason that information can't be separated from the infrastructure that displays and interacts with it.
Indeed. Software forges centralized what was a Decentralized VCS.
For one reason: profit.
I agree that centralization is probably not a hard requirement for these properties, but currently no one has come up with decentralized equivalents that provide users what they actually want - and remember, you don't get to tell them what they want!
You just described https://en.wikipedia.org/wiki/Embrace,_extend,_and_extinguis...
We developed distributed VCSes and the ability to do smart things with branches and combining contributions from people all over the world at different times. Then we somehow still ended up with centralised processes where everyone has to merge to trunk frequently and if GitHub goes down then the world stops turning. It turns out that just using git as git wasn't such a bad idea after all.
We were told that dynamic languages like JS and Python are so much more productive and that this was essential for fast-moving startups to be competitive. Today the industry is moving sharply from JS to TypeScript and even Python now has similar tools. It turns out that static typing was better for building robust software at scale than relying entirely on unit tests after all.
Just wait until lots of developers who have only ever worked with JS or Python learn what a compiled executable is. If you tell them that xcopy deployment even to a huge farm of web servers is no big deal when all you're copying is one executable file and maybe a .env with the secrets, it'll blow their minds...
Well, technically it doesn’t have to. It needs to be „a“ CI server that the team agrees upon, not necessarily a GitHub one.
I think the most important factor is to keep in mind that this is something teams decide to do.
You can certainly argue that there are other advantages to using those systems. However it's an inescapable fact that if those teams had been using git-as-git and had a good local environment for each developer then most of the developers would have been able to carry on with most of their work during all of those outages. Sometimes single points of failure fail.
Seriously, a well set up gutlab is pretty much undownable. And even if you do, what really is the impact? If the pipelines don't run you cannot integrate, sure if the core service breaks down you will probably resolve to exchanging patches the old school way. But the good part is, in order to deploy those decentrally sourced changes you still will go through centralized CI and gain its quality assurances by doing so. Where's the drawback?
Given that many companies now tie their entire deployment process into their source control and CI/CD systems that means you really are reduced to exchanging patch files. For any non-trivial change in a large system that quickly becomes impractical and so development slows to a crawl or everyone literally gives up and goes to the pub.
A different view on this is that my CI system is a kind of colleague, responsible for a lot of what testers, ops, and build eng people toiled at before. It is another aspect of the decentralized system where a robot can also check out and work with the code even while humans continue to write more, which is considerably more painful under e.g. the SVN model. My human colleagues and I still share plenty of code directly, via git pulls and other tools.
Common examples of this are multi-million line C++ codebases (e.g proprietary game engines) and monorepos in any language.
Running tests on my computer uses up valuable CPU cycles that I can use to work on something else while the CI servers are running my tests.
However I can see this working for small to medium sized codebases that really don't need CI for testing.
We have a project where the team split the tests into chunks and eventually I figured out the reason why is because they had coupling between tests, and running them all together ran into problems. I worry about other people opening that door, because it’s damned hard to close again.
How well does it work? Well .... it sorta mostly works. Sometimes a build will fail for inexplicable reasons because Gradle/Kotlin incremental builds don't seem fully reliable. You re-run with a clean checkout and the issue goes away.
How much time does it save? Also hard to say. Most of the time goes into integration tests that by their nature are invalidated by more or less any change in the codebase. That's not exactly incorrect. Some changes in some modules do avoid hitting the integration tests, though.
TC can also do test sharding in the newest versions when you use JUnit. We don't use this yet though.
Capacity planning applies to tests. As your test count goes up your budget per test goes down. Every CI tool should plot build and test duration over time and only a couple do.
There’s a reason most test frameworks can mark slow tests. You need to not only use that but ratchet down over time. Especially when you get new hardware.
I haven’t run benchmarks lately but my old rule of thumb developed over many projects and with several people better at testing than I, was a factor of eight for each layer in the testing pyramid.
That certainly puts a lot of runtime pressure on the top of the pyramid, but that’s by design. You don’t want people racing to the top, because that’s how you get a cone.
In theory you could get something reasonably close to this working locally, but the it’s a serial process so that’s pointless.
I’m assuming you are only submitting this as theoretical, because in the real world (where I’ve directly experienced this workflow) it’s a nightmare:
* Someone merges main which deletes working files in your feature branch. Automatic pushing/rebasing of main onto your feature branch creates needless work the moment you try to push upstream when git tries to ascertain state of local main.
* countless times I’ve hit bugs where if I try to merge “patch-a” from a common “feat-1” branch (meaning patch was cut from feat, not from main), but then main is updated by the auto-updater, I then have a messy working directory in which main’s new files are treated as unknown orphans and I have to spend time deleting these by hand.
I’m all for having my feat branch be up to date at merge time. But making it a rolling target is something git (from my perspective) hits the boundary of what git can reasonably do, and creates more pain than any type of positive DX
Git may or may not be part of that process.
>Ideally starting from a blank machine/VM you should be able to run a single script that get’s the latest, builds it locally, runs it successfully, and passes all tests.
This isn't automatic, and is a practice that is fairly standard today (And one I agree with). You have placed a requirement/line in the sand that says, "Only when I have ideal state should this pipeline run in linear time and output the final result, which are the return codes from tests." Your example has the user initiating the update of the upstream main, not some other process that runs git fetch on the Developer's behalf.
Your original comment, as I understand it, is contradictory to this point:
>or code to function on CI machine will also be automatically loaded into production and other developers machines.
All of us (I think) agree that auto-deployment to production is a desirable goal. But we all (I think) know that broken commits are routinely delivered to production, where "production" represents the sum of all production environments in the world. So while we can have a reasonable assumption that "Production is, or should be deployable all the time," that doesn't mean the state that is represented by Production is safe to run locally in my environment, unless I *specifically* request it. Since git doesn't have file-locking, some other team/PM/developer can decide it's time for <MASSIVE REFACTOR> that blows away my work/branch mistakenly (Or maybe even intentionally, especially if I work in an org that is terrible with communication), creating unnecessary merge conflicts/mental load. This happens in short-lived and long-lived feature branches.
In no setup, do I think it's ever safe to take away the developer's agency and let some other process keep my local machine in "sync." There are so many variables to account for that some daemon/service can't be aware of, to allow for automatic updating (and again automatic updating != user running `git fetch`).
That’s not actually involved here. The actual process of coding can take place on a separate standalone project or even a whiteboard. But somehow all that code and everything associated with it needs to be packaged up for the team or it’s never making it to production. Further your process needs to minimally interrupt other team members or team efficiency tanks.
How that’s done is up to the team but it needs to happen somehow and automation avoids headaches. I briefly worked on a project where we kept passing around updates to a VM, slow and bandwidth intensive but it did actually work.
> "Production is, or should be deployable all the time,"
That’s a separate question. I am saying the code should be bundled with any needed environment configuration required to run that code. CI is a direct test of the process.
> In no setup, do I think it's ever safe to take away the developer's agency and let some other process keep my local machine in "sync." There are so many variables to account for that some daemon/service can't be aware of, to allow for automatic updating (and again automatic updating != user running `git fetch`).
Capacity is not a requirement.
For developers it’s about being able to hit a big red button and get your local environment working rather than something you automatically do day to day. Onboarding, or coming back after 3 months on another project, etc shouldn’t involve someone trying to piece together all the little environment changes that people need to apply since the last time someone updated the onboarding document.
In practice you might be reading a diff of some script and just apply that manually. But at least the sharing process is automated.
A red local build probably means a red CI build, but a red CI build doesn’t necessarily mean a red local build. Now you’re fucked because there’s a build failure you can’t reproduce.
Reducing variance helps. Having the person who set up the CI system also copy edit the onboarding docs helps a lot with this.
It’s a matter of scale. As the team grows and particularly as you hit the steep part of the S curve of development, you’re going to have lots of builds and that 1:50 error is going to go from every two weeks to once a day.
Humans interpret every day as “all the time”. Some do this for every week, especially if it coincides with their most important commits. It’s not the ratio of failures that bothers people. It’s the frequency, and the clusters.
On who's system? Your? Mine? Someone else's? The purpose of CI is consistent continuous integration in a same like manner, not relying on developer A, B or C's systems which may vary greatly.
This falls apart for things like integration tests, which may be too large/complex/interconnected to work on a local machine, but most of the time this would be more than sufficient.
Distributed signing of artifacts is only effective if you've got fully reproducible builds. If you don't, because almost no one does because it is a huge effort, then all you have is attestation. If a broken/malicious artifact gets deployed the damage gets done and you only know who to blame afterwards.
Local tests are useful for not knowingly committing broken code insofar as your tests can determine. Outside of that a full test suite run by a beefy cluster with ready access to assets and network resources is better suited to test for deployment.
It’s unlikely that verifying such a hash is possible without rerunning the tests, and not at the same time enabling someone to trivially compute that hash without having run the tests in the first place.
In trunk based development you can do a conditional commit, where TC runs your code and only pushes it to trunk if the build is green. This allows you to push something before lunch or a meeting for someone who is blocked without coming back to angry faces because you ding-dong-dashed.
Breaking trunk is fundamentally a problem of response time. If you’re not in the office you can’t fix a red build in a timely fashion. He did this regularly and got put under house arrest.
I started using it on myself so people didn’t have to wait two hours for me to get out of a meeting and fix their api bug.
CI itself is just continually testing before continually merging. Doesn't matter where.