Faster and more flexible pipelines with a Directed Acyclic Graph
about.gitlab.com
about.gitlab.com
They are rediscovering the basic features of Make et al.
The part that saddens me is that they all end up looking like a glorified distributed make.
The issue is that "make" made sure to put enough traps so that people with less knowledge of the pitfalls tend to fall through. How many times I have seen recursive make calls without the $(MAKE) variable lead to funny races or lost compilation flags? How many paralelizable targets running in single thread because somebody did not understand some strange out of order execution(due to wrong Dependency mapping)? How many programs which are only partially built? How many projects were not actually built for huge periods of time because somebody forgot to put PHONY target? How many Frankenstein build systems exist which create Meta languages around make(kbuild) ? My latest adventure was a dormant fork bomb for 4 years in a daily updated project in the clean target! Whole development server farms being resource starved because the clean target started being executed as part of the normal build. How did the fork bomb happen? The usual recursive make.
Make is perfect actually,but we are not perfect in its usage.
Modern CI feels like a shell script using a bunch of external APIs got tossed into a blender and poured into a YAML config. Then everyone claims builds are repeatable because they run in a container, but ignore the fact that if your Dockerfile has a single command that isn't working with exact versions for resources (apt-get anyone?) you can't rely on the build being repeatable between jobs in the same pipeline.
A big problem I think we'll face in the future is there's no value in giving developers a product they can control. GitLab CI will evolve so you need a top tier plan to do anything interesting and they'll eventually own your build environment. GitHub will push things like Codespaces and they'll own your development environment. The idea that anyone would even touch GitHub Actions blows my mind.
What happened to the developers that want to control their environments? Do you actually own your codebase if it's useless without a bunch of paid subscriptions you're using as part of your development workflow? What happens when your CI provider triples the price or shuts down entirely?
The Mother of all Demos was amazing 5 decades ago. And yet we're not stuck using the exact same UI toolkit as Douglas Engelbart.
- You can't have a DAG of stages. Sometimes, especially with a monorepo which this touts, the maintenance of the "needs" becomes burdensome and you want to just block on a stage rather than explicitly named jobs
- The visualization of this is subpar.
So far my favorite implementation of this feature is Azure Pipelines. No idea if its coped into Github Actions yet or not (haven't switched over). I hate that is the case because of my underlying caution about Microsoft after the 90s and early 2000s though supposedly they are better now.
-Stage1 needs A, B, C
-D needs Stage1That helps but the "definition" of a stage is still far away from the job definition when you have enough of them and people won't know to update this when they copy/paste a job definition for adding a new test.
Plus, visualization is still terrible :)
I've used CircleCI, Travis, Drone (I don't think I've used Concourse actually) -- and I by far prefer GitLab CI, it is more featureful and more easy to use.
Hope that is helpful. If so, feel free to add a comment so the team can see.
https://gitlab.com/gitlab-org/gitlab/issues/198570#note_2936...
I have had a really negative experience reporting bugs to GitLab in the last 12 months. I have spent tons of time investigating bugs in new features only to have the devs ignore it. It feels like they are sweeping things under the rug so they can hit deadlines for officially announcing new features.
https://gitlab.com/gitlab-org/gitlab/-/issues/39534#note_283...
GitLab Senior Backend Engineer - Verify (CI) here - Thanks for reaching out! I'll try to explain current situation of these issues.
1. https://gitlab.com/gitlab-org/gitlab/-/issues/198570
We know that this bug with DAG/needs is quite annoying. That's why we've been actively discussing about this bug in other issue: https://gitlab.com/gitlab-org/gitlab/-/issues/213080
Can you please take a look at the discussions there?
2. https://gitlab.com/gitlab-org/gitlab/-/issues/39534
Unfortunately we had no choice but to give an option to fix this behavior, instead of fixing it. That's because, if we fixed this bug immediately, it would probably break existing pipelines.
That's why we allowed users to fix this behavior themselves (https://gitlab.com/gitlab-org/gitlab/-/merge_requests/24605).
Hope this is helpful and please feel free to contact us if you have any questions/feedbacks. Thanks!
For the new features part, going fast is a core part of our strategy, as is focusing on breadth over depth (https://about.gitlab.com/company/strategy/#breadth-over-dept...). Not only does this provide an easier path to collaborate and contribute for the wider community, it allows us to shorten the feedback loop with everyone that uses GitLab so that we can invest our time and effort in the areas that matter most to the wider community.Hope this is helpful.
Here they're talking about a DAG as a model for the relationships between computation rather than data.
Many systems that model computation use DAGs. For example it's how many compilers understand your code when you compile it.
The lisp people would like a word...
In seriousness, computation is data, and this is perhaps especially obvious in a CI context where your “computation” is increasingly bash scripts in some string field inside of a YAML file.
Why do you feel the need to say things in such a snarky way as this, as if I'm ignorant?
That was the point - that computation can also be represented in a 'data' structure like this. I already gave that exact example - a compiler representing computation as data.
My confusion was not on what they were doing or why they chose to use a DAG, but why they're presenting it this way and why they weren't using a DAG to begin with.
Like I've written task runners in various forms before. They've always been based on DAGs and topological sorting from the beginning. I wouldn't advertise it as such because the choice of model is obvious to me, I'd be more interested in reading about alternatives that provide other benefits.
I agree that your example illustrates the "computation is data" concept; I was just reconciling that with the apparent contradiction ("computation rather than data") earlier in your post.
I wasn't trying to call "gotcha" on you or anything; sometimes I say things that are unclear and I appreciate it when others help me clarify.
Anyone want to share their experience with git and monorepos these days?