Advantages of monolithic version control
danluu.com
danluu.com
This is the major thing I miss about Subversion, and the fact that in Subversion a subdirectory in a repository can be checked out on its own.
At the top level of our Subversion tree, there were 'web', 'it', and 'server' directories, reflecting the departments in the company (at least those departments that dealt with source code).
In the 'server' directory, there were things like 'payments', 'reports', and 'support', for things like payment processing, reporting, and stuff to help the customer support people.
So lets say we had programmer in the server department working on the credit card storage system, and on a script to make a quarterly tax report. That programmer would just have to check out /server/payments/cc_storage and /server/reports/quarterly_tax from the company Subversion repository. When he checks the history in either of those, he only sees commits that affected those directories or their subdirectories. It really is like they are separate repositories.
Suppose another programmer is working on the whole payment system. He can check out /server/payments, and automatically gets cc_storage under that, but also order_processor, paypal_callback, cc_updater, and subscription_biller.
I was in charge of the whole server department. So I could check out /server and have a copy of everything we did. The other programmers would usually save work in progress to a personal development branch, and so every morning I could update from the server, and then see what everyone had done the day before. I could do a quick review to check the less experienced programmer's work, and it also gave me what I needed to write a short note to my manager letting him know what the server department was up to and who we were progressing.
I also get the ability to use the depo explorer to poke around at absolutely everything that folks are working on. The search function in gitlab is basically a joke compared to this.
Unfortunately for perforce, git/gitlab seems better in basically every other way. It took me 2 years to have anything nice to say about perforce, my regular development flows we're just that painful.
I kinda think you might be able to convince git to let you check out a subdirectory. But I’m not sure if any of the plumbing exposes that ability or if it would take significant surgery.
If this does do all that, I think this functionality needs some SEO love because this pretty much never comes up when I search for the latest ways to grab part of a repo. All I find are conversations where people are trumpeting the wrong tools for the job.
Edit: Also, this doesn't seem to let me check out a subtree the way people mean "check out a subtree". When I check out 'just foo/bar/baz' I expect to have a directory called baz as my project root. Not a directory named foo with a single grandchild named baz.
A problem we saw with perforce and the clientspec: when people see a directory structure, they forget they don’t have all the bits. They make errors of judgment based on bad info.
I'm not saying I prefer the subversion architecture, but with subversion the pattern described is quite natural and requires no additional technology or design beyond a directory structure.
Which is weird, considering it was designed to be used as the Linux repository control system, which is a monolithic repo.
When he checks the history in either of those, he only sees commits that affected those directories or their subdirectories.
And in git, you just enter the directory and do "git log ."
It's good we now have (at least) Google, Facebook and Microsoft as examples of companies using monorepos. Those are names that carry some weight when thrown into a conversation.
https://git.archlinux.org/svntogit/packages.git/refs
https://git.archlinux.org/svntogit/community.git/refs
Orphaned branches allows you to have multiple independent trees in the same git repo, si it's basically a way to stuff many "repos" (as in history) in a single one (as in object storage).
What you were linked is the "svn to git" mirror which has to somehow translate that aspect of the setup to git.
Incidentally, the way Arch svn is currently set up works really well for arch devs and it is very hard to find a good replacement to it using git. Having a monorepo split up in lots of microrepos is straight up not possible in git.
Not saying you should, but if your goal is to have a single repo (VCS) per repo (Arch) within which all packages are stored, and replicate the {core/{foo,bar,baz},extra/{qux,tor,meh}} tree, there could be many ways to do it depending on the exact needs, by leveraging GIT_DIR, GIT_WORK_TREE, GIT_OBJECT_DIRECTORY, GIT_ALTERNATE_OBJECT_DIRECTORIES, clone --single-branch and checkout --orphan.
Single monorepo with orphan branches, single clone, multiple work trees:
# create monorepo
mkdir core
cd core
git init --bare .git
export GIT_DIR=$(PWD)/.git
# add package in its own branch isolated
export GIT_WORK_TREE=$(PWD)/bash
mkdir bash
cd bash
git checkout --orphan bash
git add PKGBUILD
git commit
# switch package
export GIT_WORK_TREE=$(PWD)/readline
mkdir readline
cd readline
git checkout --orphan readline
git add PKGBUILD
git commit
# back to bash
export GIT_WORK_TREE=$(PWD)/bash
cd bash
git checkout bash
But each time you switch you have to switch the branch too since HEAD is in core/.git, so to work around that you can share the object dir.Single monorepo with orphan branches, multiple clones but shared object dir, multiple work trees:
mkdir core
cd core
export GIT_OBJECT_DIRECTORY=$(PWD)/.git_objects
mkdir bash
cd bash
GIT_DIR=$(PWD)/.git git init # otherwise .git will be created next to .git_objects
git checkout --orphan bash
touch PKGBUILD
git add PKGBUILD
git commit -m "Add package: bash"
cd ..
mkdir readline
cd readline
GIT_DIR=$(PWD)/.git git init
git checkout --orphan readline
touch PKGBUILD
git add PKGBUILD
git commit -m "Add package: readline"
# back to bash
cd ../bash
git branch # just to check it's "bash", not "readline"
# clone existing package
cd ..
git clone some.where:core.git --single-branch --branch libarchive
cd libarchive
That's from the top of my head, and well it very much depends on what you want to achieve as a workflow.That's a trade-off I'd be willing to make though given other advantages git offers. I wonder how mercurial would fare on that regard (possibly enhanced with a bespoke extension). Not that there's any reason for your tooling to change since it fits the bill so well :)
Anyway for archmac I went the boring way and stuffed everything as subdirs inside a single monorepo as I reckoned it'd just be easier for contributions. Definitely not the same scope as Arch Linux though.
I often think about this issue of monorepos... git is such a wonderful tool, but the problems of subrepos and monorepos comes up so often it's surprising there isn't a definitive answer in git.
My understanding of Google's monorepo is that search, mail, maps, etc. all live inside one repo.
[1] https://blogs.msdn.microsoft.com/bharry/2017/02/03/scaling-g...
This basically sounds like CVS.
OTOH, I think the common CVS workflow actually matches the modern "we don't do stable branches" workflow a lot better than git does. Basically, if you had upstream CVS branches for more than released versions of software in maintenance mode you were doing it wrong.
I also tend to yearn for the days when I didn't spend 20% of my time rebasing and merging patches together, or rewritting dozens of patches worth of git history in order to move a couple minor commit hunks between patches for some reviewer. Or just juggling 20 different -next style remote repos.
Git is one of those tools that let you endlessly play with your tools rather than getting the job done.
I vehemently disagree with this sentiment. In fact, I find the opposite to be true. Git is the first VCS I used that is useful during coding instead of after it, when it's time to publish the final product, ie. a changeset. I can commit, merge, branch, rewrite and share changes freely and effortlessly whenever I need to without committing a bunch of crap to the shared repository that is of no interest to anyone.
After I'm finished getting shit done, I can then spend some time reviewing and thinking about the logical progression of changes so that the commit log will be readable to other people and older me. This is often just as valuable as writing the code itself.
There are pros and cons to each - do you want to have a hugely churning "I always have to rebase/merge" repo under you, or multiple repos and trouble keeping them in sync?
Having done both, I'm not sure which is better - it's probably very project specific.
That very well may be, but I think we're at a point where the majority of developers started after git came out (since the industry experienced explosive growth just in the last few years). They kind of forget (or didn't know) that it's not that long ago we did monorepos because we HAD to. The tooling to do a system in multi-repos was just not there. It wasn't practical.
Companies like Microsoft, Google and Facebook predates the days where building your company on top of 3000 repos was practical, and they certainly were not going to convert everything if they could help it. Thus, they built an enormous amount of tooling to make it work. It certainly has benefits (and tradeoffs). With similarly advanced tooling to support you, multiple repos also has a lot of very nice properties and scale quite nicely.
To each their own.
That was never ever the case. We did centralized repos because we had to, but every single place I worked at, whether they were using SCCS, RCS, CVS, SourceSafe or SVN, had multiple repositories for different projects. No place had more than 20 repos, but then, the largest of those was about 250 people with VCS access.
Multiple repositories make for a forgiving structure for your code base. You can tailor them however you like.
But once you have a _lot_ of code, they become hard to manage. I see the utility of a monolithic repository there—now you know exactly where all your code is: it's in this one repository!
Package managers mitigate a lot of the trouble with pulling in internal dependencies from other repositories. Nowadays, most languages have a package manager that can work with a private codebase, so monorepos aren't necessary to help with that. But monorepos can help if you have a ton of versions floating around and you don't want to support version 1.5.x of libjohnny when it's now on 4.3.x. Your code either works with libjohnny as it is right now, or it doesn't. (Which in turn makes it very clear to you how important it is to manage API-breaking changes!)
This feels a little bit rambling, but my thought is that there is some analogy between monorepos and microservices; don't use them until you _need_ to use them! You'll know it when you get there.
It might seem like a pain to do it this way, especially if you're rapidly iterating on a library—it might be that your library is not really mature, or even used by more than one repository, so it may not even make sense to have that library in a separate repository to begin with! But once you do have a mature code base, semantic versioning is a really sane way of managing dependency updates for the N number of other projects which use your library.
That and resectioning code to split or combine responsibilities in different ways. Something a monorepo makes trivial.
If that is the case, then you can isolate the likelihood of an issue as either in the library (because unit tests fail there), or in the project consuming the library (because unit tests succeed in the library).
Without testing, it's hard to have a ton of confidence in where the problem lies—which is exactly the problem you cite. And while a monorepo (or just a standalone repository with no separate libraries, which is frankly an easier setup to manage than monorepos or multirepos!) may make debugging a bit easier, it's not going to give you much more confidence in your code.
Once your organization outgrows the paradigm where you just have N standalone projects with N completely separate code bases, and you do need to commit to either a multi-repo or a mono-repo configuration, it's really, really helpful to have unit testing to allow you to isolate where issues are.
When the tests that matter cross version control boundaries you pay for it. Whether the costs outweigh the benefits is something you have to think about.
Testing is, of course, no silver bullet. Tests are written by humans, and humans make mistakes—and it's pretty difficult to achieve 100% test coverage in a production system. The goal of testing is to have confidence in the code you've written.
Tests often don't need to cross version control boundaries. You can use mock data—like would be produced by the library—on the consumer side, because you can delegate responsibility for testing of that library to the library repository itself. If your tests work great with the mock data, but things are still failing, then you can infer that the mock data and the actual data are different, and your bug is in the library.
Or so I thought. The first time I turned it on it preprod I couldn't turn it back off again because some piece of data that came from five function calls away was being shared, and nobody who participated in the PR recalled that fact.
Most of the code I'm dealing with is in a single module. I have been chipping away at fixing the insane ball of mud as I can. My coworkers often aren't that lucky. They come to me for advice on how to deal with this sort of problem but crossing 2 or three modules.
There's no low-friction way for them to fix any of this. They can't just refactor because of the coordination costs, and also the loss of historical information when you move a block of code across module boundaries or try to change module boundaries. This is the prime argument for monorepos in the literature - not making irreversible decisions on Law of Demeter problems. It's not my biggest reason, but it's sufficient for most people.
I think there's two ways you can look at your choice of configuration: ease of debugging, and ease of organization. When Google lays out why they use a monorepo, they are doing so because it simplifies their organization—there are no longer so many versions of so many libraries and apps they need to support; there's only one version of anything to support. Either everything works or everything fails.
But in your case, you're looking at it from the debugging point of view. It's easier to play around with the code in a monorepo. And that's totally fair point of view to have, particularly in your predicament.
That choice of a monorepo doesn't necessarily improve the quality of your code organization and interoperability. It's still going to be a bad bug to fix. It's just a little bit easier to debug.
I take it as a rule of thumb that a more realistic test is better than a less realistic test.
I can't think of any reason why you would want to mock anything if using the real thing is cheap and easy, building mocks is expensive, and testing against the real thing will increase the chances of detecting real bugs.
I also prefer it when my tests detect bugs in other libraries which my code depends upon because as far as users are concerned, bugs in libraries my code depends upon are bugs in my code.
I have in the past written a bunch of functional tests which check out / pull and build code from other repos to run with my code.
What you should do is test your library independently from the consumer that uses it. If that's not possible, ask yourself why that is; maybe this code should not be sequestered into a library after all, or maybe it needs to be designed a bit differently to remove some of the coupling between the library and app that seems to be at issue.
Ideally, your testing in your app should not be testing that the library works; it should be testing that your use of the library works. Too bad the real world is much messier than the ideal world.
Good luck either way!
You write custom scripts to reïmplement everything you'd get for free with a monorepo! Having worked at organisations with a monorepo and with many repos, I can confidently say that any team which is using multiple repos is very probably wrong — and the more repos, the more likely wrong they are. If you have more repos than team members, you are almost certainly wrong. You end up spending far more time managing cross-repo dependencies and changes than you would merging changes in a monorepo.
Multiple repos: not even once.
Disadvantage: it's not a package manager; it doesn't read all the dependencies from each library, resolve the duplicates and install them centrally. Instead of that, we only had submodules in the top repo (the app) - a bit like having a single requirements.txt file.
https://landing.google.com/sre/
[disclaimer: I work at Google]
https://cacm.acm.org/magazines/2016/7/204032-why-google-stor...
> At Google, we have found, with some investment, the monolithic model of source management can scale successfully to a codebase with more than one billion files, 35 million commits, and thousands of users around the globe.
so each commit adds on average over 28 new files?
Also, it depends on how much you squash - my latest 10 commits turned into just one when merging to master.
In the open source world, I have found some Unix distros use the same model. I know it's not as extreme, but the principle is quite similar. For example, in Nixpkgs all package definitions (which are actually code in the functional language Nix) are in the same repository and thus they can all depend on each other in a very easy and transparent way.
I remember the days when monorepo was the norm, and distributed version control was the weird, kooky idea. Mainstream programmers had knee-jerk notions that all managed environments were too slow.
For game development, monorepo is simpler. If one is using git, one needs to use some other software to turn the part of your repository for media into a monorepo, otherwise the asset files become a burden. (gitannex, for example)
I think this has more to do with how git handles diffs more than monorepo vs distribution. As you said, git lfs solves the issue with centralization but that's not the same as a monorepo. You can still split all your libraries out in such a system without issue.
On the other hand, I've seen the other sides of this in monorepos:
1. It's too easy to depend on code, so there is dependency bloat when something simpler would work just as well.
2. It's relatively hard to depend on things not in the repo, reinforcing not-invented-here culture.
And Google's approach, I'm sure, requires a bit of standardization and tooling investment. It's not clear to me that equivalent conformance and investment in monorepos and a package manager wouldn't work just as well.
It shouldn't, that's one of the big gains of a monorepo, is that there's only one version of everything. You don't need to version your dependencies within the repo, which means you only need to maintain one version of any external dependency.
With multi-repo environments, btw you still need to do that kind of dependency version management for certain upgrades. Essentially you can desync but you have to occasionally re-sync everything. I na monorepo environment you can just prevent desynchronization.
As for JavaScript, the hassle of importing the entire transitive closure of an NPM library you want into third party means it's much more attractive to go with NIH syndrome. I looked at importing ESLint, but it has something like 110 dependencies.
But if you don't, then a monorepo will generally slow you down because it will require coordinating changes across a much bigger group of people.
Monorepos are great for very small companies with a low communication overhead, and very large companies with the resources to build the tooling to make it work.
For everyone in between, I feel that small repos and microservices give the best developer velocity.
I've seen more than one company now that has had the same problem: how do they patch atomic cross-repo changes onto their multiple git repos? The reasons for this can vary, but the core problem is always that. As far as I see it, there are two solutions:
- Use a monorepo
- Create some external database that ties multiple hashes together for use in your ecosystem. This also requires re-inventing bisect on top of this database. I'm sure the intelligent people of HN can come up with the multitude of other tools you need to modify to make this work, but its not trivial.
If you're willing to manage that overhead somehow, that's fine, but I can't imagine its fun.
The thing is, there are already many mature well understood tools that you're probably already using in your organization whose goal is to tie multiple hashes together. Back in the day they were called package managers and had names like dpkg and rpm, and we'd use something like "pkg-config" to link to a specific one. These days we have docker and nix and a hundred language-specific dependency managers to solve every little variation on "tie these specific hashes together".
Monorepos still require tooling to manage effectively, but that tooling is at a disadvantage because it's not already being ubiquitously used. And unless you're bringing your entire dependency tree, recursively, all the way down to your OS, into the monorepo then you're still going to have to be dealing with those external dependencies one way or another.
And if you're not versioning the world, then it's really just an argument of the appropriate size of a unit of functionality in a multi-repo, and in that case "small enough to be well supported by existing tooling" isn't a terrible upper-bounds to pick for most people.
You're free to PR changes gradually, making sure things work a couple of repos at a time, until you eventually get everything. If you can tolerate temporary inconsistencies, it allows you to scale to infinity, essentially for free.
I don't see how you can do this any better in a multi-repo than a monorepo though, unless you mean to the extent of simultaneously having multiple versions of the same library in your transitive deps (and thus kind of kludgily sidestepping the diamond dependency issues). Would you mind elaborating?
>You're free to PR changes gradually
This is possible in a monorepo too, by much the same means, I'd expect: you define an adaptor that you slowly migrate everyone onto, deprecate the old thing, and then optionally remove the adaptor and deprecate it too. Am I missing something?
Leaf repos (projects nothing depends on, like apps) can do literally whatever they want at any time without affecting anyone else. Right there is a big win. Dependencies can have multiple versions and the leaf can depend on whichever version they want at any given time. You can do the same thing in a monorepo if everything is in independant folders, but that's just the worse of both worlds.
> you define an adaptor that you slowly migrate everyone onto
No no. The way we do it is: make breaking change, people upgrade to it whenever (with gentle pushes so that we're eventually all on the same version sooner than later). No adapter, no transient state within a service. Just upgrade repos one by one until you got them all, no magic involved. This assumes that your repos represent loosely coupled components (micro services or micro apps).
There longer the chain of deps is the worse this is. Even if everyone takes just couple days to upgrade, which ime is generous, your leaves end up being forced to wait weeks to upgrade in the worst case.
If the change is breaking for B, then yeah, you have to wait until B upgrades (or you can upgrade it yourself!). During that time, other projects that don't depend on B can go on using the new A.
The alternative is "stop the press, everyone is upgrading to the latest A NOOOOOOOOOW", which if the change is not automatable, might either be non-realistic, be pretty large in scope, or require you to never make drastic changes in A (which is tricky if A is a 3rd party you don't control). You can also just have these long-running transient state where somehow A is always compatible with everything no matter via compatibility wrappers.
It's tradeoffs. I like our world where we don't have to migrate everything at the same time all the time, even with drastic changes. Works quite well for us, with thousands of repos and 10s of millions of lines of code. We enjoy the flexibility. It makes certain cross-project efforts harder. That's the tradeoff.
I think the problem is that people misuse their version control as a package manager. If you want that type of behaviour, with semver and all to manage compatibility, just use a package manager to manage your dependencies, not a multirepo git contraption. You can still store your packages in one repo each if you like but at least you now get some control over interface compatibility which is a requirenent when you have multiple repos.
No. The whole point is that they don't have to think about it. Aside for micro services used machine to machine with breaking changes that cannot exist in parallel, the point is that they don't have to worry about this. Do whatever you want, other projects do whatever the hell they want, and we just slowly go toward consistency as it becomes convenient until everyone's the same. Repeat. I can upgrade a dependency with a breaking change to my repo. The other team isn't ready yet so they keep using the old version (which works) until they are.
What ends up happening in practice is that all these seemingly independent components have very strict—and sometimes even unspecified—dependencies between each other.
People create "base" packages that all other packages have a strict or loose dependency on, so when that package changes, it's a guessing game if it introduced breaking changes downstream.
With a monorepo, these dependencies are tracked and always visible, and integration testing between all dependent components becomes much easier.
It's very unlikely that you'll find truly decoupled code bases within an organization. It goes against the point of grouping people to work on a common goal to begin with.
Why? Again, the point is eventual consistency. If I make a breaking change, apps/services that are ready to migrate do so. The ones that aren't keep using the old version. Eventually, everyone's on the same page. The whole point is that it requires a lot less discipline.
If you need atomic cross repo changes, then you're doing multi-repo wrong. You need to have an upgrade path, so that you support both old and new in parallel, so you can upgrade one repo at a time.
You'd need that anyway during deployment.
The thing people need to remember is the whole labor force of a normal company is a rounding error at Google and Facebook.
Tooling isn't the issue here, you're basically arguing for kicking the can down the road and encouraging a buildup of technical debt.
When you write code, do you put everything into a single function, or do you use separate functions for different parts of the code? Multi-repo is just the extension of that.
Software quality is orthogonal to your development methods. You can have good quality multi-repos or bad ones, and you can have good quality mono-repos, or bad ones.
If you really want to take this analogy to its logical conclusion for separate repos, you're gonna have to do something more like compile every function as a dynamically loaded library. And then hope you notice a function signature change because the dynamic linker won't tell you.
Tools like Pants[1] or Lerna[2] solve a lot of the issues related to builds and dependency management.
What exactly are you missing _today_ that prevents your organization from adopting a monorepo approach?
[1]: https://www.pantsbuild.org/
[2]: https://lernajs.io/
Now, you might be referring to the fact that Facebook and Google have built up lots of tooling to help git/hg scale since their repos are too large and operations take too long. But that is not a problem you're going to have until you have at least a thousand engineers. At that size you need tooling for everything anyways.
If you look at monorepos from a package management standpoint, they are usually just high-level graphs of dependencies. Or, more accurately, sources of dependencies. How these are rolled up, shipped, and ultimately deployed is a function of the operational culture and business needs more than it is source control or even language choices, in my opinion. Business needs impact source control in any sufficiently complex, source-controlling org.
That's not to say that, for example, small companies benefit from monorepos, while large ones benefit from small repos and packages (or the opposite). I think the pros and cons are entirely decoupled from codebase size and complexity. In order to do one or the other well, you need the right business needs, operational parameters, engineering culture, and resources. So, I always find it interesting to read about Google or Microsoft leveraging one, the other, or both approaches with their own codebases.
I liken the monorepo vs small packages approach to be a little bit like rendering a frame on a CPU vs a GPU. Do you build/test/deploy each "frame" (iteration) as a top-down, more-or-less-discrete block of work, or can you parallelize it and "ship" multiple compatible streams at once?
I suggest it depends almost entirely on the problem space and the "hardware" (business needs), far more than it does the actual code or volume of code.
Perhaps Conway's law here applies here in a sense, i.e. any organization that manages source code will produce a source control management scheme that is representative of how they deploy to downstream consumers.
Once you switch to this model, you can do convenient things like landing straight from the review system (by a "Ship it" button), landing after some checks were successfully performed, and so on.
Incidentally, this is similar to how merging pull requests by rebasing works on GitHub.
When you 'hold the lock', you get to rebase and push, and 'release the lock', and if you've any sense the system will automatically ensure you haven't broken the build.
Of course, it's important not to waste time, as this serializes the commit process, as it were.
But having things split across many repos also takes a lot of tooling too. So really, for each approach you're looking at which portions of it are covered by the version control system itself, and where the gaps are that have to be plugged by auxiliary stuff. And of the auxiliary stuff, how much of it is standard enough to be things you can inherit pretty much off the shelf from your distro or other ecosystem vs. something you need to actually build and maintain in-house.
Other organizational priorities come into play too, like the importance of open source in your codebase, and what your relationships are to your upstreams, if you have them. It actually surprises me that there aren't more/better tools out there that help with synchronizing commits (or portions of commits) in and out of external standalone repos. The main patch-management tool I'm aware of is Debian's quilt, which pretty much just boils down to a handful of bash scripts and arcane conventions. Why isn't there more stuff in this space?
Now, our upstream doesn't _actually_ ship very much, and what they do ship is relatively slow moving. So we've had to extend the supplied tooling in various ways to truly meet our needs.
That said, I understand the worth of getting things done at the expense of rigor so I chalk this topic of discussion up to personal taste. It's akin to the dynamic vs static typing debate.
That's the theory, but in practice designing robust, future proof APIs has proven to be really hard in a lot of cases. You're then left with the option of supporting old APIs forever or migrating dependent code to new APIs, both of which are difficult in their own ways.
Upgrading other projects to support the new dependency version can come in a different commit. You only run into trouble when you've set up your projects to always use the latest version of their dependencies. That's a recipe for disaster.
I work for a fairly large company where a monorepo doesn’t make a whole lot of sense because each team runs several services that get released or patched independently. If you have a large product composed of many components that need to function together as a cohesive whole, go ahead, use a monorepo.
Let's say I want a small application with flask and angular.
I create a single repo for both flask and angular. I put everything flask in one sub folder called backbend and I put everything angular in another folder called frontend
WIP here: https://github.com/kusl/flaskexperiment or https://git.sr.ht/%7Ekus/flaskexperiment/
Now the problems are just starting: how do I set up ci for all my projects? Travis ci expects a single file at the root of the project and so does gitlab ci (hi sid, big fan)
I am sad I can't talk to the experts at Google about how they navigated these problems. I understand Google has many enemies that are constantly trying to exploit whatever but I still wish we could have a more open conversation here.
GitLab has a related issue to support "Several .gitlab-ci.yml for monorepos" (https://gitlab.com/gitlab-org/gitlab-ce/issues/18157).
If not, which of the "usual" VCS are best suited for monorepos? CVS? Subversion? Darcs? Bazaar? Mercurial? Git?
Should one use Mercurial simply because Facebook uses (and patches) it, or are the better choices for small-to-midsize organizations?
There is a lot of developer movement between the companies cited and it's not surprising that people take the practices they're familiar with to their new employer.
Companies incentivize people to deprecate feature X and replace it with feature Y and celebrate the win. Much harder without a monorepo.
The counter argument is open source, where the development follows the distributed model and the difficulty of syncing the monorepo with a custom build system. Figuring out a way to leverage the QA work distro people do in coming up with a consistent cut + patches would benefit everyone regardless of repo structure.
(BTW, I was slightly confused not to find a single date on this blog post, and the HTTP headers were also useless. But at least there are rough timestamps at the main site.)
This makes it sound much easier than it is. You still need to get approvals from all the teams whose code is touched. In the meantime, code may be changing out from under you. And the more code you touch, the more tests you need to run.
For anything but the most obviously risk-free changes (where a global owner can approve it), splitting up a large change into independent pieces and sending out a bunch of changelists in parallel will make more sense. There are tools to do that too.
I mean do we simple use 1 GIT repo with multiple folders for each Project/Library?
I think the benefits that massive corporations derive from monorepos demonstrates how massive corporations are a net negative to society. Imagine if, instead of an increasingly centralised and closed culture in technology, companies were small enough and interdependent enough that a decentralised model became a net positive.