2) Data model brings speed that was not possible before. git won VCS space because of sheer performance.
3) Scale just fine. Just windows have a really bad filesystem and Windows codebase is the pathologic case.
2) Data model brings speed that was not possible before. git won VCS space because of sheer performance.
3) Scale just fine. Just windows have a really bad filesystem and Windows codebase is the pathologic case.
* Pushing changes to a central repo requires including upstream commits. With 1 commit/s to that central repo, all developers are stuck in a loop until their push succeeds. It is a human spinlock with high contention.
* Some algorithms scale linearly with the number of server branches, such as pull-without-specifying-a-branch, which becomes too slow with 100K branches (a consequence of central repos).
* Some algorithms are linear with the number of files, like git status.
* Binary files don't compress nor deduplicate well, slowing pull and clone.
Those issues apply to fossil and Mercurial.
cf. https://docs.microsoft.com/en-us/azure/devops/learn/git/tech...
Can you provide a reference? I was searching a bit and only things I found was bugs in windows[1] for git lfs.
> You can call this "pathological" but this throws a lot of shade on monorepos without much critical examination of how or when they might be useful.
Windows codebase has 3.5 million files and its repo is 300GB in size. It is not normal. This is google or MS type of problem and not average git user. MS instead changing workflow decided to create GVFS[2]
[1] https://github.com/git-lfs/git-lfs/issues/2434 [2] https://blogs.msdn.microsoft.com/bharry/2017/05/24/the-large...
Apologies, I hastily mistyped, I meant 500 GB, not 5. (5 GB is about the size of my repository, which is not really so big at all and certainly something git can cope with on its own).
This series of articles should illustrate some of the issues that VFS for Git tries to address. ("GVFS" is now called "VFS for Git".)
https://docs.microsoft.com/en-us/azure/devops/learn/git/tech...
And this is a series of articles from an engineer who's been working on improving perf in large repositories in general, not strictly related to the Windows repository:
https://blogs.msdn.microsoft.com/devops/2018/06/25/superchar...
> Windows codebase has 3.5 million files and its repo is 300GB in size. It is not normal. This is google or MS type of problem and not average git user. MS instead changing workflow decided to create GVFS[2]
I didn't say it was normal. Indeed it's uncommon. I said it wasn't pathological.
That's 3x the source line count of Google's entire monorepo. [1]
So if you're using git for source code, 500GB is beyond pathological.
If you're using git for other purposes, then yes you might need something like Annex/LFS/GVFS.
[1] https://m-cacm.acm.org/magazines/2016/7/204032-why-google-st...
The problem is that the edge cases that come up have solutions which need to be looked up- not derived from understanding. And when you're scared of data loss, its a very frustrating situation
Can you specify a version control system which doesn't make it easy to lose uncommitted changes?
This makes it easy to compare, say, the state of the file now with the state of the save from 3 hours previous.
https://en.wikipedia.org/wiki/Versioning_file_system points out "Subversion has a feature called "autoversioning" where a WebDAV source with a subversion backend can be mounted as a file system on systems that support this kind of mount (Linux, Windows and others do) and saves to that file system generate new revisions on the revision control system."
Quoting http://svnbook.red-bean.com/en/1.4/svn.webdav.autoversioning... :\
> the use case for this feature can be incredibly appealing to administrators working with non-technical users: imagine an office of ordinary users running Microsoft Windows or Mac OS. Each user “mounts” the Subversion repository, which appears to be an ordinary network folder. They use the shared folder as they always do: open files, edit them, save them. Meanwhile, the server is automatically versioning everything. Any administrator (or knowledgeable user) can still use a Subversion client to search history and retrieve older versions of data. ...
> however, understand what you're getting into. WebDAV clients tend to do many write requests, resulting in a huge number of automatically committed revisions. For example, when saving data, many clients will do a PUT of a 0-byte file (as a way of reserving a name) followed by another PUT with the real file data. The single file-write results in two separate commits. Also consider that many applications auto-save every few minutes, resulting in even more commits.
It adds that Clearcase supported a similar feature.
I have never used that combination.