Scaling Git’s garbage collection
github.blog
github.blog
That actually seems kinda small.
Git’s lack of good support for large files means there’s probably an exabyte of data that, imho, should be source control but isn’t.
They do exactly what you're looking for! They provide both a DVC remote to push and pull DVC-tracked objects, as well as a UI showing them integrated alongside git-tracked files in the repository view.
(Disclaimer: I work at DagsHub)
(I was in fact asked a long time ago in an interview to estimate how much disk was needed to store Google's search index.)
Unlike my sibling saying it's "pull out of thin air", I think this is an opportunity to actually do the opposite: Ask questions back!
I have no idea how to guess that. I don't know what kind of information Google stores in it's "index". What exactly do they mean by "index" even? I need more input. I can ask the interviewer questions and start a discussion!
Of course it's possible that the interviewer really just wanted to be a smart-ass and expected me to "just have an answer" and else fail me. But then I don't want to work there anyway. Having an actual discussion, as if we were just working together and had to 'solve' whatever problem at hand is great.
My entire point was basically that it actually isn't about assumptions at all. If you just assume when I ask that question this tells me you're out as a developer I want to hire. Nobody can build proper software by just assuming the first thing that comes to mind and bullshitting.
You have to show that you can take a totally ambiguous and open question and show me that you can systematically try to find out as much information as you can in a reasonable amount of time to verify your assumptions or at least make them less assuming. Of course (a d this can happen when solving a real problem too) it might be that you cannot accurately measure some input you need and you will have to make assumptions. But you will want to try and you will want to have divided your problem into lots of smaller parts such that the part where you still have to assume is only one small part and you want to be very explicit that this part is still an assumption only which you will have to verify or improve upon as you actually implement something. You do have to start at some point, otherwise it's analysis paralysis.
And no this has nothing to do with a PM. They are notoriously bad at the above approach and instead will just assume the first best thing and bullshit their way through with that for as long as they don't get caught (no /s there but my experience with 80+% of PMs)
1 googol!
Chromium is also well over the size limit.
That makes sense for GitHub, because what they really care about is the hard drive space you use up, and a repo containing a commit they are already storing for some large project is no additional disk space.
There is not a technical limit of 4GB per repo on github.com. maybe a private repo size limit if you are not a paying customer, but it is not a technical limit of the platform, I assure you.
I am surprised they didn't directly use time_t, so that they wouldn't have to deal with this (since some platforms have already gone to 64 bit time_t)
Since uint32_t is unsigned, wouldn't it be the Y2106 problem instead?
> I am surprised they didn't directly use time_t, so that they wouldn't have to deal with this (since some platforms have already gone to 64 bit time_t)
You mentioned the problem yourself without noticing: some platforms have gone to 64-bit time_t, but others haven't. This is a file format, which can be shared by multiple platforms, so it cannot use types which change size depending on the platform.
But for this use case it's not really an issue though. FTA it sounded like they always write the mtime as now. It's unlikely they wouldn't GC the repo in 68 years to make wraparound an issue.
https://github.com/git/git/blob/e188ec3a735ae52a0d0d3c22f9df...
https://github.com/git/git/blob/e188ec3a735ae52a0d0d3c22f9df...
What would that something be for example?
If I recall correctly it has to do with the way they cache git. If you request shallow clones they have to run a git process each time to make a custom download with only the data you need. Full clones get served from a cache.
> If I recall correctly it has to do with the way they cache git.
It was not actually the shallow clone, the shallow clone is fine.
The problem was shallow fetches afterwards, as it made computing the minimum set of changes to fetch much harder for git (during the fetch operation the client and server actually try to negotiate a minimum set of changes, and the server creates an ad-hoc pack for this).
Not only that, but they'd hit a git edge case which ended up causing disproportionate CPU usage and converting the shallow clones to near-full clones but very, very inefficiently.
This issue was compounded by the very git-unfriendly layout of the repository: one of the directories had >16000 subdirectories, something Git's tree-processing code apparently was not much tested for, leading to significant inefficiencies.
[1] https://discourse.pijul.org/t/signup-on-pijul-nest-gives-for...
* forgive the caps, I know I have sinned but I do wanna be able to use this VCS.
When you have to write a web browser, you do have to fix C++ before you can actually start. I'm digging a bit deeper by avoiding blackouts so my French servers keep running (the Nest has two other locations), then I'll fix the server code.
"What real world situation does git barf at that Pijul would handle better". I'm sure there is one, but I have yet to see it.
CRDTs aren't common in real-world distributed applications because they are hard to design. But when the problem is important enough (like in this case) I think the approach is worth it.
One way for Git to do what you said would be to import the repos in Pijul, commit by commit since the last common ancestor, instead of doing a 3-way merge. Technically feasible, but the performance would be terrible.
- Cherry-picking in Pijul actually works.
- Conflicts: you don't need rerere, conflicts don't come back once you've solved them. Conflicts are the most confusing situations, this is where you need a good tool the most.
- So-called "bad merges", where a merge or a rebase goes completely wrong and shuffles your lines around. Git users rarely only call the lack of associativity a "bad merge" and don't look further, but using 3-way merge is the wrong way to merge things, because there isn't enough information to merge correctly in all cases.
- Depending on your pace of work, you might find yourself working on different things at the same time. Pijul allows you to do that without worrying too much about how you'll eventually organise your work. If I worked with Git for example, I'd spend a lot of time organising my branches, they'd never be right, and I'd spend a lot of times rebasing afterwards. Pijul frees me from that work.
- Large files, but Git doesn't treat them well for historical reasons, not for reasons inherent to its design (unlike the other points).