I contribute to git.git, and it would be interesting to know if there's inherent issues stopping you from doing that, or if it's implementation problems in some cases (e.g. missing plumbing commands or features). There's definitely interest from upstream in reviewing patches / helping if there's missing or inadequate plumbing.
You'd get upstream features for free as they come along. E.g. presumably you haven't implemented the new MIDX format, but that speeds up pack file access by a lot for some use-cases, and presumably the boring bits of low-level git operations aren't much of a selling point in and of themselves.
Aside from whether you'd use "git" itself, such a trick of using a slave process you'd talk to over IPC of some sort would cover some of the issues you wrote about, e.g. issue with sharing global state with libraries like Breakpad.
In terms of getting the right data, one example is that we need to know the full set of non-ignored sub-directories in the working directory, so we can watch them for changes. It's easy enough to generate this ourselves as we calculate the status output, but I don't believe that git will emit it.
In terms of performance, we rely on being able to read objects efficiently. For example, to show a commit, we can't just use the output of "git diff", as we need the full file contents to be able to calculate syntax highlighting correctly. You could go a long way with "git cat-file --batch", but there are plenty of contexts where you can't practically batch requests, and process creation costs + the lack of caching across requests (which can be quite significantly due to the delta encoding of objects) would be quite significant.
There's going to be cases where it sucks, e.g. what you point out with wanting both raw blobs and their diffs, you'd need to do that in two plumbing commands now.
But just on that example: Having poked at some of the diff code recently I can tell you there's no big technical hurdle to just exposing that sort of thing. I.e. spewing out machine-readable raw blobs and their diffs, it just happens not to be exposed now.
I think what a program like Sublime Merge would want/need short of C API access (which is unlikely to happen) is a git version of an open-ended "plumbing" IPC protocol of the sort that Common Lisp VMs tend to expose. I.e. being able to have one (or few) "git command-server" processes spawned, and ask them questions like "look up this blob" or "diff these two blobs" (where the previous blob lookup would be cached).
Obviously patching/coordinating/upstreaming those sorts of changes is going to take work, but so is duplicating and keeping up-to-date with the diff, pack, status etc. code.
I'm not trying to tell you what to do, just saying that the git project is definitely friendly to "we're a commercial product and need this missing plumbing for our editor" (unlike say, GCC).
The plumbing that's there now is mostly in the state it's in because it's what git itself needed in the past when it was more of a collection of shellscripts, as well as being biased towards what git server operators like GitHub needed (because they sent more patches), which is why plumbing for say batch blob operations tends to be better than the one for "status".
In any case it would be very interesting to have some post about the sort of read-only operations Sublime Merge is doing with its own custom git code.
(Also, IPC and fork+exec has overhead that mmap or thread in the same program does not.)
The libgit2 code is GPL with a linking exception, so you can use it (unlike "git" itself) as a C library in a proprietary commercial product.
> IPC and fork+exec has overhead[...]
The "git cat-file --batch" command is something you'd invoke once, and then as your program runs you keep feeding it SHA-1s on stdin and it spews out their content on stdout. So even on Windows the overhead of that should be fine.
It's clear from ben-schaaf's other comments (which I read later) that one concern was the simplicity of downstream APIs being able to read the data using a normal C variable.
But that just leaves more questions. People in this thread are mentioning pack files, assuming that a multi-GB "git object" must be in a pack, but I notice the original post doesn't say anything about it.
If they're reading packs with this they'll need to parse it, resolve deltas etc. So likely the code that deals with the mmap()'d variable is small in any sane codebase (they're surely not doing delta resolution repeatedly all over the place...).
If they're very large loose objects those will most likely be zlib compressed, so wouldn't this need to go through some intermediary API layer anyway? I guess if SM itself is adding them it could add them uncompressed.
Since ben-schaaf mentioned this not being about performance, but about saving memory I thought this might be something like wanting to extract a small part of a 1GB object from git for display. That seems like a thing an editor might want to do.
In that case "git cat-file --batch" would suck, but not for some intrinsic reason. An API could be added that could take the start/end of an object to print out.
You might also have lawyers who are lazy about it and don't want to deal with the liability, "we heard Apple banned GPL code..." or "the FSF sued Cisco...".
But there's no license reason for why you can't use that GPL code in some way, and everyone from Google with Android to Oracle with Oracle Linux and their DB bundles GPL code that's directly used by some accompanying proprietary piece of software.
But what I was more going for is that there's also a non-legal aspects to it that go beyond the license, which is that some maintainers of free software are actively hostile to their software being used as a smaller component in some proprietary product.
The GCC project is probably the most famous example of this, I think this has changed somewhat in recent years with LLVM+Clang, but they used to jealously guard things like their AST format. So e.g. someone with a proprietary editor (or Emacs for that matter...) could never hope to use GCC for spewing out parsing information for some C code.
I think it's fair to say that the Git project isn't like that. If someone maintaining proprietary software needs some plumbing interface to hook their stuff up and is willing to submit patches it'll be received as well as any other change (subject to review, maintenance & backwards-compatibility concerns etc.). If they find it useful it's likely that other people will too...
The effects of oh-noes-my-file-is-gone can be somewhat mitigated by using the heuristics built into NSData (instead of using mmap directly).
For example, you call NSData’s `dataWithContentsOfFile:options:error:` with the `NSDataReadingMappedIfSafe` option [1]. The framework will then transparently mmap the file unless it believes there’s an elevated risk of the file going away.
Apple doesn’t disclose how NSData exactly makes that decision; however, I’ve found a few reports that say it uses mmap internally when the file is on the root filesystem, and fall back on an in-memory copy otherwise.
It’s a rather dumb heuristics though, and may not solve the issue entirely.
[1] https://developer.apple.com/documentation/foundation/nsdatar...
"Memory mapped files work by mapping the full file into a virtual address space and then using page faults to determine which chunks to load into physical memory. In essence it allows you to access the file as if you had read the whole thing into memory, without actually doing so."
I feel like this could be done in c++ directly, by maintaining an internal cache for each file that keeps track of which parts of the file are loaded and uses read() to load chunks on demand. Error handling would be a lot simpler (no signals, just a failed read()) and there would be less OS-specific code.It totally would have been simpler overall, but each incremental step we made was significantly less work than the refactoring required for pread.
Not necessarily. With O_DIRECT, pread() doesn't put pages into page cache: it just DMAs them directly into your process. Using O_DIRECT and the process-private caching we've been discussing, sophisticated programs (like databases) can (and do!) implement their own "page cache" systems. And because databases have access pattern information that the generic kernel VM subsystem doesn't, such a database can frequently do a better job doing this caching on its own.
Question.
In 10 years will you be saying this about the next incremental problem that you run into? If you think this likely, then the next incremental problem is an excuse to do it right.
Unless you actually need to read the file multiple times (compared to looking at the parsed in-memory data multiple times), this should be fast enough.
Showing the scope of change within the editor is a rather nice touch. Visualization of complexity, if you will.
Are you sure the mechanism used by thread_local is safe to use in a signal handler?
Relevant bug from Rust: "TLS accesses aren't async-signal-safe", https://github.com/rust-lang/rust/issues/43146
I'm thinking even if it works on certain OSes, it's not guaranteed, because a signal handler's context is not a thread context - or is it?
E.g. the thread_local mechanism might depend on compile-time options (affecting how thread_local, whether it allocates memory on demand, and how it relocates the memory block when loading a shared library), whether it's main program or an -fPIC shared library, and the type of thread library (different ways of implementing pthreads).
Is there any chance that merging and rebasing via drag-n-drop is coming to Sublime Merge? For me that's the one big feature which keeps me from switching from Gitkraken to Sublime Merge.
> The signalfd mechanism can't be used to receive signals that are synchronously generated, such as the SIGSEGV
> Limitations
> The signalfd mechanism can't be used to receive signals that are synchronously generated, such as the SIGSEGV signal
This is because synchronous signals are fired at the thread that caused them, and signalfd read() calls can't be used to read signals fired at other threads.
This needs a stronger justification. mmap allows reading and writing large data structures without copying, which can be a huge benefit depending on the use case.
They are saying that if they somehow knew up front what the performance gains would be, and what the cost in bugs and complexity would be, they wouldn’t have used mmap at all.
You said above that “mmap allows reading and writing large data structures without copying, which can be a huge benefit depending on the use case.”
Yes, of course that’s true, and the Sublime Text authors are clearly well aware it’s true. That’s why they decided to use mmap in the first place. They agreed with you.
This is them reporting, with hindsight, that for their use case mmap introduced a lot of tricky bugs that required complex platform-dependent fixes, and that the performance gains were real but modest. Therefore, in hindsight, it probably wasn’t a good choice.
Which part are you arguing with?
I’m pretty sure most programmers are capable of writing a signal handler that sets a flag (volatile sig_atomic_t) for “parsing failed,” and a loop that checks that flag in addition to checking whether the loop is finished for other reasons. Signal handling doesn’t have to be complicated.
I think you’ve mixed up synchronous signals - SEGV, BUS, ABRT - with asynchronous signals like QUIT, INT, USR1. Asynchronous signals can be handled easily with a flag and a loop as you mentioned; synchronous signals are much trickier and much more complicated.
While it's true that memory bandwidth is sometimes a limiting factor in performance, I've found that much more frequently people overestimate the cost of memory operations and don't check their estimates against benchmarks.