GitMounter: A FUSE filesystem for Git repositories
belkadan.com
belkadan.com
> A file is further divided into blocks with variable lengths. We use Content Defined Chunking algorithm to divide file into blocks.
> This mechanism makes it possible to deduplicate data between different versions of frequently updated files, improving storage efficiency. It also enables transferring data to/from multiple servers in parallel.
I use it on old PC without issue. Drawback: since the files are not stored in clear, in case of data corruption of the Seafile repositories, I need backup (never happened to me).
https://news.ycombinator.com/item?id=37573679
This lets you:
- Mount a Git repo + branch state to a folder: `git xet mount https://xethub.com/XetHub/LAION-400M.git --prefetch 0`
- Analyze parquet files using DuckDb:
`import duckdb`
`duckdb.query("select COUNT() from 'data/.parquet'")`
Regardless, if you are asking whether you can check out commit that doesn't have a branch pointing at it, yes you can.
You can have a work tree for every commit in your repository if you like.
You might have to wait a moment at startup for it to make a list of commits, but after that git is very well designed for this sort of browsing.
Fair, it'll likely break most tools written in the last 10-15 years.
> (e.g., commits) that require a lot of work to find.
That's solvable with a cache. I'm surprised git doesn't seem to have one for those, at least going by how long it takes to generate a full "shorthand log" on a large repo.
PS, the single biggest problem with million-entry directories is the propensity of tools like ls(1) to want to sort the darned things, so one has to remember to use `ls -f`, or to at least use the C locale to get memset() collations instead of much much slower Unicode collations. Another problem is that the POSIX stat(2) family of functions combine reading metadata that could come from just the directory (e.g., a file's inode number) (and which contents has already been read) with reading metadata that requires [possibly much] extra I/O to get, so if you're doing `ls -l` on a million-entry directory you might as well go on vacation (but make sure to send the output to some file, cause your scrollback buffer just won't do).
Also you just have to use ls -f. Edit: Oh, you even mention this yourself in another comment. Serious non-problem then. If you're worried about the filesystem side, it could also load the list of commits gradually.
This is also going to be a great prank on the next engineer who's never gonna figure why he can't see any of the files the build jobs are seeing in his shell.
What did they "patent"? Object databases? Versioning files? Mounting file-systems?
Anyway, that are at most some super weird software patents; so you don't have to care outside of the US, I guess.
But it's 90s IBM enterprise business model to the core and the rest of the Rational product suite sucks.
But there are good uses for mountable "git like" repos. For example for backup systems.