Why SQLite Does Not Use Git
sqlite.org
sqlite.org
It's always a bit frustrating when working with a team because everyone understands a different part of git and has slightly different ideas of how things should be done. I still routinely have to explain to others what a rebase is and others have to routinely explain to me what a blob really is.
In a team of the most moderate size, teaching and learning git from each other is a regular task.
People say git is simple underneath, and if you just learn its internal model, you can ignore its complex default UI. I disagree. Even just learning its internal model leads to surprises all the time, like the blobs that I keep forgetting why aren't they just called files.
What git needs is a chair lift up that hill. A way to easily get people there. But I have no idea what that would look like. Lots of people try, few do very well at it.
https://git-scm.com/book/en/v2/Git-Internals-Plumbing-and-Po...
They mix multiple only slightly related commands in one.
In fact, one of the things in common among the ORMs that have left a bad taste in my mouth is that they all tried to abstract away SQL without leaking enough of it.
Picking the ideal interface to abstract is critically important (and very hard).
In the case of ORMs, available solutions abstract the schema (tables, rows, fields), the objects, or use templates. My solution abstracted JDBC/ODBC. The only leak in my abstraction was missing metadata, which I was able to plug (with much effort!).
My notions for interfaces, modularity, abstractions are mostly informed by the book "Design Rules: The Power of Modularity". http://a.co/hXOGJq1
So given a mental model is inevitable, seems reasonable that that model should be the actual model.
I think the longevity of SQL has proved there's value is non-leaky abstracted interfaces.
How is sql non-leaky? To be proficient with sql you have to understand how results are stored on disk, how indexes work, how joins work, etc. To debug and improve them you need to look at the query plan which is the database exposing it's inner workings to you.
You have to know about the abstractions an sql server sits on as well. Why is it faster if it's on an SSD instead of an HDD? Why does the data dissapear if it's an in memory DB?
No, you don’t. As far as I know, the data is stored in discrete little boxes and indexes are a separate stack of sorted little boxes connected to the main boxes by spaghetti. This is the abstraction, it works, and I don’t need to know about btrees, blocksizes, how locks are implemented, or anything else to grok a database.
Have you created an index? Was it clustered or non-clustered? That's not a black box, that's you giving implementation details to the database.
There’s no question that knowing more will get you more, but I think for the question of “when will things go sideways and I need to understand internals to save myself”, one would be able to use a relational database with success longer than git, getting by on abstractions alone. Running a high-performance installation of either is really outside the scope of the original point.
Yes, most of us will have to do both at some point, but they can be thought of as discrete skills.
If, on the other hand, you are working to create said history (and devise/use an advanced workflow for that), it's very helpful if you understand the underlying concepts. Which also goes for designing database layouts - someone who doesn't understand the basics of the optimizer will inevitably run into performance problems, just as someone who doesn't understand Git's inner workings will inevitably bork the repository.
You may need to go deeper and understand the underline model if you want performance but sticking to normal form can make unnecessary for a lot of people a lot of the time.
You can have a useful separation of work between a developer understanding/using sql and a DBA doing the DDL part and the optimization when needed.
That is, I am not proficient with relational databases, and I can handwave why an SDD is faster, and why data may disappear from an in-memory DB.
But I couldn't do an outer join without help. Nor do I know when I would want to do one.
Bob Martin wrote the essay at http://blog.cleancoder.com/uncle-bob/2017/12/09/Dbtails.html , in which he writes:
> Relational databases abstract away the physical nature of the disk, just as file systems do; but instead of storing informal arrays of bytes, relational databases provide access to sets of fixed sized records.
This isn't true. SQLite does not use fixed size records.
This suggests to me that a lot of people who consider themselves proficient with SQL don't know how the results are stored on disk, nor the difference between the SQL model and the actual implementation details, making them not proficient under your definition.
Because you know that information for other reasons as most people would. Just because the information is gained for other reasons does not make it irrelevant when using a database though.
> This isn't true. SQLite does not use fixed size records.
It's actually true of most/all modern databases these days. The point isn't knowing the exact structure the database uses to store it's information (even though it can be useful) but knowing how efficiently it can find the information for any given request. Knowing when a database is doing an index lookup or a full table scan is very important and I wouldn't consider someone that can't make a reasonable guess to be proficient in sql. Many of these details are even exposed in the sql, when you create an index and decide if it's clustered or non-clustered your giving the database specific directions about how the data will be physically stored.
The fact that you need to know anything about how they do their work internally to be reasonably competent at using them makes them a leaky abstraction.
Also, SQL has well-established processes and formalisms to design schemas which generally result in solid performance by themselves. That's what RDBMS are around for, after all: enabling efficient and consistent record-oriented data manipulation. This is quite difficult to do correctly in reality; for example, if you write your own transaction mechanism for disk/solid-state storage, you are going to do it wrong. This is genuinely difficult stuff.
There is a ton of internals that SQL abstracts so well that very few DB programmers know or (have to) care about them. Things like commit and rollback protocols, checkpointing, on-disk layouts, I/O scheduling, page allocation strategies, caching etc.
Certainly. My comment, however, concerned what you meant by 'proficient', and not simple use.
You used the qualifier "all modern databases". Was that meant to imply that SQLite is not a modern database?
My point remains that there are many people who are proficient in SQL, and would do very well with SQLite, even without knowing the on-disk format.
That is why I disagree with your use of the term "proficient".
Same exact thing above applies to so many things in software development, from IDEs, to code editors (Vim/Emacs/Sublime/etc), to programming languages, to deploy tools, the list goes on. There’s a reason software development is classified as skilled labor and not a low end job generally. You’re expected to have knowledge of, or be willing to learn a lot, to do your job.
I recall the time when mp3 was to demanding for many CPUs, so you had to convert to non-compressed formats. Today you do need to know that downloading non-compressed audio will cost you a lot of network traffic. Once performance is a concern, all abstractions have to be discarded.
That said, I do feel some "porcelain" git commands are poorly named and operate inconsistently -- compared to the plumbing of the acyclic graph concepts which is good but limited.
Half the time, when I know what I want, I both keep forgetting git flags and sub-commands - and struggle to find them in the man pages.
Like the fine:
git diff --name-only # i list only the files that are changed. But I'm not --list-files.
git diff --stat git rev-list --all --parents | grep "^.\{40\}.*<PARENT_SHA1>.*" | awk '{print $1}'
whereas in fossil, you use fossil timeline after <COMMIT>
I mean, one of these looks just a little more straightforward than the other, doesn't it?Also, a cursory test in a local git repo just now showed that command seems to print out only immediate descendants--i.e., unless that commit is the start of a branch, it's only going to tell you the single commit that comes immediately after it, not the timeline of activity that fossil will--and all it gives you is the hash of those commit(s), with no other information.
I use git myself, not fossil, but if this is something you really want in your workflow, fossil is a pretty clear win.
How many other ways of looking at commits or trees are there, that are hard in git but impossible in fossil because the author didn’t feel like it?
You could alternatively use:
git log --graph git log <COMMIT>.. git log --all --ancestry-path ^<COMMIT>I don't understand HN's hardon for hating Git.
There are about 5/6 fundamental operations you do in git/hg. If that's too much then again, there's not an abstraction that is going to help you out.
> git/hg
Mercurial was a great solution to the same problem that Git set out to tackle, virtually free of Git's foibles. The tradeoff was a few minor foibles of its own, but a much better tool. It's a fucking shame that Git managed to suck all the air out of the room, and we're left with a far, far worse industry standard.
I just said you can't give specifics on what to change, because there isn't much too change.
>And you act as if source control were invented with Git
No I'm not?
>and we're left with a far, far worse industry standard.
Yeah, we definitely should have gone with the system that can't do partial checkouts correctly or even roll things back. Branching name conflicts across remote repositories and bookmark fun! Git won for a reason, because it's good and sane at what it does.
The master can do no wrong.
Lacked a few dubious features such as merging multiple branches at the same time too.
It has improved but git is still noticeably more efficient with large repositories. (Almost straight comparison is any operation on Firefox repository vs its git port.)
Those dubious features are so relevant to daily work that I didn't even knew they existed.
Instead Mercurial uses additional cache file which instead is slower on Linux with big repos. But happens to be faster in Windows.
And the octopus merge is used by kernel maintainers sometimes if not quite a lot. That feature is impossible to add in Mercurial as it does not allow more than two commit parents.
No it doesn't? People use octopus merges all the time, every single day.
To emphasize that even more: Try to explain the concept of an ML-style sum type (i.e. a discriminated union in F#) to someone who only knows languages with C++-based type systems. You'll have a hard time to even explain why this is a good idea, because they will try to map it to the features they know (i.e. enums and/or inheritance hierarchies), and fail to get the upsides.
No, Mercurial's design is fundamentally inferior to Git, and practically the entire history of Mercurial development is trying to catch up to what Git did right from the start. For example having ridiculous "permanent" branches -> somebody makes "bookmarks" plugin to imitate Git's lightweight branches -> now there are two ways to branch, which is confusing. No way to stash -> somebody writes a shelve plugin -> need to enable plugin for this basic functionality instead of being proper part of VCS. Editing local history is hard -> Mercurial Queues plugin -> it's still hard -> now I think they have something like "phases". In Git all of this was easy from the start.
Another simple thing. How to get the commit id of the current revision. Let's search stack overflow:
https://stackoverflow.com/questions/2485651/print-current-me...
The top answer is `hg id -i`.
$ hg id -i
adc56745e928
The problem is, this answer is wrong! This simple command can execute for hours on a large enough repository, and requires write privileges to the repository! Moreover, it returns only a part of the hash. There's literally no option to display the full hash.The "correct" answer is `hg parent --template '{node}'`. Except `hg parent` is apparently deprecated, so the actual correct way is some `hg log` invocation with a lot of arguments.
Also, on the git/hg debate, I feel I've had problems (like the stash your modification and redownload everything) more often with git that hg. I mean perhaps it tells something about my capability to understand a directed acyclic graph, but hg seems less brittle when I'm using it.
What I don't like in git is the loss of history associated with squashing commits, I would prefer having a 'summary' that would keep the full history but by default would ne used like a single commit.
The DAG is already powerful enough to handle both the complicated details and the top-level summaries, it's just dumb that the UIs don't default to smarter displays.
(I find git stash essential given that `git add --interactive` is a painful UX compared to darcs and git doesn't have anything near darcs' smarts for merges when pulling/merging branches. Obviously, your mileage will vary.)
That seems like you made an assertion as well. I think there are counter-examples.
For example, the point of gitless is (quoting http://gitless.com/ ):
> Many people complain that Git is hard to use. We think the problem lies deeper than the user interface, in the concepts underlying Git. Gitless is an experiment to see what happens if you put a simple veneer on an app that changes the underlying concepts
Some commentary is at https://blog.acolyer.org/2016/10/24/whats-wrong-with-git-a-c... .
Many HN discussions as well, including https://news.ycombinator.com/item?id=6927485 .
I think that's an exaggeration. For example, Darcs and Pijul aren't based around a "graph of commits" like Git is, they use sets of inter-dependent patches instead. I'm sure there are other useful ways to model DVCS too.
Whilst this is mostly irrelevant for Git users, you mentioned Mercurial so I thought I'd chime in :)
> The only thing Git can really fix is changing it's command flags to be consistent across aliases/internal commands.
I mostly agree with this: Git is widespread enough that it should mostly be kept stable; anything too drastic should be done in a separate project, either an "overlay", or a separate (possibly Git-compatible) DVCS.
I said graph, I didn't say which graph. Both systems still use graphs. And still a graph you have to understand how to edit with each tool. The abstraction is still the same, and if you have problems with Git, you're going to have problems with either of those tools as well. The abstraction is not the problem, it's the developers inability to conceptualize the model in their head.
Where is the exaggeration?
You said "that graph" which, in context, I took to mean the git graph.
> Both systems still use graphs
True
> The abstraction is still the same
Not at all, since those graphs mean different things. Each makes some things easier and some things harder. For example, time is easy in git ("what did this look like last week?"). Changes are easy in Darcs ("does this conflict with that?"). Both tools allow the same sorts of things, but some are more natural than others. I think it's easy enough to use either as long as we think in its terms; learning to think in those terms may be hard. For git in particular, I think the CLI terminology doesn't help with that (e.g. "checkout").
> if you have problems with Git, you're going to have problems with either of those tools as well
Not necessarily. As a simple example, some git operations "replay" a sequence of commits (e.g. cherrypicking). I've often had sequences which introduce something then later remove it (bugs, workarounds, stubs, etc.). If there's a merge conflict during the "replay", I'll have to spend time manually reintroducing those useless changes, just so i can resume the "replay" which will remove them again.
From what I understand, in Darcs such changes would "cancel out" and not appear in the diff that we end up applying.
> Where is the exaggeration?
The idea that "uses a graph" implies "equally hard to use". The underlying datastructure != the abstraction; the semantics is much more important.
For example, the forward/back buttons of a browser can be implemented as a linked list; blockchains are also linked lists, but that doesn't mean that they're both the same abstraction, or that understanding each takes the same level of knowledge/experience/etc.
What I'm getting at is that if you don't understand what the graph entails, and what you need to do the graph, any system is going to be "hard to use." This idea that things should immediately make sense without understanding what you need to do or even what you're asking the system to do, is just silly.
I've never seen someone who understands git, darcs, mercurial, pijul, etc go "I totally understand how this data is being stored but it's just so hard to use!" I don't think that can be the case, because any of the graphs those applications choose to use have some shared cross section of operations:
* add
* remove
* merge
* reorder
* push
* pull
I see people confused about the above, because they don't understand what they're really asking the system to do. I don't think any abstraction is ever going to solve that.
Git does have a problem with its command line (or at least how consistent and ambiguous it can sometimes be), but you really should get past it after a week or two of using it. The rest is on you. If you know what you want/need to do getting past the CLI isn't hard. People struggle with the former and so they think the latter is what's stopping them.
I posted a bald statement. He replied directly with snide remarks and fallacies. Look at the timestamps and edits. I have every right to be annoyed and make it known that I am annoyed in my posts when the community refused to consistently adhere to guidelines.
Enforce guidelines that keep discussions rational, not because people don't want to be accosted in public for their misleading, emotionally bloated statements.
They are currently not being applied. Is it fair for me to point out how inconsistently the posts are being treated?
https://news.ycombinator.com/newsguidelines.html:
>Don't say things you wouldn't say face-to-face. Don't be snarky.
"Every day humans make me again realize that I love my dogs, and respect my dogs, more than humans. There are exceptions but they are few and far between." [2]
>Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize.
"Which sort of doesn't matter since everyone thinks GitHub is source management." [1]
>Please don't post shallow dismissals
"You all lost out on "the most sane and powerful" as a result." [1]
"Calling it a sane and powerful source control tool is just not supported by the facts, calling "the most ..." is laughable." [1]
"Calling Git sane just makes it clear that you haven't used a sane source management system." [1]
"Lots of people are too busy/whatever to know what they are missing, maybe that's you. It's not me" [3]
>When disagreeing, please reply to the argument instead of calling names. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."
"Arguing with some random dude who thinks he knows more than me is not really fun." [3]
"Dude, troll much?" [4]
[1] https://news.ycombinator.com/item?id=16806588
[2] https://news.ycombinator.com/item?id=16807652
[3] https://news.ycombinator.com/item?id=16806877
[4] https://news.ycombinator.com/item?id=16807763
At least I had the decency to correct dishonest statements with vanilla citations in my posts, and they still got ignored.
I didn't insinuate that people are worth less than pets that I bought and own (who can't even choose who to be dependent on) because they don't agree with my perspectives over a piece of software. In what context would this be an acceptable statement to make face to face or in a public setting and you go "well, you know, it's kind of okay to say!"
I'm exceedingly interested in where I crossed that line in a considerable manner because that's one distant line to cross. Next time someone says something I perceive to be incorrect, or they get on my nerves for continually disagreeing with me, I'll be sure to tell them my dog is worth more than them since that's actively being allowed and has a precedent of moderator support.
And for the record, my tone is probably "abrasive" in this post because the above actions and outright blind eye towards outright lies and uncalled for statements is aggravating. I have a feeling you're not doing anything just because of who he is, and not because what he is saying is warranted or even accurate (it's definitely not, as I demonstrated across several different posts).
I've archived this thread so people are free to review my actions and moderator actions at a later date: https://web.archive.org/web/20180411062201/https://news.ycom...
I've said my piece.
My current team are mostly controls engineers, working on PLCs. But the software we're now working with has its configurations tracked in git. These aren't dumb people, they're quite talented, but their education wasn't in CS, and "directed acyclic graph" is not a thing they have a mental model for.
So, my advice is to try to write some instructions for yourself for all the common cases you might run into during your work. It will not only help you realise what you actually need from git, but also will serve as a good cheat-sheet.
https://leanpub.com/learngitthehardway
I start with simple examples and work up from there. It's based on training I've conducted at various companies, and avoids talk of Merkle trees or DAG.
This was a big problem that bugged me too, so for every team I've worked with I've created a few scripts for the team's most common version control operations.
Most devs, including me, are pretty lazy so they'd all rather run this script than go to Stack Overflow to figure out git arcania.
This helps standardize conventions too: Feature branches/linear DAGs/topic branches/dev branches/prod branches/whatever weird thing a team does they all just do that using the script so it's standardized.
I just know 5 basic commands; pull, push, commit, branch, and merge. Never ran into any issues. People who run into issues are usually editing git log or doing something fancy with “advanced” commands. I have a feeling that these people get into trouble with git cause they issue commands without really knowing what those commands do or even what they want to achieve.
What's the alternative? Managing all dependencies by an external dependency manager does not exactly reduce complexity (if you're not within a closed ecosystem like Java + Maven that has a mature, de-facto standard dependency manager; npm might count, too).
It's absolutely not feasible for C++ projects; all projects that do this have horrible hacks upon hacks to fetch and mangle data and usually require gratuitous "make clean"s to untangle.
My mental model is basically that they're separate repos, and the main repo has a pointer to a commit in the submodule. Do your work that needs to be done for the submodule, push your changes, and then check out that new commit. Make a commit in the main repo to officially bump the submodule to that new commit. Done.
The annoying part is when you do a pull on the main repo, you have to remember to run git submodule update --recursive.
Git subrepo or subtree are some of a solution but not quite complete and easy to use.
In some other scms (P4 and SVN, partly hg) the answer is don't do that, which had a whole lot of its own problems.
If, for example, you add a submodule with the wrong url, then want to change the url, then you instinctively change .gitmodules. But that won't work, and it won't even nearly work.
If you add a submodule, then remove it, but not from all of those places, and try to add the submodule again (say, to a different path), then you also get wierd errors.
If you add a submodule and want to move it to another directory then just no.
Oh and also one time a colleague ran into problems because he had added the repo to the index directly - with git add ..
Oh and let's talk about tracking submodule branches and how you can mess that up by entering the submodule directories and running commands...
But seriously, the fact that there is a .gitmodules file lulls you into a sense that that file is "the configuration file". If you don't know about these other files, then it's natural to edit .gitmodules. When you make errors, the fixing those errors are pretty hard. There is no "git submodule remove x" or "git submodule set-url" or "git submodule mv".
For example, do you know how, on the top of your head, to get an existing submodule to track a branch?
How do you think someone who does not quite understand git would do it? Even with a pretty ok understanding of git infernal, you can put yourself deep in the gutter. (case in point, if you enter the submodule directory and push head to a new commit, you can just "git add submodule-directory" to get point the submodule to the new commit. But if you were to change upstream url or branch or something else in the submodule, you're screwed. That's not intuitive by a long shot)
Edit: git submodule sync is not enough by the way... You can fuck up your repo like crazy even if you sync the two configuration files.
> I'll be darned if I understand the fear of the merge commit
I apologize in advance for not adding much substance in this reply, but I agree too much to just upvote alone.
Here to 'present' feature branches, we take a feature development branch will all the associated crud... Once it's ready to merge, the dev checkouts a new 'please merge me' branch, resets (or rebase -i --autosquash) to the original head, and re-lay all the changes as a set of 'public' commits to the subsystems, with proper headings, documentation etc.
At the end, he has the exact same code as the dirty branch, but clean... So he merges --no-ff the dirty branch in (no conflicts, same code!) and then the maintainer can merge --no-ff that nice, clean branch in the trunk/master.
What it gives us is a real, true history of the development (the dirty branch is kept) -- and a nice clean set of commits that is easy to review/push (the clean branch).
My point is, while your basic commands do the work, your habits and knowledge keep you from losing code like this without you knowing.
When using Git daily we never really did anything complicated, just a few feature branches per developer, commit, push, pull-request, merge. Basic stuff. We had Git crap out all the time. Never something that couldn't be fixed, but sometimes the fix was: copy your changes somewhere else, nuke your local repo, clone, copy changes in and then commit an continue as normal.
Not even once I lost I code worked on with git. stash is a reliable companion across branches and large timespans.
Same is true of merge-based pull.
In other words, shouldn't git just fix ux for branches and rip out stash?
# git stash:
prev_ref="$(git rev-parse --abbrev-ref HEAD)"
git checkout -b wip-stash
git add .
git commit -m 'wip stuff'
git checkout "$prev_ref"
# git stash pop:
git checkout wip-stash -- .
git checkout -D wip-stash
It's quite a considerable saving. I suppose by "fix UX" you mean make it so the saving would be less anyway, but I think really they're just conceptually different: - branch: pointer to a line of history, i.e. a commit and inherently its ancestors
- stash: a single commit-like dump of patches
If stashing disappeared from git tomorrow, I think I'd use orphan commits rather than branches to replace it.You have a working directly/checkout - that can be: identical (apart from ignored files) to some version in git; or different.
If it's different ; some or all changes can be marked for storing in the git repo - most commonly as a new commit.
It's a bit unfortunate that the repo typically is inside your work directory/checkout - under '.git' along with some files like hooks, that are not in the repo at all...
Haven't used reset personally though but only when trying to fix someone's repo.
Switched to mercurial from svn and workflow was painless for the team. Interestingly, we slowly started adopting more distributed techniques like developer merges being common. With svn, I think I was the only one who could merge and it would be rare and added product risk.
Then after about a year of mercurial we switched to git and our brains had adapted. Our team was small, 5-10 people.
Somewhat relatedly, in 2002, I worked in a large team of 75 people or so with a large codebase of a few hundred thousand lines of active dev. It used Rational ClearCase had “big merges” that happened once or twice a release with thousands of files requiring reconciliation. There was a team who did this so it was annoying to dev in, but largely I didn’t care.
Company went through layoffs and the team was down to one. He quit, the company couldn’t merge, so couldn’t release new software versions.
There was a big crisis so they went to the architects and pulled a few out of dev work. It turns out I was the one who could figure it out and dumb enough to admit it.
That sucked. It took a few weeks to sort out and modify our dev process to make merges easy and common. But it was not fun. Upside is we ended up not having any “non-programmer” op/configuration management people since the layed off/quit team were ClearCase users, who didn’t code.
Moral- don’t let people know you can do hard, mundane tasks.
Honestly, its because a lot of it comes down to preference and what value you gain from using version control. It is very much like code style standards -- it doesn't matter what is in the standard so much as your teammates all using the same one.
If part of the blocker for your team is that no one is experienced enough with git to have a strong opinion, I'd be happy to brainstorm with you for an hour to learn about your current process and offer a tailored opinion.
While I have no opinion on git, I can’t abide by all the precious chaotic mutant misuse, like git-flow.
I’d happily accept a subset of primitives, if only to disallow bad ideas. Kinda like Git vs SVN, C/C++ vs Java, flamethrower vs peanut butter.
It's pull then rewrite all your personal commits to be based on the latest tip from that pull.
This confusion happens because many popular SCMs historically have the "commit" and "push" operation in a single step. Git keep them separate.
Not using local branch is another confusion caused by the perspective of historical/traditional SCMs (people thinking branches are the domain of a centralized server and are outside of their control.)
Keeping "local feature branches" just on your dev machine is bad for many many reasons:
- you want to encourage low barrier cooperation in your team -> sharing changes
- you want changes to the CI pipeline early so the potentially slow testing machinery works in parallel with the developer
- you want to keep the team up to date on what changes you make
- you don't want to lose work if the machine/OS dies, or the developer leaves/becomes sick/goes on a 4 week vacation during which they forget their disk crypto password
So, in practice you can try to use rebase opportunistically, when out of chance your WIP work is still unpushed because the change was only made very recently. This is error prone. Or you can rebase published branches explicitly, by destroying the original branches in the PR merge phase. But all this is big bother if the purpouse is to just beautify history and at the same time hide the real trial and error that went into making the changes.
But maybe to you, 'publish' means 'publish to master'? In that case I can assure you, they are not necessarily the same thing. I regularly work on a local feature branch, publish that branch to the shared repo, rebase it on top of master, then force-push to the shared tracking branch. When I'm done I merge it into master and don't rebase master on top of anything.
Basically it makes it so that all of the local-only commits are sequenced after any remote changes that you have not seen yet.
[edit]
YZF is correct. In the context of pulling (i.e. "git pull --rebase") my description is correct. However in general rebasing branch X to Y that diverge from commit C is:
rewind branch Y to commit C; call the old tip of Y Y'
play all commits from C -> X on Y
play all commits from C -> Y' to branch Y.
When explaining to others, I should probably say 'pull, reapply, then push'.
Perhaps 'rebranch' is a better word choice than 'rebase', to conceptually more closely match what's actually happening under the hood.
Last time our devop did 20 commits to get something on elasticbeanstalk right, I squashed it all into just one clean commit that got merged into master branch.
It will help you to commit more often without worry until the moment you have to hand in your work.
What's annoying is that git is just expected knowledge these days and having a github account is enough to claim it. There's not a good way to sell the fact that you're a bit more into it than that.
I've even said to git "experts" that branches should really be called refs and their eyes glaze over. It's difficult for me to understand what git is in their heads.
I know you can target commits through them - which utilizes the ref syntax... But they're still not really referencing anything directly.
They're completely arbitrary and are just a feature to improve gits workflow.
But just as you wouldn't call a symlink to a zip archive a zip file itself, you also shouldn't call a branch a ref.
A branch points to anything you want it to point to. It can be any ref you want and can be changed at will.
right? Seeing as you can git update-ref branches, but you need to git symbolic-ref HEAD.
A branch points to the tip(last commit) of a particular timeline.
In that case, they were thinking the git was you.
If you don't want the tutorial, you can go straight to the sandbox here: https://learngitbranching.js.org/?NODEMO
The distributed nature of git led to the simple and secure contribution model of everyone working on their own repos and not needing to give write access to anyone else. This pretty directly led to an explosion of open source software.
I discovered this supporting SVN servers for whole bunch of developers.
The error messages are clearer, it is multiplatform, all the advanced functionalities are there, a nice graphic interface exists.
I really do not understand why git won, apart from github.
I’d say that git is fine for 90% of development (or some arbitrarily large number), but so is fossil. I don’t even think that SQLite-in-git would necessarily be a deal-breaker that couldn’t be worked around (drh ‘sqlite can chime in here). The whole space (from personal projects to global collaboration) is diverse enough that there’s no talking about “better” without qualifying the situation, either.
Fossil is good for a large subset of work that can benefit from source control management, regardless of git.
What git definately has is
1) scaleabilty, which is probably of no consequence for 99% of the cases it is employed
2) network effect, for better AND worse
He already has
> With Git, it is very difficult to find the successors (decendents) of a check-in ... This is a deal-breaker, a show-stopper.
This operation is not easy in any DAG. It involves:
- find all or desired branch tips - walk backwards until hitting tge desired checkin - memoize already seen parents to not walk them multiple times
The network effects are there, though.
For the large number of files one https://code.facebook.com/posts/218678814984400/scaling-merc... is one source. There's some earlier discussion of the same issue at https://news.ycombinator.com/item?id=3549679 that goes into some technical details.
There has been some work on git since then to address some of those issues (e.g. see https://blogs.msdn.microsoft.com/bharry/2017/02/03/scaling-g... ) but it's not clear to me that it helped enough to catch up to where Mercurial is for large repos.
For large numbers of changesets, just try running "log" or "annnotate" on any file with a long history in git. I just did this simple experiment:
1) hg clone https://hg.mozilla.org/mozilla-central/
2) git clone https://github.com/mozilla/gecko-dev.git
3) (cd mozilla-central && time hg log dom/base/nsDocument.cpp)
4) (cd gecko-dev && time git log dom/base/nsDocument.cpp)
It's not quite apples to apples because the git repo there has some pre-mercurial CVS history in it. But note that I'm not even using --follow for git and the file _has_ been renamed after the mercurial repo starts, so git is actually finding fewer commits than mercurial is here.
Anyway, if I do the above log calls a few times to make sure the caches are warm, I end up seeing times in the 8s range for git and the 0.8s range (yes, 10x faster) for mercurial.
That all said, most repos do not have millions (or even hundreds of thousands) of changesets or files. So the scalability problems are not problems for most users of either VCS.
Yes. CVS and SVN. The old and ugly ones.
If we saw the Clearcase kernel module you can be sure that that was going to be the root cause of the crash. That thing seemed to really terrible, and it wouldn't surprise me if the rest of the product was as bad.
I don't know if things have changed now, but to use Clearcase you needed a kernel module that provided a special filesystem that you did your work against.
That... Augh, that's actually kind of a good idea, even if it was before its time. But FFS...
In their respective times they were a big improvement. I believe CVS was the first client/server revision control system (why that feature was added was a horror story)
Which version control you use matters a lot less than having a sane development process.
(RCS only worked on single files, moving to CVS allowed you to maintain a tree. Before RCS I was using SCCS!)
So, not recent. But version control is one of the few infrastructural components that's allowed to have decades of churn, in my book. Software lives that long.
If not for my yammering there wouldn't be any git in use by any team there. (Only a year ago)
All the code progressed fine. git would still have improved it but nobody had bothered to switch yet.
I still use SVN from time to time with my old repos, but if given an opening, I migrate it to git without hesitation.
Git is an incredibly powerful tool for managing a set of files over time, but if you just use the handful of basic commands, then I agree that the immediate big win is branching.
Personally, I found that moving from Subversion to Git fundamentally changed my work habits (for the better). I was a lone developer at the time, so the collaboration aspect wasn't really important.
I noticed that Git made it so easy to create a repository that I put everything into version control: not just application code, but random scripts and notes.
The other gain was that I learned to work in small, focused commits, because Git is so fast that commiting often is not a burden. Once I made that change, the commit history became meaningful and useful in a way that that Subversion never was: I could quickly revert code, and look back at individual commits for information.
This!
My company uses SVN. SVN works fine so there isn't really a reason to spend the man hours migrating a crap ton of projects to to Git. Before SVN existed we used CVS and we migrated to SVN from CVS about, idk, 15-ish years ago?
That's more or less what I was getting at.
https://pijul.org/manual/why_pijul.html
A bit like DARCS (also very hipster, in Haskell and has some math behind it), but then fast.
https://pijul.org/model/#efficient-algorithms
Oh and it uses a cool hi-perf storage lib (also in Rust, by the same devs):
Pijul lets you describe your edits after you’ve made them, instead of beforehand.
Pardon my French, but about fuckin time.On a big product, forensics matter. Not day to day, but often enough and if your metadata is rotten then you’re left with the oral history of the project as your only guide. And even that may not exist, depending on project structure.
When selecting technology I look for “a rising tide lifts all boats” situations and opt-in tools have limitations in that regard.
There’s a big gap between ‘can do’ and ‘will do’ and I feel like we downplay that frequently in our industry, and to our own peril.
In that section's context, it sounds like naming a branch after having already started on it. In which case, that seems to me the tiniest bit less useful than git's ability to rename branches (git branch -m oldname newname).
What am I missing?
Darcs was magical -- in both senses of the word. It was incredible to see it figure out which patches depended on which, allowing a fluid exchange of changes between branches in a way that quickly becomes a nightmare in git. But it was also magical in that nobody really understood the internals. Not in the sense of git where the underlying data model is pretty simple, and the "version control" aspect is a (thin!) UX veneer on top, but in the sense that it was like quantum physics. When something went wrong, it was almost always impossible to fix. And with Darcs, things did go wrong, because it had bugs, specifically a certain dreaded "exponential conflict" edge case where, if it encountered an identical line change in two patches from different branches (or something like that, it's been more than 10 years), computation time went through the roof and the merge command almost never finished. At several points we had to start history from scratch to avoid spending an entire day fighting the conflict problem. Another thing with Darcs (and presumably Pijul) was that since it tracks patch inter-dependencies, you can rarely cherry-pick individual patches -- pulling out one patch tends to pull with it a whole string of related patches, all connected. Which is often what you want (git just fails horribly in such cases), but sometimes you do want to "forcibly cherry-pick" and manually fix, change identity be damned. I don't know if Pijul supports this.
It looks like Pijul fixes the conflict problem, but it still seems to keep the "quantum theory of patches" that requires an above-average developer to understand. If it has no bugs, then maybe the problem is moot, but in our industry, transparent, "self-repairable" tech seems to win in the long run over the esoteric, opaque and magical.
That said, it's clear the Darcs/Pijul has a vastly better UX, which I'm all for. Git's data model works remarkably well for what it does, but it's always been obvious to me that its "record snapshots and try to make sense of them after the fact" philosophy is a bit flawed. The article mentions branch history. And rename detection doesn't work well with how most people work, for example; it's a clever kind of lazy evaluation, but probably designed for Linux kernel devs, so not clever enough. Darcs had a patch type specifically for renames, and it worked very well.
Another thing I wish version control systems had was what you might call a high-level changelog. It would let you group and annotate commits after the fact, but without changing them. For example, you might want to group a bunch of patches as a single "feature" commit. Then you could make a "release" group that groups a bunch of feature commits. In other words, several levels of nesting, with each commit containing child commits and so on. Viewing the log should show only the highest-level groups, with the option to expand them visually so you can see what they contain. You should be able to group things like this after the fact without changing commit order, and you should be able to annotate the log (e.g. add more information to a commit message) without mutating the underlying patches. Git was on the verge of ventured into this territory with its (now discouraged) "merge commits" -- a high-level commit that represents a single logical merge but encapsulates multiple physical patches -- but that didn't go anywhere. The nice thing about a high-level history like this is that you could use it to drive release notes and change logs, and it would greatly aid in project management and issue tracking, because you could manage entire sets of commits by what issues or pull requests or milestones or whatever they relate to.
Quoted for truth.
The patch theory is complex, but it isn't that complex. Especially since there is plenty of alternate implementations out there of Operational Transforms (OTs) and Conflict Free Replicated Data Types (CRDTs), it's relatives/cousins/descendants. In theory, any developer than can grok a blockchain or a Redis cache should be able to grok the patch theory.
Darcs suffered much more from being written in Haskell, I think, than from the actual complexity of its patch theory.
Pijul being written primarily in Rust maybe has a chance of also getting over that hump a bit easier than Darcs had. Though now it also has the uphill climb of competing against git's inertia.
> Git was on the verge of ventured into this territory with its (now discouraged) "merge commits"
Discouraged only by people that don't know `--first-parent` exists as a useful `git log` and other command arguments. The useful thing about a DAG is you can very easily slice it to create arbitrary "straight line" views. You don't have to constantly smash and squash history to artificially force your DAG into a straight line.
And a summary of an alternative proposed there: http://endoflineblog.com/oneflow-a-git-branching-model-and-w... "As the name suggests, OneFlow's basic premise is to have one eternal branch in your repository. This brings a number of advantages (see below) without losing any expressivity of the branching model - the more advanced use cases are made possible through the usage of Git tags. While the workflow advocates having one long-lived branch, that doesn't mean there aren't other branches involved when using it. On the contrary, the branching model encourages using a variety of support branches (see below for the details). What is important, though, is that they are meant to be short-lived, and their main purpose is to facilitate code sharing and act as a backup. The history is always based on the one infinite lifetime branch."
O_o I'll find my way out.
On Windows, that fact that fossil comes as a single statically-linked executable that works without any special installation procedure[0] is really nice.
Also, I have to appreciate the builtin Wiki. I use it to keep a kind of diary of what I did and why, as well as gather helpful links I have come across over time.
[0] Other than putting the executable somewhere on your %PATH%, of course.
1. It's unclear to me what he means. Yes git doesn't store anything like a doubly linked list of commits, and thus finding the "next" commit is more expensive, but you can do this with 'git log --reverse <commit>..', and it's really snappy on sqlite.git.
It's much slower on larger repositories, but git could relatively easily grow the ability to maintain such a reverse index on the side to speed this up.
2. Yeah a lot of the index etc. is complex, but I wonder how something like "git add -p" works in Fossil. I assume not at all. Is there a way to incrementally "stage" merge conflicts? Much of that complexity comes with significant advantages.
3. This is complaining about two unrelated things. One is that GitHub by default isn't showing something like 'git log --graph' output, the other is that he's assuming that git treats the "master" branch magically.
Yeah GitHub and other viewers could grow some ability to special-case the main branch and say "..and this was merged into 'master'" at the top of that page, but in any case all the same info exists in git as well, so it's just a complaint about a specific web UI.
Now, I don't know how Fossil answers that question. Maybe it's got some clever trick (like a separate "universal ID" vs. "per-branch ID" for each commit, maybe). Maybe SQLite doesn't need that and doesn't care. But it's not like this is a simple feature request. Git was designed the way it was for a reason, and some of us like it that way.
For the record, the initial version of git was developed in haste as a replacement for bitkeeper, after Larry McVoy (who shows up here on HN once in a while, and always has great posts) got a little aggressive about the licensing.
BitKeeper was open-sourced a year or two back: https://github.com/bitkeeper-scm . I haven't had a chance to use it extensively yet, but I hear it still has a good selection of features that git either chose not to implement or didn't properly understand.
Don't get me wrong, I think git is pretty great and have been one of its primary advocates over SVN, etc., at most companies I've worked with (I think 2013 was the first time I came on board a team that was already using it), but its history is informative.
EDIT: upon scrolling, I see that Larry has already dropped by. Read his posts! https://news.ycombinator.com/item?id=16806588
I.e. in some DAG implementations the branch would be a fundamental property at write time, in git it's just a small bit of info on the side.
What I'm referring to is that for the common case of something like the SQLite repository that uses branches consistently it's easy to extract info from git saying "this commit is on master, and on the LHS of any merge to it", or "this commit is on master, but also the branch-off-point for the xyz branch".
The branch shown in the article is a perfect example of this. In that case all the same info exists to show the same sort of graph output in git (and there's even options to do that), it's just not being done by the web UI the author is using.
The "--cached" option would go away; if you have a WIP commit on top, then "git diff" does "git diff --cached", and if you want the diff to the previous non-WIP commit, you just say so: "git diff HEAD^".
stashing wouldn't have the duality of saving the index and tree. It would just save the changes in the tree. Anything added is in the WIP commit; if you want to stash that, you just "git reset HEAD^" past it, and later look for it in your reflog, or via a temporary tag.
Everyone thinks that until they need to use it for something. If all you do is a bunch of linear small changes with obvious implications and two-line commit messages, then the index is nothing but an extra step.
But at some point you're going to want to drop a thousand-line change (from some crazy source like a contractor or whatnot) on top of a giant source tree and split it up into cleanly separable and bisectable patches that your own team can live with. And then you'll realize what the index is for.
"Index" is probably not a helpful term. I think of them as simply "staged changes", that is, changes that will be committed to the repository when I run `git commit`, as distinct from local changes that will not be committed when I run `git commit`. With a Git repository checked out, just editing a file locally will not cause it to be included in a commit made with `git commit`. Rather, `git add` is how you describe that you want a change to be included in the "staged changes" that will be committed. You can add some files and not others, or even parts of a file and not other parts.
The need for this doesn't come up especially often, but it's really helpful when it does. One common case where this can come up is when you've been developing for a while locally, and you realize that your changes pertain to two different logical tasks or commits, and you want to break them up. Maybe one commit is "upgrade these dependencies" and the other is "add feature X". You started upgrading dependencies while building feature X, but the changes are logically unrelated, and now the dependency change is bigger than you expected and deserves to be reviewed on its own.
So with all of these changes in your workspace, you'll stage just the changes for "feature X" or "upgrade dependencies" and then run `git commit`. At this point, maybe you'll move this commit into its own branch or pull request in order to code review and ship it separately. (You might use `git stash` to save the remaining uncommitted changes while you do this.) Then you'll return to the remaining changes, which you will stage and commit as well. You've just gone from a bunch of unstructured, conflated changes to two or more separate commits on different branches/PRs (if that's what you want), that can be reviewed and shipped independently. You've gone from a massive change that's too big to review, to multiple bite-sized pieces.
These tools are also especially helpful if you, for any reason, need to manipulate source control history, such as breaking up one already-made commit into several commits, or simply modifying an existing commit. To do this, you would take that commit, apply it to the local workspace as if it's an unstaged change, and then, starting from the point in history before that commit was made, stage parts of the changes again and check them in. At this point, you can push the changes as a new branch, or even rewrite history by replacing your existing branch.
To give a use-case for this last capability, imagine that a developer accidentally checks in sensitive data as part of a big commit. Before (or even after) shipping the change, you realize this, so you want to go back and edit that commit, to remove the part of the change that checked in data while leaving the rest of the changes. You would describe these manipulations with the index as described in the previous paragraph.
Everywhere you said stage you could say commit (or amend) and then not need the extra step of committing afterwards.
How would you describe this? One massive `git commit` with a ton of parameters? I don't see how it could work.
How would you describe "commit the first hunk of fileA (but not the second), and the second hunk of fileB (but not the first), and all of file C?". How do you "just commit those changes"? I believe you are missing how to actually describe this on the command line or with an API.
The index is absolutely needed. It's what allows you to build up a commit through a series of small, mutating commands like `git add fileC`, `git add -p fileA`. The value of the index is that you can build up your pending commit incrementally, while displaying what you've got with `git status`, then adding to it or removing from it.
However, unlike today, those commands would automatically also do the equivalent of:
if HEAD is marked as WIP
then git commit --amend
else git commit --special-wip-flag
As somebody who understands and uses the git index, I would wholly approve of this change.You can build a commit using multiple small "commit --amend --patch" commands. These use the index, but only in a fleeting, ephemeral way; changes go into the index and then immediately into a commit. They go into a new commit, or if you use --amend, into the existing top-most commit.
The index is a varnish onion.
Git has too many kinds of objects in its model which are all bags of files.
I do not require a staging area that is neither commit, nor tree.
Look at GIMP. In GIMP, some layer operations get staged: you get a temporary "floating layer". This gets commited with an "anchor" operation, IIRC. But it's the same kind of thing: a layer. It's not a "staging frame" or whatever, with its own toolbox menu of operations.
$ git commit --patch
... pick out changes: commit 1 ...
$ git commit --patch
... pick out changes: commit 2 ...
$ git commit --amend --patch
... pick out more changes into commit 2 ...
$ git commit --patch
... pick out changes: commit 3 ...
There, now we have three commits with different changes and were never aware of any "index"; it was just used temporarily within the commit operation.Oops, the last two should have been one! --> git rebase -i HEAD~2, then squash them.
The index is too fragile. You could spend 20 minutes carving out very specific changes to stage into the index. And then you do something wrong and that staging work is suddenly gone. Because it's not a commit, it's not in the reflog.
You want changes in proper commit objects as early as possible, then massage with interactive rebase, and ship.
Suppose HEAD points to a huge commit we would like to break up. One simple way:
$ git reset HEAD^
Now the commit is gone and the changes are local. Then just do the above procedure: commit --patch, etc.That's largely outdated. For years now, git's commit command has been able to stage changes and squirrel them into the commit in (apparently to the user) one operation.
Only people who learned git ten years ago (and then stopped) still do "git add -p" and then a separate "git commit" instead of just "git commit --patch" and "git commit --amend --patch" which achieve the same thing.
Situations like integrating a big blob of messy changes happen all the time in real software engineering, and that's the use case for the git index.
I would convert the change to a unified diff, remove the change, and then apply the selected hunks out of that diff with patch. ("Selected" means making a copy of the diff, in which I take out the hunks I don't want to apply. Often I'd just have it loaded in Vim, and use undo to roll back to the original diff and remove something else.)
Using reversed diffs (diff -uR) I used also to selectively remove unwanted changes, similarly to "git checkout --patch"
This is basically what git is doing; it doesn't require the index. The index is just the destination where these selective hunks are being applied.
At some point I am ready to craft this code into multiple commits. After my first git add -p and git commit, I don't know if HEAD is in a state where it even compiles. It takes further work and discipline to produce a whole series of good commits.
Imagine I have a file with lines 10 and 150 changes. How do you commit just one without some form of index (or alike)?
Well, that's the weasel word. In my grandparent comment, I proposed the "alike", didn't I?
Nowhere did I say, just remove the index from Git, but don't replace its functionality with any other representation or mechanism.
In git, we can do that today in such a way that the index is only temporarily involved:
$ git commit --patch
... interactively pick out the change you want ...
$ # now you have a commit with just that change
It is not some Law of Computer Science that the above scenario requires something called an "index", which is a big archive holding all of the files in the repo, where these changes are first "staged" before migrating into a commit.1. typing `git commit`, checking the output, then typing `git commit -a`, or
2. typing `git commit`, then moving on with your life, and realizing minutes, hours, or days later that the changes you meant to include were not actually included, so you have to go back and add them if you're lucky, untangle them from whatever subsequent changes you were trying to make and/or do an interactive rebase if you're unlucky, and maybe face the prospect of doing a `git push --force` if you're really unlucky
Scale that up to several days or weeks to match the learning period and repeat for every developer who has to sit down and interact with it. That's the overhead we're talking about.
The article got it right; this is a monumental waste of human effort.
> Every developer has a finite number of brain-cycles
But even when we use "git add", we are not aware of the index. The user can easily maintain a mental model that "git add" just puts files into some list of files to be pulled into version control when the next commit takes place. That is, until that silly user makes changes to the file after the git add, and those changes do not make it in because they forgot the "-a" on the commit.
I won't argue there is a relatively long learning period with git ... It helps if you have some experienced mentors in this area. But you get a lot of power for this...
Where "it" is just the snapshot of that file as it is now, not as it will be at commit time; then you have to add it again!
..I actually use it with almost every commit, so I don't add reminder comments and debugging statements.
So sure: if you want a not-index which works like the index and has a bunch of 1:1 operations that map to the existing ones, then... great. I still don't see how that's much of an argument for getting rid of the index.
Remember that git, initially, didn't hide the index in the add + commit workflow! You had to "git add" and then "git commit". So the fact there is only "git commit" to do everything is because they realized that the index visibility is an anti-pattern and suddenly wanted to hide it from view.
Since the index is already hidden from view (and largely from the user's model of the system) in the add + commit workflow, we are not going to optimize the command set by turning the index into some other representation. That's not what this is about.
The aim is consistency elsewhere.
For instance, if the index is an actual commit, then if we abandon it somehow, like with some "git reset", it will be recorded in the reflog.
Currently, the index is outside of the commit object model, so it gets destroyed.
It's possible for a git index to have content which doesn't match the working tree; in that case when the index is lost with a git reset, that content is gone.
If the index is a commit, it can have a commit message. It can be tagged, etc.
The way it would work is that when I realize I want to do a partial commit I'd stash, pare down the changes in my working tree (probably using an editor with visual diff), test what's in my working tree, commit, and then pop the stash.
I had hoped that this would already be doable with git, but it isn't, at least not in a straightforward way. The problem shows up when you try to apply that stash. You get loads of merge conflicts, because git considers the parent of the stash to be the parent of HEAD, and HEAD contains a bunch of the same changes.
I'm sure there's some workaround for this, but every time I've asked people always tell me to not bother testing before committing!
$ git stash
$ git checkout stash@{0} -- ./
$ $EDITOR # pare your changes down
$ make runtests # let's assume they pass
$ git add ./
$ git commit
$ git checkout stash@{0} -- ./
$ make runtests # let's assume they pass
$ git add ./
$ git commit
*(In case it appears otherwise, this isn't actually supposed to be a defense of Git. Originally, this was 15 steps, but I edited it into something briefer and more straightforward.) git commit --patch # Just commit the bits you want
git stash # Stash the rest
<do your tests>
git stash pop
<continue developing>You're probably going to say I shouldn't test before every commit, but I rarely work in a branch where the bar is as low as "absolutely no testing required". I generally at least want my build to pass, or some smoke tests to pass, and I can't reliably verify either of those with a dirty work tree. And actually, the fact that all of the commits on my branch are effectively going to end up in master (unless I squash) makes me want to to have even my feature branches fully tested.
Ah, I see what you want to do now.
> You're probably going to say I shouldn't test before every commit
If have no business telling you what you should do. If you want to test before committing, your wish is my command
git stash --patch # Stash the bits you don't want to test
<do your tests>
git commit <options> # Commit the rest when the tests pass
git stash pop
<continue developing>I never use `--patch` (even with `git add`). I prefer to use vim-fugitive, which lets me edit a diff of the index and my working tree. It looks like being able to do something similar with stashes is a requested, but not yet implemented, feature for vim-fugitive: https://github.com/tpope/vim-fugitive/issues/236
git commit --patch
(and therefore `add --patch` is never even required!) git checkout --patch
to selectively revert hunks, and git stash --patch
to selectively stash. I couldn't be happier with this workflow!I.e. you get it.
Importantly, the semantics is available through a common interface rather than a different design and implementation of the semantics for the index versus commits.
> you now have a "commit" that doesn't behave like a commit
Well, now; literally now you have a commit that doesn't behave like a commit: the index.
If a real commit is used for staging, it behaves much more like a commit. It's just attributed as do-not-publish so it doesn't get pushed out. Under this model, all commits have this attribute; it's just false for most of them. Thus, it isn't a different kind of commit.
> I strongly believe that things that behave differently should be named differently.
Things that do not behave completely differently can use qualified names, in situations when it matters:
"work-in-progress commit; tentative commit; ...."
For instance we use "socket" for both TCP and UDP communication handles, or both Internet and Unix local ones.
Too hard, they said. They wanted something which just needed to be pretty, not something with tough human interface problems.
I know about git log (and the millions of options it has), but it's ugly as a sin, difficult to just quickly browse, and is just so busy that it's hard (for me at least) to quickly get information on if employee x merged branch y into z or if v has the latest commits from w, etc...
So I ended up writing a user script to blow up the small window on the GitHub page to a larger size, and get rid of some stuff that we don't care about and made it a dashboard of sorts.
I spent some time one day trying to find something like it, butcame up with pretty much nothing that was easily setup and maintained and that I didn't need to build my own application around.
There is also gitg, which also includes primitive commit capabilities and uses the full-fat GNOME widgets, making it pretty.
It's the same overly busy UI, the same lack of easily seeable branch names (yes I know this isn't how git works internally, but it's extremely useful), the same vertical layout on our widescreen monitors, and they are still ugly to me (although this is honestly one of the lowest on my list of priorities).
I'm looking for something like [0] but with a better UX (Github's version doesn't let you scroll with the mouse wheel for instance), is fast enough to pull up at a glance to quickly see the state of the whole repo, and won't get disabled if the repo has too many forks (and I'm guessing branches, although I've never seen it happen from then alone) like Github's does.
It's easy to glance at and see where a branch is (what was merged into it, what it was merged into, etc...), who did the work (in the example it shows in a "tooltip" when you hover over the dot), roughly when the work was done, and what that branch includes.
I recently aliased `git history` to `git log --graph --all --simplify-by-decoration`, and have `glog` as `log --graph` (and so seldom use `log`).
If you’re earlier than Git 2.12 or so, you’ll need to add --decorate to get the branch and tag names shown.
Otherwise they had very comparable features including being able to display the histories, merges, etc.
I hardly ever have to use cmd git.
The visualization features are amazing though but could use a little bit more design.
If you're talking about locally modified files, I use the integration with Visual Studio to manage those and it is doing an OK job. The only gripe I have with it is that you can't change an already staged file or those changes won't get checked in. I don't remember it working like that a year ago so something must have changed.
I use them a lot and love how it makes some advanced stuff much easier (committing hunks, interactive rebases). They're also cross platform, which is great.
Still, it's incredible how crappier it gets over time. It feels like every year there's some kind of new major refactor or UI framework refresh that just destroys whatever little stability the application had, as if a new lead comes in and decides to start everything from scratch.
As an engineer I kinda understand the reasoning (using a nuclear bomb to get rid of tech debt), but it's pretty baffling from a product lifecycle standpoint.
I wish they had a more stable model and were charging for registration really. I'd gladly pay them $50 or whatever for some modicum of stability.
I mostly only ever used it for branches, cherry picking and conflict resolution, and it was awesome for that.
I don't use it now, and I don't use branches or cherry picking. My personal philosophy is that I am not keen on branches: never have been, never will be. I see repositories as a two pizza team concept. If you need branches, you should split the repo, make more iterative changes, or look at establishing or improving a test suite. It keeps things simple and cognitively low overhead for everyone (as per the sqlite team's comments).
I work on a branch, raise a merge request to master, let my colleague do a check and point out any issues and then merge it once we're both happy with it. He does the same thing for me to check.
Other than that, no branches.
It is not immediate, needs to wait for another human to be around and mentally present, demands some sort of QA standards are created and commonly understood, and creates the need to make and manage a forum for listing and discussing merge requests.
You could get some of the way and retain instant feedback with automated tests, yet remove a lot of the overhead.
Downsides is it's not well known, the principle dev is only doing bug fixes, and it while it's blindingly fast on small to medium repos, it doesn't scale well for larger repos.
On the command line? Hg log --graph
https://svn.apache.org/repos/asf/subversion/developer-resour...
ClearCase is 20 years old or so, but no git gui comes close.
The reason is git's peculiar object model. It has nodes (commits), edges between nodes (parent references) and references (pointers to commits) but no branches. That means it is not in general possible to tell which branch a particular commit belongs to. Asking a question like that just doesn't make sense i git's world. Therefore visualizing complicated branching scenarios, in ways that makes sense to users, becomes almost impossible. This is why people advocate rebasing https://blog.carbonfive.com/2017/08/28/always-squash-and-reb... The idea is to "lie" to the vcs because otherwise your history becomes a hodge-podge mess of merges. :)
Btw, all other features of ClearCase were just horrible and brain-damaged. But the branch visualizer kicked ass.
That's because commits can belong to multiple branches, which is by design.
> Therefore visualizing complicated branching scenarios, in ways that makes sense to users, becomes almost impossible.
There a reason users need to visualize complicated branch scenarios?
> This is why people advocate rebasing https://blog.carbonfive.com/2017/08/28/always-squash-and-reb.... The idea is to "lie" to the vcs because otherwise your history becomes a hodge-podge mess of merges. :)
Rebasing/cherry-picking isn't lying, it's avoiding the mess of a merge at the expense of being a tad more difficult.
Seems you have a very concrete idea of what you think a VCS is, and git disagrees with you, thus they "all suck".
I could be wrong, but I think what they mean is the fact that once you merge/rebase your dev branch back into master, all of those commits you made in dev are now commits in master.
You could conceivably have a source control system that still allows commits to belong to multiple branches, but bakes the branches of a commit into the commit. (And I believe many/most other source control systems do just that.)
The fact that the branches of a commit can change later on always struck me as being completely inconsistent with the commonly held idea about git that outside of master it's ok to forego testing. That really only works if you squash into master (and test that squash commit). If you don't squash, and especially if you fast-forward, then all of those commits are now "in" master, and so you've dumped a bunch of untested stuff in master.
Just shows you how lazy some people are!
Seriously, I find git very non-sane if not quite insane. I really liked darcs and still prefer Mercurial but sometimes the value attained by adopting standardisation overcomes the extra value in a "better" alternative, and so… here we are.
Most sane? That’s a matter of perspective. I’m still a little shocked that git “won out” over mercurial. Even as Subversion was eating CVS, everybody knew distributed revision control was going to ultimately prevail. I was pretty sure Darcs wasn’t going to achieve popularity, but I’d have bet anything that Mercurial would be the successor to Subversion. It was far more natural / similar for anyone who's ever worked with CVS/SVN.
I’d have bet on the success of more friendly CLI UX and comparatively easier time on Windows even with the disamenity of Python.
Calling Git sane just makes it clear that you haven't used a sane source management system.
Git has no file object, it versions the repo, not files. There is one graph for all files, the repo graph. So the common ancester for a 3 way diff is the repo GCA, which could very well be miles away from the file GCA if you have a graph per file (like BitKeeper does).
No file object means no create event recorded, no rename event recorded, no delete event recorded. If you ask for the diffs on "src/foo.c" all Git can do is look at each commit and see if "src/foo.c" was modified in that commit. That's insanely slow in a big repo. And it completely ignores the fact that src/foo.c got moved to src/libc/foo.c years ago and there is a different src/foo.c that is completely unrelated. There is an option to try and intuit the renames when you are spitting out diffs but noone uses that because it's even more insanely slow.
Git is basically a tarball server. Calling that a source management system is an enormous stretch. Calling it a sane and powerful source control tool is just not supported by the facts, calling "the most ..." is laughable.
Yeah, I get it, Git won. You all lost out on "the most sane and powerful" as a result. Which sort of doesn't matter since everyone thinks GitHub is source management.
As someone who has used (in a professional setting) version control systems ranging from RCA to SVN to the current crop (git, mercurial, even darcs), the fact that git versions the repository as a whole instead of individual files is a godsend.
You've not known hell until you've had to deal with an RCA/CVS repository with 15+ years of history and thousands of files, each that maintains their own version history (and associated version number!).
I'd gladly take the comparative "slowness" of git when dealing with large repositories.
> Git is basically a tarball server. Calling that a source management system is an enormous stretch.
What, in your opinion, is the definition of a version control system then?
In BitKeeper files work like they do in Unix, there is a (globally) unique name for each file object. Where the object lives in the repository is an attribute of that object, as are the contents, the permissions, who has changed etc.
Here's a super common workflow that's easy in BitKeeper and miserable in Git. I'm debugging an assert. I want to see who added the assert. I pop into the gui that shows me the per file graph and contents, search for the assert, hover over the rev and see that it was done a long time ago. I look in the area above the assert and I see a recent change, hover over that, see the comments and go "hmm, maybe this". Double click that line and I pop into a different gui that shows me the whole commit that contains that suspect line.
Note that because I have a graph per file, I have checkin comments per file. More work for you poor committers but a godsend for us debuggers. More breadcrumbs are more better.
In Git, less breadcrumbs, single commit message. Git wants to go from the rev to the commit, it's miserable to look around in a file and then go backwards to the commit.
When I was supporting BitKeeper our average response time to bug report or a crash report was 25 minute. 24x7. The only reason it was that long was because we were all in North America so there was a window where we were all asleep. Response time 6am-6pm PST was typically under 5 minutes. And I credit the fact that the tool accurately recorded everything and you could find the history really easily.
Oh, and it didn't slow down as the repo got big. Git is fine in little repos but it sucks pretty hard in big ones. Sucks even worse if you are on NFS. I can dig up benchmarks, we built up a synthetic 4M file repo and ran a bunch of tests on it (it was a modified version of the facebook repo builder, the facebook one had some stuff in it that made Git look incredibly bad, we looked at that and decided that wasn't real world or fair, we took that part out).
Less, better breadcrumbs, yes.
I get why, as a dev, you want git commit -m'Fixed bug' but as the debugger guy, the reviewer gui, anyone who reads the code, that's a horrible thing to do to those readers.
Who, if you wait long enough, will be you. And I'll laugh my ass off at all the lazy committers who really could use more breadcrumbs when they have to debug their code later. Been there, done that, I haven't worked with people that lazy in decades.
edit: since HN won't let me extend the thread, let me reply to the comment below because BK does do something special.
The GUI for checkins presents you with a list of files, a place to type comments, and a big pane that shows the diffs. You type in comments for the first file, go to the next, type in comments (yes, there is a way to say use the previous comments). As you move from file to file, the bottom pane shows the diffs for that file so you can see what changed in that file.
The special sauce, that Git most definitely does not have, is when you get to the last file, which in BK is the ChangeSet file, this is where you would type the commit message. What are the diffs? There aren't any so we stuff all the comments you just typed on individual files. What does that do? Well, on the files you are usually typing in details of how you did this or that, when you get to all of those comments, you naturally uplevel and type in why you were doing all that.
It dramatically increases the usefulness of commit messages. That's why Intel pretty much mandated the use of the gui checkin vs the command line checkin.
Here's what he's taking about: http://www.bitkeeper.org/man/citool.html and http://www.bitkeeper.org/man/templates.html
Here is git's analogous feature: https://robots.thoughtbot.com/better-commit-messages-with-a-...
It's possibly the least interesting and least unique selling point an SCM can have. It's really funny that you keep bringing up this example around the thread.
EDIT: To respond to the above edit:
Again, he's Proving The Point. Commits should be atomic. If you have to individually comment on file changes then the correct thing to do would be to put those in their own commit, no? I'm not really sure what's being described is necessarily a cool feature, but rather a way to avoid making sure your changes are truly related. I honestly don't see the point. This seems like a feature that was written because of the decisions that were made into how BitKeeper works internally, not because it's a fantastic idea. You can get the same thing with atomic commits in git. You comment per file, because BK tracks changes per file. Git does not do this. You should be making your commits atomic because git is tracking the actual content. Atomic commits will accurately describe what's being changed, and then of course those all get lumped together in a patch/PR.
Git isn't lacking the feature you're describing, it just kind of is there without any extra data tracking required, because it's not making up for technical design decisions.
I may be misunderstanding that particular workflow (I'm sadly unfamiliar with bitkeeper), but this seems like a workflow that I accomplish relatively often with the use of tig[1].
On a separate note: thank you for the wonderfully detailed reply. It's such a pleasure to have an opposing view be so thoroughly explained.
[1]: https://jonas.github.io/tig/
Edit: I apologize that my first, original comment seems to have brought out the incorrigible tech trolls. Sometimes the internet is the worst.
Every day humans make me again realize that I love my dogs, and respect my dogs, more than humans. There are exceptions but they are few and far between.
Edit: Never mind, I see you answered this below. https://news.ycombinator.com/item?id=16807319
This. A thousand times this. A change isn't a single file, it's likely a changeset of multiple, sometimes hundreds or thousands of files. Most often, you need to know the entire changeset, not just what happened with one file. This is especially true of large-scale refactoring where you're changing the public interface of something. Depending, that can have a far-ranging impact and you want to see that history all together.
But lots of times you want to look at the file view, find the line of code that looks like the problem, and then zoom out to the changeset. BK makes that trivial, Git makes that miserable to impossible.
[1] Actually there was a little known system, Aide-de-camp, that one of my people told me about that had changesets so I didn't invent it, but I reinvented it. And made the world aware of the concept. Back when you could search Usenet via dejanews you could search for "changeset" and date limit it to before me talking about it. There were maybe 5 hits. A few years after BK came along there were 100's of thousands of hits. So I wasn't first but I am definitely the reason that you know what a changeset is.
git log -- filename
Unless the file hasn't been touched for many years in a Linux kernel size repository, that's faster than you can blink.I do occasionally wish that git blame was faster. Building up a cache of side index information for that would be a nice addition.
In my company people simply use checkout to apply changes from another branch, but that doesn't handle well for example file removals. It also creates a completely different commit so the branches diverge more and more, making things like rebase more time consuming.
It is a huge pain, which would not exist if we would use SVN as a backed for example.
What you should probably do in this case is use git merge nevertheless, but reset all changes outside of the directory of interest before committing the merge. This way, you get the history of the merge in the DAG, which will make git's merge resolution work.
Unfortunately, I'm not aware of a built-in way of doing this.
I suspect on a HDD `git blame` would be unbearably slow on anything but the smallest of repositories, but it has been many years since I last worked with source code stored on a HDD.
However, this is something every single version control system can do, so surely you were referring to something else. Could you explain what operation you were referring to?
We built a GUI tool that lets you look at a versioned file, it shows you the graph in the top pane and either diffs between two versions or the contents of a particular version in the bottom pane.
It's the goto tool for figuring out stuff. It is not just a GUI version of "git blame". When you use it you can see the history of each line by hovering over that line, you get a popup that shows the checkin comments for that line. And it is fast, as in below human reaction time, so you use that feature.
And you can double click on any line and boom, you are looking at the changeset that introduced that line.
I'm tired so I'm probably not doing a good job explaining this, but we supported commercial customers for a couple of decades and we had just incredible response time to each issue and I credit this work flow for that. Someone would call and say I have this assert and one of would get into the gui, start looking and we would know the cause of the problem in seconds or single digit minutes and I don't me 5-10 minutes, I mean 1-2 minutes.
Maybe I'm clueless and there is a way to do this in git but I haven't found it. When I have to work with git repos I fast-export them into BK just so I can have a more sane way to look at the history. It's not great history but it's better than Git.
Edit: I didn't explain what that gui did on the ChangeSet file. So that's what gitk (gittk?) is, it shows you the repo graph. You can click on a node and see the commit, you can left click and right click and see the diffs between those changesets.
So far as I know, BK is the only system that puts the metadata in the same system as the user data.
This criticism doesn't ring true for me. What are you saying is missing in the Git experience?
In Git, I would use `git blame` to determine which commit contributed the problematic line in question. It displays the file along with the commit that most recently modified each line. At that point, I know which commit last changed the line (and the commit is the changeset). Aren't we done?
If I need more history, I can use `git log` on the filename to see what commits have changed that file over time, and I can inspect how each commit individually changed the file if needed. There are editor-integrated tools to walk back through this history easily.
A `git blame` that doesn't take forever to run. Especially in the face of file moves and renames.
Thanks for your contributions to better SCM design through BK. I agree with the potential value of tracking files and branches as well as commits and being able to easily navigate through that space -- as well as the emphasis on adding more metadata to make future understanding easier.
It's unfortunate we don't have a better funding model for FOSS development (or alternatively a basic income) because otherwise BitKeeper might have been open from the start -- and then we might have avoided the limitations of git as a hasty workaround for licensing issues.
And no one cares about it. For a reason.
>Git has no file object, it versions the repo, not files.
Which is the correct thing to do.
>if you have a graph per file (like BitKeeper does).
Which is insane to do because you're basically making things more complex than they need to be, which is why merging in git is not only SUBSTANTIALLY faster than nearly every SCM out there, it also works a good deal of the time without issue.
>No file object means no create event record, no rename event recorded, no delete event recorded.
Most people do not need these. If you're going to choose to use file GCA's for the reasons listed above then you better have a damn fucking good idea of how much these features are actually used, because the trade offs are ENORMOUS.
> That's insanely slow in a big repo.
And also not really done that often. So you know, good job choosing to optimize for things no one is going to use on an astoundingly consistent basis or even needs to be super duper speedy in the first place.
>You all lost out on "the most sane and powerful" as a result.
Your concerns are misplaced? Misdirected? Git does things a certain way and people use it because people were tired of SCM's that decided to fix problems no one really cared about, and then do the things that developers do care about very poorly. You actually demonstrated this pretty thoroughly in your own post while attempting to call ,presumably, bitkeeper "sane."
>Calling it a sane and powerful source control tool is just not supported by the facts
I'm sorry, feel free to tell every other developer in the world, including the ones that are involved in far more collaborative efforts than your work requires, that the tool they're using is just not sane. I guess being one of the most used SCM's in the world, on one of the biggest OS projects in the world aren't really relevant facts into how "sane" an SCM is. I guess that's totally why bitkeeper used to be sold and now is open source.
Your claim about merging is false though, demonstrably. Picking a repo gca when you could have used a much closer file gca is better. BK does that and automerges, correctly, more frequently and is way way faster than Git.
I've written two source management systems. I'm confident in my knowledge. Arguing with some random dude who thinks he knows more than me is not really fun. So go enjoy Git. Lots of people are too busy/whatever to know what they are missing, maybe that's you. It's not me, I kinda like my audit trail to be accurate.
Even Linus admitted to me in my kitchen that Git's audit trail is lossy. But go enjoy Git, I'm glad it works for you. Knock yourself out.
Linus wrote Git. Fanboy followed.
Demonstrate it. You said it was demonstrable so presumably your Totally Sane SCM project should have evidence to back this up.
>Arguing with some random dude who thinks he knows more than me is not really fun.
Why because you basically made a complete fool of yourself?
>Even Linus admitted to me in my kitchen that Git's audit trail is lossy.
You might describe it as FUZZY, but not "lossy." Nothing is being "lost." My random internet dude that is actually an insane assertion to make, especially the anecdote about Linus being in your god damn kitchen. Not that it really changes anything about the characteristics of any SCM and why people use it.
>But go enjoy Git, I'm glad it works for you.
Enjoy your dead project. Glad it could be surrounded by a graveyard of other Totally Sane (C) SCM's who are collectively responsible for an untold amount of wasted man hours and licensing fees.
http://www.mcvoy.com/lm/photos/2007/05/276.html
ZFS guys (Brian, Jeff and Bill) at the same party (I thought I had one with Linus and them, they were all at the same table talking file systems):
http://www.mcvoy.com/lm/photos/2007/05/257.html
Funny story, I invited Linus to the pig roast and he didn't RSVP so there wasn't anywhere for him to sleep. I ended up sticking him and his daughter in a VW popup van :)
You can kindly wander off now, random internet dude. Especially since the guy who wrote Git agrees that it is lossy, so take your FUZZY and go home.
That comment is not very nice and does not serve your otherwise interesting comments.
For what reason?
Whatever, it paid the bills nicely for almost 20 years.
And for things like this:
>I didn't want to do anything that even smelled of BK. Of course, part of my reason for that is that I didn't feel comfortable with a delta model at all (I wouldn't know where to start, and I hate how they always end up having different rules for "delta"ble and "non-delta"ble objects). [1]
[1] http://lkml.iu.edu/hypermail/linux/kernel/0504.3/1588.html
In fact popular IDEs / editors back then supported CVS, then SVN and pretty much nothing else.
You'd be amazed (or not) how little technical superiority matters. Developers just want to click stuff and get back to developing.
Which is not to denigrate git, which I love so much it hurts.
I think Git won because it was superior to CVS and SVN. I don't think it won because it was better than BitKeeper, Fossil or Mercurial.
Lastly, does anybody have experience with git merges and BitKeeper merges? You make it sound slow, but it also sounds like you've never used it.
I should write up a blog post about it, it's pretty complicated to understand because to get it, you need to understand SCCS's interleaved delta format. If you understand that format, then imagine that you put a line number in front of each data line in the weave. Check out the GCA, local, and remote versions of the file with those line numbers prefixed. Now run a 3 way merge on that.
All the complexity in smerge.c comes from dealing with the cases where that doesn't work, but man, it works great 99% of the time.
The computing world is full of absolutely terrible technologies people keep using even though much better alternatives exist. At first I considered listing a few and then I realise that it's likely most readers are actually users of one of those, but I'm sure you can think of a lot.
Yeah, I've only worded it 50 different ways.
I don't claim that Git is perfectly. But most of the problems are solved by better UI (Git does track when files are renamed, created and deleted), not by wrecking what a commit is.
Git understands commits. CVS style 'every file has it's own history' is wrong. Git's not perfect and I would not be too bothered if Mecurial beats it long term, but my god if I had to go back to CVS-style flow with concept of multi-file checkin, I would quit my job.
Actually, no, it doesn't. Git tracks just the before/after state: before the commit these files existed, after the commit those files existed. It infers creation/deletion/rename, when necessary, by comparing these two (or more) states.
Which doesn't always work. Not that it matters.
It would be like the Myspace founder/creator badmouthing Facebook and claiming Myspace was still superior and that Facebook's way of doing things is insane. Whether or not the claims are legitimate is irrelevant in light of the fact that your competitor squashed you and may have made you bitter.
I'm not a fan of using a DB to store versions. It's just not the right tool. Before we open sourced, we jealously guarded "the weave" which is how the history data is stored. The weave gives us so much, bk blame is instant, bk grep is instant, there is a "bk grep -R" that will look in all versions of a file that is instant, or you can do "bk grep -R<revs>" and look in just those revs, all instant.
The weave is compact, fast, merges better, it's just a better storage format than a DB. Here's an example. In most version control systems, lets say there is a 100 line file. I clone that repo and I modify the first 51 lines of that file. You clone the same thing I cloned, and you modify the bottom 50 lines of that file. So we have 99 unique lines and 1 line that we both modified. Now Joe Merge clones my repo and merges your repo. He's the guy that closed the DAG. He had to manually merge the one line that we both changed so when you do $SCM blame the correct answer is the top 50 lines are me, the bottom 49 lines are you, and the manually merged line is you, right?
That's what happens in BK. It's not what happens pretty much anywhere else. Either the entire top chunk or the entire bottom chunk will look like it was done by Joe. Why? Because everyone else passes data by value, BK (the weave) passes data by reference. Everyone else copies the data across the merge point. BK does not, the only new data that will be in the merge node is the one line that joe had to merge by hand.
This can have some space savings implications, which can be a big deal for big files, but in my opinion the far bigger implication is blame. Joe merged in your stuff and now your stuff looks like he wrote it. Someone is tracking down a bug and they should be talking to you but they are talking to Joe.
I'm not sure I understand your example. At least, the way I understand it, both git and mercurial deal with it the way you say BK does. They attribute the 50 first lines to you, the 51th line to joe if it looks like neither what was in your or the other's version, and the last 49 lines to the other.
You establish a naming convention on top of that to use it as a VCS. For example, one common such convention is to have directories named "trunk", "tags", and "branches" at the top level, with your projects living in subdirectories under "trunk". Under this convention the way you represent a tag is by simply copying your project directory from trunk/my_project to a new directory named tags/my_project/tagname. Similar for branching...just copy to branches/my_project/branch_name.
Don't like that convention? Develop your own that fits your work better.
Oh my yes. Binary diffs were lovely. Also, support for large assets.
For reproducible data science work, you’re going to need at least code and data. One of those things is a total PIA with git if the data aren’t trivially small.
I wonder how fossil does in this regard...
[0]: yes, I know this has been debunked. It's an analogy.
Please stop.
> Or that I used the generic term rather than the git-specific term?
In the context of the discussion, generic term refers to "check in", and git-specific term refers to "commit" (which is the only possible alternative term to be using in that context). The obvious reading is that you're calling "commit" a git-specific term.
So I wouldn't say I'm putting words into your mouth, unless you were unaware that "commit" is the correct alternative term, which is of course a possibility that I didn't consider.
Also unless your organization's project is open source, you don't really need DVCS. In my organization we use git nearly exclusively, but in the end everything relies on the central repo. We actually would do much better if we used SVN instead of git and have less issues. For example we already had significant mistakes such as someone delete main branch or performed a force push (yes, you can restrict it, but with git you need to know what to expect before blocking it). We also have repo for CMS, where we would greatly benefit from ability to merge by directory. I also see people trying to checkout latest version from git for just specific subfolder, but with that is also difficult, but trivial and extremely lightweight on SVN.
Yes, for still has very valuable tools for the local developer: stash, staging changes, bisect, local history (great if you work without internet access), but you can actually use git with other SCM and get the best out of both worlds.
Unfortunately when I mention that we could have our main repo under SVN, people look at me like some kind of dinosaur that is proposing it, because I get confused by git.
Just because git is a great tool for Linux kernel development, doesn't mean that it is the best SCM for for your organization. And if your organization uses something else than git, it doesn't mean you can't use git, in fact the tool of my choice is still git, I just think that most companies don't need DVCS for their main repo.
> There is no significant way in which I found Pascal superior to C, but there are several places where it is a clear improvement over Ratfor. Most obvious by far is recursion: several programs are much cleaner when written recursively, notably the pattern-search, quicksort, and expression evaluation.
While the author goes about comparing Pascal to C (and Ratfor), he fails to disclose his affiliation with being the author of C.
What Kernighan did do is write the book on C. (You might argue that it's a conflict of interest, too, but it's not. It makes a lot of sense that someone would like a programming language so much that they'd decide to write a book on it.)
Now, the reasons why you created Fossil, that interests me. But this is not the case. It's an ad.
In the 90s we had cvs and perforce... and then svn. Then there was BitKeepr. From what I can read of history, 2002 to 2005 Linux was under BitKeeper and then in 2005 that relationship soured and Linus went off to write Git. At the same time there was a large growth of other version control systems. The graphic shows bazaar, darts, hg, plastic and then Fossil is also in there.
Even at 2006 when Fossil was released, there wasn't significant mindshare on any of those platforms yet. Why use git? It wasn't even a year old when Fossil was released.
The migration of SQLite to Fossil (from CVS) was done in 2009.
That's several years after the software was written that SQLite switched to it. I wouldn't exactly put that in the place of dogfooding. It kind of is - but Fossil was a mature project when SQLite shifted its codebase.
I wrote Fossil specifically to support SQLite development. If Fossil does nothing else other than support SQLite, then it is a success. Any other use of Fossil is just gravy. That we were conservative in moving the main SQLite source code into Fossil does not negate that fact.
We do also use Fossil for dogfooding SQLite. See, for example, item 15 on the release-testing checklist: https://www.sqlite.org/checklists/3230000/index
Other selling points from the Fossil site:
* Integrated Bug Tracking, Wiki, and Technotes
* Built-in Web Interface _and_ Self-Contained (I combined these)
* Simple Networking (no git://, just HTTP and SSH)
* CGI/SCGI Enabled
* Autosync - "Fossil supports "autosync" mode which helps to keep projects moving forward by reducing the amount of needless forking and merging often associated with distributed projects."
* Robust & Reliable - "Fossil stores content using an enduring file format in an SQLite database so that transactions are atomic even if interrupted by a power loss or system crash. Automatic self-checks verify that all aspects of the repository are consistent prior to each commit."
It would be easy to get into comparing git and Fossil feature by feature. That's interesting. But it's more interesting to compare the philosophies between the two tools.
edit: for grammer.
So what are their philosophies? Fossil is for small things and git is for scalability?
> Git is a content-addressable filesystem. Great. What does that mean? It means that at the core of Git is a simple key-value data store. What this means is that you can insert any kind of content into a Git repository, for which Git will hand you back a unique key you can use later to retrieve that content.
That's a great thing for matching the mindset of someone writing operating system software.
Fossil is on top of a relational database... which is also a great thing for matching the mindset of someone writing a relational database.
Git is a version control system; it’s not a complete project management solution. You can build integrated solutions around it, like GitHub, but if GitHub’s issue tracking is too primitive for you and you want to use Jira instead, that’s super easy.
I actually favor that approach myself. If I wanted to use Fossil but their issue tracker didn’t work for me, the best case scenario is that I just integrate Fossil with Jira or whatever and just haul around this vestigial issue tracking system that I don’t use. Likewise, if I wanted to use Fossil’s issue tracking but not their version control, what then?
> Built-in Web Interface _and_ Self-Contained
Is this different from `git instaweb` ?
> Simple Networking (no git://, just HTTP and SSH)
So out of 3 methods git supports they support 2?
"Open-Source, not Open-Contribution: SQLite is open-source, meaning that you can make as many copies of it as you want and do whatever you want with those copies, without limitation. But SQLite is not open-contribution. The project does not accept patches. Only 27 individuals have ever contributed any code to SQLite, and of those only 16 still have traces in the latest release. Only 3 developers have contributed non-comment changes within the previous five years and 96.4% of the latest release code was written by just two people. (The statistics in this paragraph were gathered on 2018-02-05.)"
> In order to keep SQLite completely free and unencumbered by copyright, the project does not accept patches. If you would like to make a suggested change, and include a patch as a proof-of-concept, that would be great. However please do not be offended if we rewrite your patch from scratch.
The working directory
The "index" or staging area
The local head
The local copy of the remote head
The actual remote head
Git contains commands (or options on commands) for moving and comparing content between all of these locations.In contrast, Fossil users only need to think about their working directory and the check-in they are working on. That is 60% less distraction."
I don't think about any of this when developing. I check out a branch, work on it, commit to it, and push it back up. If the fix is larger than a few commits i'll make a feature branch. What's so hard about that?
Also, everyone using git isn't a bad thing. It means we finally have at least one standard in development.
I'm sorry, I don't buy this at all. Git is one of the most simple source control tools there is. If you can't understand a DAG then there isn't much else you probably can understand in the development world.
>This manifests as frequent unintended results
No it doesn't. Every time I've seen people complain about "unintended results" it's literally been because of the above, and they've been complete morons so I'm never surprised when these people have "trouble" with git. You're building a graph, and you're doing pretty basic manipulation of that graph.
>I never had any such problem with Perforce
Uh what? Permission issues? Terrible branch performance? Merging between branches is basically a gamble -- it's actually insane how much this used to mess up over the most basic of merges. Having to "upgrade" the system? Ever done that. I'm going to guess no.
If you haven't learned about urbit yet, you could try that for the same "I don't know what you're talking about" experience in your life.
Of course? I'm also free to call them idiots online.
>There's a practical issue with mapping every command to each copy of the dag
That is certainly the most valid criticism of git, in my opinion. And certainly the biggest learning curve for most people. But that's not where people tend to have true, blue operational issues in my experience.
There isn't a way around this. With any tool, especially when dealing with SCM. You have to understand how your selected tool is going to perform certain operations because it's going to dramatically effect your workflow with that tool.
>The biggest thing with Git is you have to invest time into understanding it
Just like you have to do with every other SCM, or just generally any piece of software?
You can, and will, argue they are morons. That's an excellent way to deal with a poorly designed system.
It does not "actively" lose work. Losing work in Git is a pretty specific operation and generally requires you to spell out what's going on (either via 'reset' or 'checkout'). And also less likely to happen, since you are regularly using the index to stage changes and should be commiting often (since it's cheap to do, unlike trying to commit in BitKeeper, let's say, where locking problems sometime even require to change the window you're committing in [1]). The moment things are in the index/repository, it's pretty unlikely you're going to lose work.
Say, in other systems, me accidentally clicking something is a surefire way to just wipe it out without a second thought. In TFSVC I can easily undo my active changes with a simple click of a button. Again, the contention was that you actively lose work while working with git. It's hard to do, as deleting work in git requires some fairly specific commands.
Even removing a file requires you to stage the deletion, then commit it. And even THEN, even if you amend some other previous commit with that deletion you can role the entire thing back with reflog and get your file back. If I took that amended commit and rebased onto another branch, rolled that all into one commit via squashing, I'd still be able to go back to where I was with the reflog.
If however, you edit some lines in a file and then tell git to checkout that file again and they're not staged, then sure, you're going to lose work. But that's not something that's common and it certainly isn't a recipe for having git "actively" delete work.
[1] http://www.bitkeeper.org/man/citool.html
> If the repository is locked, and you try to bk commit, the commit will fail. You can wait for the lock to go away and then try the commit again; it should succeed. If the lock is an invalid one (left over from an old remote update), then you can switch to another window and unlock the repository. After it is unlocked, the commit should work.
If you want to live in a world where two people can be wacking the repo at the same time with undefined results, be my guest, you seem like that sort of person. We are not. We like atomic commits.
As for BK being expensive compared to Git, you are so right. 10 years ago. These days we do quite well and we do it correctly.
I dunno why you have a hardon to smear BK, but bring it dude, I'm happy to make you look foolish.
Excuse me? You brought up BitKeeper. You're the one that claimed I've never used a sane SCM, because you think BitKeeper is sane and I clearly didn't include that in my list of the World's Most Sane SCM's list. Added bonus for the claim we're all "missing out" on sanity.
Here I'll even link you to the post you made: https://news.ycombinator.com/item?id=16806588 And a screenshot https://imgur.com/a/JO8IE
>I'm happy to make you look foolish.
You're actively on this forum basically demonstrating why BitKeeper and other SCM's have lost. You bring up silly things like templated commit messages, random anecdotes that don't technically make any sense, and claim annotating/blaming history in files is hard to do in git.
>These days we do quite well and we do it correctly.
Some guy got so pissed off at your SCM and made a new for one free, without wasting untold amounts of man hours and capital. And he wasn't the only one (Mercurial). He did it so well that other companies now use his SCM as a cornerstone for their platforms (GitHub and BitBucket).
I think this is your fundamental misunderstanding. Git would be another esoteric tool for crazy kernel devs without GitHub. Git won because of GitHub. Very simple. Out in the real world of teams of 12 developers rewriting the same business logic over and over until they retire, Git is GitHub. I've worked with several developers who don't understand the difference, and think that they're using the GitHub client whenever they interact with Git locally (and, if they use GitHub Desktop, they are). All your discussion about git's dominance being a testament to the tractability of the git UI is a false equivalence.
GitHub could switch to BitKeeper under the covers overnight, and as long as they branded it in a non-scary way, very few people would know the difference.
When is it ever the case that general acceptance means "objectively the best" instead of "obviously the path of least resistance"?
No it didn't. GitHub was built because of git's popularity. That's a really silly claim to make and doesn't even logically make sense. The tool was popular, someone built a hosted service for it.
Like think about the insanity of what you're saying. A company was built around some "esoteric" tool, despite plenty of alternatives existing at the time and you think people just whimsically invested in this company because of... what exactly?
You're literally just talking from historical ignorance and making up shit. It's so easy to verify any of the claims you decided to type out yet you chose not to anyway. GitHub believed in Git's popularity, and it paid off.
>I've worked with several developers who don't understand the difference
That's a really neat anecdotal story.
>GitHub could switch to BitKeeper under the covers overnight
No they couldn't. They couldn't even switch to it over a two year time frame. You are making absolutely ridiculous claims with nothing to back it up. BitBucket, owned by Atlassian, has hosted Mercurial repositories. Don't see anyone lining up to switch over to Mercurial.
A really easy claim that is obviously negated by a practical example and you still chose to make it? Legitimately silly.
>All your discussion about git's dominance being a testament to the tractability of the git UI is a false equivalence.
That's not even what that word means.
>When is it ever the case that general acceptance means "objectively the best" instead of "obviously the path of least resistance"?
Don't know if you were around for when git was released but there is now a whole graveyard of SCM's that people actively dropped to switch to git. And BitKeeper was definitely one of them. But I guess those were all dropped because of a future platform that no one knew about at the time. Heavy sarcasm by the way.
Indeed I was, and FWIW I distinctly remember everyone throwing their weight behind Mercurial, in large part for its superior cross-platform support.
I stand by my position. git usage would be minor without GitHub. GitHub could switch off git if they wanted to and they'd take most of the git user base with them.
Except for nearly every single open source project.
>Mercurial, in large part for its superior cross-platform support.
And then that got side stepped by the obvious problems with repository size and branching issues (that hg would go onto address later).
>I stand by my position.
If you want to ignore a pretty basic historical timeline that's totally up to you.
It's sad, because there are other useful work flows, but GitHub is SCM at this point. I agree with what someone said elsewhere, they could swap out git for bitkeeper and nobody would care (well the people that are still butthurt over the licensing would whine but it's apache v2 now, that should be good enough).
I'm sorry, what dev conferences are you going to? You're kind of just claiming that these same people are too stupid to understand what git allows you to do out of the box so they wouldn't mind BitKeeper's (or any other SCM like it) problems and limitations as long as GitHub hosted it for them with a nice logo (which again, isn't true because other hosted SCM's solutions lost as well). That's just incredibly tone deaf and doesn't make sense from a historical timeline perspective. This is just straight up denial at this point.
>well the people that are still butthurt over the licensing would whine but it's apache v2 now, that should be good enough
It's not just a licensing issue and you know it. You are being dishonest with everyone here and yourself. There is a historical record in the lkml archives that you're choosing to ignore.
So what's so special about GitHub if not Git? Slightly better UI?
And I think, might be wrong, but I think github was first.
Right now today, there are developers out there who are putting in the work to learn git. They are smart and dedicated to their craft. They are also losing work because git makes it easy to shoot yourself in the foot and because it doesn't easily surface to you how to recover when you do so. It's not because the developer is an "idiot". It's because git's UI is user hostile and does a poor job of reflecting the core concepts in operation when the user is using it.
To be fair to Git they've put a lot of effort in fixing this. It used to far worse than it is now. But due to backwards compatibility and other concerns there is still a lot confusing and unintelligible options out there and that is even after you get the whole concept of a DAG and the index vs working tree.
This is a pretty silly assertion to make. You're actively ignoring the SCM's that were chosen by far larger org's. Git took the dev world by storm because it was better. Not because an open source OS used it. Other SCM's had years worth of developer training and interaction with their systems and STILL lost. Years worth of sales connections, demos, networking, you name it and STILL lost.
>They are also losing work because git makes it easy to shoot yourself in the foot
Example?
Most people can get the quickie commands down pretty fast, and as long as there are no complaints or errors, it's all OK. But as soon as you need something that takes several steps, your average git user falls apart. Rebases are the most common example. If you don't block force commits from the get-go on a project used by more than two devs, you'll learn that you need to do that pretty quick.
- https://stackoverflow.com/questions/10099258/how-can-i-recov...
- https://www.clearlyagileinc.com/blog/recovering-commits-from...
Everyone of these is an example of the git UI/UX confusing you about what is happening. Reflog allows you to recover if you know about it but if you don't you will feel betrayed by the tooling. Blaming the developers for not knowing what Git was going to do is ignoring git's bad ui/ux decisions. Mercurial for instance doesn't have this problem. In fact it tells you whenever you do something that might lose data that it made a backup and where it put the backup at. Everything the developer needs to know right there after they accidentally shot themselves in the foot trying to do their job.
Other SCM's had years worth of developer training and interaction with their systems and STILL lost. Years of worth sales connections, demos, networking, you name it and STILL lost.
You are comparing Git to high cost options like Perforce which lost because they were expensive and OpenSource doesn't like expensive proprietary solutions in a field where OpenSource rules. I'm talking about solutions like mercurial which have a similar underlying model and a better user experience out of the gate. Git beat hg not because it was better architecturally. They have the same underlying datamodel. It won because the linux kernel used it and that gave it the necessary cachet and authority by association. This is not bad git is worlds better than perforce or cvs or svn in many ways. I'd rather use git than most of those in most situations. But I also would rather use hg than git any day because it's the same datamodel with a ui that doesn't mislead me.Um, I don't see it. Your first 2 links are the same question. All "3" examples basically give you a single command to run to get your work back. What's confusing?
> Mercurial for instance doesn't have this problem.
WHAT? Mercurial actively changes the underlying data objects it stores and there is NO WAY to get them back. Rollbacks actively delete content and you CAN'T restore them, shuffle them around, or anything like that. You do have the Journal extension but it's not apart of the core product. And people definitely haven't had problems with extensions before ;). See: https://book.mercurial-scm.org/read/undo.html
>In fact it tells you whenever you do something that might lose data that it made a backup and where it put the backup at.
So does git. Every operation that Mercurial gives you a warning about git does as well. Even better, git still gives you a way out even if you do the thing you shouldn't have done. It warns you in the prompt, it warns you with comments in the text editor of your choice that pops up for input.
Is there some operation you can give an example of that git lets you do that Mecurial somehow magically prevents or gives you a way out of?
>You are comparing Git to high cost options like Perforce
And what about the other opensource SCM's that lost?
>They have the same underlying datamodel.
They do not.
>I'm talking about solutions like mercurial which have a similar underlying model and a better user experience out of the gate.
It definitely does not. Not sure what you want but I'll take practical solutions like git being able to handle partial checkouts, branches, and blames over a nice gui. Git may have inconsistent command flags but it still does what you need it to.
>use hg than git any day because it's the same datamodel
You keep saying this but it's definitely not true, hence hg's limitations on certain operations. https://stackoverflow.com/a/1599930
However, if you want to send a patch to someone else's project on Github, you have to create a remote fork, download the fork, and push it to your remote fork, then send the pull request. (I have tried other ways but they don't seem to work.)
This is fairly annoying since it's extra steps and it clutters many people's Github accounts with forked versions of projects they contribute to. There's no good reason for these forks to exist.
But this is a problem with Github and not git itself.
So, it doesn't take long to realise that you need to understand the index after all. You need to learn git add -N, or use "git diff HEAD" instead of git diff.
In particular, I'm interested in the last section "4.2 Features found in Git but missing from Fossil".
Maybe this is some deficiency of my workflow, but both those things make Fossil sound extremely unappealing to me. You're telling me I have to push all local changes every time I want to push anything? What if I was just debugging or experimenting with one branch, and then I had to switch to a "serious" branch to push some critical changes?
And rebasing is fantastic when you have a feature branch workflow.
I use git professionally (of course), but while I've used rebase to fit into workflows with github though I dislike the history garbling it does, I've never bothered with rebase with my fossil projects though I use feature branches. Just normal merges. I let it fork and merge at an appropriate point. (What I actually miss now and then is "git rerere" and a useful GUI like gtk.)
Some notes on this choice focusing on the differences I actually use -
1. Fossil gives me peace of mind that I have everything (all code/notes/bugs/tags/branches) synced to the server with a single command "fossil sync". If you have autosync, then even better. I never get this peace of mind with git.
2. There're only two files to deal with - the fossil repo file and the fossil executable - which functions as the command line tool, a server for cloning and syncing, and minimal GUI.
3. I like the tagging system in fossil more. Since you can reuse tags unlike git, I just use a single "release" tag to mark code points pushed out .. with the commit containing details. I similarly use tags to mark points in the commit tree to revisit later (for example) - across branches. In fact, branches are just recurrent tags in fossil, so one less concept.
4. More peace of mind 'cos I can't leave a "dangling commit" that will be "garbage collected". Since I can't leave an unreachable commit, if I want to stash something, I just commit it with full notes, update to an earlier commit and "let it fork". (fossil does have stash, but I don't bother with it as .. less peace of mind).
5. I can customise the bug system for different uses.
The same is about the rebase - the difference is that while git keeps the history "as the developers want it to be", fossil keeps it "as it actually happened".
Git preserves history as it happened on the remote, which is what matters for collaboration. Why foul up the canonical history up with what amounts to scratch paper? Typically, people only amend local history if they feel it would easier to follow later. Why take that option out of their hands?
The way I do work and the way I submit it are two separate problems with overlap. I use git for both and it works beautifully. Even for personal non-code projects I never intend to collaborate with others, I still use git because it promotes a workflow I'm comfortable with.
On the team I'm in, the convention is to include code changes in the same commit as any test changes. There are advantages and disadvantages, but it's the convention, and it's best we follow it or else it might lead to CI problems. But, in my personal workflow, I tend to try to change the tests, commit, code changes, commit, and then iterate. With git, this is as simple as doing it however which way I want, and then squashing the commits together before pushing. What do I do with Fossil?
I'm trying to figure out what problem it's trying to solve with this, and it just comes off as hollowly idealistic.
Git does not track branch history.
This makes review of historical
branches tedious.
This isn't true. It's just the Fossil example has a better UI than their github example. Pop open gitk or anything with a nicer UI and you can easily follow branch history. $ cat .git/refs/heads/master
170ec1365f9fc0ca281e72e6789d1df281d168ae
From it, you can infer the commit history which gitk shows, but if you delete the branch (and discard the reflog), it's gone.In Fossil, each check-in (commit) actually records the name of the branch it belongs to: http://fossil-scm.org/index.html/doc/trunk/www/fileformat.wi...
That's not a proper data structure, and you can't point to a random commit after merging and deleting a branch, and ask git to tell you which branch it came from.
- Eve has a repository, Alice and Bob have access to it. Eve goes on vacation and nobody commits to "master"/"trunk".
- Alice makes a branch "alice-fixes" and creates commits on it.
- Bob comes and creates a branch "bob-features" from some point from "alice-fixes".
- Bob then merges "bob-features" into "master" and deletes it.
- Alice gets fired and her branch is deleted without merging.
In git, you can't see that some of the commits in the history came from "alice-fixes". Fossil, on the other keeps track of branch names and _changes_ in commits:
When Alice created her branch: -trunk +alice-fixes When Bob created his branch: -alice-fixes +blob-features When Bob merges his branch: -bob-features +trunk
You can't delete this information, because it's recorded in the commit artifacts. You can't delete branches — you can only close and hide them (you can also apply edits to commits by adding "edit" artifacts [don't remember what the are called], but the history is preserved, not modified.)
If you really want that behavior from git, you can have it, just create a tag for each branch HEAD. Actually I think this is all that fossil does as well.
> If you really want that behavior from git, you can have it, just create a tag for each branch HEAD. Actually I think this is all that fossil does as well.
Fossil has branch/tag name directly in the checkin manifest (I've linked to the document describing file format somewhere above.) Tag and branch names are embedded directly in the history, they are not separate references to commit hashes like in git.
That is not true. You get a branch diagram just like the fossil example in the OP.
Merges in git actually have a "direction", i.e. git records which branch was merged into which. So that information is recorded.
merge
/\
commit-hash1 commit-hash2
When branching, Fossil creates a branch commit with the branch name (similar to git tag object), and also records changes to branch name in commits.Pull actually works the wrong way around – it merges the remote branch into the local branch and then pushes that as the new remote branch. So if you follow the first parent, you land in the pullers local part and bypass everything he merged.
http://fossil-scm.org/index.html/doc/trunk/www/branching.wik...
In fact, I'd say Subversion, Git, and Fossil call different things "branches".
Note: I'm not saying that one way is superior to another, I'm just saying that Fossil does track branch history, but Git doesn't.
the annoying part of getting it back was that you can't just do "checkout {name}" you have to do it by the SHA ID and then commit again by the original name, but before all that, you can 100% follow any since-deleted branch by its name. there's just a weird disconnect between viewing its history and having it directly in your hands again.
> The principle maintainer of SQLite cannot function effectively without being able to view the successors of a check-in.
I wish it went into more detail about why this is critical to his workflow.
If I were to hazard a guess, this ability to follow descendents helps follow the evolution and intent of current state of code, given some previous state? (how did we get here from there and what else changed on the way).
In git that'll kinda sorta be possible with git bisect.
What does make me chuckle however is seeing git used as a drop in replacement for svn without using any of the advanced branching/merging features it offers. I find it funny to hear a devops youngster eschewing the benefits of git after hearing that <n> hip opensource projects use it, just to find that they use using it with a single tree (or bunch thereof), just continually committing everything to the master branch.
Just like people used to use cvs and svn.
Are you going to tell me you've never lost work because you didn't know git well enough to avoid putting your branch into a bad state?
I just wrote the referenced article last night. Normally, I takes months or years before something like this gets picked up and discussed on HN, and I have more time to refine the text. This one snuck up on me. Come back in a month or two and the article will probably be much improved. You are reading an initial draft.
On the other hand - it is a funny cartoon, don't you think? And it does kind of capture how most people use Git in a snarky kind of way, doesn't it? :-)
Any system that has consistent-ish rules can be learned. The more rules, and the subtler they are, the harder it is to learn. Git's problem is that it is necessary to know how it really works, to avoid getting into a bad state. It's complicated enough that it is easier than it should be to get into a bad state. The extent of the rules you have to know to avoid trouble does not match the simplicity of the day to day operations you do on the repo.
It's the only thing I've tried since SVN days, but I can totally see the room for alternatives...
Ok I'll bite. Why?
I can see it's a nice idea - but at some point you add function X, then ten check ins modify that function. what's the difference between looking back at ten and looking forward at ten?
Perhaps I need an example.
If I remember rightly fossil was written by the Sqllite people? I tried it out for a while back in the day - I think it had this store the tickets in the branch alongside the code - it was an appealing idea, but got complicated quickly.
git tag --contains <commit>
or git branch --contains <commit>> Fossil, in contrast, strives to keep all changes from all contributors mirrored in the main repository (in separate branches) at all times.
From https://www.fossil-scm.org/index.html/doc/trunk/www/fossil-v...
Semi interesting full circle side note: One of those projects was actually a rudimentary version control system for a report writing app, which used SQLite as the VC data store.
Why? There was a time not too long ago where this was true about git. Companies switched to git anyway.
But when you know git and try bzr, it’s mostly the same thing but with different commands. Any gain is small and the pain is instant.
Despite that, the Git command line interface is horribly inconsistent. Commands take separate words, `--options` or one-character flags without a scheme. Lots of synonyms such as staging area, index, or cache all meaning the same thing in the terminology does not help to make learning Git easier.
For new users I would recommend Gitless, which tries to create a better interface to git and to solve the ambiguity in commands. As Gitless is a frontend to libgit2, it works with any git repository and you can also use normal git commands. The downside is that documentation usually only shows how to do stuff with 'git'. If you only learned 'gl', you have no idea how to reproduce it.
But it ticks the two boxes
- powerful enough (to e.g do proper merging, which cvs and svn never could)
- tooling support, meaning it’s supported out of the box in bug trackers, build systems etc.
All others (hg, perforce, svn, cvs, pijul, fossil, ...) fail one or both of the above.
Now, I hope that one day git will be replaced by something nicer. But for the time being it’s what we’re stuck with.
My pet peeve: sequential revision numbers. Why not have an arbitrary numbering of commits?? Saying “I have bug X in rev 1234 but it’s not in 1230” is fantastically powerful compared to “I have the bug in a1b34h but not in 3ae452”. These would be a sequence for a particular centralized branch - typically mainserver/master.
Because this is the reality: it’s distributed version control but we almost all use centralized version control. This is also why it’s so odd that Git LFS took years to make it into git. Why would I want all past revisions of a binary (Yes, binaries must often be in version control, whether anyone thinks it’s a bad idea or not)?
Is a distributed bug reports etc even desirable? How would they work?
Not thinking about the actual remote head and/or the local copy of the remote head seems bad? I'm worried that Fossil is just obscure information, rather than presenting it.
Is it possible to `fossil rebase origin/master` without an internet connection?
// edit:
Reading more of the docs, what I quoted at the top is definitely, ah, misleading.
> When autosync is turned off, the changes you commit are only on your local repository. To share those changes with other repositories, do: fossil push URL
> When you pull in changes from others, they go into your repository, not into your checked-out local tree. To get the changes into your local tree, use update:
https://fossil-scm.org/index.html/doc/trunk/www/quickstart.w...
We have confirmed that Fossil has
The working directory
The "index" or staging area
The local head
The local copy of the remote head
The actual remote head
So yeah, point #2 in the OP is lies.Looks like the love is mutual
> With Git, it is very difficult to find the successors (decendents) of a check-in.
"git log --children" and "git log --reverse" are very difficult indeed.
> Fossil users only need to think about their working directory and the check-in they are working on. That is 60% less distraction.
Since Fossil is a distributed version control system, there is also remote state to keep in mind. I don't know Fossil, so I don't know the details, but simply pretending the remote state does not exist seems at least misleading.
> Setting up a website for a project to use Git requires a lot more software, and a lot more work, than setting up a similar site with an integrated package like Fossil.
Setting up GitLab with Omnibus takes about five minutes. It's not Git itself, but rather a third-party package, but why should I care? (And in general, I wouldn't set up anything at all – I'd just use GitHub or GitLab or Bitbucket for my open-source software.)
It's fine for different people to prefer different tools. Git can be annoying at times. However, it's a blessing that the open-source community is moving towards a standard everyone can work with, and Git is Good Enough to be that standard. Using an obscure alternative and justifying it with a list of downright wrong claims doesn't seem to create a welcoming atmosphere for new contributors.
(And don't get me wrong – SQLite is a great piece of software, and I'm thankful people put in their time to create it.)
We tried git too and lost our changes inevitably. Our git knowledge is to be blamed but blame git CLI too! It is not easy to learn at all. After 12 years of git, I see articles with tips and guides on using git. That kind of hints how esoteric git CLI can be.
For folks like me, who spent their evenings in high-school screwing around with obscure linux distros and have been scouring badly written man pages trying to fix their computer for half their life, the little stuff is forgivable, and a lot of the big stuff is pretty good. But it's a shame how many inessential hurdles there are. See also:
https://git-man-page-generator.lokaltog.net/
Biggest piece of advice for not losing data: learn about git-reflog and git-reset before doing anything that modifies history. Nothing is destroyed, even by "scary" operations like rebase, so it's always possible to recover something, but you need to know how to find stuff.
I do some contract work for a university, where we have lots of student interns going through. We use GitHub, and we rarely rebase or cherry-pick, basically because while I could spend a bunch of time trying to teach git to new interns that will be gone in 4 months, It seems like a more efficient use of time to just occasionally put up with a messy history.
1.Revert 3 git commits. This is one-click revert in wiki.
2.Diff 2 arbitrary commits. One-click in wiki.
3.change 2 lines of code. Edit->Submit in wiki.
Why git has to be so arcane, especially reverting to specific point in time? The first code versioning control system that emulates the wikipedia UI model will win the market.
He's basically looking for
git branch --contains $commit_or_branch
isn't he?
I do think it's a thoughtful write-up, hopefully git maintainers and github read it.
# begin quote
Every developer has a finite number of "brain-cycles". Fossil requires "fewer brain-cycles to operate", thus "freeing up intellectual resources" to focus on the software under development.
# end quote
It is useful to have a 'lingua franca' of open-source version control, but I'm not sure git is necessarily the best choice for that (not that network-effects usually result in the 'best' being chosen)
That said, that doesn't make the points SQLite makes invalid.
But, now that I do know git the biggest change from I noticed from previous VCSes is how much I work on multiple issues in the same repo. Something that was extremely hard with CVS, SVN, P4 (10yrs ago).
A friend was struggling with git recently and ranting about it. He didn't get it and didn't understand why anyone would use it compared to what he was used to (non DVCS). I wrote him this analogy
> Imagine some one was working with a flat file system, no folders. They somehow have been able to get work done for years. You come along and say “You should switch to this new hierarchical file system. It has folders and allows you to organize better”. And they’re like “WTF would I need folders for? I’ve been working just fine for years with a flat file system. I just want to get shit done. I don’t want to have to learn these crazy commands like cd and mkdir and rmdir. I don’t want to have to remember what folder I’m in and make sure I run commands in the correct folder. As it is things are simple. I type “rm filename” it gets deleted. Now I type “rm foldername” and I get an error. I then have to go read a manual on how to delete folders. I find out I can type “rmdir foldername” but I still get an error the folder is not empty. It’s effing making me insane. Why I can’t just do it like I’ve always done!”. And so it is with git.
> One analogy with git is that a flat filesystem is 1 dimensional. A hierarchical file system is 2 dimensional. A filesystem with git is 3 dimensional. You switch in the 3rd dimension by changing branches with git checkout nameofbranch. If the branch does not exist yet (you want to create a new branch) then git checkout -b nameofnewbranch.
> Git’s branches are effectively that 3rd dimension. They set your folder (and all folders below) to the state of the stuff committed to that branch.
> What this enables is working on 5, 10, 20 things at once. Something I rarely did with cvs, svn, p4, or hg. Sure once in awhile I’d find some convoluted workflow to allow me to work on 2 things at once. Maybe they happened to be in totally unrelated parts of the code in which case it might not be too hard of I remembered to move the changed files for the other work before check in. Maybe I’d checkout the entire project in another folder so I'd have 2 or more copies of the project in separate folders on my hard drive. Or I’d backup all the files to another folder, checkout the latest, work on feature 2, check it back in, then copy my backedup folder back to my main work folder, and sync in the new changes or some other convoluted solution.
> In git all that goes away. Because I have git style lightweight branches it becomes trivial to work on lots of different things and switch between them instantly. It’s that feature that I’d argue is the big difference. Look at most people’s local git repos and you’ll find they have 5, 10, 20 branches. One branch to work on bug ABC, another to work on bug DEF, another to update to docs, another to implement feature XYZ, another working on a longer term feature GHI, another to refactor the renderer, another to test out an experimental idea, etc. All of these branches are local to them only and have no effect on remote repos like github (unless they want them to).
> If you’re used to not using git style lightweight branches and working on lots of things at once let me suggest it’s because all other VCSes suck in this area. You’ve been doing it so long that way you can’t even imagine it could be different. The same way in the hypothetical example above the guy with the flat filesystem can’t imagine why he’d ever need folders and is frustrated at having to remember what the current folder is, how to delete/rename a folder or how to move stuff between folders etc. All things he didn’t have to do with a flat system.
> A big problem here is the word branch. Coming from cvs, svn, p4, and even hg the word "branch" means something heavy, something used to mark a release or a version. You probably rarely used them. I know I did. That's not what branches are in git. Branches in git are a fundamental part of the git workflow. If you're not using branches often you're probably missing out on what makes git different.
> In other words, I expect you won’t get the point of git style branches. You’ve been living happily without them not knowing what you’re missing, content that you pretty much only ever work on one thing at a time or find convoluted workarounds in those rare cases you really have to. git removes all of that by making branching the normal thing to do and just like the person that’s used to a hierarchical file system could never go back to a flat file system, the person that’s used to git style branches and working on multiple things with ease would never go back to a VCS that’s only designed to work on one thing at a time which is pretty much all other systems. But, until you really get how freeing it is to be able to make lots of branches and work on multiple things you’ll keep doing it the old way and not realize what you’re missing. Which is basically way all anyone can really say is “stick it out and when you get it you’ll get it”.
> Note: I get that p4 has some features for working on multiple things. I also get that hg added some extensions to work more like git. For hg in particular though, while they added after the fact optional features to make it more like git go through pretty much any hg tutorial and it won't teach you that workflow. It's not the norm AFAICT where as in git it is the norm. That difference in base is what really set the two apart.
Sorry that was so long but my question for the Fossil guys would be "which workflow does Fossil encourage?" Lots of parallel development like git or like many other VCSes not so much parallel dev. Are branches light and easy like git or are they only meant for marking versions like the were in SVN, P4, CVS. Do branches even need to be related or can they be completely unrelated like gh-pages and the VCS won't complain that you're "off master" as hg does (did?)
> 1. Git data model only lets you see ancestor commits, not descendants. Maintainer needs to find descendents.
It's technically true about git's data model, but it's not too hard to reconstruct forward history based on all refs (branches for this purpose). (see "git rev-list").
They don't justify the use-case that demands this information be trivially available, and I've never personally needed it, so I find this a weak argument. I have needed to find what branches contained a given commit (either a regression or a bugfix), and that's available with "git branch --contains <SHA>"
> 2. The mental model of git is complex (working dir, index, local head, local copy of remote head, remote head)
Granted. It's especially always complicated trying to explain to beginners the distinction between local head, local copy of remote head, and remote head. Though it makes sense to me when working out technical implications of git's distributed data model, it's obvious that many users are overwhelmed. (I think the mental model is actually great, but I won't impose my opinion on others)
> 3. Fossil's branch history display is better than GitHub's, and git doesn't tell you if a branch has been merged
No mention of "git log --graph" (or any GUI alternative) which goes a long way of solving their complaint. I agree that GitHub's graphless history is frustrating, but that's a GitHub issue, not Git.
However, Git does "swallow" merged branches, as the merge commit may be a property of the receiver branch, not the original branch (depending on your development model). Git doesn't enforce something there, so I agree with the issue.
> 4. Git lacks essential wiki/bug-tracking. If you use 3rd-party tool for these, they're centralized
Granted, though I'm not quite convinced in how "essential" it is to have that tightly integrated in the version control system, and thus have it distributed, but that's probably an artifact of me being used to existing git-based systems.
A counterpoint to this is that Git (now) has a healthy ecosystem of alternatives for you to choose from. If you're not satisfied with Fossil's offering, how easy is it to change?
> 5. Git requires administrative support for the extra web tools that Fossil otherwise integrates
This follows from the previous point, so granted. I'd add to this that Git lacks a good access-control/code-review system, hence the rise of so many alternative portals (GitHub/Gitlab, Gerrit...)
> 6. No-one really understands git [w/ XKCD link]
Now you're just trolling.
Overall, I'm surprised at how many points I actually agree with the author(s), but the main difference is that git purposefully aims at a more limited feature-set that excludes project-management (ie wiki/bugtracking), which I guess is the consequence of its Torvalds/kernel-originated development history.
The branch name is not preserved in commits. I am curious of what you are thinking of. You can not reconstruct a branch, it's impossible.
This is technically correct, the best kind, but it's not the whole story.
The default merge commit message does contain the branch names. Unless it's a fast-forward, in which you maintain linear history, so there are no problems.
If your team uses a common application to write their code which manages Git, it will align everyones method of use.
It's funny watching an instructor training a group on Git. Everyone looks lost because there are too many scenarios being explained and how to handle them depending on what your intent might be. Before there is a solid grasp on the code management workflow the lecturing about commit comments begins, shifting clear over to the other side of the brain to guarantee none of it is collectively understood.
Presenting too many options to the collective is counter-productive.
You love them more than humans because they're fundamentally incapable of calling you out on your fud?
I engage because I want to help. BitKeeper is mostly dead, I've offered to go to facebook and transfer as much as I can to their SCM.
We detached this subthread from https://news.ycombinator.com/item?id=16807652 and marked it off-topic.
Wow, we have gone full circle. Can’t wait for Linus to announce the git ICO!
Gitcoin is the name of a real blockchain project that let's people put ethereum bounties on their github issues.
(Slightly off topic, sorry)
Because Fossil-- like Git-- doesn't solve, attempt to solve, or even advertise itself as solving the problem of decentralized consensus.
All existing blockchain technologies at least claim to be a solution to the problem of decentralized consensus.
Therefore Dr. Richard Hipp is only "technically" right in the ways that do not matter, in the same way that I'm "technically" doing functional programming any time I write a javascript function.
I assume he knows this and is being satirical.
On the command line:
git for-each-ref --contains <commit, branch, etc.>
Lists all refs (branches, tags) that contain the passed argument.I'd be curious to know the author's need for this. I've used something similar on GitHub to determine when a given commit, typically a bugfix, is released.
> There is no button in GitHub that shows the descendents of a check-in.
There is: go to the commit's URL (https://github.com/$ORG/$PROJECT/commit/$COMMIT) and below the subject & body of the commit, but above the author/commit time, there's a section that will have all branches, tags, and pull requests that the commit is a part of.
For example: https://github.com/aio-libs/aiohttp/commit/d7f0511ead6d05cf6...
That commit is part of the "master" branch, and the "v3.1.2" tag.¹
¹Note that more tags will likely appear in the future, if you're reading this comment some time from when I'm writing it.