What's New in Mercurial 3.0
hglabhq.com
hglabhq.com
Here are some of git's real problems:
* Performance issues with multi-GB git repos * Handling of large binary files * Submodules - Mercurial has subrepos, but I don't know how they compare
It's really disappointing to see 3.0 announced with no solution to this.
And what kind of data corruption do you see? It tracks files by their SHA1 value and downloads them as needed. Did you report any corruption issues you hit? Or have links to anything specific?
what client do you use?
I'd like to see a second class of files which are only checksummed (or timestamped) on 'status', 'diff', 'add' and then binary-diffed on commit for possible compression (or perhaps deduped with checksums of blocks) with features like 'git/hg binary-add somefile.jar' to differentiate.
Why do you say it's bad practice? Game assets come to mind, these aren't derived artifacts and you can't build the game without them.
This is one of the cases where git's inability to handle rewriting a public history is a pain. The art and programming departments should be able to develop along separate branches, and the programmers should be able to pull a tree from art that has everything between tags squashed (but in a reversible fashion in case there's a need to bisect something later).
Historically I've dealt with the problem by keeping code and assets in separate repositories (using different version control systems). While it's worthwhile to do so in order to use git for code it really is a pain.
To answer a couple of levels up, it's not bad practice because it isn't useful... there are many circumstances where having big binary blobs versioned and associated with code. It is bad practice because the tools are all written _for_ code and can get nasty when filled with binaries.
So there's a problem and a solution, but nobody seems to have put together the right magic yet.
It's only a bad practice because our tools don't support it properly.
There seems no reasonable argument that there is some predefined size limit that all assets in version control must, by natural law, fall under.
The criteria I have for version controlled assets is they must 1) be versionable, 2) have a comparison tool, and 3) be strongly connected to the other assets under control.
Yes, exactly.
It operates under a similar principle (committing hashes rather than file contents) but stores the actual file contents elsewhere rather than using a compression/deduplication within the repo itself.
It's neat, but it's more about replacing something like Dropbox than it is replacing git (despite being based on git). It also isn't particularly stable, though I have nothing but respect for the developer.
I backed Joey's kickstarter, but to my shame I have yet to give git-annex-assistant a spin.
Using this workflow with mercurial is really frustrating when I do it - the occasional pull request for a python-based project. A git branch as a concept is really simple, the mercurial ways I just cant wrap my head around (granted I only use it occasionally).
As a platform I like how all the porcellain in mercurial is implemented in a high level python. I can only wonder how productive writing custom porcellain commands in mercurial is given that interface.
Some people mean that Git's model is complicated, which what you're talking about. I actually find its model very simple, some people find it complicated, but at any rate, I certainly think Git and Mercurial have comparable complexity in the model.
What most people mean, though, is that Git's UI is complicated. To be blunt, I think this is simply objectively correct. For example, "git checkout foo" might mean go to the foo branch, or might mean revert a file called "foo". There is no way to know. If you want to be sure to revert a file called foo, you can do "git checkout -- foo", I believe, but I don't know a branch equivalent (and it at any rate won't be symmetrical). Want to create a new branch? That's "git checkout -b newbranch". Want to delete it? That's "git branch -d newbranch". There are tons of things like this in the Git UI, where commands have basically arbitrary parameters in different contexts.
Way back when Git was first created, the plan was for what is now called Git to be the underlying implementation of a higher-level UI. I really, profoundly wish that had actually panned out. Instead, Git's low-level commands gradually grew more user-friendly until it hit the "good enough" zone. That's what people usually mean when they say Git's complicated.
Want to revert your changes? That's unpacking from the repo to the working dir. Want to switch to a different branch? That's also unpacking files from the repo to the working dir. The only difference is that in one case the current working branch stays the same and in one case it changes (or you could look at it as in one case you set it to the current value and in the other you set it to something different).
This is why people say you need to understand the underlying model of git to grok it. The commands make perfect sense from the perspective of the data model.
Also, want to create a branch? "git branch new-branch". Delete it "git branch -d new-branch". The "checkout -b" is an optimization, just like "rm -r x" is an optimization of "find x | xargs rm" (why do two commands when you can do one?).
The only thing I don't like about git UI is that it can't decide what to call the "index". It's called "index", "staging area", and "cache" in various places (including command line flags), which is confusing.
Making people understand the internal data structures in order to understand a UI is... not good.
The "underlying model of git" that David refers to (blobs, trees, commits, tags, branches, HEAD, etc.) aren't seen as internal data structures to be glossed over by a UI. They're the very essence of git.
Users don't need to know the 'very essence' of a tool inorder to use it. Do you think that people who drive cars know in detail about how car works? They just need to know about starting a car, making it go forward, backward, parking etc and that's what most car drivers care about.
Git is like driving a manual transmission. In a manual car you have to understand that the clutch disengages the motor from the drive shaft and how the different gears work in general. It's not rocket science and most people can pick it up.
Git is nothing like driving manual.
> In a manual car you have to understand that the clutch disengages the motor from the drive shaft and how the different gears work in general.
No, you most definitely don't, that's complete lunacy. The vast majority of (manual) drivers[0] have no idea how things work and they don't give a fuck. Different gears are for "go faster" and "go slower", and the clutch is for "change gear". People learn to do it right because the alternative is to stall or get a horrible grinding sound and pay top bucks to get stuff fixed, not because they understand how things work under the hood.
[0] in countries where it's the norm, not in countries like the US where manual is for nerds and passioned
Sorry, that's exactly what I meant by "the different gears work in general". You need to understand that underlying model before you can make it work, even if it's only intuitive, and not cerebral.
> the clutch is for "change gear"
That, and it's also "don't stall when I stop".
I think most people understand that it disconnects the engine. I don't think that's as crazy a leap as "understanding the details of how a gearbox works" (which I don't even know, since I've never looked in one).
Anyway, I maintain that my analogy is apt.
No, you only need to understand what the final effect is. The vast majority of drivers neither understand nor care that gearbox speeds change the conversion ratio between the engine's rotating speed and the axle's, if you did and had to gear speeds would be labelled by their conversion ratio not 1-6 (and then you'd have to include the axle's conversion ratio in the mix). And I don't doubt that a Git-based gearbox would do exactly that, and that going in reverse would require either using an inverter or would require using a completely separate reverse transmission.
> Anyway, I maintain that my analogy is apt.
And I maintain that it is not.
I think professional drivers do indeed know how their car works. They certainly should.
We are talking about professional developers/programmers/devops here, right?
As to git's basic model - a dag of commits, with branches being little more than auto-updating labels for commits - that fit into place in the first 5 minutes of usage. My learning time with git was shorter than any previous source control I used, except for sourcesafe, in so far as that is a SCM.
If you want to revert/reset a file, why don't you just use the reset command? That seems like a more direct way than using checkout.
If you want to create a new branch, why not use "git branch branch-name"? As another said, the "git checkout -b branch-name" is just shorthand for convenience.
I've been learning git over a mere 2 months (I came from svn), I'm finding its UI to be very pleasant.
As someone else said, that index/staging/cached situation is bad, though.
* reverts one file
* reverts all files
* removes changes from the index, but doesn't otherwise change anything at all
* moves your branch, without changing any files
* moves your branch and changes all files
This rather makes my point."git checkout -b", meanwhile, is not a shorthand for "git branch something", but rather "git branch --track something origin/something", unless there is no "origin/something", in which case you are again correct.
Because that's what git tell you to do when you run "git status"
# (use "git checkout -- <file>..." to discard changes in working directory)
Some time around year 1960, the plan was to develop a language with syntax ("M-expressions") on top of the Lisp S-expressions.
We are still waiting, I guess.
Has anyone attempted to actually implement this? Building a 'VCS' on top of git but with a simple 'UI'?
[1] https://gist.github.com/russelldavis/e5173ce1269fae67baaf
You don't have to use Python to add more commands on top of Mercurial. Its CLI is also its API, just like git, and you don't typically have to use Python to write tools on top of hg, just like you don't typically need to write C to hack on top of git.
The stdout of hg is guaranteed to remain stable, so you can script it by just parsing stdout. The options are guaranteed to remain stable, so your old tools won't need to be updated in case hg's CLI some day changes, because its CLI never changes.
There are also a bunch of hg CLI commands that start with debug (e.g. hg debugparents or hg debugdag) that you can use to directly manipulate hg's internal revlog data structure and do fun things like create a corrupt repo, if that's what you want to do. :-)
Any first hand experiences with this?
http://inversethought.com/hg/revset/file/2ddbf1893f3b/lol.py
http://jordi.inversethought.com/blog/on-gitology/ explains git's complexity extremely well (without even getting into horribleness of the command line UI). The following is a key quote, though you should read the whole thing.
The following gitological concepts are not particular to git:
* repositories (repos)
* commits or changesets (csets)
* directed acyclic graph (DAG)
* branches
* pushing and pulling changes
* whole-repo tracking, not individual files
* rebasing csets
* pushing and pulling csets
The following are purely gitological and add unnecessary complexity, in addition to eventually being unavoidable:
* Exposing the index/staging area
* Exposing other implementation details: blobs, trees, commits, refs
* refs and refspecs
* Branches are refs
* Detached HEADs (a.k.a “not on a branch”)
* Distinguishing remote and local tracking branches
* Choosing which branch to pull onto
* Bare repos
* Hard, soft, mixed resets
* Porcelain vs plumbing
But after a while of working with both, you start seeing git and hg as relatively equal products, just with different UI quirks.
I eventually moved most of my projects to git, though. Not because I thought it was better, but because it was more widely used. And being the collaborative type, using git was the path of least resistance.
$ git init --bare --shared foo.git
Initialized empty shared Git repository in foo.git/
$ cat foo.git/config
# ...
[receive]
denyNonFastforwards = true[receive] onlyFastForwards = true
However, we are a small team consisting of mostly Linux kernel developers, so that may influence the level of trust we put in not screwing things up. We also work pretty much independently; were that not the case, this would get ugly.
What team is that, and what are you working on?
Amazingly useful for making sure patches are easy to apply while following a remote branch.
My biggest problem with hg is the lack of real topic branches and how they become impossible to delete -- and having to use things like quilt on top to try to make local branches more sane -- but it's frustrating because it's not the same thing.
I consider Mercurial Queues as one of Mercurial's youthful mistakes. It was ok in 2005, but we have much better things with bookmarks, histedit, and rebase and even with hg commit --amend.
Evolve is basically the last nail in MQ's coffin.
We were discussing the advantages of Mercurial having publishable rebases.
That can be a feature or a bug.
This would be wildly useful for a public branch that needs periodic rebasing, because unlike a git rebased branch, you'd have a history of the rewrites.
On the other hand, most users who locally use git rebase -i to transform a local series of WIP patches into a sensible patch series for submission do not want any record of the intermediate commits (which may not bisect, or even build, and which may have commit messages like "WIP: try fixing it again"). git makes it easy and sensible to commit early and often, and then sort out a sensible patch series from the result.
For the case when you want to push out to the world… Well, that's mostly for catastrophic things—oh, no I accidentally checked in the private key! In that case I also don't want the history of that kept around.
So I don't know. I like the idea, but I'm not sure when it's applicable.
That's just a habit you acquired because right now rewriting public history is a "problem". It shouldn't be a problem. In fact, it's something people do, e.g. how about being able to edit a pull request as it's being discussed and it being ok if that pull request gets merged as it's being discussed?
There are all these different levels of "saveyness" and the lower levels are just not as interesting.
In particular, I don't want stupid untested typos in the commit history because they just aren't interesting or helpful.
But I'm willing to say that I'm probably missing something… I just don't see what it is yet. :-)
Those hidden, obsolete revisions are not shown on your DAG and they are generally not pushed nor pulled. In most respects they behave as if they were not even there. It is only when you need them that you can show them or go back to them (by using the --hidden flag of some of mercurial's command such as hg log or hg update). This gives you a nice safety net (since rewriting history is no longer a destructive operation) that you can use _if you want_. It also makes it possible to rewrite revisions that you have already shared with other users (since when you push a successor revision you also push the list of revisions that it is the successor of).
I think evolve is a significant step forward on the DVCS paradigm as it enables safe, distributed, collaborative history rewriting. This is something that, AFAIK, was not possible up until now.
http://git-scm.com/blog/2010/03/17/replace.html
It requires some manual setup on all checkouts for the changes to propagate automatically.
With Evolve, there is something similar to .git/refs/replace, called obsolescence markers, which may or may not indicate which commit replaces the obsolete commit (some commits are replaced, others are just pruned). These markers are created automatically every time you rewrite history. They don't have to be created manually like with git replace. Moreover, the obsolete commits are hidden from the UI unless you pass the --hidden argument to commands. Lastly, these obsolescence markers are propagated with push and pull operations. It doesn't seem to me like git replace can work over the wire?
I thought that being in the /refs/ namespace would make them eligible for easy synchronization once set up, but on second thought it doesn't seem like it. Git examines parents of refs to determine when something needs to be updated, but would use the parent of the object replacement in this case.
I think a mechanism similar to "git notes" would be better, where the ref points to a history of commits with each tree containing files for each replaced object. I've hacked git to do this at one point so we could retroactively edit git commit messages, but abandoned the effort after discovering "git notes".
Since it's really a commit that changes what history looks like, it's safe to push to other users.
More details here: http://mercurial.selenic.com/wiki/ChangesetEvolution
This isn't really a nitpick though, since this means that similar porcelain could be implemented on top of git fairly easily. The underlying data-model supports it.
Changeset evolution has some similarities with a distributed reflog. Like Git, commits are immutable in Mercurial and we can only "change" a commit by creating a new version and then hide the old version. Mercurial "hides" the old version today by stripping it from the repository — the old version is then stored in a bundle in the .hg/strip-backups folder.
This is far from optimal, so a first step was to add a concept of hidden commits. Hidden changesets have been part of core Mercurial for some time now. The evolve extension enables it and actually changes commands to use it. So "hg commit --amend" will normally strip the old commit, but when evolve is enabled, it will instead hide it. This is both faster and safer.
The next step is the introduction obsolete markers. These are small markers that tell you when a commit is succeeded by a better version. When you amend a commit, an obsolete marker will be created that say "the new version obsoleted the old version". This information is something that Git doesn't store, and it is by distributing these markers that we can make Mercurial more intelligent. As an example, if I amend a commit that you have already based work on, then evolve will know that it should rebase your work onto the successor I created. It will tell you about this when you pull from me and get the new version along with the obsolete marker. Your commits will be called "unstable" as that point, meaning that they are descendants of a commit marked obsolete (they descend from the commit I amended and thus marked obsolete). You can run "hg evolve" and it will figure out that it should run "hg rebase" behind the scenes.
Seen like this, I would say evolve is similar to what happens in Git when you edit history, but with some extra meta data that will allow you to edit shared history with confidence.
In hg "rebase" just means "change the base" not "rewrite commits". So I assume you mean "rewrite" in general.
With evolve, the obsolete commits stay around foreverish, but they slowly fade from history as new people clone or pull, since obsolete commits don't get pulled or pushed by default.
> If I accidently commit "the keys to the kingdom" how do I get them out of the history?
Mercurial never actually removes any functionality, since it's got the deepest commitment to backwards compatibility I've ever seen. Thus, you will delete commits the same way you do now: hg strip --no-backup. That deletes commits with extreme prejudice, locally. Now you just have to run this in every copy of your repo, including remote ones, but the genie-out-of-the-bottle problem is one you can't avoid with a DVCS.
rebase doesn't mean "rewrite commits", it means "create new commits based on these ones, based of a new base". Your original commits are still there, and are pushed to the remote, but are GCed after a certain period (default 30 days?) if they are not referenced from anywhere. Since unreferenced commits are pushed to the remote, you can easily restore those commits within the GC period.
My point is that git says "rebase" even when the underlying base of the commits affected is not changing. This is an artifact of the UI, since the command to rewrite in git is typically git rebase -i.
Mercurial will also (very helpfully) create a backup bundle of the changeset you strip, so you will need to securely erase that as well.
Now if only Atlassian's bitbucket was as popular as github!
[1]: http://mercurial.selenic.com/wiki/Phases [2]: http://mercurial.selenic.com/wiki/PublishingRepositories
For git there is ticgit, which is a tracker that lives as a git branch.
Its definitely great to have several tools available and see a really productive evolution for good version control systems.
Identity, Authentication, Privacy, Subscriptions, Notifications, Contacts, Invites etc.
That's what we're working on at Qbix
The Bitbucket devs are nice guys and have said that they do listen to issue feedback like this.
[paths]
default = http://live.hglabhq.com/hg/hgsharp/hgsharp
default-push = http://live.hglabhq.com/hg/hgsharp/hgsharp
When you create a new repository with `hg init`, the `[paths]` section is, naturally, empty, hence the complaining.