How to undo almost anything with Git
github.com
github.com
http://justinhileman.info/article/git-pretty/
Just another way of presenting similar information. No affiliation, just a satisfied consumer of the info :)
The thing that ended up saving me was our CI -- we autodeployed passing builds to our staging env; so we were able to ssh in and `git push -f` back from staging to our repo.
You either had git configured on the server to git-gc after every push, or were unable to ssh into the server?
If neither one of those is true, then -IIRC- you could have either:
1) Logged in to the server, rewritten the affected repo's branch to point to the pre-disaster commit hash.
2) Pull down the repo's .git directory from the server, rewrite the branch, and force push that.
Would either one of those have been more work than working with your CI system?
git reflog show remotes/origin/develop
This will show the reflog for the branch, then you can find the hash just before your push, check it out, then force push it back.
I've recovered from other developer's accidental force pushes in less than 5 minutes, with no commits lost.
edit: This should be in reply to jordigh's comment above
alias gitpushf='git push --force-with-lease'
In fact, let me ask HN. Does anyone work in a place where devs commit directly to master? I can't even imagine that workflow. Everything I've seen or heard of is PullRequests, code-review, +1, then merge to master.
That won't help you when you think you're pushing to and from your own branch, but are accidentally on develop and force-pushing to the remote develop branch.
Most people only know github and I don't think they expose this locking feature in their UI which is a shame.
Branches also encourage a workflow [in my experience] where the devs implement multiple unfinished features on different branches. That's another form of technical debt. If everyone works right on master, it is more natural to finish a feature before moving onto the next.
I've also had to rewrite history on a git repository to remove things like binary files.
In the top level comment for this thread, I'd say the situation could have been avoided if that developer were cherry-picking to the git repo on his computer, then pushing that branch to the server. The history would be present on both that dev's computer as well as the server, so losing the history on the server would not affect the branch on that developer's machine. Not sure if the comment implies the developer had SSH'd to a live server & performed his cherry-pick there, but that's how it sounds.
The only use I've found for [feature] branches in my day to day work is a place for a junior developer or un-trusted OSS contributors to push changes for me to review before I merge them in. If I'm on an experienced team, I find things go smoothest if we all stay on master.
You can solve that by regularly merging or rebasing master onto your feature branch (as described in Fowler's article).
I think this is probably common in places that have used SVN/Perforce/Whatever since forever and migrated to Git because it's now a "best practice"
A little while ago, I accidentally did a `git clean -xdf` on my home directory (wrong tmux tab). I index my home directory: just the most important config files, among other notes and text files. That `git clean` call wiped half my home directory before I realized what was happening and frantically tried to ^C and ^\ it. I had to find other ways to recover my files.
The deleted files weren't essentially important, which is why I didn't back them up frequently or index them, but they were moderately important. That was a bad day.
I realize it's probably better to dump everything in ~/.config and index that instead, while maintaining symlinks in ~/. It was just the way I had it set up.
Second why ever use force merge?
That said, as the article points out, you need to consider them compromised once they've been pushed and rotate the creds.
To paraphrase from my own comment last month [1]:
"Some time ago I published my blog to GitHub, with my MailGun API key in the config file (stupid mistake, I know). In less than 12 hours, spammers had harvested the key AND sent a few thousand emails with my account, using my entire monthly limit.
Thankfully I was using the free MailGun account, which is limited to only 10,000 emails/month, so there was no material damage. And MailGun's tech support was awesome and immediately blocked the account and notified me, reseting the account after I had changed the keys."
If you're wondering how they were able to harvest GitHub commits so quickly, just read the article linked in that thread [2]. Basically you have bots drinking from GitHub's events firehose and the GHTorrent project. Every commit is monitored and harvested for passwords on the fly.
[1] https://news.ycombinator.com/item?id=8818035
[2] http://jordan-wright.com/blog/2014/12/30/why-deleting-sensit...
I suppose one of the key features of Git is the ability to rewrite history. It makes a lot of sense in the context of an open source project pulling in changes from the wild. For most of us such utilities aren't just useless they're actively harmful.
Never leave me P4. Please God never leave me.
The "if" clause here happens rarely enough that I can quote just the first four words of the sentence: "Git is really bad."
There are 2000 words here on how to undo.
(I'm not so sure what the right English word for "head" is, I hope it still makes sense even if I use the wrong one)
Or another example. Cars are a highly specialised tools. Nobody is allowed to use one if they haven't shown that they have learned to use it properly via standardized exams. Maybe if a company uses a lot of git the mistake is that they don't require people to learn it properly first, either by offering courses or by filtering in the hiring process.
And since I've spent month learning it and now can use it even in potentially destructive situations like "git rebase -i" I can tell you that for me it's way more useful than SVN or CVS ever were.
I have a directory tree full of test data. As the project goes along, the test data will evolve, and thus should go under revision control.
Testing needs to start with known files, so, hey!, git checkout test_data - except that means my latest code revisions need to go into the test_data branch even before they're tested :-(.
Then the tests make their changes to the data, which the tests check, and which I then want to throw away. So: "git checkout test_data -f; git clean -f" -- except that cleans out the source code area as well as the test data area.
I'm thinking the test data should be separate repository. Is that a mistake?
[Edit] I've tried looking at stackoverflow.com, but searching for "git testing" returned ~7000 articles, the first few hundred of which didn't look relevant.
does that help?
git checkout test_data 'testdata/'
to only grab the 'testdata/' directory from that branch. You can do this with files with a pattern, like '*.c', as well.
A workflow example would be doing a rebase after a push is almost universally seen as naughty so why permit it without some kind of UI like --I-really-know-what-i-am-doing=yes or something?
WRT the UI itself, I've read a couple people claiming the emacs magit package is easier to use than the CLI itself, which would isolate the problem to the GUI. I have not personally invested the time into magit and would find comments on that theory by people who have experimented to be interesting.
Heck, the just the contrast between `git add <file>` and `git reset HEAD <file>` is terrible.
Not sure where you're going with `add` versus `reset` though. `git rm --cached ${file}` would presumably do what you want with parallel syntax.
If you are interested in using Mercurial on your own project and are simply worried about hosting, here's a list of providers: https://mercurial.selenic.com/wiki/MercurialHosting
If you are thinking of using Mercurial when the rest of your team uses Git, then I would just suck it up and learn Git (or convince your team to switch). It is not, by any means, the most complicated thing you have to learn for your job ;-).
Back in the bad old days when I was forced to use Visual Source Safe for version control, we used to maintain our code in multiple repositories -- using CVS for development and simply pushing to VSS when features were complete. That kind of tactic is still open to you, but the overhead is rather large for the minimal difference between Git and Mercurial.
Sounds interesting. So was this basically a team of 'bandits' secretly using a non-endorsed VCS to collaborate and get shit done, and then pushing to VSS to keep the pointy-haired boss happy?
As you might imagine (the use of GPL software banned, the use of MS software enforced), the rules were politically rather than technically oriented. There were rumours that we had an agreement with MS to follow these rules and I wouldn't say it was out of the question. By and large as long as you followed the letter of rules nobody cared much after that. Using CVS was rather a major coup, though, given the GPL licence.
Possibly younger people will be amazed at the ridiculous conditions some people worked under back then. After I left that position I swore I would never agree to work with restrictions like that. In fact, the ability to work on free software is one of the first things I bring up as important to me in a job interview. Even a slight hesitation is enough to make me walk.
Honestly, I also think github's defacto status is a major obstacle to adoption.
There are actually projects that let you work with an hg/git repo and push to the other type. http://hg-git.github.io/ for instance. No idea if these things are production ready or not though.
Once you grok that it's more or less an immutable DAG with diffs for edges and hashes for node names, that tags are read-only labels for hashes and branches are read-write labels for hashes, all possible operations are obvious; you just need to find the right incantation.
The concept of git being so simple is what makes working with it much easier than something like svn or cvs, where doing the equivalent of a rebase, cherry-pick or merge of a diff into multiple different branches is sufficiently difficult that I developed my own tools and workflow to get around them. When I had to work with svn writing bug fixes or doing development, instead of committing my work, I saved my work to patch files which I saved / reapplied when I switched branches. I developed scripts to do 3-way merges. I haven't had to do any of that crap with git. Git is far more logical. It's just missing a consistent command-line UI.
That's problem #1. More or less immutable = mutable. Other SCMs limit mutations to additions and use a separate command (svnadmin) to do things that may permanently lose information. The svn repository may be ugly, but it can be relied on to store history.
"you just need to find the right incantation."
Incantation is the right term. As you admit, it's "just" missing a consistent command-line UI.
The combination of these two makes me very weary whenever I do anything remotely difficult in git.
A third thing that scares me is the ease with which people talk about things still being there "as long as git hasn't garbage collected them". To me, that sounds like having a memory allocator with an 'unfree' call that you can use to try and recover accidentally freed memory.
Your GC fears sound like superstition, sorry. Nothing to do with manual memory allocation, and the problems of manual memory management are irrelevant. GC just collects nodes no longer reachable from branches or tags. Very unscary once you understand the dag nature.
https://git-scm.com/book/en/v2/Git-Tools-Rewriting-History:
and you can rewrite commits that already happened so they look like they happened in a different way. This can involve changing the order of the commits, changing messages or modifying files in a commit, squashing together or splitting apart commits, or removing commits entirely – all before you share your work with others.
That surely looks like changing the graph, not just its attributes.
> "Your GC fears sound like superstition, sorry. Nothing to do with manual memory allocation, and the problems of manual memory management are irrelevant. GC just collects nodes no longer reachable from branches or tags. Very unscary once you understand the dag nature."
Thanks for triggering me to reread the documentation. I thought that gc would (potentially) collect all unreachable roots, but rereading https://www.kernel.org/pub/software/scm/git/docs/git-gc.html, I find:
The optional configuration variable gc.reflogExpire can be set to indicate how long historical entries within each branch’s reflog should remain available in this repository. [..] It defaults to 90 days.
The optional configuration variable gc.reflogExpireUnreachable can be set to indicate how long historical reflog entries which are not part of the current branch should remain available in this repository. [...] This option defaults to 30 days.
So, it seems that they work hard to prevent collection of nodes that you may want to refer to.
That makes this lack of documentation:
--auto With this option, git gc checks whether any housekeeping is required; if not, it exits without performing any work. Some git commands run git gc --auto after performing operations that could create many loose objects.
waaaaaaay less of a problem. I have looked hard, but cannot figure out what those 'some commands' are that may do a gc. The best I could find is http://stackoverflow.com/questions/5137447/list-of-all-comma.... That's 5 years old, greps the git source code, and not the official documentation.
It's a powertool for power users. It sks that it became the major tool of our time, but that doesn't change it's design decisions. (And I'm a huge git fan. But I'm also a power user who is happy to spend weekends to learn all the details about such a tool)
git rebase HEAD~N --onto master
Maybe --squash will fix it? Something to look at, I guess.