Git In Two Minutes (updated after 8 years)
garyrobinson.net
garyrobinson.net
Saw this though for switch/restore:
"THIS COMMAND IS EXPERIMENTAL. THE BEHAVIOR MAY CHANGE."
Update: based on other comments, I WILL seriously look into replacing the use of checkout. It sounds like these newer features are probably stable enough for the simple things that are being done in the blog post, and that updating the post will more it a bit more useful as a starting point for learning git.
I'm thinking of posting the modern syntax as an alternative rather than a replacement to checkout, but I have yet to investigate how stable things are now... will do when I have time.
But meanwhile, the checkout method is fine, it works, and it's well-established for many years.
"So if you're writing a script that needs to work with dozens of past and future Git versions, use git checkout. If you need to teach humans how to talk to Git, use git switch. Some details of some flags may change in the future, but I'd argue that that'll be a smaller mental challenge than trying to teach which parts of git checkout do what."
I'll try to use more switch/restore
I've actually used "git stash" more than it should be (I'm probably applying real bad "p4/g4/svn" like ideas in my head to the development. As soon as I go into project with few more people, and I'm lost, though I was able to make few PR's in github for things - but everytime had to le-learn the process).
Some dedicated/stubborn devs also used to (maybe still do?) manage local history in a git-based tool with pushes on demand to a g4 changelist for review.
git stash
git pull
git stash pop
# running away to hide myself from the git gods....https://stackoverflow.com/a/30209750
https://leosiddle.com/posts/2020/07/git-config-pull-rebase-a...
See also the reddit thread linked by a nephew comment.
Switch seems straightforward but I don't understand restore. Is it possible to describe restore in terms of checkout?
Changes not staged for commit:
(use "git add <file>..." to update what will be committed)
(use "git restore <file>..." to discard changes in working directory)
...
Another example where it recommends "switch" git checkout f9b45dd
Note: switching to 'f9b45dd'.
You are in 'detached HEAD' state. You can look around, make experimental
changes and commit them, and you can discard any commits you make in this
state without impacting any branches by switching back to a branch.
If you want to create a new branch to retain commits you create, you may
do so (now or later) by using -c with the switch command. Example:
git switch -c <new-branch-name>
Or undo this operation with:
git switch -Sometimes, they are really stupid, and they will checkin passwords.
With the RCS archives, I can use vi (or nano, or any other reasonable editor), and remove this foolishness.
When I run "git cvsconvert" any foolishness is ENGRAVED IN STONE.
Removal is possible in git, but not easy.
This is my problem. THERE ARE SO MANY IDIOTS. What can I do?
EDIT: For Windows-centric users of git, you need to run this in every and all repositories RIGHT NOW.
git grep -i password $(git rev-list --all)
[actually, everybody should try it]
edit: This is interactive rebase, where you can rewrite history. Instructions are in the file you are editing once the command is executed. As a rule of thumb, don't rebase on production, however with such a SNAFU you probably need to.
Still, a rebase is a critical part of fixing this. If you rebase every branch that has the offending data in their history, and recreate any tags that have the data, then the offending data would be present only in unreachable objects (which can still be found through the reflog or knowing a commit hash). Eventually, git's garbage collection will clean up such unreachable data. You can force this behaviour earlier by using the `git gc` command. By default git will not remove recent objects (which is why you can reflog your way out of a botched rebase), but you can override this behaviour as well.
Of course, all of the above assumes you are working on your own repository. Given git's decentralized design, you need to clean up and garbage collect on every clone that did a fetch since the offending data was checked it. Worse, each clone also keeps track of remote branches, and considers those to be reachable as well, so a garbage collection will not work correctly until a given clone fetches from all of its remotes after those remotes removed the offending data [0].
Further, there is not good tooling to check that you actually did this correctly, so when you are done, you need to hope you fully purged the data.
The plus side of all of this is that once something is checked in, it is extremely difficult to accidently delete it. The downside is that it is also difficult to deliberately delete it. Also, it is fairly easy to make it difficult enough to find to negate the day to day benifets of it still existing, which does very little to protect against a motivated attacker.
[0] Fourtuantly, you do not need all of the remote to have garbage collected, so you could have every make the data unreachable, fetch on all remotes, then garbage collect.
If there is some foolishness, git propagates it EVERYWHERE.
Every git repository is equal. I have no control of proliferation.
Sigh.
If I stapled your HR record, including your bank account and routing number, to our shared refrigerator so I could remember it, would you think that appropriate?
Really, this was the context of our meeting.
Now I am left with the aftermath of his departure, and knives at my sides.
edit: mistakes happen always it is better to change the workflow in such a way that a the consequences can be contained.
The other reply comments are saying the same thing, but your (parent) comment needs to be called out.
If possible, switch from using passwords or tokens without expiry to using ephemeral or time-limited tokens such as machine or pod identities, JWT tokens, IAM service accounts, public/private key pairs (if you can get by with only a public key in the repo) or two-factor authentication. Consider distributing time-limited passwords with Hashicorp Vault.
Some teams might use a cloud-hosted secrets manager or password manager like 1password to distribute passwords and then have code load the password as it runs on a developer machine. GitHub has a secrets scanner, if you pay them enough money for “Advanced Security” such that secrets can never be pushed upstream if recognized as such.
Also, converting from any repo format to another repo format requires care and multiple repeated attempts. A lot can go wrong. Reposurgeon (from a sibling comment) is highly recommended but also not easy to use. It takes a lot of attention to detail to really get the details right.
We had the same problem in TFS.
I found them, I warned them, my riot got us kicked off our shared server.
We migrated to our corporate TFS cloud, and they kicked us off for precisely the same reason.
This just won't/can't stop.
Here’s a brief précis on getting started: https://gist.github.com/ryfactor/f70529438f254d44c0077176508...
It's such a common problem that there's been many tools and workflows that have been set up to mitigate it. This situation often results from not having those tools and workflows in place to give even the "idiots" a "pit of success" to fall into.
e.g., pre-commit hooks installed on dev machines by team policy, and pre-receieve hooks on central repos, both along with CI jobs running truffleHog etc. And even that is downstream of proper code design of minimizing credential use, and focusing/localizing it into files which are already .gitignore'd
I find I don't ever bother with managing local branches for anything in my own workflows. I just git-reset --hard to bounce between fetched copies of upstream branches as needed. Stuff I need to work on for more than a few commits in an afternoon gets pushed to a remote branch regularly anyway as part of general safe development hygiene.
It sounds like switch and/or restore are probably stable enough that they could be used instead of checkout for the specific tasks I'm trying to do in the blog post, so yet another update will be warranted when I get a chance!
It's completely usable as-is, so I hope no one is scared off by thinking it should incorporate switch and/or restore. It's usable now, and if you do feel like saving a link to it for future reference, I expect that there will be another way in the future to do a couple of the tasks using the more modern syntax.
A couple times since then, I've noticed there was something I was using that wasn't in it, but that could be added without making it significantly longer or more complicated. So there were a couple of updates.
I did another one today, and had the thought that since this guide now includes the benefit of 8 years of practical experience in actually using it, while still essentially being "Git In Two Minutes", it might be worth posting to HN. So here it is.
It doesn't say anything about github. It really is for a solo developer who wants to start using git in a very painless way. With the additional goal that you can profitably use git for years without going beyond the described features.
https://www.garyrobinson.net/2014/10/git-in-two-minutes-for-...
* Since it's an introduction maybe consider "--oneline" instead of "--pretty=oneline" for memorability
* Undoing a bad commit - since it mentions a commit message typo, maybe instead redo it with "git commit --amend"
There are other cases where git reset is useful, but generally not for the reasons given.
Once vice I have for solo development that I won't do at work is really crappy commit messages. Like "ditto" for a commit that is similar to the one before or "typo" if I am just correcting a spelling mistake.
You want meaningful history, not unfiltered history.
The harm is that the history is no longer as useful if it's not curated -- whether that's for tracking down the origin of a bug, or just using it to remember what you did a few weeks ago.
It's easy to see the flaw in "keep everything" by taking the argument to its logical conclusion, which would be recording the history of every keystroke in your editor.
If you have false starts or alternate approaches that you gave up on but think you might still want, by all means save them, just not in "main."
Yeah that's what I meant by save them.
> Reverts are not too bad becuase they are pretty meaningful.
I was responding to: "I still revert rather than reset and don't force push".
Say you are working privately on a branch, saving your work, etc. You realize you made a silly error, and need to undo a commit. Creating a revert commit here is a pretty serious error of git organization imo. Typically you do not want to save this work... it is just your private mess as you were working through the problem, and when you are done and have solved it, you want to save the final result, as clean history, perhaps as multiple commits, perhaps as one.
After you have published your work publicly, then yes, you need to do a revert commit if you made a mistake.
If I went down a direction then tried something else, I prefer revert so I don't lose that. That said, tagging then resetting is probably just as good and might have the best of both worlds. You get to keep what you did and you keep the history clean. I only just thought of this now!
Tagging and resetting is of course a bit like stash, but I see stashes as very temporary. It is easy to lost them due to stash pop.
I don't see how that's logical in any way.
I like this guide a lot as a cheatsheet. However, when it comes to beginners I fear it is one of those things that make sense only when you already know what the guide is talking about.
It would take me more than two minutes just to explain a completely new developer what a commit is and why they would want one. And God help us if I throw the output of "diff" at them without warning...
The sad truth is, you cannot explain git in two minutes. I nonetheless admire the author for giving the problem a fair fight.
Or am I just an idiot?
Git is simple, but it takes a lot of experience to appreciate its simplicity. So don't beat yourself up, git is hard and it's okay to be lost.
Yet another example of how, as so often in programming, the "KISS principle" ("Keep It Simple, Stupid") is deceptive: Simple isn't (at least not always) equal with easy.
But yeah, with a more well-thought-out set of commands it would have been a lot easier for a lot of people. I think it's just simply (heh!) that it was released a tad too early, before anyone had thought more deeply about keeping the syntax orthogonal, and the world's been stuck with those choices (largely, choices not made in the first place) ever since, for backwards compatibility.
Git _is_ confusing; it became the de-facto standard because it was the first free DVCS (and DVCSs solve lots of problems that were common to old non-distributed VCSs), not because its UI is particularly well-designed.
Many professional software devs I work with still really have no idea how git works, they've just memorized the 3 or so commands they absolutely need and maybe how to recover when something out of the ordinary happens.
Git is basically a linked list (if you have one branch) and a graph (if you have more than one branch). Then all you are really working with are pointers to certain nodes in that graph.
In git terminology these nodes are called objects and there are different types of objects such as tree objects and commit objects etc. Objects are compressed with a library called zlib and are stored as files within the .git/objects directory, the file name is the hash of the decompressed object.
But, back to the graph and the pointers to certain nodes. A branch itself is nothing more than a pointer either. You can look at the files in .git/refs/heads to see where they point.
Now, what you do with git checkout is changing the node you currently view in the graph. If you make a commit you create a new node. Merging is nothing more than taking two nodes and creating a new node that is connected to both of them.
The main problem i think is that concepts like branch, tag, commit etc are actually just overcomplicated abstractions of the much simpler graph nodes and pointers and are thus confusing.
Git is designed for collaboration. Git's UI is absolutely unintuitive, especially coming from other version control software. But that's irrelevant. Everyone wants to use it because of the network effect.
In terms of _why_ the UI is unintuitive, I believe it's because it intentionally exposes its internal representation. Unlike its predecessors like subversion, Git has infinite flexibility in terms of the workflow that teams of people can adopt. But the flexibility comes at the cost of learning how to use a distributed graph of commits.
Personally, I never learned git by reading recipes of commands, like the OP. I only really grokked it by understanding how it works internally. This article helped me:
Git for Computer Scientists https://eagain.net/articles/git-for-computer-scientists/
IMO flexibility is just a meaningless buzzword used to describe overly complicated and poorly designed products. If a product has crappy UX, they just call it “developer oriented” and “flexible”.
When I read about “flexibility” of a software product, I think that this product is of low quality. And that I will have to spend some time to make it work for me (hours? days? weeks? years?). To me as a user this flexibility does not matter at all.
The product either works or it doesn’t. If it works, I don’t want to spend a career to learn its internals. I just want it to support my use case and give clear step by step instructions in its documentation.
I better spend my time on something more important than exploring how this or another crappy software product is flexible.
I wish the world had standardized on mercurial instead, but here we are.
[Edit: tpyo.]
Locally, git almost makes sense. I make a change, I commit it to the tree of changes. I can roll it back, or branch it off, or merge branches, or restore from a previous point, and so on.
It's when some kind of external repository gets involved where the model breaks down for me, especially if it's public/open source.
Someone wants to help me fix a bug? Cool! They got my code from GitHub and modified it. Now they want their changes to be merged into my version. They submit a pull request. A what? They want me to pull their code into mine? Why isn't it called a push request, where they request permission to push their changes into my maintained repository?
And so on.
If you haven't given them push privileges to your repository, they can't push anything to it. That's why they have to request you to do it... And then it isn't a push, now is it? They're asking you to get code from their repo; that's a pull.
The larger issue is of course that many people never use the git CLI with remote repositories: Far, far too many think that "GitHub is git." Dunno if that's what's happened in your case: You write, "They got my code from GitHub..." If they'd got it from a git repo on your server directly, their "pull request" would consist of an e-mail or something asking you to check out their code, and then you would do that in your repository by typing a command literally starting with "git pull ..."
I use it so often that I have an alias in my shell for it: `gis`.
It shows all the relevant information and generally suggests commands that would perform what you might want to do in the situation you are in (new untracked files, tracked files changes, staged changes that you may want to unstage, merging conflicts, rebase in progress, etc.). It really helps with Git discoverability, which is not negligible when you set yourself to present such a complexe tool in two minutes!
Regarding git there is:
Lecture 6: Version Control [0]
git checkout HEAD -- <filename>
I find it hard to remember things unless I understand their purpose. It looks like it's specifying an argument but without an argument name (e.g. --verbose), unless it's similar to a pipe | symbol and <filename> is being passed to the checkout command as some special kind of argument?https://unix.stackexchange.com/questions/11376/what-does-dou...
This is not git specific, but a common convention for CLI tools. Try running:
> touch -- -i foo
> rm -- -i foo
And compare what happens without the "--".docker run -it nginx —- ls la
startx /path/to/client --with --client-options and arguments -- --server-options -go --here
Here, the double dash is used to separate the client options from the server options (the path to the server binary is compiled-in, IIRC). rm -- -f
Idk the word, just minus minus.I've defintely had this sort of problem in the past :)
Before you downvote, ask yourself how do you list tags, branches and remote
You point is what, that these are needlessly asymmetric? It's true. But they're in my head because I do them every day, and it's not like I'm suffering under the burden of remembering a handful of flags. That's a pretty far cry from "not actually usable", so maybe your hyperbole is a little misplaced?
Anyone with half a brain can understand that if they can't keep something as simple as printing a list consistent you better believe nothing else will be straightforward. Which is my point. NOTHING is straightforward and I haven't met a single person who likes the CLI if they do anything more complex than a commit, push and pull. I know people who still refuse to use rebase and don't understand bisect or blame. They use a GUI to restore files
> git branch --all
Tags are different, since they are actual objects in git, but using ls-remote isn’t tough, and one could create an alias.
As I understand it, they're attempting to make it "actually usable" by adding more consistent (and maybe even "intuitive") options to the syntax. In order to avoid a breaking change, though, they're also leaving the old confusing stuff in. Sure, it's regrettable that it's there in the first place, but understandable that they don't want to break people's scripts from, by now, decades back.
Sure, that makes it a little harder to learn, but it can be done, by just making a conscious decision to ignore those crufty old bits.
> This option can be used to separate command-line options from the list of files, (useful when filenames might be mistaken for command-line options).
This is a common idiom, e.g., to grep a file for a pattern that matches a grep argument,
> grep -- -e file
Edit: fixed dash weirdness.
For instance, "git checkout test.c" will check out the file test.c unless you happen to have a branch called "test.c". (See also people complaining that checkout is overloaded".
Simmilarly, "git add --verbose test.c" will add test.c, and log that to stdout. "git add -- --verbose test.c" will add test.c and --verbose, where "--verbose" is the name of a file.
This isn't git specific, most CLI tools use "--" to restrict the parsing on arguements that follow.
That would be nice. It is a pity most man pages are elitist and lack good useful examples.
It can be really difficult distilling all you know about a certain topic into something short and concise like this.
Because it's about using git as a single-user, single-machine version control tool.
https://jwiegley.github.io/git-from-the-bottom-up/
git's the kind of tool where understanding how your commands operate on the underlying data structures will probably make it a lot easier to use efficiently. (And they're beautifully simple despite how powerful and flexible they are)
As a tool you might be using for hours per week for a few more decades, it's worth the investment going beyond the "in x minutes". I tend to provide both to juniors.
Obligatory: https://xkcd.com/1597/
You don't learn git in two minutes if you don't talk about rebase and remotes.
You want to learn a dvcs in 2 minutes? try fossil.
You don't learn anything in two minutes if you don't even read the headline:
> Git In Two Minutes (For A Solo Developer)