Git's initial commit
github.com
github.com
http://selenic.com/hg/rev/0#l10.1
The revlog data structure from then is still around, slightly tweaked, but essentially unchanged in almost a decade.
Nowadays git has references (branches), and hg has bookmarks which are the same, plus hg also has the option to label every commit with a permanent branch name. They also still have branching-by-cloning, and if you listen to Linus's original Google code talk about Git, you can see that he conflates "branch" and "clone" because that's what he originally envisioned! Even in 2007 he was still thinking in bitkeeper terms too. I bet that branching with references was Junio Hamano's idea, after Linus did the code hand-off.
I find branching-by-cloning a bit more natural in hg, because you can push to any repo. It's useful for quick, throwaway, local, easy testing out of ideas. In git, you can only push if your push doesn't modify HEAD, which typically translates into only being able to push to bare repos.
My usual development routine is to make a ton of small commits that add up to a small set of good commits, to promote bisect-ability. I do dozens of rebases, squashes and amends when working on a topic branch. I have to use Mercurial for one of my clients, and it's a nightmare doing my development model in an SCM where I can't toss commits around willy-nilly like I can in Git.
Yes you can. `hg histedit` is a lot like `git rebase -i`, and `hg rebase` is like `git rebase` without -i and `hg commit --amend` is a lot like `git commit --amend`.
There are also some really cool things that we're working on with hg:
hg just starts out more user friendly, and puts the rest in extensions. I like it more!
ok, hg is a bit slower
(I've also spent some time thinking about how it's kind of a hack, and what we can do to make it better: http://akkartik.name/post/wart-layers)
I enjoy diving into reddit every now and again. But I use github for work (and code for fun, although it's 'serious' fun). Although open-source collaboration is a fundamentally social activity, I think that mixing source control with a social network does inevitably leads to these kinds of comments. And I wouldn't dream of mixing that up with my professional identity.
Maybe it's just a marker of how versatile github is, and the community of people who write programs and put them in source control.
git rev-list --max-parents=0 HEAD | tail -1> git rev-list --reverse HEAD
Yeah, I think you're being a little mean. If you browse to that user's GitHub page, it looks like it's just somebody new who's excited about software. Good for them.
The comments are pointless, sure, but also harmless. Similar comments might crowd out productive discussion if they were on (say) the head of the master branch, but I doubt that any serious development is happening on git's initial commit anyway. Let the new people have their fun.
As far as newbie disruptiveness goes, it could be far worse. When I was getting started with Linux, I posted this cringeworthy gem to LKML, now enshrined in the archives for all eternity: https://lkml.org/lkml/2000/10/22/69 If newbies today are merely posting "yay, git!" and "thank you!" to a secondary forum where it doesn't disrupt development, I'd say they're doing pretty well in comparison. :)
As far as disruption, it did occur to me later that somebody may be getting notification emails about these comments. But it's not too bad, as I assume they could just send the emails to /dev/null, since Github is not the official host of git. (As a tangential note, I sort of wish Github would handle this better. So many Github-mirrored projects end up with something like "don't submit pull requests or open issues here, they will be ignored" in their repo description.)
Thus, successfully self-hosting a version control system is some measure of evidence that the developers know what they are doing and can manage the changes. (And thus they understand change management and we can trust them to be working on version control software.)
http://en.wikipedia.org/wiki/Self-hosting
"Other programs [than compilers] that are typically self-hosting include kernels, assemblers, command-line interpreters and revision control software.
[1] https://en.wikipedia.org/wiki/Magic_number_%28programming%29...
The readme is the best explanation of git I've seen.
-> % git cat-file -p 8c48d1a36c3d11db44c75a431d4f09cb0035222f
tree 288c2d5379768f685f391bdbffd31b8965318c63
parent 002ae35061beef02453b7fb1045a50fa2f7f30f8
author Denis Bilenko <denis.bilenko@gmail.com> 1246939605 +0700
committer Denis Bilenko <denis.bilenko@gmail.com> 1246939605 +0700
MANIFEST.in: include libevent.h and libevent-internal.h
-> % git cat-file -p 288c2d5379768f685f391bdbffd31b8965318c63
100644 blob 6e543dc13df1b556fd95530061ac0c77a9178309.hgignore
100644 blob 79c7beb2227ce149c7a71e58e2f7379071b7a189MANIFEST.in
100644 blob 0d05178544942a035a82599900bec27fbac1c9c5README.eventlet
040000 tree edb8f37fa622315dcf7bf4f7316d5e85c48cfdbdexamples
040000 tree 64cf252d77a4162099442bb0153985fc20ed5ba3gevent
040000 tree 261052e04b4aece469b2e767e394aafbc9d88a32greentest
100644 blob 488e805c563dfeeb6af5e7a1a8953b706d9676e3setup.py
-> % git cat-file -p 6e543dc13df1b556fd95530061ac0c77a9178309
syntax: glob
*~
*.pyc
*.orig
dist
gevent.egg-info
build
htmlreports
results.*.db
gevent/core.so
And yeah it's still very similar though it currently doesn't store the objects individually but rather packs them together.http://alblue.bandlem.com/2011/08/git-tip-of-week-trees.html
I'm not sure when that change was made but it must have been very early on, because the repository format has been basically stable for many years now.
It's actually really great to see that the model hasn't changed much (there must have been a long phase of thinking before though)
If you want to go deeper, you can check out this page:
https://chrome.google.com/webstore/detail/hackbook/logdfcelf...
* +Side note on trees: since a "tree" object is a sorted list of +"filename+content", you can create a diff between two trees without +actually having to unpack two trees. Just ignore all common parts, and +your diff will look right. In other words, you can effectively (and +efficiently) tell the difference between any two random trees by O(n) +where "n" is the size of the difference, rather than the size of the +tree. *
Um, What?
Hence diffing arbitrary commits with git is always O(N) in the number of changed files, regardless of the number of interstitial commits.
I'm no git internals expert, but I suspect for a flat list of files the complexity is still O(n) where n is the number of files (not changes) because at very least you must check that n checksums are the same.
Sure. The constant factors make a huge difference though - even if you've cached all the data in memory walking all those structures and diffing the actual file data is going to be enormously slower than simply walking a list of hashes, so you're really saying that the total time is big * O(number of files changed) + small * O(number of files). If small*N ~ big then it's reasonable to just disregard that cost - it's going to be lost in the noise.
If you want to see the commits going forward from here.
echo https://github.com/git/git/commit/$(git log --pretty=format:%H | tail -1)http://jcooney.net/post/2011/06/22/First-Check-in-Comments-f...
Or am I over-thinking it?
In any case, I believe he was joking. The odds of a sha1 collision are very very low.
I wonder if there are any git sha1 collisions out there in aggregate, say across all of github. Would they even notice if there were?
Despite the incredibly high number of all commits there must be, I think the chance of a collision is still very unlikely. 2^160 is a pretty big number.
On the other hand, 2^80 is "only" approx. 1.2 * 10^24. Still, good luck colliding with that without big effort.
In hindsight, it's good that git didn't choose MD5, since collisions for MD5 can be generated almost trivially now. However, the decreasing security of SHA-1 could be a concern for the future.
> Source control management systems such as Git and Mercurial use SHA-1 not for security but for ensuring that the data has not changed due to accidental corruption. Linus Torvalds has said about Git: "If you have disk corruption, if you have DRAM corruption, if you have any kind of problems at all, Git will notice them. It's not a question of if, it's a guarantee. You can have people who try to be malicious. They won't succeed. [...] Nobody has been able to break SHA-1, but the point is the SHA-1, as far as Git is concerned, isn't even a security feature. It's purely a consistency check. The security parts are elsewhere, so a lot of people assume that since Git uses SHA-1 and SHA-1 is used for cryptographically secure stuff, they think that, OK, it's a huge security feature. It has nothing at all to do with security, it's just the best hash you can get.
Whether or not it's sloppy is up for debate and just a matter of personal preference.
I personally like being able to do it since it allows me to do away with the 2 extra lines auto indent puts in if i add brackets. That's a 50% reduction for a 4 line if. Maybe I should just buy a bigger monitor.
https://github.com/owen2/little-braces
Also, it is available from the online add-in manager if you search for "little braces." The VS2013 community edition is just in time :)
I couple it with the indent guideline plugin for best effect (braces are super small, light lines to track indent level, 2 space indent...).
Not _line_, _statement_. Consider
if(flag)
foo(); bar();
and if(flag)
foo =
bar +
baz;
That first example always calls bar().Warning: I haven't tested this, and am beginning to doubt a bit. It must be correct, but why, then, don't I remember seeing this in underhanded C contests? Combining that with macros allows you to hide the semicolon.
int main() return 0; int main(argc, argv)
int argc;
char **argv;
{
int local;
} Do not unnecessarily use braces where a single statement will do.
if (condition)
action();
[0] https://www.kernel.org/doc/Documentation/CodingStyleI think a good language shouldn't have braces to mark blocks in first place. Given indentation,they are redundant most of the times and they just contribute in clunk. This is exactly the case with Python and hence this is essentially a default style and people hadn't be complaining about it's causing bugs.
C:
if (condition)
statement_1();
statement_2();
Python: if condition:
statement_1()
statement_2()
Personally, I always use braces in C and C++, even though it is more clunky. I want the assurance. I also frequently have to make changes to code that does not use braces, and then I have to add the braces in because I am adding statements to a conditional. To me, that is more clunky.Wrt. using { }, I omit them if I put the block to be executed on the same line, and usually I do that only with special cases, e.g.
if(expr1) continue;
if(expr2) throw new RuntimeException();From this day forward, let us proclaim to always use brackets; so that intent is more obvious to the reader.
If you're worried about bugs, there are other things in C/C++ to criticize first ;)
if expression1 expression2
it can be fiendishly hard to determine the boundary between expression1 and expression2. stupid. contemptible and despicable.
That sums it up quite well. Every day I pay thanks to The One Who Programmed Me that my workflow doesn't put me in need of that shitload of crap that is git. I pity those who do need git.