Linus Torvalds: Git proved I could be more than a one-hit wonder
techrepublic.com
techrepublic.com
Can you or someone else reflect on my file system? I work for the government doing statistical analysis of healthcare data, and there is no VCS where I code, other than how you name the files and where you put them in folders and how you back them up manually.
I am facing a major data-branching event where I'm going from ~40 scripts (R, SQL, STATA) on one dataset, to then three overlapping but different datasets and having ~100 scripts. I just don't know if my brain and wits are up to the task of maintaining these 3 branches with 3 different languages, given all I have is folder/file names and my knowledge reservoir and memory...
I know this is a perfect use case for git, but I've never used it before and no one else in my department uses it. I don't know if I have the time and energy left at this job to implement a new system of VCS AND reproduce my code for 3 different-but-similar projects.
Burnout approaches...
The GUI is stunningly beautiful and functional, and there are more than enough keyboard shortcuts to keep things snappy once you're in the flow. I used to live and die by the terminal, now I am in love with sublime merge.
I used the portable version for a job where I didn't have install rights to the corporate laptop, and it preserved my workflow and kept me sane during my dev work. The portable version can run a little slow, but it's a pretty good solution.
Some truly locked down environments may not allow it but if the poster has other open source tools like R they can probably run .exe files.
Get your department to pay for you and ~2 colleagues to go on a git training course for a few days. As well as teaching you how to use git, it'll give you some time with an expert to look at your problem, and give you some relaxation time helping the burnout, and with 3 of you on the course, you'll likely get buy-in for a new setup.
Beware that git isn't a silver bullet. While it solves a bunch of issues, it causes many new ones - especially when you have lots of people who aren't knowledgeable about version control using it. I wish git had better integration with 'regular files' - ie. so that Mary in the marketing department can update a readme file without having to learn a totally new way of working. I wish you could "mount" a git repo as a drive in Windows, and all changes would be auto-committed to a branch as soon as a file is saved, and that branch were auto-merged to master as long as tests pass. Then people without git knowledge can work as before.
Cool idea for a project
Does it only check files passing tests? I read quickly and didn't see that
git works by creating its own .git directory wherever you create a new git repository, but doesn't touch the files and directories outside of it until you tell it to.
So you can have a directory of old code and you just cd to it and run 'git init', and now you have a git repository in the same directory. It won't be managing any of the files yet, but it will technically be there.
Because git is just a bunch of extra data in a .git directory, and because git is also built as a distributed VCS, the "make a copy of a directory to back it up" methodology actually works pretty OK with git. Ideally you should be using 'git clone' to copy your directories and 'git pull' to keep them in sync, but if you just Control-C Control-V your source code directory, git will actually be just fine with that, and you can still later use git to sync the changes between those two directories.
I'm not going to put a full git tutorial into this post, about how you add files to the repository and make commits, but I just want to convey that while git has a justifiable reputation for sometimes devolving into arcane incantations -- it's actually low effort to get started and you only need to learn three or five commands to get like 95% of the value from it.
Once you learn those three or five commands, you'll find yourself running 'git init' in nearly every directory you make -- for your random scripts, for your free time coding projects, for your free time creative writing projects -- and you'll even find it easy to use on horrible "27 directory copies of the source code with 14 file renames" projects where none of your teammates use git; you can use git yourself in such cases without adding any real friction, and it still helps you even if your teammates just copy your code directory or send you copies of their code directories.
EDIT: One other note: git can also go away easily if you decide you don't like it. You don't need to run git commands to create, edit, copy or otherwise modify the files in your code base, like you do with some other source control systems, so if you can just forget it is there if you are busy and don't want to worry about it, and then later go ahead and add or commit all of your work. If you really don't like it, you just stop running git commands and you're no longer using it: you don't need to 'export' or 'ungitify' your code base. So it's pretty low-risk in that way as well.
- you can put the git directory somewhere other than in your working directory, if you really want to. Or reference a bunch of .git directories in a series of commands without having to change your current directory. Sometimes this is handy (usually for automation or something like that).
- If you're nervous about some command you're about to run—something that might screw up your git tree—just copy the .git directory somewhere else first. You can copy it back to entirely restore your state before the command, no need to figure out how to reverse what you did (assuming it's even possible).
My brain is basically overloaded with stress and I'm headed for burnout...only 18 months into this position. I just can't handle the tech stack, the shitty office, the commute, the feelings of being the worst analyst and the worst researcher in every single room I'm in. It is totally wearing me down. Management said new employees can get work from home after 12 months, then at 18 months I asked, and they revoked their verbal agreement and said they'd reconsider their decision if I made an article and let someone else be first author on it (unethical).
Outside of my complaints...I'm just not a great worker. I just feel that the whole team and department would be better off without me, that I can not handle this tech stack and QoL and its frustrations...govt is a very very restrictive environment and I feel like a circle being jammed into the square hole. I can't implement most of what these comments stated because I can not install anything onto my computing environment...even Python, I have to go through red tape and request special access to use Python instead of R and STATA.
I'm sorry to vent but all of these shortcomings are seriously burning me out.
I sincerely doubt that you. You sound like a conscientious employee in an environment not set up for the kind of work you were hired to do. You also sound like you want to leave your job - which can give you leverage. Not that you should threaten to quit, but that since you are so unhappy, you are willing to quit. That means you can start saying what kind of computing environment you need. Not want, but need.
Personally, I think that having source control is basic table-stakes when writing code as a part of a job.
Otherwise, another poster commented that a git training course paid for by the company could help (+ give you some relief from burning out).
Since the internet is also for acting like you know what you are talking about and offering unsolicited advice, I'll also drop some here. Feel free to ignore it, and I hope you situation gets better, either at your current job or a new one.
I won't speak too much to your work skills, because I don't know you; but feeling like and worrying that you're terrible at your job is a pretty normal experience. You pretty much have to rely on whether other people think you are doing a good job because people in general are garbage at judging their own skill. It's pretty hard to tell the difference between "I think I'm doing poorly and am" and "I think I'm doing poorly and am actually doing fine", without a lot of feedback from people you trust (ideally, your coworkers).
If your coworkers think you're doing fine, well, you can't stop worrying about it, but you'll at least have some evidence against your feelings; if your coworkers think that you're under-performing, they might at least be able to offer some advice on how to do better.
The burnout advice I have to give is in three parts: first, focus on making some small, incremental progress every day; second, avoid the temptation to overwork; third, make sure to invest time in your life outside of work.
The first is both about positive thinking and also about developing good work habits. The second is because it doesn't usually work (you end up doing less with more time, which is even more depressing than feeling like you aren't getting enough done in 8 hours). The third is because you will feel better and be more resilient if your entire identity isn't invested in your job. It's easier to both to avoid burnout and to recover from it when it does happen if your job is only one part of your life.
And remember git doesn't save directories that are empty.
And if git seems too difficult to start with, Subversion can also "host" a repository on the file system, in a directory separate from your working directory.
I understood and was comfortable with SVN within a few minutes (using the TortoiseGit front-end, which I highly recommend).
I wrestled with git for months and at the end still feel I haven't subdued it properly. I can use it reliably but SVN is just so much friendlier.
So my suggestion is go with SVN + TortoiseGit. SVN is your butler. Git is a hydra that can do so much, once you've tied it down and cut off its thrashing heads.
It's not just me, our whole (small) company moved to it it burnt too much of our time and mental resources.
Edit: after learning TortoiseGit, learn the SVN command line commands (it's easy), and learn ASAP how to make backups of your repository!
Getting started with SVN is very quick, but once you need to peek under the hood, you'll find out it's super complicated inside.
Git is just the other way around: the interface is a mess, but the internals are simple and beautiful. Once you understand four concepts (blobs, trees, commits, refs), the rest falls into place.
Recommended four page intro to git internals: https://www.chromium.org/developers/fast-intro-to-git-intern...
Could you explain what you mean by svn being super complicated inside? I presume you mean from a user's not a programmer's perspective; I never found it confusing, ever.
It has it's flaws (tags are writable, unless that's been cured) but it's really pretty good, and far better than git for a beginner IMO.
In Git, no matter how strange the situation, everything is still blobs, trees, commits, and refs. There are very few concepts used in Git, and they're simple and elegant.
SVN to Git is like WordPress to Jekyll - WordPress is easier to use than Jekyll, but Jekyll is simpler than WordPress.
SVN's concepts are straightforward - commit stuff, branch, branches are COW so efficient, history is immutable unlike git (for better or worse) erm, other stuff. Never got confusing.
Yes, I did. Thanks.
However, branching may not be the ideal solution given how you describe your issue. With git branches, we typically dont want to run something then switch branches then run something else. I would say branches are primarily for organizing sets of changes over time.
If you have multiple datasets with similarities, what you may need more than git is refactoring and design patterns. To handle the common data in a common way, and then cleanly organize the differences.
That said I would still definitely want all scripts in git. It is not that hard to learn, lean on someone you know or email me if you need to.
In particular, dvc carefully handles large binary files.
I rarely use functional programming but I certainly see its appeal for certain things.
I think the concept of immutability in functional programming confuses people. It really clicked for me when I stopped thinking of it in terms of things not being able to change and started instead to think of it in terms of each version of things having different names, somewhat like commits in version control.
Functional programming makes explicit, not only which variables you are accessing, but which version of it.
It may seem like you are copying variables every time you want to modify them but really you are just giving different mutations, different names. This doesn't mean things are actually copied in memory. Like git, the compiler doesn't need to make full copies for every versions. If it sees that you are not going to reference a particular mutation, it might just physically overwrite it with the next mutation. In the background "var a=i, var b=a+j", might compile as something like "var b = i; b+=j";
Over time someone discovered that the number of repositories and usage was much greater than they expected. What they found was that non engineering folks who had contact with engineering had asked questions about how they manage their code, what branches were, and etc. Some friendly engineering teams had explained, then some capable non engineering employees discovered that the server was open to anyone with a login (as far as creating and managing your own repositories) and capable employees had started using it to manage their own files.
The unexpected users mostly used it on a per user basis (not as a team) as the terminology tripped up / slowed down a lot of non engineering folks, but individuals really liked it.
IT panicked and wanted to lock it down but because engineering owned it ... they just didn't care / nothing was done. They were a cool team.
I highly recommend git-annex. It is like git-lfs but a bit less mature but much more powerful. Especially good if you don't want to set up a centralized lfs server.
Binary files don't cause the issue, but because binary files don't deltify / pack well significant use of them makes repos degenerate much faster.
(I checked this out the day before I left for vacation, so to be fair, my research might have not been thorough enough to find each and every implementation - but I think it is comprehensive enough to make some preliminary judgement)
My god, the things I've seen in repos. vim .swp files. Project documentation kept as Word documents and Excel spreadsheets. Stray core dumps and error logs in random subdirectories, dated to when the repo was still CVS. Binary snapshots of database tables. But the most impressive by far was a repo where someone had managed to commit and push the entirety of their My Documents folder, weighing in at 2.4GB.
(for those with a severe sugar hangover, I'm being a little bit sarcastic)
https://news.ycombinator.com/threads?id=miohtama
"Git is what a version control UX would look like if it were written by kernel developers who only knew Perl and C"
Back in a day we had Subversion, Mercurial, Bazaar, some others. I used all of these. All of them were more coherent than Git. However they were slower - but not much - and they were not used by the most popular software project in the world. Then, GitHub popularized git and Github become well funded enough to take over the software development world.
Bitbucket, now Atlassian, started as a hosted Mercurial repos. Bazaar was DVCS for Ubuntu, developed by Ubuntu folks.
Will we see another DVCS ever again? I hope yes. Now all software developers with less than 5 years of experience are using Google and GitHub as the user interface for Git. Git's cognitive burden is terrible and can be solved. However Git authors themselves are not priorising this.
Software development industry could gain a lot of productivity in the form of more sane de facto version control system, with saner defaults and better discoverability.
I don't think we should underestimate how much performance differences can color opinions here, especially for something like CLI tools that are used all the time. Little cuts add up. At work I use both git and bazaar, and bazaar's sluggishness makes me tend to avoid it when possible. I recall Mercurial recently announced an attempt to rewrite core(?) parts in Rust, because python was just not performant enough.
In my experience, saving a little on runtime doesn't make up for having to crack open the manual even once. UI is a big cut.
From an implementation POV, it's also generally easier to rewrite core parts in a lower-level language, than it is to redesign a (scriptable, deployed) UI.
But only if the data structure is simple and works well for the problem domain. Bazaar (a DVCS from Ubuntu; I mean bzr, not baz) had a much simpler and consistent UI, but it had several revisions to the data structures, each one quite painful, and it was slow; they were planning the rewrite to a faster language but never got to it. (Mercurial also used Python and wasn't remotely as slow as bzr - the data structures matter more than the language).
Please don't ask me to give up features so you don't have to read the documentation. It is the most basic step of being a good citizen in a software ecosystem.
Maybe, but I used to work in a multi-GB hg repo, and I would have given up any amount of manual cracking to get the git speed up. Generally you only open the manual a few times, but you can sync many times a day. I'd give a lot to get big speeds up in daily operations for something I use professionally.
If one cannot figure one of the most common use case of a version control system without Googling a StackOverflow answer then we have a problem somewhere.
I will say from experience that it's not hard to use git productively with a bit of self-study and only a few of the most common commands. You still have to understand what those commands actually do, though.
man git
The problem is people are unwilling to read the documentation. I have little patience for them demanding I change my workflow to accommodate their sloth.Fortunately, I don't have to worry because the overlap of 'people who don't RTFM' and 'people who are capable of articulating how they want to change git' have so far failed to write a wrapper that's capable of manipulating git trees without frustrating everyone else on the same repo.
And of course they can't: version control[1] is not a trivial problem. So I see no reason for us to demand that someone knows how to do it without studying when we don't expect the same for other auxiliary parts of software development such as build systems or containerisation or documentation.
[1] As opposed to the backup system the link wants to use it as: asking better questions is another important step. There's little reason to checkout an older commit as a developer unless you want to change the history, in which case it's important you understand how that will interact with other users of the same branch. If you don't need it to be distributed, you already have diff or cp or rsync or a multitude of other tools to accomplish effective backups.
(one of my favorite, for example, is that `git checkout` causes silent data loss while every other git command will print out giant errors in that scenario)
I don’t agree bazaar was a UX panacea over git, and it was not just “not by much” slower. Subversion was a piece of shit full stop (especially if you had the misfortune of using the original bdb impl), bested in this regard only by VSS. I think slower “not by much” is the understatement of the century for a repo of any substantial size for all but mercurial on your list.
You don’t even mention perforce, leading me to think most of your experience is skewed by the niche of small open source projects.
Mercurial was a contender... great windows support too. I think it was less kernel that killed it and more github.
I think it was perceived performance that led git to besting Mercurial, which the Linux Kernel team certainly contributed to that drama, including the usual "C is faster than Python" one-upmanship, this especially funny because it was despite most of git at the time being a duct taped assortment of nearly as much bash, perl, awk, sed scripts as C code.
For the same reason that BD won over HD-DVD: «Greater capacity tends to be preferred to better UX», except in this case it's performance rather than capacity.
As just the first example off the top of my head, Pijul [1] is "new" compared to the others you listed. There will likely always be folks exploring alternatives.
> Now all software developers with less than 5 years of experience are using Google and GitHub as the user interface for Git. Git's cognitive burden is terrible and can be solved. However Git authors themselves are not priorising this.
I have the opposite impression, that the Git Team is finally getting serious about the UX, whether its just all the fresh blood (thanks at least partly to Microsoft moving some of their UX teams off of proprietary VCSes to converge on git), or that Git's internals are now stable enough that the Git team feels it is time to focus on UX (as that was always a stated goal that they'd return to UX when all the important stuff was done).
Clear example: The biggest and first announcement in the most recent Git release notes was about the split of `git checkout` into `git switch` and `git restore`. That's a huge UX change intended to make a lot of people's lives easier, simplifying what is often people's most common, but most sometimes most conceptually confusing git command given the variety of things that `git checkout` does.
The Git UX is better today than it was when it first "beat" Mercurial in the marketplace, and there seems to be at least some interests among git contributors to make it better.
git is not the same as those other pieces of software mentioned.
git's default workflow encourages lots of parallel work and making tons of branches (which because of bad naming are confusing because git branches are not what other software calls branches) .
it's a fundamental difference and has increased my productivity and changed my work style for the positive in ways what would never have happened with CVS, svn, p4, hg, etc... all of which I used in the past for large projects.
If you're using git and your mental model is still one of those other systems you're doing it wrong or rather you still don't get it and are missing out.
I'm not suggesting the UX couldn't be better but when you finally get it you'll at least understand what it's doing and why the UXs for those other systems are not sufficient.
CVS was much much much slower; multiple branch handling was horrible until ~2004 (and even on a single branch you did not have atomic commits). Also, no disconnected operation.
SVN was only a little slower than git, but didn't have disconnected operation, and horrible merge handling until even later (2007 or 2008, I think)
Bazaar 2 was, at the time, while comparable in features, dead slow compared to git. But it also sufferend from bazaar1 (branched from arch=tla) being incompatible with bazaar2 and an overall confusing situation.
Mercurial and Git were a toss-up. Git was faster and had Linus aura, Mercurial had better UI and Windows support. But all the early adopters were on Unix, and thus the Linus aura played a much bigger part than the Win32 support.
Github became externally well funded after the war was over. But it was self well funded, because git was more popular (in part because github made it so ...)
Really, I think the crux of the matter is that Git's underlying data model is really simple, and the early adopters were fine with UX ... mostly because those adopters were Perl and C people. So the UX was not a factor, but speed and Linus aura were.
However I believe in long run Hg caught up in the speed and Bazaar was getting a lot of better as well.
SVN merge was nightmare. People avoided doing work that would result a merge as it hurted to get it executed nicely.
The killer feature though was that git didn't put my data at risk while I worked. With the normal SVN workflow, your working directory was the only copy of your changes. And when you sync'd with upstream it would modify the code in your working directory with merge information. Better hope that you get that merge right, because there's no second chances. Your original working directory state is gone forever, and it's up to you to recreate it with the pieces SVN hands you.
So write it.
I'm sorry to be so dismissive but it seems notable that those who like git get along with using it while those who complain about it just throw peanuts from the gallery. If it's obvious to you where git's flaws lie, it should be easy to write an alternative. If saner defaults and better discoverability are all you need, you don't even have to change the underlying structure, meaning you can just write a wrapper which will be found by all the competent developers whose productivity is so damaged that they do what they do when they encounter any problem and search the internet for a solution.
It seems notable this hasn't happened.
Depends, we went from CVS to Git and nightly jobs tagging the repository went from taking hours to being almost instant.
Otherwise I’m hoping for pijul to somehow gain popularity (and a bit of polish) and become mainstream. I guess a motto for it could be “the high quality user interface and semantics of darcs without the exponential time complexity”
But, having said that, I made my peace with the git command line years ago, in part by learning to appreciate aliases:
co = checkout
ci = commit
dt = difftool
mt = mergetool
amend = commit --amend
pfwl = push --force-with-lease
(The first two are my personal hangovers from Subversion.) I also have a "gpsup" shell alias which expands to git push --set-upstream origin $(git_current_branch}
The latter is taken from Oh My Zsh -- which actually has dozens of git aliases, most of which I never used. (When I realized "most of which I never used" applied to all of Oh My Zsh for me, I stopped using it, but that's a different post.)tl;dr: I used to have a serious hate-on for git's command line, but one of its underestimated powers is its tweakability.
EDIT: I found a couple of interesting references for folks who may be curious about this as well. I especially like [2] for its diagrams.
[1] https://stackoverflow.com/questions/20151158/using-git-repos...
[2] https://www.kenneth-truyers.net/2016/10/13/git-nosql-databas...
What it also doesn't give you for free is sensible search indexes. I do think though that combining it with a search index could be very powerful.
"I will, in fact, claim that the difference between a bad programmer and a good one is whether he considers his code or his data structures more important. Bad programmers worry about the code. Good programmers worry about data structures and their relationships."
--- Linus Torvalds, https://lwn.net/Articles/193245/
But computers don’t “compute”, they don’t do math. Computers are simplifying, integrating machines that manipulate symbols.
Data (and its relationships) is the essential concept in the term “symbolic manipulator”.
Code (ie a function) is the essential concept in the term “compute”.
Not trying to start a flamewar, I just found the distinction you drew interesting.
If one were to say computers do math, they would be saying computers reason. Reason requires free will. Only man can reason; machines cannot reason. (For a full explanation of the relationship between free will and reason, see the book Introduction to Objectivist Epistemology).
Man does math, then creates a machine as a tool to manipulate symbols.
> Reason requires free will
isn't it still kind of an open question whether humans have free will, or what free will even is? How can we be sure our own brains are not simply very complex (hah, sorry, oxymoron) machines that don't "reason" so much as react to or interpret series of inputs, and transform, associate and store information?
I find the answer to this question often moves into metaphysical, mystical or straight up religious territory. I'm interested to know some more philosophical approaches to this.
If objective reality doesn’t exist, we can’t even have this conversation. How can you reason—that is, use logic—in relation to the non-objective? That would be a contradiction. Sense perception is our means of grasping (not just barely scratching or touching) reality (that which exists). If a man does not accept objective reality, then further discussion is impossible and improper.
Any system which rejects objective reality cannot be the foundation of a good life. It leaves man subject to the whim of an unknown and unknowable world.
For a full validation of free will, I would refer you to Chapter 2 of OPAR. That man has free will is knowable through direct experience. Science has nothing to say about whether you have free will—free will is a priori required for science to be a valid concept. If you don’t have free will, again this entire conversation is moot. What would it mean to make an argument or convince someone? If I give you evidence and reason, I am relying on your faculty of free will to consider my argument and judge it—that is, to decide about it. You might decide on it, you might decide to drift and not consider it, you might even decide to shut your mind to it on purpose. But you do decide.
It's not that I reject the idea of objective reality–far from it. However I do not accept that we can 1) perfectly understand it as individuals, and 2) perfectly communicate any understanding, perfect or otherwise, to other individuals. Intersubjectivity is a dynamical system with an ever-shifting set of equilibria, but it's the only place we can talk about objective reality–we're forever confined to it. I see objective reality as the precursor to subjective reality: matter must exist in order to be arranged into brains that may have differences of opinion, but matter itself cannot form opinions or conjectures.
I'll assume that book or other studies of objectivity lay out the case for some of the statements you make, but as far as I can tell, you are arguing for objectivity from purely subjective stances: "good life", "improper discussion"... and you're relying on the subjective judgement of others regarding your points on objectivity. Of course, I'm working from the assumption that the products of our minds exist purely in the subjective realm... if we were all objective, why would so much disagreement exist? Is it really just terminological? I'm not sure. Maybe.
Some other statements strike me as non-sequiturs or circular reasoning, like "That man has free will is knowable through direct experience". Is this basically "I think, therefore I am?" But how do you know what you think is _what you think_? How do you know those ideas were not implanted via others' thoughts/advertisements/etc, via e.g. cryptomnesia? Or are we really in a simulation? Then it becomes something like "I think what others thought, therefore I am them," which, translated back to your wording, sounds to me something like "that man has a free will modulo others' free will, is knowable through shared experience." What is free will then?
"free will is a priori required for science to be a valid concept" sounds like affirming the consequent, because as far as we know, the best way to "prove" to each other that free will exists is via scientific methods. Following your quote in my previous paragraph, it sounds like you're saying "science validates free will validates science [validates free will... ad infinitum]." "A implies B implies A", which, unless I'm falling prey to a syllogistic fallacy, reduces to "A implies A," (or "B implies B") which sounds tautological, or at least not convincing (to me).
I apologize if my responses are rife with mistakes or misinterpretations of your statements or logical laws, and I'm happy to have them pointed out to me. I think philosophical understanding of reality is a hard problem that I don't think humanity has solved, and again I question whether it's solvable/decidable. I think reality is like the real number line, we can keep splitting atoms and things we find inside them forever and never arrive at a truly basic unit: we'll never get to zero by subdividing unity, and even if we could, we'd have zero–nothing, nada, nihil. I am skeptical of people who think they have it all figured out. Even then, it all comes back to "if a tree falls..." What difference does it make if you know the truth, if nobody will listen? Maybe the truth has been discovered over and over again, but... we are mortal, we die, and eventually, so do even the memories of us or our ideas. But, I don't think people have ever figured it all out, except for maybe the Socratic notion that after much learning, you might know one thing: that you know nothing.
Maybe humanity is doing something as described in God's Debris by Scott Adams: assembling itself into a higher order being, where instead of individual free will or knowledge, there is a shared version? That again sounds like intersubjectivity. All our argumentation is maybe just that being's self doubt, and we'll gain more confidence as time goes on, or it'll experience an epiphany. I still don't think it could arrive at a "true" "truth", but at least it could think [it's "correct"], and therefore be ["correct"]. Insofar as it'll be stuck in a local minimum of doubt with nobody left to provide an annealing stimulus.
I will definitely check out that book though, thanks for the recommendation and for your thoughts. I did not expect this conversation going into a post about Git, ha. In the very very end (I promise we're almost at the end of this post) I love learning more while I'm here!
I’ve enjoyed this discussion. It has been civil beyond what I normally expect from HN. From our limited interaction, I believe you are grappling with these subjects in earnest.
This is a difficult forum to have an extended discussion. If you like, reach out (email is in my profile) and we can discuss the issues further. I’m not a philosopher or expert, but I’d be happy to share what I know and I enjoy the challenge because it helps clarify my own thinking.
At least for certain actions and situations, the "direct experience" of free will is measurably incorrect.
Doesn't mean free will doesn't exist (or myabe it does), but it's been established that that feeling of "I'm willing these actions to happen" often times happens well after the action has been set into motion already.
More basically and fundamentally, I'd suggest that no, numbers aren't symbols: numbers are numbers (i.e. they are themselves abstract concepts as you suggest), and symbols are symbols (which are much more concrete, indeed I'd say they exist precisely because we need something concrete in order to talk about the abstract thing we care about). We can use various symbols to represent a given number (say, the character "5" or the word "five" or a roman numeral "V", or five lines drawn in the sand), but the symbols themselves are not the number, nor vice versa.
This all scales up: a tree is an abstract concept; a stream is an abstract concept, a compiler is an abstract concept — and then our business is finding good concrete representations for those abstractions. Choosing the right representations really matters: I've heard it argued that the Romans, while great engineers, were ultimately limited because their maths just wasn't good enough (their know-how was acquired by trial-and-error, basically), and their maths wasn't good enough because the roman system is a pig for doing multiplication and division in; once you have arabic numerals (and having a symbol for zero really helps too BTW!), powerful easy algorithms for multiplication and division arise naturally, and before too long you've invented the calculus, and then you're really cooking with gas...
There is also informática/computación; both Spanish words to refer to the same thing but used in Spain/America.
I guess that literally they'd be something like IT and CS.
https://www.dictionnaire-academie.fr/article/A9O0665
A search for "computer" does not find anything; though I suspect many French actually use computer not ordinateur.
To use one is colloquially 'to data'; as in, a verb form of data :)
In German the subject is called "Informatik", translating to information science. I find that quite elegant in contrast.
" Before we had all this high falutin' opinions of ourselves as programmers and computer scientists and stuff like that, programming used to be called data processing.
How many people actually do data processing in their programs? You can raise your hands. We all do, right? This is what most programs do. You take some information in, somebody typed some stuff, somebody sends you a message, you put it somewhere. Later you try to find it. You put it on the screen. You send it to somebody else.
That is what most programs do most of the time. Sure, there is a computational aspect to programs. There is quality of implementation issues to this, but there is nothing wrong with saying: programs process data. Because data is information. Information systems ... this should be what we are doing, right?
We are the stewards of the world's information. And information is just data. It is not a complex thing. It is not an elaborate thing. It is a simple thing, until we programmers start touching it.
So we have data processing. Most programs do this. There are very few programs that do not.
And data is a fundamentally simple thing. Data is just raw immutable information. So that is the first point. Data is immutable. If you make a data structure, you can start messing with that, but actual data is immutable. So if you have a representation for it that is also immutable, you are capturing its essence better than if you start fiddling around.
And that is what happens. Languages fiddle around. They elaborate on data. They add types. They add methods. They make data active. They make data mutable. They make data movable. They turn it into an agent, or some active thing. And at that point they are ruining it. At least, they are moving it away from what it is."
https://github.com/matthiasn/talk-transcripts/blob/master/Hi...
This encapsulates it for me and informs my coding everyday. If I find myself having a hard time with complexity, I revisit the data structures.
But based on this, I always take greatest care about the data structures. Especially when designing database tables, I keep all the important aspects of it in mind (normalization/denormalization, complexity of queries on it, ...) Makes writing code so much more pleasurable and it's also key to make maintenance a non-issue.
Amazing how far-sighted this is, when considering that most Web Apps are basically I/O - i.e. data - bound.
One of the things he pounded into everyone's heads back then was "The most important decision you have to make is how to represent the problem. Do that well and programming will be easy. Get it wrong and there isn't a force on this world that will help you write a good solution in any programming language."
In APL data representation is of crucial importance, and you see the effects right away. It turned out to be he was right on that point regardless of the language one chose to use. The advise is universal.
I believe it was the first major CS book that emphasised data structures.
https://en.wikipedia.org/wiki/Algorithms_%2B_Data_Structures...
"Show me your flowchart and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won't usually need your flowchart; it'll be obvious."
But I don't think Brooks was trying to suggest it was an original idea to him or his team, either. I imagine there were a decent number of people who reached the same conclusion independently.
Coined by Peter Naur (of BNF-"fame"), by the way.
Id rather split up independent structures.
There’s a reason why equivalent Clojure code is much much shorter than comparable programs in other languages.
(Apologies for linking to esr but it's a good book)
Imagine Einstein alive and denying climate change. Would you apologize every time when you are referring to the theory of relativity?
P.S. Sorry, if you don't agree with the apologising comment and were just informing about possible reasons.
The chapter I linked to was just a summary of ideas put forth by others - though admittedly written well.
My problem with esr is more his arrogance and conceit than politics (which I also find distasteful)
I've read and liked his book, btw, but I had to ignore all his stupid Windows-bashing where he attributes every bad practice to the Windows world and every good one - to the Unix world.
Sometimes I quote HP Lovecraft and sometimes I feel like apologizing for his being racist (and somewhat stronger than just being a product of his times). But most of the time, also not. But it does usually cross my mind and I think that's okay and important. In a very real "kill your idols" way. Nobody's perfect.
And that's just for being a bigot in the early 20st century, which, as far as I know, is of no consequence today.
However if Einstein were alive and actively denouncing climate change today, I would probably add a (btw fuck einstein) to every mention of his theories. But that's just because climate change is a serious problem that's going to kill billions if we would actually listen to the deniers and take them seriously. This hypothetical Einstein being a public figure, probably even considered an authority by many, would in fact be doing considerable damage spouting such theories in public. And that would piss me off.
What I mean to say is, you don't have to, but it's also not wrong to occasionally point out that even the greatest minds have flaws.
Also, a very different reason to do it, is that some people with both questionable ideas and valuable insights, tend to mix their insightful writings with the occasional remark or controversial poke. In that case, it can be good to head off sidetracking the discussion, and making it clear you realize the controversial opinions, but want to talk specifically about the more valuable insights.
And this IS in fact important to keep in mind both, even if you think it is irrelevant. Because occasionally it turns out, for instance, through the value of a good deep discussion, that the valuable insights in fact fall apart as you take apart the controversial parts. Much of the time it's just unrelated, but you wouldn't want to overlook it if it doesn't.
And then he accused women in tech groups of trying to "entrap" prominent male open source leaders to falsely accuse them of rape.
And then he claimed that "gays experimented with unfettered promiscuity in the 1970s and got AIDS as a consequence", and that police who treat "suspicious" black people like lethal threats are being rational, not racist.
Basically, he's a racist, bigoted old man who isn't afraid to spout of conspiracy theories because he thinks the world is against him.
But in this case the "consequence" to esr was somebody apologizing for linking to him. Methinks the parent protests too much
> I see it used as a reason to limit speech because this speech that I disagree with is insidious and sinister.
Limiting speech is a very nuanced issue, and there's a lot of common misconceptions surrounding it. For a counterexample, if you're wont to racist diatribes, that can make many folks in your presence uncomfortable; if you do it at work or you do it publicly enough that your coworkers find out about it, that can create a toxic work environment and you might quickly find yourself unemployed. In this case, your right to espouse those viewpoints has not been infringed -- you can still say that stuff, but nobody is obliged to provide audience.
And as a person's publicity increases, so do the ramifications for bad behavior -- as it should. Should esr be banned from the internet by court order? Probably not. Does any and every privately owned platform have the right to ban him or/and anybody who dis/agrees with him? Absolutely: nobody's right to free speech has been infringed by federal or state governments. And that's the only "free speech" right we have.
> Should esr be banned from the internet by court order? Probably not.
Where's the uncertainty in this?
> Does any and every privately owned platform have the right to ban him or/and anybody who dis/agrees with him?
Those that profess to being a platform and not a publisher should not be able to ban him, nor anybody else, for their views, whether expounded via their platform. That's why they get legal protections not afforded to others. Do you think the phone company should be able to cut you off for conversations you have on their system?
It should be noted that the basic premise of Domain-Driven Design is that the basis of any software project is the data structure that models the problem domain, and thus the architecture of any software project starts by identifying that data structure. Once the data structure is identified then the remaining work consists of implementing operations to transform and/or CRUD that data structure.
It really isn't. DDD is all about the domain model, not only how to synthesize the data structure that represents the problem domain (gather info from domain experts) but also how to design applications around it.
I think I found all these quotes on SQLite's website, https://www.sqlite.org/appfileformat.html
The first thing is to understand FSMs and State Transition Tables using a simple two-dimensional array. Implementing a FSM using a while/ifthenelse code vs. Transition table dispatch will really drive home the idea behind data-driven programming. There is a nice explanation in Expert C Programming: Deep C secrets.
SICP has a detailed chapter on data-driven programming.
An old text by Standish; Data Structure Techniques.
Also i remember seeing a lot of neat table based data-driven code in old Data Processing books using COBOL. Unfortunately i can't remember their names now. Just browse some of the old COBOL books in the library.
Tsk. Now I'll never know if I'm a good programmer. I do all my programming in Prolog and, in Prolog, data is code and code is data.
That same comment was made to a class I was in by a University Professor, only he didn't word it like that. He was discussing design methodologies and tools - I guess things like UML and his comment was he "preferred Jackson, because it revolved around the the data structures, and they changed less than the external requirements". (No, I have no idea what Jackson is either.)
Over the years I have come to appreciate the core truth in that statement - data structures do indeed evolve slower than API's - far slower in fact. I have no doubt the key to git's success was after of years of experience of dealing with VCS systems Linux hated, he had an epiphany and came up with the fast and efficient data structure that captured the exact things he cared about, but left him the freedom to change the things that didn't matter (like how to store the diff's). Meanwhile others (hg, I'm looking at you) focused on the use cases and "API" (the command line interface in this case). The end result is git had a bad API, but you could not truly fuck it up because the underlying data structure did a wonderful job of representing a change history. Turns out hg's API wasn't perfect after all and it's found adapting difficult. Git's data structure has had hack upon hack tacked onto the side of it's UI, but still shines through as strong and as simple as ever.
Data structures evolving much more slowly than API's does indeed give them the big advantage of being a solid rock base for futures design decisions. However they also have a big down side - if you decide that data structure is wrong it changes everything - including the API's. Tacking on a new function API on the other hand is drop dead easy, and usually backwards compatible. Linus's git was wildly successful only because he did something remarkably rare - got it right on the first attempt.
https://www.percona.com/sites/default/files/hipp%20sqlite%20...
"Show me your flowcharts and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won't usually need your flowcharts; they'll be obvious." -- Fred Brooks, The Mythical Man-Month, pp. 102-103
The Linus comment is about designing your programs around a data representation that efficiently models your given problem.
Doesn't it show more of the drawbacks of this functional data structure?
As for the working directory - yes, there could be more management around that. I'm not sure why the git community went for nested repos / submodules rather than partial checkouts. It's a different question than the data structure of the repo itself, though. Compared to other VCS it still seems miles ahead.
Large repos: It seems one could alleviate that by limiting the pull history (and LFS if needed), right?
browser version of this is datascript (oss).
https://www.youtube.com/watch?v=fHSZz_Mx-Uo
To become a Git power user, it is far more beneficial to learn its underlying content-addressable-store data structure rather than explore the bazillion options in its command line interface. It is surprisingly easy to create a repository manually and then to start adding "blob" files to the store!
The whole impetus for git (someone correct me if I'm wrong):
1. Linux source was hosted on BitKeeper before git. It basically was one of the first distributed source control systems (or the first? not sure if anything predated it).
2. Linux devs got into conflict with the BitKeeper owner over open-sourcing and reverse engineering, so Linus realized they needed a new system because no other source control system had the features they needed like BitKeeper had (mainly, I understand, the distributed repos).
So basically, Linux is to Unix like git is to BitKeeper (roughly).
SCCS is a single-file revision management system, which inspired RCS (but they are not compatible IIRC; one keeps changes forward and the other backwards). CVS was a management system on top of RCS to provide "repository-wide" actions and branches. SVN was supposed to be "CVS done right".
git is a blockhain/dag content addressable storage, inspired and borrowing a lot from monotone (a system created by Graydon Hoare who later went to create rust; Linus credits monotone and graydon for inspiring the design of git, which is essentially monotone but much simplified).
McVoy and Linus did, in many ways, collaborate on making BitKeeper more comfortable for Linus and Linux development; but the design and implementation of BitKeeper predates that and goes back to SCCS through TeamWare, and maybe even farther.
> Git development began in April 2005, after many developers of the Linux kernel gave up access to BitKeeper, a proprietary source-control management (SCM) system that they had formerly used to maintain the project. The copyright holder of BitKeeper, Larry McVoy, had withdrawn free use of the product after claiming that Andrew Tridgell had created SourcePuller by reverse engineering the BitKeeper protocols.
But as the BitKeeper article makes clear, the relationship between BitKeeper and OSS had been frosty for a while.
Larry McVoy comments here occasionally as https://news.ycombinator.com/user?id=luckydude
People have noted multiple times that its current build system is fairly hostile to packaging on *NIX, but nobody seems to be putting in the work and the time. I'm somewhat curious about BitKeeper, but not enough to make my first goal to go down into its guts (complete with a custom stdio from NetBSD with modifications!) to make it play nicely with my system.
[1] https://users.bitkeeper.org/t/thoughts-after-a-few-days-with...
It has a much cleaner interface, easier to use, better formalism, code signing is built-in and required rather than tacked-on and optional, and insanely good documentation (I still sometimes recommend people to read the monotone docs as an introduction to git/DVCS).
It was also basically finished and stable when the BitKeeper drama happened. It was one of the few alternatives Linus looked at and publicly evaluated before writing git. But unfortunately, its performance was pretty poor for a project the size of Linux at the time, and a combination of human failings + the well architected abstractions which hid the underlying monotone data structures convinced Linus that it was less effort to write git than to fix monotone.
In a literal sense that was true since git was basically written in a weekend.. but more than a decade on we're still basically stuck with all the design short-cuts and short-falls of a weekend project. The world would be a better place if we used monotone instead :(
Edit: these days I'm more a fan of the formalism of darcs, and really looking forward to a stable pijul. But there is a soft spot in my heart for monotone, and more than once I've considered forking it to modernize with compact elliptic curve crypto instead of RSA and a faster non-relational database or filesystem backend instead of SQLite.
I use Arch (2006) btw.
There were several systems which predated it. The most arcane one I know of is Sun's TeamWare[1] which was released in 1993 (BitKeeper was released in 2000). It also was used to develop Solaris for a fairly long time.
Larry McVoy (the author of BitKeeper) worked on TeamWare back when he worked at Sun.
> 2. Linux devs got into conflict with the BitKeeper owner over open-sourcing and reverse engineering, so Linus realized they needed a new system because no other source control system had the features they needed like BitKeeper had (mainly, I understand, the distributed repos).
That is the crux of it. Linus does go over the BitKeeper history in the talk he gave at Google a decade ago on Git[2].
[1]: https://en.m.wikipedia.org/wiki/Sun_WorkShop_TeamWare [2]: https://www.youtube.com/watch?v=4XpnKHJAok8
There must be hundreds of thousands of me, hungry to try to understand what makes an OS tick.
Maybe I didn't read the code, but I learned so much by compiling it. Knowledge I wouldn't have bothered with had I just downloaded a binary blob.
Edit: I mean, their value propositions to internal teams would disappear. They would still have value as social networks and as centralized hosting.
I believe OSX ships with Python2 by default, right?
https://docs.python.org/2/library/cgihttpserver.html
Hmm...
Implying that it's a mistake because people don't know about it is very odd. Clearly, I know about it, as do others.
It's been possible for 10+ years to serve git repos using gitweb, and mercurial repos using hgweb. We did this in like 2007-2009ish, locally, in our LAN, among developers... because our code couldn't be pushed to a third party for... reasons.
Eventually I did setup our own internal SSH server to serve the repos, but for quick browsing of a team's repo state, using the built in HTTP server is just fine.
How are people here saying they've never heard of this? Oof... Typical developers, not even familiar with their tools.
Social networks are popular because of their userbase and the ease of discovering other users and code. There is nothing you can build into a piece of client software that enables that.
what do you think are their value propositions? I think they biggest part of their value proposition has to do with a centralized git repository as a service.
The centralized part is really important for most companies. To the point that many git users don't really understand it's decentralized nature.
I guess since my team uses a self-hosted instance of GitLab, I'm biased and don't put any value on the social network aspect or the hosting aspect.
Git is good in some ways, terrible in others. I've used it for years and still don't feel really comfortable with it, but I've never had something as multiheaded as the linux repo.
Git is overkill for so many projects, I hate being forced into for everything.
For years probably 99% of my interactions with git, or any other VCS, is through editor extensions like magit or fugitive.
And of course it's remote, every action in SVN is remote since it's centralized (except for shelving).
http://svnbook.red-bean.com/en/1.7/svn.ref.svnadmin.c.create...
Why would I [re]learn those tools if I already know git?
If I'm going to move to a new VCS, it's going to be one that actually gives me something I didn't have before, like Fossil. Not some other VCS that captures the same concepts with a slightly different cli UX (which hardly even impacts me at all, since I rarely interact with such systems on the command line rather than through porcelain provided by an editor extension.)
If you are forced into using it to everything but still haven't taken the steps necessary for understanding it, why is that my problem?
Git costs nothing to use, you add it to a project and then it sits there until you do something with it. If you want to use it as a "super save" function, it'll do that. If you want to use it to track every change to every line of code you've written, it'll do that too.
But those other systems were the whole point of the post you replied to ;)
Mercurial? Similar DVCS concepts, but you no longer have to worry about garbage collection or staging areas...
And no staging area is strictly simpler than having a staging area, which is contrary to your assertion.
Sometimes, I try a few different approaches to make something work. Each of these attempts is a different branch--I might need to revisit it, or pull stuff out of it. Good luck staring at a branch name and working out if it's landloop or landloop2 that had the most working version of the code.
I've been using git for a few years, and staging has been all cost with zero benefit so far.
Well yes, but the GP is claiming that git is the most simple thing above file storage.
Staging may be a feature, but it adds complexity. Perhaps useful complexity, but complexity nonetheless.
Not making any larger comparison here, what I'm saying is that a single file is simpler than a single directory.
If you use git at all, you may as well use it for everything. If you have control over which version control system to use, there's no good reason to actively use multiple ones at the same time.
If I ever need to manage the codebase for a huge, multi-level project involving large numbers of geographically dispersed developers then I'm sure I'd use git. For simpler projects, not so likely.
both fork and the intellij ide are great for that, handling the common cases solidly and building up so many convenience functions I can't live without them now, like whitespace aware conflict resolution or single line commits.
Bottom line, git is the VSC equivalent of C, it is quite powerful, but its got a lot of foot-gun moments and requires a lot of human rigor/experience to make it work well. Rigor that even the best users frequently mess up.
So yes, if you make the hypothesis that things went out so we would all be using BSD, then we would. And yes, successful projects and people always come from a part of luck. But so what? What happened in reality is what happened, and if they went lucky good for them, but this does not really removes anything from their achievements.
The same way it doesn't stop gamblers, stock pickers, actors, and entrepreneurs from mistaking survivorship bias for talent.
That said, I don't think git was purely luck either.
BSD-style licenses tend to be a better fit for business models that sell proprietary extensions to free software. These form lock-in moats that inhibits the growth of any deeper ecosystems. We've seen this over and over again with things such as graphical subsystems for non-free UNIX, but also more recent examples with firewalls and storage boxes. Those are great for what they do, but work on the free parts are seen more like a gift to the community than a money maker.
The tit-for-tat model of the GPL enables those ecosystems to form. By forcing your competitors to free their code in exchange for yours, game theory dictates that moats cannot form, and when everyone stands on other's shoulders development is faster.
I'd say that's pretty much experimentally proven by now. Of course, reality is not as black and white, especially when GPL-style companies contribute to BSD-licensed software and vice versa. Perhaps PostgreSQL is a prominent example of that. There are however traces of these patterns there too, for example in how the many proprietary clustering solutions for the longest time kept the community from focusing on a standard way.
That's an interesting take, but I'm not sure I understand.
Are you saying Linux would've lost (presumably to proprietary OSs?) if it used a permissive license?
If the GPL has been proven to be more suitable for business, why is the use of GNU GPL licenses declining in favor of permissive licenses?
I don't have a hog in this pen and so I'm not trying to provoke. I'd just like to hear thoughts on why it looks so different from where I sit.
It isn't, at least not exactly. It's declining in favor of a combined permissive/commercial license model. And it's only doing that for products that are meant to be software components.
The typical model there is that you use a permissive license for your core product as a way of getting a foot in the door. Apache 2.0 is permissive enough that most businesses aren't going to be afraid that integrating your component poses any real strategic risk. GPL, on the other hand, is more worrisome - even if you're currently a SAAS product, a critical dependency on GPLv2 components could become problematic if you ever want to ship an on-prem product, and might also become a sticking point if you're trying to sell the company.
But it's really just a foot in the door. The free bits are typically enough to keep people happy just long enough to take a proper dependency on your product, but not sufficient to cover someone's long-term needs. Maybe it's not up to snuff on compliance. Or the physical management of the system is kind of a hassle. Something like that. That stuff, you supply as commercial components.
I used to follow LLVM development. There was lots of mailing list traffic of the form "I'll send you guys a patch as soon as management approves it..." followed by crickets.
Basically, RMS was exactly correct about the impact that loadable modules would have on GCC's development.
The last thirty years in a nutshell.
The size and quantity of companies working on things related to a project determines whether a strong copyleft license or a permissive license makes the project more successful
Say you want to make a business around a FOSS project. Which license should you choose for that project?
If your business starts gaining traction, people may realize it's a good business opportunity, and create companies that compete against you.
I'll simplify to two licenses, GPL and MIT. Then there's two options, based on which one you chose originally:
1) If you chose the GPL, then you can be sure that no competitor will get to use your code without allowing you to use theirs too. You can think of this as protection, ensuring no other company can make a product that's better than yours without starting from scratch. Because everyone is forced to publish their changes, your product will get better the more competition you have. However, your competitors will always be just a little behind you because you can't legally deny them access to the code.
2) OTOH if you chose MIT, a competitor can just take your project, make a proprietary improved version of it and drive you out of the market. The upside is if you get to be big enough, you can do exactly that to _your_ competitors.
You can see that when you are a small company the benefits of GPL outweigh the cons, but for big ones it's more convenient to use MIT or other permissive licenses. In fact, I think the answer to your question "why is the use of GNU GPL licenses declining?" is because tech companies tend to be bigger than before.
Now say you want to make a business around some already existing software. And say there's two alternative versions of that software, one under the GPL and one under the MIT license (for example, Linux and BSD). Which one should base your business on? And contribute to? Well, it's the same logic as before.
As with most things, I think Linux succeeded because it was a worse-is-better clone of existing systems that happened to get lucky.
I really wish he'd collaborated with some more people in the early stages of writing git to come up with an interface that makes sense, because everyone is constantly paying the cost of those decisions, especially new git learners.
You're underestimating how much inertia is created simply by being the out-of-the-box default, and how hard that inertia is to overcome even by better alternatives.
Built another tool similar but not same, using the same storage method underneath.
But the existence of something even worse doesn't excuse something that is merely bad. And git is so much more widely used that its total overall harm on developer productivity is worse.
I started in the late 90's when cvs was popular. Then we moved to svn. You had productivity issues of all sorts, mainly with branching and merging.
Have you ever worked with Rational ClearCase? It's a true horror show.
Does it happen often at your workplace? What kind of issues are we talking about?
It's great being able to change a ClearCase config file to choose a different branch of code for just a few files or directories, then instantly get that branch active for just those specific files.
These are problems that every single person learning git has to figure out and then come up with their own solutions for.
So much of this drama seems propped up on things that just aren't that difficult.
While there are a metric ton of things which are confusing about git, this was perhaps not the greatest example.
So perhaps it's a great example if you've gotten it wrong?
In the case of the branch, the correct result is something like "what has changed since my branch diverged from its parent"--basically what you see in a PR. I think this is unnecessarily obscure in Git because a "branch" isn't really a branch in the git data model; rather it's something like "the tip of the branch".
I don't think I've ever wanted to compare my workspace against a branch, but clearly diffing the branch is useful (as evidenced by PRs). Similarly, I'm much less inclined to diff my workspace against a particular commit, but I often want to see the contents of an individual commit (another common operation in the Github UI).
In essence, if Github is any indicator, Git's data model is a subpar fit for the VCS use case.
It's unfair insofar as it does what I expect it to, which is to diff between what I'm curious about, and where I am.
In other words, if you elide the second argument, it defaults to wherever HEAD is.
The point being, this is not something I personally need to look up. I'd venture a guess that your familiarity with hg is interfering because the conventions are different.
By contrast, Mercurial's UI makes the repository history a more first-class citizen, and it is very easy to answer basic questions about the history of the repository itself. If you're doing any sort of source code archaeology, that functionality is far more valuable than comparing it to the current state: I don't want to know what changed since this 5-year-old patch, I want to know what this 5-year-old patch itself changed to fix an issue.
Git users also need to answer questions like "What changes are in my feature branch?" (e.g., a PR) and "What changed in this commit?" (e.g., GitHub's single-commit-diff view). These aren't Mercurial-specific questions, they're applicable to all VCSes including Git, as evidenced by the (widely-used) features in GitHub.
Even with Git, I've never wanted to know how my workspace compares to another branch, nor how a given commit compares to my workspace (except when that commit is a small offset off my workspace).
> In other words, if you elide the second argument, it defaults to wherever HEAD is.
Yeah, I get that, but that's not helpful because I still need to calculate the second argument. For example, `git diff master..feature-branch` is incorrect, I want something like `git diff $(git merge-base master feature-branch)..feature-branch` (because the diff is between feature-branch and feature-branch's common ancestor with master, not with HEAD of master).
One of the cool things about Mercurial is it has standard selectors for things. `hg log -b feature-branch` will return just the log entries of the range of commits in the feature-branch (not their ancestors in master, unlike `git log feature-branch`). Similarly, `-c <commit>` always returns a single-commit range (something like <commit>^1..<commit> in git). It's this consistency and sanity in the UI that makes Mercurial so nice to work with, and which allows me to recall with better accuracy the hg commands that I used >5 years ago than the git commands that I've used in the last month.
git show <commitish>
will show the log message and the diff from the parent commit tree.
`git branch [branchname]` creates a branch without switching to it.
`git checkout -b [branchname]` creates a branch and checks it out.
And `git reset --hard` will also discard changes. (Arguably, this is better than `git reset` discarding local changes, as it is more explicit.)
git is a tool. Different tools take different amounts of time to master. People should probably spend some time formally learning git just as one would formally learn a programming language.
And I have already explained why `git reset --hard` makes more sense in my opinion.
I agree that Git can be hard to wrap your head around, and that the commands could be more intuitive. But Git is complex in large part because the underlying data structure can be tricky to reason about -- not because the UI on top of it is terrible.
The commands johnmaguire2013 listed are the ones usually recommended for beginners and I have found them easy to understand.
"git branch [name]" is for creating branches; it tells you if the branch already exists. Pretty easy to understand.
"git checkout [name]" is for checking out branches; it tells you if you're already on that branch.
You can run these sequentially and it works fine; there's no need for `git checkout -b [branchname]`.
I think there is sometimes some productivity porn involved in discussions of git, where people feel really strongly that everything should be doable in one line, and also be super intuitive. It's a bit like the difference between `mkdir foo && cd "$_"` on the command line, vs just doing mkdir and cd sequentially. IMO the latter is easier to understand, but some experienced folks seem to get upset that it requires typing the directory name twice.
I've actually taught how to use git to many teams, and I always start with merkle trees. They are actually easy to grasp even for designers and other nontechnical people, and could be explained in 10 minutes with a whiteboard. And then suddenly git starts to totally make sense, and I'd dare to say, become intuitive.
I can never seem to guess what things do in Git, and I consider myself fairly comfortable with the core concepts of Git. Having written many types of append-only / immutable / content address data systems (they interest me), you'd think Git would be natural to me. While the core concepts are natural, Git's presentation of them is not.. at least, to me.
edit: formatting with the * .
so, that says, copy all of the files out of the current branch at the current commit into the local dir. What this will do in practice is "discard current changes to tracked files". So If i had files foo, bar, baz, and I had made edits to two of them, and I just want to undo those changes, that's what checking out * from HEAD does. It doesn't however delete new files you have created. So it doesn't make the state exactly the same.
So why not just git checkout HEAD? Well, you already have HEAD checked out (you are on that branch), so there's nothing for git to do. You want to specify that you want to _also_ explicitly copy the tracked file objects out also. It's kind of like saying "give me a fresh copy of the tracked files".
The confusing thing is that in practice it is "reverting" the changes that were made to the tracked files. But `git revert` is the command you use to apply the inverse of a commit (undo a commit). One of the more confusing aspects of git is that many of the commands operate on "commits" as objects (the change set itself), and some other commands operate on files. But it's not obvious which is which.
git reset --hardIt's just that I've been using git CLI for so long, and know exactly which commands to use in any circumstance without having to look them up, that I don't benefit much from switching to something new, whereas someone who hasn't yet put in that time to really learn git would stand to benefit more.
In comparison to an average CLI program's usability, I think git's got a very good one. It's not perfect, but I think saying it's "poor" is really exaggerating the problems.
In particular, I love how well it does subcommands. You can even add a script `git-foobar` in your $PATH and use it as `git foobar`. It even works with git options automatically like `git -C $repo_dir foobar`.
> It's really hard to explain to new users why you need to run `git checkout HEAD * ` instead of `git reset` as you'd expect
Why would you ever do `git checkout HEAD * ` instead of `git reset --hard`? The only difference is that your checkout command will still leave the changes you've done to hidden files, and I can't think that's ever any good.
> why `git branch [branchname]` just switches to a branch whereas `git checkout -b [branchname]` actually creates it
If you think those behaviors should be switched, good, because they are.
EDIT: How did you manage to add the asterisk to the checkout command in your post so that it's not interpreted as italics without adding a space after it?
In an alternative timeline where BSD would be dominant, would we have e.g. free software AMD drivers? Would we have such big variation in containers, VMs, and scalabe system administration as we do on Linux? I wonder. No doubt that world would also be prettier than what we have now - in line with ways in which the BSDs are already better than Linux - but who knows.
This might be a controversial opinion but: Linux likely "won" because it was better in the right areas.
You're right that those law suits were settled long before Linux gained momentum though. FreeBSD and NetBSD were released after Linux and their predecessor (386BSD) is very approximately as old as Linux (work started on it long before Linux but it's first release was after Linux). As far as I can recall, 386BSD wasn't targeted by lawsuits.
Also wasn't BSD used heavily by local ISPs in the 90s?
In any case, I think Linux's success was more down to it being a "hacker" OS. People would tinker with it for fun in ways people didn't with BSD. Then those people eventually got decision making jobs and stuck with Linux because that's where their experience was. So if anything, Linux "won" not because it was "better" than BSD on a technical level but likely because it was "worse" which lead to it becoming more of a fun project to play with.
(I think it's a pity that the useful innovations that happened in Linux cannot be moved back over to FreeBSD because of licensing -- the computing world would be better off if it could.)
IMO it's Linux that should want the features from FreeBSD/Solaris. I want ZFS, dtrace, SMF, and jails/zones. Linux is basically at feature parity, but the equivalents have a ton of fragmentation, weird pitfalls, and are overall half baked in comparison.
For example, eBPF is a pretty cool technology. It can do amazing things, but it requires 3rd party tooling and a lot of expertise to be useful. It's not something you can just use on any random box like dtrace to debug a production issue.
systemd
There were a couple of times early on when I wanted to try both Linux and one of the BSDs on my PC. I had CDs of both.
With Linux, I just booted from a Linux boot floppy with my Linux install CD in the CD-ROM drive, and ran the installation.
With BSD...it could not find the drive because I had an IDE CD-ROM and it only supported SCSI. I asked on some BSD forums or mailing lists or newsgroups where BSD developers hang out about IDE support, and was told that IDE is junk and no one would put an IDE CD-ROM in their server, so there was no interest in supporting it on BSD.
I was quite willing to concede that SCSI was superior to IDE. Heck, I worked at a SCSI consulting company that did a lot of work for NCR Microelectronics. I wrote NCR's reference SCSI drivers for their chips for DOS, Windows, Netware, OS/2, and Netware. I wrote the majority of the code in the SCSI BIOS that NCR licensed to various PC makers. I was quite thoroughly sold on the benefits of SCSI, and my hard disks were all SCSI.
But not for a sporadically used CD-ROM. At the time, SCSI CD-ROMs where about 4x as expensive as IDE CD-ROMs. So what if IDE was slower than SCSI or had higher overhead? The fastest CD-ROM drives still had maximum data rates well under what IDE could easily handle. If all you are going to use the CD-ROM for is installing the OS, and occasionally importing big data sets to disk, then it makes no sense to spring for an expensive SCSI CD-ROM. This is true on both desktops and servers.
The second problem I ran into when I wanted to try BSD is that it did not want to share a hard disk with a previous DOS/Windows installation. It insisted on being given a disk upon which it could completely repartition. I seem to recall that it would be OK if I left free space on that disk, and then added DOS/Windows after installing BSD.
Linux, on the other hand, was happy to come second after my existing DOS/Windows. It was happy to adjust the existing partition map to turn the unpartitioned space outside my DOS/Windows partition into a couple Linux partitions and install there.
As with the IDE thing, the reasons I got from the BSD people for not supporting installing second were unconvincing. The issue was drive geometry mapping. Once upon a time, when everything used the BIOS to talk to the disk, sectors where specified by giving their actual physical location, specifying what cylinder they were on (C), which head to get the right platter (H), and on the track that C and H specifies, which sector it is (S). This was commonly called a CHS address.
There were limits on the max values of C, H, and S, and when disks became available that had sectors whose CHS address would exceed those limits, a hack was employed. The BIOS would lie to the OS about the actual disk geometry. For example, suppose the disk had more heads than would fit in the H field of a BIOS disk request. The BIOS might report to the OS that the disk only has half that number of heads, and balance that out by reporting twice as many cylinders as it really has. It can then tranlate between this made up geometry that the OS thinks the disk is using and the actual geometry of the real disk. For disks on interfaces that don't even have the concept of CHS, such as SCSI which uses a simple block number addressing scheme, the BIOS would still make up a geometry so that BIOS clients could use CHS addressing.
If you have multiple operating systems sharing the disk, some using the BIOS for their I/O, and some not, they all really should be aware of that made up geometry, even if they don't use it themselves, to make sure that they all agree on which parts of the disk belong to which operating systems.
Fortunately, it turns out that DOS partitioning had some restrictions on alignment and size, and other OSes tended to follow those same restrictions for compatibility, and you could almost always look at an existing partition scheme and figure out from the sizes and positions of the existing partitions either what CHS to real sector mapping the partition maker was using. Details on doing this were includes in the SCSI-2 Common Access Method ANSI standard. The people who did Linux's SCSI stuff have a version [1].
I said "almost always" above. In practice, I never ran into a system formatted and partitioned by DOS/Windows for which it gave a virtual geometry that did not work fine for installing other systems for dual boot. But this remote possibility that somehow one might have an existing partitioning scheme that would get trashed due to a geometry mismatch was enough for the BSD people to say no to installing second to DOS/Windows.
In short, with Linux there was a good chance an existing DOS/Windows user could fairly painlessly try Linux without needing new hardware and without touching their DOS/Windows stuff. With BSD, a large fraction would need new hardware and/or be willing to trash their existing DOS/Windows installation.
By the time the BSD people realized they really should be supporting IDE CD-ROM and get along with prior DOS/Windows on the same disk, Linux was way ahead.
[1] https://github.com/torvalds/linux/blob/master/drivers/scsi/s...
However, I also agree with @wbl that the lawsuits were ultimately the decisive factor. The hardware requirements situation of BSD was a tractable problem; it just needed a flurry of helping hands to build drivers for the wide cacophony of PC hardware. The lawsuit-era stalled the project at just that critical point. By the time that FreeBSD was approaching an acceptable level of hardware support Linux already had opened up a lead... which it never gave up.
As to the C/H/S low-level-format, NCR could read some Adaptec formatted drives, while Adaptec couldn't read NCRs. Asshole move. Never mind.
As for the BSDs being behind? Not all the times. I had an Athlon XP 1800+ slightly overclocked by about 100Mhz to 2000+ in some cheap board for which i managed to get 3x 512MB so called 'virtual channel memory' because dealer thought it was cheap memory which ran only with via chipsets. Anyways 1,5GB RAM about twenty years ago was a LOT! With Linux of the times i needed to decide how to split it up, or even recompile the kernel to have it using it at all. No real problem because i was used to it, and it wasn't the large mess it is today.
Tried NetBSD. From a two or three floppy install set. I don't remember the exact text in the boot console anymore, just that i sat there dumbstruck because it just initialized it at once without further hassle. These are the moments which make you smile! So i switched my main 'workstation' from Gentoo to NetBSD for a few years, and had everything i needed, fast and rock solid in spite of overclocking and some cheap board from i can't even remember who anymore. But its BIOS had NCR support for ROM-less add-on controllers built in. Good times :-)
Regarding the CD-ROM situation, even then some old 4x Plextor performed better than 20x Mimikazeshredmybitz if you wanted to have a reliable copy.
As to sharing of Disks by different OS? Always bad practice. I really liked my hot-pluggable 5 1/4" mounting frames which took 3,5" drives, with SCSI-ID, termination, and what not. About 30 to 40USD per piece at the time.
And empirically hasn't this been the case in our history in tech? Windows being popular over Mac, IE being more popular than netscape, etc etc.
Some would argue that his impact wasn't even net-positive. He might have done it in good faith, but it didn't really work out well.
indeed
Back in the 90’s, my sister asked what “Linux” was and I explained it to her as a free replacement for Windows (I know, I know, but that was the right description for her). She asked why it was free while Windows was expensive enough to make Bill Gates the richest man in the world. I told her that the guy who wrote it gave it away for free. She said, “wow, I’ll bet that guy feels really stupid now.”
Gaining user attention at a global scale is always extremely competitive, even if you give it all away for free.
Some people like Bill Gates got extremely lucky thanks to excellent social connections but others like Linus who were not so lucky had to go to extreme lengths to break through all the social and economic obstacles imposed on them by the incumbents.
I don't think there were many (any?) free UNIX or UNIX-like OSs at that time. MINIX wasn't albeit Tanenbaum wanted it accessible to as many students as possible so teh licence fee was relatively cheap compared to other UNIXes of the time. 386BSD (the precursor to FreeBSD and NetBSD) wasn't released until around a year after Linux albeit they started quite a bit before.
I guess there was lots of free OS's in the hobbyist sense but nothing that actually competed with MINIX.
Gates was smart enough, though, to negotiate a non-exclusive license with IBM.
Furthermore, his version worked on modern hardware
- Father was a wealthy attorney.
- Mother worked for IBM.
- His first OS demo for IBM worked the first time even though they had only tested it on an emulator before. This is extremely unusual; there are a lot of factors which can make the emulator behave differently from the real thing.
- IBM did not see the value in software and did not ask for exclusivity (they could easily have demanded it).
Sure he is a very smart guy, but mostly he is a ridiculously lucky guy.
https://en.wikipedia.org/wiki/Haber_process#Cause_of_populat...
Maybe Norman Borlaug as having the most wealth creation?
About 50% of the world population depend on crops produced with artificial fertilizers. They enabled billions of people to live at all. In my opinion they are setting the bar quite high.
https://en.wikipedia.org/wiki/History_of_the_Haber_process
First extraction was 110 years ago, then it was industrialized 106 years ago.
That means they are out for
"anybody in the last hundred years" just barely.
It's amazing to think of how much change they enabled in the last 103 years.
Roughly as I understand it Haber-Bosh process enabled us to get to the billions (e.g. 1B & 2B). Norm Borlaug, & the green revolution which he helped start, built on top of fertilizer & enabled the next couple billions.
https://en.wikipedia.org/wiki/Green_Revolution
Norm Borlaug is credited with saving over a billion people from starvation & famine.
Years of successful project management are many small hits, not just one big hit.
It sounds like maybe "one trick pony" might be closer to what Linus was getting at. Being a one trick pony is better than being a one-hit wonder, but Linus has shown he's got at least a couple of tricks. :-)
And I, for one, have not forgotten the rather impressive work he did at Transmeta.
And for the sake of not making this sound like hero worship, I still side with Tanenbaum when it comes to the monolithic kernel vs. micro-kernel debate...
EDIT: corrected nonsensical double mention of ‘microkernel’.
For others that are not familiar, you mean the microkernel vs monolithic kernel debate.
Not sure why this is always overlooked.
The real-time code translation approach where you have a cisc front end but the code executs on a risc core was immediately copied by Intel (and sued for that). Without that technology we would still be doing 200 mhz at 200 watt.
What transmeta did was to keep the x86 instruction set (CISC) but internally convert them to a simpler RISC-ish instruction set and run it in a much simpler and power effective RISC core.
Intel copied this idea which allowed P4 (maybe already PIII?) to make a giant performance leap. Nowadays all high performance CISC CPUs from AMD and Intel do this.
Wow! Even Linus has some form of 'impostor syndrome' and self-doubt after all the technical achievements.
The authors of Pijul[1] found a solution for the exponential merge problem, and from what I've heard there have been discussions that darcs might in the future switch to Pijul's algorithm.
[0]: http://darcs.net/FAQ/Performance#is-the-exponential-merge-pr...
[1]: https://pijul.org
It just so happens that what I usually want in practice is to make copies by-value. There are other integration workflows that make copies by-reference with dedicated merge commits. But my team's most common use case for long-lived branches is to track bugfixes for deployment to remote hardware separately from mainline development. So, by-value copies of small changes is much more common.
IIRC, the mercurial project had already started when Linus started working on Git. (In fact, very early versions of Xen used BitKeeper, following suit with Linux; when the license changed, Xen moved over to mercurial because git wasn't ready yet, and stuck with it for a number of years.)
The main reason Linus wasn't happy with mercurial was the performance -- Linux just has far far more commits and files to deal with than nearly any other project on the planet, and even at the time, operations in mercurial took just a bit too long for Linus.
Not really. I work with Mercurial every day, in a FAANG company's monorepo which is many many times the size of the Linux kernel. It's not always pleasant, but it's not clear git would be any better. Mercurial's performance issues are solvable, have been solved to a large degree, but IIRC when git started that was not the case. It's a shame really. The fragmentation is annoying sometimes, even if it's the result of historical accident rather than any bad decisions made at any particular moment in time.
Mercurial had a better interface and early on was equal or better (as far as I know) for most things. Everyone went with Git because Linus made it. I think it also might have been faster for huge code bases... but that pretty much only affects the kernel team and a handful of others. Your standard CRUD app could easily use a slower VCS without noticing.
Instead of going on merits, everyone just followed Linus. Which is what always happens in technology communities. Everyone just does whatever some guy at the front is doing.
Was GitHub that much better? Or was it network effect of everyone on GitHub? (sincere question, I don't know)
Git started winning people over because mercurial was atrociously slow. The real nail in the coffin was GitHub which truly was revolutionary. Nothing else really played a part.
What do you mean by "there"?
If you mean you've been working that long my first VCS experience was with MS SourceSafe and quickly moving to CVS. So...
The first I heard of Git (as best I recall) I definitely knew it was from Linus.
Mercurial slowness was never a problem for me, and these were the times before SSD! It probably was slower but I used it on some pretty large codebases (NetBeans) and it was fine.
BitBucket was much better for me than GitHub because it offered free private repositories with 5 users, which just happens to be enough for a small team/company/startup.
Sun Microsystems picked Mercurial for OpenJDK/NetBeans/etc, Mozilla was on Mercurial.
I'm still puzzled how Git won because in my bubble it was a tool with much worse commands and 'metaphors' compared to Mercurial.
I think it's a big loss for the industry that we are all (me included) on git.
How much earlier was Github than Bitbucket?
Bitbucket supported both Git and Hg.
You don't think creating Linux and git are "merits"? Most people I know that adopted git had no clue who Linus was, or that he was even involved in git.
But creating Linux should not be a merit toward Git. Obviously the experience would make Git better. But why did a better UI lose to an inferior one?
I'd love to be wrong because now Mercurial is sort of dead and I have to use Git every day if I want to work with others. I'd love to have a better attitude about it. So far it just looks like another thing that one because X popular guy made it or Y big company made it.
Because it wasn't better? Git won because it's overall a superior VCS, had better support for complex workflows, was performant, and had GitHub which means teams could circumvent IT.
I think from the perspective of a repository maintainer git was always the better choice. And those are the ones who choose the VCS. For most devs, that just want to commit, mercurial had the friendlier user interface. But I am glad git won, because now as a developer I use a lot of it's functionality that I never expected from an VCS.
"Linux itself got its name from Ari Lemmke who ran the FTP server the original Linux Kernel was uploaded to. Linus Torvalds, the creator of the Linux kernel, wanted to name the kernel Freax, but Ari instead gave him a folder called “linux” to upload his kernel to"
IDK what "dick" sounds like. A man's first name?
I was pointing out "what it sounds like" to me, as "git" is an uncommon word in American English.
I was unhappy to see Linus looking over weight and out of shape in that picture. I am 68 so I understand getting old, but I consider Linus to be a ‘world resource’ and I wish him well, health wise. His wife used to be a karate champion so he at least has an exercise expert at home.
1. https://neil-gaiman.tumblr.com/post/160603396711/hi-i-read-t...
If your first hit is sufficiently large you never need another.
It's different in founding startups..., once you're lucky, twice you're good.
* it always was fast (written in C vs Python for Mercurial);
* Linus and Linux are behind it.
That's a pretty bold statement. Git might be almost ubiquitous these days, but if you erased git from the world, the ripples would be much less due to the plethora of alternatives out there. There's no reason to think that things like Github wouldn't have evolved with alternative VCSs. Don't get me wrong, it's a great tool and improves the QoL of developers, but it's hard to think about any dev not being able to do their job without it.
What else could we take away from Trovalds in order to push him to his next project? :)
ex post facto
It has support for most dive computers and does full technical dive planning with mixed gasses.
I then went on to use subversion for a long time in different settings. Up until the day I finally tried git. This is not about centralized vs. distributed, it's simply superior software, period. Nothing else I've tried even comes close.
And GitHub chose Git because of two reasons:
1. It was affiliated with Linus Torvalds.
2. It had a catchy name.
As far as I remember Mercurial had a much better UX.
Compared to Mercurial, Git is over-complicated and a lot of the commands don't make sense (e.g. 'git checkout -b mybranchname' to make a new branch WTF?). In a way, it shows how superficial we are as a society; even among software developers.
Edit: thanks webarchive https://web.archive.org/web/20081218003732/http://whygitisbe...
A bit lower are other multi-hit inventors. Gavin King, inventor or Hibernate and Seam, is in this list.
Everybody else comes third. It's great to invent something useful. (A singleton.) As a software maintenance engineer, I recognize great maintenance engineers in this category, too.
I am deeply impressed by Torvalds. Git and Linux are huge.
Pike, Thompson and Ritchie are in the same league. Not many others, maybe none.
A level below that, people like Gavin King. He invented Hibernate and Seam, not too shabby. (But not like Linux.)
A level below that, the rest of us mortals. I know quite a few one-hit authors (they're good) and some maintenance wizard, but it's nothing like inventing multiple super-hits. Hats off to Linus.
Hence, don't respect/admire such people too much, they are the same avg frustrated chump from time to time like you are. With as many worries.
I honestly think that Git has the potential to be relevant a lot longer than Linux.
i once had a job at a state agency that was conservative enough to use CVS and subversion, so i got to see what the bad old days were like. i'm so glad git exists.
“I’m an egotistical bastard, and I name all my projects after myself. First ‘Linux’, now ‘Git’” --https://websetnet.net/microsoft-now-using-linus-torvalds-ope...
He would never admit though ...
Almost nothing substantial was ever one person's doing.
Anecdotally...
When I first began my git adventures, nearly 10 years ago, I was extremely apprehensive. I knew SVN. I liked SVN. I had no reason to go elsewhere. And then I tried git (albeit using a GUI at the time), and it didn't click. So I went back to SVN. Later, after being forced to use git for a project at work, I started to understand what makes it special. I finally felt I had a version control system/tool that I could trust without being so hands-on like I was with SVN. Put simply -- git _just_ works. The concepts take time to really grasp (rebase versus merge, reflog, etc), but once grasped, it becomes very easy to see why it is so popular.
I now use git locally as much as I use it with a hosted repository (GitHub, GitLab, BitBucket, etc). I use both command line and GUIs, and I thoroughly enjoy being able to trust git so completely.
Man, has he gotten fat.