What git needs is a chair lift up that hill. A way to easily get people there. But I have no idea what that would look like. Lots of people try, few do very well at it.
What git needs is a chair lift up that hill. A way to easily get people there. But I have no idea what that would look like. Lots of people try, few do very well at it.
They mix multiple only slightly related commands in one.
In fact, one of the things in common among the ORMs that have left a bad taste in my mouth is that they all tried to abstract away SQL without leaking enough of it.
Picking the ideal interface to abstract is critically important (and very hard).
In the case of ORMs, available solutions abstract the schema (tables, rows, fields), the objects, or use templates. My solution abstracted JDBC/ODBC. The only leak in my abstraction was missing metadata, which I was able to plug (with much effort!).
My notions for interfaces, modularity, abstractions are mostly informed by the book "Design Rules: The Power of Modularity". http://a.co/hXOGJq1
So given a mental model is inevitable, seems reasonable that that model should be the actual model.
I think the longevity of SQL has proved there's value is non-leaky abstracted interfaces.
How is sql non-leaky? To be proficient with sql you have to understand how results are stored on disk, how indexes work, how joins work, etc. To debug and improve them you need to look at the query plan which is the database exposing it's inner workings to you.
You have to know about the abstractions an sql server sits on as well. Why is it faster if it's on an SSD instead of an HDD? Why does the data dissapear if it's an in memory DB?
No, you don’t. As far as I know, the data is stored in discrete little boxes and indexes are a separate stack of sorted little boxes connected to the main boxes by spaghetti. This is the abstraction, it works, and I don’t need to know about btrees, blocksizes, how locks are implemented, or anything else to grok a database.
Have you created an index? Was it clustered or non-clustered? That's not a black box, that's you giving implementation details to the database.
There’s no question that knowing more will get you more, but I think for the question of “when will things go sideways and I need to understand internals to save myself”, one would be able to use a relational database with success longer than git, getting by on abstractions alone. Running a high-performance installation of either is really outside the scope of the original point.
Yes, most of us will have to do both at some point, but they can be thought of as discrete skills.
If, on the other hand, you are working to create said history (and devise/use an advanced workflow for that), it's very helpful if you understand the underlying concepts. Which also goes for designing database layouts - someone who doesn't understand the basics of the optimizer will inevitably run into performance problems, just as someone who doesn't understand Git's inner workings will inevitably bork the repository.
You may need to go deeper and understand the underline model if you want performance but sticking to normal form can make unnecessary for a lot of people a lot of the time.
You can have a useful separation of work between a developer understanding/using sql and a DBA doing the DDL part and the optimization when needed.
That is, I am not proficient with relational databases, and I can handwave why an SDD is faster, and why data may disappear from an in-memory DB.
But I couldn't do an outer join without help. Nor do I know when I would want to do one.
Bob Martin wrote the essay at http://blog.cleancoder.com/uncle-bob/2017/12/09/Dbtails.html , in which he writes:
> Relational databases abstract away the physical nature of the disk, just as file systems do; but instead of storing informal arrays of bytes, relational databases provide access to sets of fixed sized records.
This isn't true. SQLite does not use fixed size records.
This suggests to me that a lot of people who consider themselves proficient with SQL don't know how the results are stored on disk, nor the difference between the SQL model and the actual implementation details, making them not proficient under your definition.
Because you know that information for other reasons as most people would. Just because the information is gained for other reasons does not make it irrelevant when using a database though.
> This isn't true. SQLite does not use fixed size records.
It's actually true of most/all modern databases these days. The point isn't knowing the exact structure the database uses to store it's information (even though it can be useful) but knowing how efficiently it can find the information for any given request. Knowing when a database is doing an index lookup or a full table scan is very important and I wouldn't consider someone that can't make a reasonable guess to be proficient in sql. Many of these details are even exposed in the sql, when you create an index and decide if it's clustered or non-clustered your giving the database specific directions about how the data will be physically stored.
The fact that you need to know anything about how they do their work internally to be reasonably competent at using them makes them a leaky abstraction.
Also, SQL has well-established processes and formalisms to design schemas which generally result in solid performance by themselves. That's what RDBMS are around for, after all: enabling efficient and consistent record-oriented data manipulation. This is quite difficult to do correctly in reality; for example, if you write your own transaction mechanism for disk/solid-state storage, you are going to do it wrong. This is genuinely difficult stuff.
There is a ton of internals that SQL abstracts so well that very few DB programmers know or (have to) care about them. Things like commit and rollback protocols, checkpointing, on-disk layouts, I/O scheduling, page allocation strategies, caching etc.
Certainly. My comment, however, concerned what you meant by 'proficient', and not simple use.
You used the qualifier "all modern databases". Was that meant to imply that SQLite is not a modern database?
My point remains that there are many people who are proficient in SQL, and would do very well with SQLite, even without knowing the on-disk format.
That is why I disagree with your use of the term "proficient".
Same exact thing above applies to so many things in software development, from IDEs, to code editors (Vim/Emacs/Sublime/etc), to programming languages, to deploy tools, the list goes on. There’s a reason software development is classified as skilled labor and not a low end job generally. You’re expected to have knowledge of, or be willing to learn a lot, to do your job.
I recall the time when mp3 was to demanding for many CPUs, so you had to convert to non-compressed formats. Today you do need to know that downloading non-compressed audio will cost you a lot of network traffic. Once performance is a concern, all abstractions have to be discarded.
That said, I do feel some "porcelain" git commands are poorly named and operate inconsistently -- compared to the plumbing of the acyclic graph concepts which is good but limited.
Half the time, when I know what I want, I both keep forgetting git flags and sub-commands - and struggle to find them in the man pages.
Like the fine:
git diff --name-only # i list only the files that are changed. But I'm not --list-files.
git diff --stat git rev-list --all --parents | grep "^.\{40\}.*<PARENT_SHA1>.*" | awk '{print $1}'
whereas in fossil, you use fossil timeline after <COMMIT>
I mean, one of these looks just a little more straightforward than the other, doesn't it?Also, a cursory test in a local git repo just now showed that command seems to print out only immediate descendants--i.e., unless that commit is the start of a branch, it's only going to tell you the single commit that comes immediately after it, not the timeline of activity that fossil will--and all it gives you is the hash of those commit(s), with no other information.
I use git myself, not fossil, but if this is something you really want in your workflow, fossil is a pretty clear win.
How many other ways of looking at commits or trees are there, that are hard in git but impossible in fossil because the author didn’t feel like it?
You could alternatively use:
git log --graph git log <COMMIT>.. git log --all --ancestry-path ^<COMMIT>I don't understand HN's hardon for hating Git.
There are about 5/6 fundamental operations you do in git/hg. If that's too much then again, there's not an abstraction that is going to help you out.
> git/hg
Mercurial was a great solution to the same problem that Git set out to tackle, virtually free of Git's foibles. The tradeoff was a few minor foibles of its own, but a much better tool. It's a fucking shame that Git managed to suck all the air out of the room, and we're left with a far, far worse industry standard.
I just said you can't give specifics on what to change, because there isn't much too change.
>And you act as if source control were invented with Git
No I'm not?
>and we're left with a far, far worse industry standard.
Yeah, we definitely should have gone with the system that can't do partial checkouts correctly or even roll things back. Branching name conflicts across remote repositories and bookmark fun! Git won for a reason, because it's good and sane at what it does.
The master can do no wrong.
Lacked a few dubious features such as merging multiple branches at the same time too.
It has improved but git is still noticeably more efficient with large repositories. (Almost straight comparison is any operation on Firefox repository vs its git port.)
Those dubious features are so relevant to daily work that I didn't even knew they existed.
Instead Mercurial uses additional cache file which instead is slower on Linux with big repos. But happens to be faster in Windows.
And the octopus merge is used by kernel maintainers sometimes if not quite a lot. That feature is impossible to add in Mercurial as it does not allow more than two commit parents.
No it doesn't? People use octopus merges all the time, every single day.
To emphasize that even more: Try to explain the concept of an ML-style sum type (i.e. a discriminated union in F#) to someone who only knows languages with C++-based type systems. You'll have a hard time to even explain why this is a good idea, because they will try to map it to the features they know (i.e. enums and/or inheritance hierarchies), and fail to get the upsides.
No, Mercurial's design is fundamentally inferior to Git, and practically the entire history of Mercurial development is trying to catch up to what Git did right from the start. For example having ridiculous "permanent" branches -> somebody makes "bookmarks" plugin to imitate Git's lightweight branches -> now there are two ways to branch, which is confusing. No way to stash -> somebody writes a shelve plugin -> need to enable plugin for this basic functionality instead of being proper part of VCS. Editing local history is hard -> Mercurial Queues plugin -> it's still hard -> now I think they have something like "phases". In Git all of this was easy from the start.
Another simple thing. How to get the commit id of the current revision. Let's search stack overflow:
https://stackoverflow.com/questions/2485651/print-current-me...
The top answer is `hg id -i`.
$ hg id -i
adc56745e928
The problem is, this answer is wrong! This simple command can execute for hours on a large enough repository, and requires write privileges to the repository! Moreover, it returns only a part of the hash. There's literally no option to display the full hash.The "correct" answer is `hg parent --template '{node}'`. Except `hg parent` is apparently deprecated, so the actual correct way is some `hg log` invocation with a lot of arguments.
Also, on the git/hg debate, I feel I've had problems (like the stash your modification and redownload everything) more often with git that hg. I mean perhaps it tells something about my capability to understand a directed acyclic graph, but hg seems less brittle when I'm using it.
What I don't like in git is the loss of history associated with squashing commits, I would prefer having a 'summary' that would keep the full history but by default would ne used like a single commit.
The DAG is already powerful enough to handle both the complicated details and the top-level summaries, it's just dumb that the UIs don't default to smarter displays.
(I find git stash essential given that `git add --interactive` is a painful UX compared to darcs and git doesn't have anything near darcs' smarts for merges when pulling/merging branches. Obviously, your mileage will vary.)
That seems like you made an assertion as well. I think there are counter-examples.
For example, the point of gitless is (quoting http://gitless.com/ ):
> Many people complain that Git is hard to use. We think the problem lies deeper than the user interface, in the concepts underlying Git. Gitless is an experiment to see what happens if you put a simple veneer on an app that changes the underlying concepts
Some commentary is at https://blog.acolyer.org/2016/10/24/whats-wrong-with-git-a-c... .
Many HN discussions as well, including https://news.ycombinator.com/item?id=6927485 .
I think that's an exaggeration. For example, Darcs and Pijul aren't based around a "graph of commits" like Git is, they use sets of inter-dependent patches instead. I'm sure there are other useful ways to model DVCS too.
Whilst this is mostly irrelevant for Git users, you mentioned Mercurial so I thought I'd chime in :)
> The only thing Git can really fix is changing it's command flags to be consistent across aliases/internal commands.
I mostly agree with this: Git is widespread enough that it should mostly be kept stable; anything too drastic should be done in a separate project, either an "overlay", or a separate (possibly Git-compatible) DVCS.
I said graph, I didn't say which graph. Both systems still use graphs. And still a graph you have to understand how to edit with each tool. The abstraction is still the same, and if you have problems with Git, you're going to have problems with either of those tools as well. The abstraction is not the problem, it's the developers inability to conceptualize the model in their head.
Where is the exaggeration?
You said "that graph" which, in context, I took to mean the git graph.
> Both systems still use graphs
True
> The abstraction is still the same
Not at all, since those graphs mean different things. Each makes some things easier and some things harder. For example, time is easy in git ("what did this look like last week?"). Changes are easy in Darcs ("does this conflict with that?"). Both tools allow the same sorts of things, but some are more natural than others. I think it's easy enough to use either as long as we think in its terms; learning to think in those terms may be hard. For git in particular, I think the CLI terminology doesn't help with that (e.g. "checkout").
> if you have problems with Git, you're going to have problems with either of those tools as well
Not necessarily. As a simple example, some git operations "replay" a sequence of commits (e.g. cherrypicking). I've often had sequences which introduce something then later remove it (bugs, workarounds, stubs, etc.). If there's a merge conflict during the "replay", I'll have to spend time manually reintroducing those useless changes, just so i can resume the "replay" which will remove them again.
From what I understand, in Darcs such changes would "cancel out" and not appear in the diff that we end up applying.
> Where is the exaggeration?
The idea that "uses a graph" implies "equally hard to use". The underlying datastructure != the abstraction; the semantics is much more important.
For example, the forward/back buttons of a browser can be implemented as a linked list; blockchains are also linked lists, but that doesn't mean that they're both the same abstraction, or that understanding each takes the same level of knowledge/experience/etc.
What I'm getting at is that if you don't understand what the graph entails, and what you need to do the graph, any system is going to be "hard to use." This idea that things should immediately make sense without understanding what you need to do or even what you're asking the system to do, is just silly.
I've never seen someone who understands git, darcs, mercurial, pijul, etc go "I totally understand how this data is being stored but it's just so hard to use!" I don't think that can be the case, because any of the graphs those applications choose to use have some shared cross section of operations:
* add
* remove
* merge
* reorder
* push
* pull
I see people confused about the above, because they don't understand what they're really asking the system to do. I don't think any abstraction is ever going to solve that.
Git does have a problem with its command line (or at least how consistent and ambiguous it can sometimes be), but you really should get past it after a week or two of using it. The rest is on you. If you know what you want/need to do getting past the CLI isn't hard. People struggle with the former and so they think the latter is what's stopping them.
I posted a bald statement. He replied directly with snide remarks and fallacies. Look at the timestamps and edits. I have every right to be annoyed and make it known that I am annoyed in my posts when the community refused to consistently adhere to guidelines.
Enforce guidelines that keep discussions rational, not because people don't want to be accosted in public for their misleading, emotionally bloated statements.
They are currently not being applied. Is it fair for me to point out how inconsistently the posts are being treated?
https://news.ycombinator.com/newsguidelines.html:
>Don't say things you wouldn't say face-to-face. Don't be snarky.
"Every day humans make me again realize that I love my dogs, and respect my dogs, more than humans. There are exceptions but they are few and far between." [2]
>Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize.
"Which sort of doesn't matter since everyone thinks GitHub is source management." [1]
>Please don't post shallow dismissals
"You all lost out on "the most sane and powerful" as a result." [1]
"Calling it a sane and powerful source control tool is just not supported by the facts, calling "the most ..." is laughable." [1]
"Calling Git sane just makes it clear that you haven't used a sane source management system." [1]
"Lots of people are too busy/whatever to know what they are missing, maybe that's you. It's not me" [3]
>When disagreeing, please reply to the argument instead of calling names. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."
"Arguing with some random dude who thinks he knows more than me is not really fun." [3]
"Dude, troll much?" [4]
[1] https://news.ycombinator.com/item?id=16806588
[2] https://news.ycombinator.com/item?id=16807652
[3] https://news.ycombinator.com/item?id=16806877
[4] https://news.ycombinator.com/item?id=16807763
At least I had the decency to correct dishonest statements with vanilla citations in my posts, and they still got ignored.
I didn't insinuate that people are worth less than pets that I bought and own (who can't even choose who to be dependent on) because they don't agree with my perspectives over a piece of software. In what context would this be an acceptable statement to make face to face or in a public setting and you go "well, you know, it's kind of okay to say!"
I'm exceedingly interested in where I crossed that line in a considerable manner because that's one distant line to cross. Next time someone says something I perceive to be incorrect, or they get on my nerves for continually disagreeing with me, I'll be sure to tell them my dog is worth more than them since that's actively being allowed and has a precedent of moderator support.
And for the record, my tone is probably "abrasive" in this post because the above actions and outright blind eye towards outright lies and uncalled for statements is aggravating. I have a feeling you're not doing anything just because of who he is, and not because what he is saying is warranted or even accurate (it's definitely not, as I demonstrated across several different posts).
I've archived this thread so people are free to review my actions and moderator actions at a later date: https://web.archive.org/web/20180411062201/https://news.ycom...
I've said my piece.
My current team are mostly controls engineers, working on PLCs. But the software we're now working with has its configurations tracked in git. These aren't dumb people, they're quite talented, but their education wasn't in CS, and "directed acyclic graph" is not a thing they have a mental model for.
So, my advice is to try to write some instructions for yourself for all the common cases you might run into during your work. It will not only help you realise what you actually need from git, but also will serve as a good cheat-sheet.
https://git-scm.com/book/en/v2/Git-Internals-Plumbing-and-Po...
https://leanpub.com/learngitthehardway
I start with simple examples and work up from there. It's based on training I've conducted at various companies, and avoids talk of Merkle trees or DAG.