HNHacker News
TopNewBestAskShowJobs

wyoung2

168 karma · joined December 15, 2015

submissionscomments
wyoung2··on Fossil Chat
Yes, chat messages are stored in a table in the SQLite DB backing the repository you're chatting about, which allows you to close the browser window when you need some peace to concentrate on work, then revisit the chat room later and see what's been going on while you were away.

As for the dream of making SQLite faster or better, I doubt it. Fossil works best for small teams with projects sized reasonably for those teams, and chat privilege isn't generally given out to the masses. There's only one chat room per repo. Therefore, there simply isn't enough I/O involved to drive much in the way of SQLite changes.

Fossil proper has resulted in SQLite improvements, though, such as the recursive CTE feature added in 3.34.0. That directly supported a feature for tracing the history of files through renames across repository history using a single efficient (though complicated) SQL query: https://fossil-scm.org/forum/forumpost/5631123d66

wyoung2··on Fossil Chat
Yes. Developers working independently is best for productivity, but sometimes you need to coordinate something, which then brings up the question: what to use? The last time I tried to list it, I came up with over a dozen options, all of which either require some sort of admin setup hassle or put you into the claws of some megacorp. With Fossil chat, you've already got Fossil set up, so it becomes the low-friction path.
wyoung2··on Fossil
> one must assume those same people using fossil will also want to hide their "imperfect" code too

One must not, because if one did, one would be wrong. :)

Go read Weinberger. It's $10 on Kindle right now.

> Please read the next sentence I wrote right after where you stopped quoting.

First, don't assume that because I didn't quote your posting in full that I didn't read it in full. HN is a threaded messaging system: we don't need to fully quote everything just to maintain the flow of the conversation. The history is right there to see on the page.

Second, how does "If a group of people are collaborating and they agree that rebases are going to happen, nothing is wrong with letting them do that," argue against the article in question? That's just a blind assertion, not logical argumentation.

> not a problem with git-rebase,

Sure it is: if rebase commits a breaking change to the blockchain immediately because you weren't able to test it, you have two options: 1. Commit a fix, pushing the broken commit later, potentially breaking bisect and such. 2. Do more rebase squashing and such to fix it in place before pushing it.

Argument 2 is "We need rebase because we used rebase." :)

Fossil's alternative is to not commit anything to the blockchain automatically. If Fossil did have rebase, it would make the changes in the checkout tree only, and you'd have to commit it separately.

The argument is not "Fossil can't have rebase because Git's version is badly considered," it's "Fossil's developers don't want rebase and Git's version is badly considered anyway." Fossil could avoid the design error, but that doesn't make rebase a good idea.

> If the repo is on Github and the PR is being merged through its web interface

...then you're using proprietary software with tremendous lock-in, but okay, if you're willing...

> the merge can remember to use it.

You're really going to insist on that? Commands by the foot, instead of a sensible default?

> it's not related to rebase then, in git or otherwise.

It's an example. The argument we've received multiple times from Git fans is that developers need rebase to make the timeline "clean", but Fossil shows that you don't have to modify history to do that. You just need sufficiently powerful tools that let you preserve history while changing its presentation to the user to suit various needs.

Git's porcelain is showing here.

> derp... - foo.bar() ... + foo.bar();

So you've committed without compiling first, much less running the tests, and your solution is "I need rebase?" No, my friend, you need to compile and run the tests before committing!

Maybe you want a better example?

wyoung2··on Fossil
I don't see why you would bisect in this situation in the first place. The problem's fixed now, as of commit 5.

But okay, let's take your example as-is: you determine the "good" point is commit 1, and the ...um, other good point is commit 5?

Well, that doesn't work. I guess we have to arbitrarily ignore commit 5 and say commit 4 is bad. A bisect will show that. Then in Fossil, if you visit the /info page for commit 4, it will show its child commit 5 as fixing the problem.

Try again. Squashing doesn't solve anything here.

wyoung2··on Fossil
The whole point of those articles is to compare and contrast.
wyoung2··on Fossil
It isn't shipped on macOS, and the Homebrew version of Git has it packaged separately.

I fired it up on a Git export of a Fossil repo here, and it's missing features of Fossil's web UI timeline view:

1. Cherrypick markers

2. Diff arbitrary versions (it shows only diff from previous)

3. Branch coloring

4. Hyperlinks to produce new timeline views: click author to get list of commits by that author, click timestamp to get a new timeline surrounding that point in time, etc....

5. Integration with other web UI features. For instance, you can create a Fossil wiki article attached to a commit from Fossil UI, but you can't attach a GitHub wiki article to a Git commit from gitk.

For another, Fossil has a feature to produce zip and tarball downloads of particular versions, which have links in the commit info page, which means visitors to your project don't need to clone the whole thing and roll it back to that version manually to get a single version. That's integrated into the Fossil web UI timeline: click a version, click Zip, done.

6. A modern browser fits its platform and offers more UI affordances than the mid-1990s Motif-inspired Tk. (Better copy/paste behavior, font rendering, zooming, native controls, etc.) Web UIs get a lot of angst these days, but like them or not, they're native citizens of the host platform these days in a way that Tk is not.

wyoung2··on Fossil
You don't even need the temporary files:

    $ vim
    ...write, write, write...
    :w !fossil wiki create "The Foo Article" -
    :q!
Later:

    $ vim
    :r !fossil wiki export "The Foo Article" -
    ...write, write, write some more...
    :w !fossil wiki commit "The Foo Article" -
Wrapping all of that up into Vim macros would be a small matter for someone sufficiently interested.

Also, someone wanting to edit wiki-like docs in a text editor rather than in the web UI is probably going to find Fossil's embedded docs feature more useful anyway: https://fossil-scm.org/home/doc/trunk/www/embeddeddoc.wiki

wyoung2··on Fossil
> Windows Vista...components were being developed in private branches.

The claim isn't that Microsoft developers on the Vista project used Git and private branches. The claim is that private branches are another form of siloing which leads to the same sorts of communication problems. They're a way to purposefully hoard code so your fellow developers can't see it. It is exactly what McCarthy was warning about in his "beware a guy in a room" comment; and McCarthy was at Microsoft when he wrote the book cited.

> I can't tell if the book specifically says private branches have anything to do with ego or not.

Weinberger wrote his seminal book in 1972, so probably not. :)

Human psychology hasn't notably changed since 1972. It doesn't matter if you're developing with punched cards or with worldwide Kubernetes clusters, humans are humans.

> Single-person development happens in branches only that single person cares about, so whether the branches are private or public makes no difference to that person

If you're doing single-person development, then the concerns over siloed development don't apply at all.

This section of the document is talking about communication among developers on the same project. If there is no communication on your project, its points are irrelevant, not wrong.

>4.0...So test each commit then.

You're missing the point. If Fossil offered Git-style rebase, the commit would be pushed up to the remote repo you cloned from before you could possibly test it, because of its autosync feature.

The autosync feature and the practice of leaving it enabled as much as possible is justified here: https://fossil-scm.org/fossil/doc/trunk/www/fossil-v-git.wik...

> `git rebase --ignore-date` will reset the commit date

How often do you suppose that's done in practice?

The point in the article you're rebutting is that Fossil doesn't make you do that at all, because it doesn't create timewarps or require after-the-fact date rewrites to avoid them.

> I'm not sure what (4) is referring to by "routine display"

It refers to the "fossil amend COMMITID --hide" feature. Affected branches no longer show up in the timeline, in the default branch list, etc., but no info is destroyed. It's just a tag telling the web UI and CLI not to show that branch by default.

> I don't need to go through half-commits that got reverted afterwards.

You're assuming 20/20 foresight. The very nature of software bugs is that you don't know you're committing them at the time, so how can you prospectively know which elements of a commit are good and which bad? You can mitigate it through testing, code review, etc., but Bugs Happen. There are whole companies dedicated to that fact.

The very point of bisect is, "Given this pile of commits between points GOOD and BAD, which one caused the symptom I'm seeing now?" If you knew the answer to that, and thus were able to make the in-advance judgement you suggest, you wouldn't need to do the bisect, because you wouldn't have committed the bug in the first place.

> this directly contradicts 4.0 unless we are to assume that only merge commits get tested

Only if you assume you have 100% test coverage, both in terms of lines of code and functionality. If you are in such a happy position, and you always run your tests before committing, then yes, it is impossible to commit a bug to the repo.

I wanna see that repo, the one without bugs because it has 100% functional-test coverage.

Even SQLite hasn't got that, evidenced by the fact that its test suite continues to change, even for historical features.

> If I have to squash two commits in my WIP branch I absolutely want to be able to.

You're toggling between "have to" and "want to".

And again, you're assuming 20/20 foresight, that you will never want to come back and tease those commits apart again.

The merge point is the proper place to logically "squash" things, not within the WIP branch.

wyoung2··on Fossil
Crypto is something best delegated to others. There are a bunch of ways to host Fossil, several of which offer HTTPS proxying: https://www.fossil-scm.org/home/doc/trunk/www/server/index.h...
wyoung2··on Fossil
A lot of us Fossil users host it on $5 VPSes. It's written in C atop SQLite, so the network is almost always the bottleneck, not the server-side processing speed.

sqlite.org and all of the other repos D. Richard Hipp maintains (Fossil, Pikchr, the separate SQLite docs repo, the forums for Fossil and SQLite...) all run on a $40/month VPS. The page generation time is calculated and displayed at the bottom of non-static pages like this one: https://sqlite.org/src/

wyoung2··on Apple aims to sell Macs with its own chips starting in 2021
> not built by Apple

If you hadn't included that restriction, I'd tell you about Numbers. Not all spreadsheet type problems fit into its limitations, but for those that do, it makes you ask "Why don't all spreadsheets work this way?"

But you did make that a restriction on your criteria, so forget I said anything. :)

The first non-Apple software I thought of when you asked the question was PDFpen, which operates vaguely like what is now called "Acrobat DC" used to. Acrobat drives me up a wall; it's like they asked, "How can we make the Microsoft ribbon interface even worse? Let's do that!"

I have about three Markdown editors that are Mac-only that I like better than any cross-platform alternative I've tried. (Byword, Markdown Pro and MultiMarkdown Composer.) I have three because none of them does everything right, so I occasionally have to switch for some task or other. I'd happily switch to something cross-platform if it matched the union of those three apps' features, those I care about, anyway.

I suspect I have three solid Markdown editors to choose from because they're all based on solid OS-level rich text editing features that exceed what you get on Windows, so the non-Mac competition has to waste a lot of resources reinventing wheels. Just as one example, you get grammar checking for free in most Mac native text-editing software. Where it doesn't occur, it's usually because the app is cross-platform and is thus avoiding Mac-specific features.

A related example is the built-in dictionary. You can get dictionary/thesaurus apps on other platforms, but they generally aren't as deeply integrated as the one on macOS.

Then there are the times where you have a cross-platform app that simply works better on macOS. VLC and MacVim (as compared to gVim) come to mind.

wyoung2··on A new hash algorithm for Git
In large measure, you actually can't, since there's a good chance those repos are behind HTTPS-only these days, and those versions of Git will be linked to ancient versions of OpenSSL that won't even talk to modern TLS implementations, the two being unable to agree on a common ciphersuite.

Beyond about 10 years, you usually end up freezing old binaries in place along with old data in order to continue manipulating it anyway.

wyoung2··on A new hash algorithm for Git
> Fossil...would force me to commit the proverbial 500-line blob all at once

Nope.

If it were me doing such a thing as you describe, I'd start the work on a feature branch. If I'm working on that repo with other active developers, this lets them see what I'm up to and possibly help; and if not help, then at least be aware about where my head's at, so they can better predict what's likely to land on the shared working branch later.

If I got to a point where only part of the branch needed to be applied, I could cherrypick those individual changes, either down to the parent branch or up to a higher-level feature branch.

All of this happens in public, with the work fully recorded, so someone doesn't have to reconstruct the development history after the fact later.

This mode of development helps keep your project's bus factor above 1.

wyoung2··on A new hash algorithm for Git
> ideology conflates architectural decisions and workflow processes with individual worth

No. You start with the ideology based on your local culture and project needs, then you pick the tool that supports your project's needs.

This is why we spend so much time talking about philosophy in the Fossil vs. Git article, particularly this section: https://fossil-scm.org/fossil/doc/trunk/www/fossil-v-git.wik...

Which of the two philosophies matches better with the way your project works? That alone is a pretty good guide to whether you want Fossil or Git. (Or something else!)

wyoung2··on A new hash algorithm for Git
> If the people checking all see the valid files

...which will likely contain thousands of bytes of pseudorandom data in order to force the hash collision...

> they cannot raise any alarms

You think a human won't be able to notice that the diff from the last version they tested looks awfully funny? Code that can fool the compiler into producing an evil binary is one thing, but code that can pass a human code review is quite another.

You might be surprised how often that occurs.

I don't do a diff before each third-party DVCS repo pull, but I do diff the code when integrating such third-party code into my projects, if only so I understand what they've done since the last time I updated. Commit messages, ChangeLogs, and release announcements only get you so far.

Back when I was producing binary packages for a popular software distribution, I'd often be forced to diff the code when producing new binaries, since several of the popular binary package distribution systems are based on patches atop pristine upstream source packages. (RPM, DEB, Cygwin packages...)

Each time a binary package creator updates, there's a good chance they've had to diff the versions to work out how to apply their old distro-specific patches atop the new codebase.

Someone's going to notice the first time this happens, and my guess is that it'll happen rather quickly.

wyoung2··on A new hash algorithm for Git
If you try to use Fossil 1.37 — the last 1.x release — to clone a repo that has SHA-3 hashed artifacts in it, it says, "server returned an error - clone aborted". Since 1.37 pre-dates this feature, it can't give a more detailed diagnosis than that.

If you have an old clone made from before the transition and try to update it, I'm not sure what it says, since I don't have any of those around any more. It has, after all, been three years since Fossil began to move on this problem, so that it's largely a past issue for us now.

This transition time was indeed annoying for us over in Fossil land, but Git's going to have to go through a transition like this, too. The question isn't whether but how long we'll have to wait for it to begin and how long it'll take to complete.

wyoung2··on A new hash algorithm for Git
SQLite can be considerably faster than the filesystem: https://www.sqlite.org/fasterthanfs.html

If you think your filesystem-based Git repo is easy to manipulate, go poking around in there, and what you'll find is a bespoke one-off pile-of-files database! Given a choice between Git's DB and SQLite, I put more trust into SQLite.

> I just want a C program in /usr/bin that does version control.

...which Git doesn't provide. Git is hundreds of files scattered all over your filesystem, a large number of which aren't C binaries anyway, and of those that are, only one of them is the front-end program sitting in /usr/bin, whereas Fossil can be built to a single static executable in /usr/bin.

And if you can't build Fossil statically on your system, it's likely due to an OS limitation rather than something about Fossil itself, as on RHEL where they've made fully static linking rather difficult in the past few releases.

Getting back to Git, large chunks of Git are written in POSIX shell, Perl, Python, and Tcl/Tk. Almost all of Fossil is written in C, and the rest of the code is embedded within that binary running under built-in interpreters rather than depending on platform interpreters.

This has nice knock-on effects, one of which is that Fossil is truly native on Windows, whereas you have to drag along a Linux portability environment to run Git on Windows. Another is that Fossil plays nicely with chroot/jail/container technology.

> I'm also not interested in version control systems that are dragging along a wiki and bug tracker.

Not a GitHub or GitLab user, then, I'm guessing?

wyoung2··on A new hash algorithm for Git
> history does not matter if the change was parented in some temporary context

It does if it means a big ball o' hackage lands on the public working branch, since it complicates merges, backouts, cherrypicks, and bisects.

Git users can also hide individual commit messages behind one big combined message, losing part of the project's development history and logical progression.

When I pull your repo and build it, and I find that it doesn't build on my system, I don't want to dig through a 500-line merge commit to figure out why you changed this one line from the one that used to build last week, I want the 14-line diff it was part of so I can begin to understand what you were thinking when you committed it. If I later find out that that 14-line change was wrong but the rest of your 500-line merge was fine, I want to be able to back it out with a single command. (In Fossil, it's `fossil merge --backout abcd1234`.)

> confuse other people with irrelevant information when they try to navigate the history.

How much time do you spend navigating the project's history vs looking at the tip of the current branch?

I'd wager that the times you dig back into the history, it's because you are in fact trying to figure out why you got here, which means a trail of detailed breadcrumbs will be more likely helpful than "...and between one week and the next, something changed in commit abcd1234, but we've lost all of its internal context, so we'll be spending next week reconstructing it because Angie's on vacation now."

wyoung2··on A new hash algorithm for Git
> sometimes I don't care about history and I'm just trying to coordinate developers across timelines

The fact that Fossil preserves history does not prevent you from coordinating with people across timelines. It is rather the whole point of a DVCS.

> conflating a workflow decision with a moral failing

I think it's fairer to say that we don't think a data repository is any place for lies of any sort, even white lies.

> I've learned to be somewhat skeptical of programming/workflow heuristics advertised as rules, and to be very skeptical of heuristics advertised as ideologies.

Sure, flexible tools are often better than inflexible ones, but you also have to consider the cost of the flexibility. Here, it means someone can say "this happened at some point in the past," and it's just plain wrong.

That isn't always an important thing. Most filesystems and databases operate on the same principle, presenting only the current truth, not any past truth.

Yet, we also have snapshotting in DBMSes and filesystems, because it's often very useful to be able to say, "This was the state of the system as of 2020.02.04."

You don't need a snapshotting filesystem for everything, and you don't need Fossil for everything, but it sure is nice to have ready access to both when needed.

> You've never accidentally committed a password to repo, or had to respond to a takedown request?

Fossil has shunning for that: https://fossil-scm.org/fossil/doc/trunk/www/shunning.wiki

And no, shunning is nothing at all like rebase, which should be clear from the article.

Fossil also has the `amend` command: http://fossil-scm.org/fossil/help?cmd=amend

And no, it is also not like rebase, because it only adds to the project history, it never destroys information.

wyoung2··on A new hash algorithm for Git
When is the right time to worry? Maybe wait until someone publishes a practical attack, then wait years for the new code to get sufficiently far out into the world that you can switch to it?

I mean, I see you're expressing concern, but the first major red flag on this went up three years ago, and another big one went up last month. (https://sha-mbles.github.io/)

When we dealt with this same problem over in Fossil land, we ended up needing to wait most of three years for Debian to finally ship a new enough binary that we could switch the default to SHA-3. Fortunately (?) RHEL doesn't ship Fossil, else we'd likely have had to wait even longer.

Atop that same problem, Git's also got tremendously more inertia. Git has to wait out not only the Debian and RHEL stable package policies but also all of that infrastructure tooling they brag on. Every random programmer's editor, merge tool, Git front end... all of that which a project depends on will have to convert over before that one project can move to a post-SHA-1 future.

This is going to be a colossal mess.

wyoung2··on A new hash algorithm for Git
It only takes one person to raise the flag.

Sure, many thousands of people doing blind "git clone && configure && sudo make install" could be burned by a problem like this, but someone would eventually do a diff and see the problem on any project big enough to have those thousands of trusting users in the first place.

I'm not excusing these SHA-1 weaknesses, only pointing out that it won't be trivial to apply them to program source code repos no matter how cheap the attacks get.

For instance, the demonstration case for SHAttered was a pair of PDFs: humans can't reasonably inspect those to find whatever noise had to be stuffed into them to achieve the result.

I also understand that these SHA-1 weaknesses have been used to attack X.509 certificates, but there again you have a case very unlike a software code repo, where the one doing the checking isn't another programmer but a program.

wyoung2··on A new hash algorithm for Git
Wow! I wouldn't have guessed that Git had that vulnerability. Fossil solves it easily: creating a new repo involves generating a random project code (a nonce) which goes into the hash of the first commit, so that even two identical commit sequences won't produce identical blockchains.

Fossil lets you force the project ID on creating the repo, but the capability only exists for special purposes.

wyoung2··on A new hash algorithm for Git
It's worth noting that this attack is a property of the Merkle–Damgård hash construction, not of SHA-1 specifically, which means SHA-2 (Git's path forward) is also vulnerable:

https://en.wikipedia.org/wiki/Merkle%E2%80%93Damg%C3%A5rd_co...

https://www.reddit.com/r/crypto/comments/44p5jc/eli5_why_are...

Fossil uses SHA-3, which has an entirely different construction, which is not at this time known to have a similar weakness. SHA-3 is also much newer, with a much shorter list of known attacks.

wyoung2··on A new hash algorithm for Git
In addition to D. Richard Hipp's thoughts as HN user SQLite — author also of Fossil, so he oughtta know — I offer these:

1. Keep in mind that Fossil and Git are both applications of blockchain technology, which in this particular practical case means you must not only forge a single artifact's hash, you must also do it in a way that allows it to fit into the overall blockchain.

2. Fossil's sync protocol purposefully won't apply Dr. Hipp's hypothetical evil.c to an existing Fossil blockchain if presented it. Fossil will say, "I've already got that one, thanks," and move on. Only new or outdated clones could be so-fooled.

wyoung2··on A new hash algorithm for Git
> Furthermore, the evil.c file with the same SHA1 hash would need to be valid C code that does something evil while still yielding the same hash

...and also produce an innocent-looking diff!

I mean, you could stuff a bunch of random bytes into a C comment to force the desired hash in the output using these documented attack techniques, but anyone inspecting the diffs between versions is likely to see such an explosion of noise and call foul.

If you want an analogy, it's like someone saying they've learned to impersonate federal agent identification cards, only it requires that the person carrying the fake ID to have a thousand rainbow-dyed ducks on a leash in tow behind him.

Such attacks are fine when it's dumb software systems doing the checks, but for a source code repository where people do in fact visually check the diffs occasionally?

Well, let's just say that when someone manages to use SHAttered and/or SHAmbles type attacks on Git (or even Fossil) I expect that it won't take a genius detective to see that the repo's been attacked.

wyoung2··on A new hash algorithm for Git
> That cute rhetoric will not fool anyone.

Well, let's see, the Fossil equivalents are:

1. Do nothing at all for a conversion from the SHA-1 to SHA-3 — yes, 3, not 2 as in Git! — because it's automatic for months now and dead easy going back 3 years now. (https://www.fossil-scm.org/fossil/doc/trunk/www/hashpolicy.w...)

2. "fossil diff"

3. "fossil ci"

4. Why are you rebasing in the first place, again? https://www.fossil-scm.org/fossil/doc/trunk/www/rebaseharm.m...

wyoung2··on Is Git Irreplaceable? (2019)
Agreed. Bisecting a deep and complex version history can be a lifesaver, especially under a time crunch.

Fossil has bisecting with much the same CLI as Git, for what that's worth.

wyoung2··on Is Git Irreplaceable? (2019)
> Ammending changes the hash of the commit

Not necessarily. Fossil's `amend` command works by adding additional information to the repo that the web UI and commands like `fossil info` look at when building up information intended for direct consumption by the user.

In this way, you can edit commit messages, rename branches/tags, add/remove tags, etc. to historical check-ins without breaking the blockchain / Merkle tree of commits.

This allows Fossil to keep all of the historical information about what happened to a given file, commit, ticket, etc. while still allowing a coherent presentation of the current state of affairs to the user.

wyoung2··on Is Git Irreplaceable? (2019)
Sorry, but this is a terrible analogy.

The core problem with it is that very few people can get paid more by being better at using their [D]VCS, whereas those more skilled with their programming language(s) of choice often do get paid more to wield that knowledge.

Consequently, most people do not fully master their version control system to the same level that they do with their programming language, their text editor, etc.

To be specific, there are many more C++ wizards and Vim wizards than there are Git wizards.

In situations like this, I prefer a tool that lets me pick it up quickly, use it easily, and then put it back down again without having to think too much about it.

You see this pattern over and over in software. It is why all OSes now have some sort of Control Panel / Settings app, even if all it does is call down to some low-level tool that modifies a registry setting, XML file, or whatever, which you could edit by hand if you wanted to. These tools exist even for geeky OSes like Linux because driving the OS is usually not the end user's goal, it is to do something productive atop that OS.

[D]VCSes are at this same level of infrastructure: something to use and then get past ASAP, so you can go be productive.

wyoung2··on Is Git Irreplaceable? (2019)
It's updated now, here: https://www.fossil-scm.org/fossil/doc/trunk/www/fossil-v-git...

Thanks for the feedback!

← PreviousPage 2 of 6Next →