35,550 karma · joined July 2, 2014
Want to talk to me? Drop me an e-mail:
125-888-0378@kylheku.com
This could change without notice; check the profile.
Git repositories: https://www.kylheku.com/cgit
Mastodon: @Kazinator@mstdn.ca
LinkedIn: https://www.linkedin.com/in/kaz-kylheku-8a8b94197/
meet.hn/city/49.2608724,-123.113952/Vancouver
Socials: - kazinator.at.hn
---
Parent commits can have their own signatures, and it's true that there is an attack possible there where the same parent hash could point to two different commits, that have valid signatures of some kind (possibly from the same sneaky developer who is a bona fide project member).
Be that as it may, it's perfectly okay for the signature on a banana to validate only the banana, and not the gorilla that is holding it, or the vine the gorilla is swinging on, and the whole jungle.
It would be worth it to have better quality commit signing for SHA-1-based repos.
We can round up the bits that make up a commit in a SHA-1-based repo, and sign those bits securely; this is a thing that is possible.
Yes, I didn't understand that the GPG signing just operates on the top level object in the commit and trusts the SHA-1 hashes contained in it.
The signing process doesn't recursively traverse the bytes of the commit to pull them into GPG, like you would expect.
It's like, imagine you made a "bill of materials" of your project's files consisting of their names and CRC-32 checksums, and then signed this file, and called your project securely signed, LOL.
This aspect can be fixed without foisting new hashing scheme into the content tracker. In fact, it must be fixed; users on SHA-1-based repos deserve secure signing.
It's really sneaky that the SHA-1 business (not intended to be a security mechanism) was embroiled into the signing implementation; that GPG is demoted to the strength of SHA-1.
Was that just to save some cycles? It's certainly faster just to sign the commit object!
The "bytes passed to GPG" of course get hashed by GPG, using something better than SHA-1.
All bytes that comprise the commit should be hashed by GPG, rather than depending on the content referencing hash in the object tracking system.
This is something that is possible; it is not a logically deductive necessity that we just scan the topmost object and trust the hashes it contains.
That digest can be the SHA-256; since the infrastructure is there for it, signing should use SHA-256 regardless of what hash is used by the repository for identifying and linking content.
Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?
Oh, that would never be a problem for widely disseminated, popular, open source project, so it doesn't matter.
The GPG signature signs some kind of hash calculated over the commit, minus the GPG header, which is thereby added.
The git hash is then calculated over the whole thing. The git hash is on the outside, and not part of the signing.
A SHA-256 sum, though very good, only assures you with great confidence that you're looking at the same thing you looked at before, or that someone else is looking at elsewhere.
It is not a digital signature, and we don't want digital signatures to serve the role of content hashes.
Speaking of signatures, we have support for them in Git; you can use gpg to sign commits, and set it up to be done automatically.
Nobody is going to fake your commit such that the fake has the same SH-1 hash and your GPG signature.
The worry there is that the key holder (whether the legitimate one, or a malicious party who got a hold of the key) somehow does this: creates a new commit, signed with their key, which somehow has the same SH-1 as an existing signed commit. The git hash includes the GPG signature, so there is a significant layer of difficulty there which is likely harder than faking an unsigned SHA-256 commit.
The stewards of Git are going to do whatever they want, and there is nothing you can do about it if you don't have the clout to create a fork that takes the lead.
No amount of discussion will do anything because they've already decided that their view of the situation is correct. Git hashes are not just content identification but a digital certificate mechanism, and their collision resistance is a grave issue that must be fixed, the end.
You will be browbeaten in any discussion; it's not worth the energy in a world replete with issues.
Plain git init could fail with a diagnostic: informing to use one of the two aliases or an option.
What people don't want is making git repos SHA-256 by accident and finding out later that they made repos not compatible with older git.
Nothing else matters.
Git hashes are not supposed to be a security mechanism. If your basis for trusting that you have the right checkout is the git hash, in a situation where you have legitimate concern about untrusted parties manipulating remote repositories, then you're simply wrong.
If you need to certify the authenticity of some code, and you've decided that a Git hash of any kind is going to be your certificate, you have a problem between keyboard and chair which is not fixable by stronger hashes in Git.
I don't want instability and churn in tooling.
For instance consider that pitch class inversion of a major scale around its root results in precisely a major third key change. E.g. if we invert C major, we get C Phrygian, and C Phrygian is a mode of A♭ major: Coltrane change!
Try improvising around the two chord loop CM7, C♯M7 (or D♭M7), CM7, C#M7 by shifting between C major and its inversion C Phrygian; it is very melodic and intuitive.
You can feel/invoke similar mood changes when following the Coltrane changes.
C#M7/D♭M7 contains a Fm triad, so it evokes hint of the common parallel minor device. C Phrygian and C Aeolian/minor are in fact only one note off, Phrygian is Aeolian with the second degree flattened.
A Coltrane key change can be fibbed by the improviser by a parallel major/minor shift which is so ingrained into the pop music repertoire that it's easy to find familiar idioms that the listener will latch onto. You omit one note that doesn't work and are good to go.
Here is a cool thing about the Coltrane keys (related to the above semitone chord change). In each key you can pick a major triad such that their roots form a semitone progression. WE already identified C,C♯/D♭ from C and A♭. Then the third key, E, gives us the B triad.
The B triad in E major is at the root of the Mixolydian mode. The C triad in C is of course at the root of Ionian. The D♭ triad is at the root of Lydian. I.e. we have the I chord from one key, IV chord from another, and V chord from the third, and they end up chromatic. This can be exploited in improvisation. As the accompanist goes through changes you can just move by a semitone and then change one note in the scale at the same time. E.g. if we go from Mixolydian to Ionian, we sharpen the 7th. At the same time we go up a semitone, since the root is moving from B to C. Then for the next change, we do it again; root moves to D♭, scale shape changes one note to make Lydian. You can emphasize the chromaticism so the half note shift intent is clearly conveyed, while the changing note in the scale takes the sharp corners off it and makes it more interesting. The small chromatic movement helps the listener make sense of the weird chords that are going on (because the Coltrane chord progressions have a II-V-I in each of the three keys and whatnot).
Speaking of the II-V-I progressions in those keys: here is one more thing. It's possible to see a tritone movement there that evokes the tritone substitute.
Here is why. Consider the II-V-I in the key of C, where I is C. Now consider II-V-I in the key of E. What is the II note? It is G♭. That is exactly a tritone away from C.
So it's like we have a circle of fourths with tritones thrown in: D-G-C-G♭-E♭-A♭-... 4, 4, ♯4, 4, 4, ♯4, .. (if we do the key changes in that direction).
I use this as a trick to find the chords on the guitar fretboard.
The tritone has musical implications though, too. Suppose we alter the CM7 landing note to C7 and also change the G♭m7 II chord of E to G♭7. Then we have a note-sharing tritone root movement between two dominant chords.
If we have a compiler for language X in language X, and don't use any tools such as a parser or lexer generator, then there is nothing but code in X in the project, and that makes the developer of language X feel like they have earned major brownie points ... err, I mean, ... that they have kept their project free of cumbersome dependencies.
If language X is a one-implementation invention, then the tooling won't exist which hits these checkboxes: (1) is written in X; (2) generates code for X. By the time language X is mature enough that its ecosystem has something like that, it is long past the point where it would make sense to introduce it into its one and only implementation. There would have to be interest in writing another implementation.
E.g. the first C compilers would never have used Yacc.
Among languages that have one implementation, and that use parser generators, we will almost always see that another language is used for bootstrapping and the generator is for that language.
For some designers, that is a bruise to the ego, or else an unappetizing dependency. Even if they are boostrapping with another language, they are thinking forward to a future release where they will ditch that: they will rewrite parts that are in the boostrapping language in the new language to make it self-hosting.
If you use tooling like parser generation, which is in the ecosystem of, and oriented toward, the to-be-jettisoned-one-day boostrapping language, that throws a barrier in the path toward self-hosting. So you tend not to do it.
Like if you are boostrapping with C, and have it in the back of your mind to get rid of it, do you want to be bringing in a complicated tool with its own input language, which generates C? You think twice and are more likely to go ahead if you've resigned yourself to sticking with the C dependency.
Most of the make replacement tools are promoted by propaganda which attacks a strawman version of make, whereby it is assumed to be at the center of a shitty recursive situation.
Of course, it's not the same as having the parse tree all done from a previous pass and just walking it to do semantics.
I.e. basically not at all.