HNHacker News
TopNewBestAskShowJobs

chrishill89

87 karma · joined May 27, 2022

submissionscomments
chrishill89··on At the end you use `git bisect`
The question as stated answers itself. Some are able to make nice multi-commit PRs. So they do.
chrishill89··on At the end you use `git bisect`
I found a bug in a OSS program where an entry contained a gibberish string. This was C so apparently it was some not-properly initialized variable. I couldn’t hope to trace down where that occurred. But with time to write a bisect script and half an hour to run the bisect session I was able to find the commit.
chrishill89··on Git Notes: Git's coolest, most unloved­ feature (2022)
> You can always use git notes for yourself but it is as I called it gimmick feature. I can make a bet you will use it for a week maybe couple weeks and then just stop because in description it sounds good but in practice - no one is using it.

I'll take that bet.

chrishill89··on Git Notes: Git's coolest, most unloved­ feature (2022)
There’s plenty of data which is extra-commit (doesn’t belong in the commit messages but is relevant). Like if the tests passed[1], manual testing notes, iteration on code review, notes for code spelunking when you find a problematic 6+ month old commit.

[^1] https://news.ycombinator.com/item?id=44348438

chrishill89··on Git Notes: Git's coolest, most unloved­ feature (2022)
There are many uses for Git notes even though you might use a project management tool. Particularly all those things which are relevant for developers only, or that the developers can use as the data source for "higher-level" goals.
chrishill89··on Git Notes: Git's coolest, most unloved­ feature (2022)
By default no copying will be done. While `notes.rewrite.amend` and `notes.rewrite.rebase` are true you also need `notes.rewriteRef` which tells it what notes refs should be rewritten. And it has no value by default. (you can set it to a glob copy over all notes refs)
chrishill89··on Git Notes: Git's coolest, most unloved­ feature (2022)
"Pickaxe" most often means those options. But it's also a (apparently legacy) command alias for git-blame.
chrishill89··on Git Notes: Git's coolest, most unloved­ feature (2022)
I use Git Notes all the time. It doesn’t look like the article mentions that you can attach the notes to emails you send out with git-format-patch(1) and git-send-email(1). I use them for emails to attach comments on patches that shouldn’t go into the commit message:

    I did <insert research notes> and found no other places in the code base where this needs to be fixed.
And as the cover letter for a single patch (if needed/not cowered by the commit message).

And also like a commit message on the iterations on the patches. So for a patch series that go over three versions the note may say what updates where done in versions 2 and 3.

And other than that I use notes for:

- Private notes on how I’ve manually tested the commit

- Link to CI

- A localized changelog for customers (who are not technical)

chrishill89··on Git Bug: Distributed, Offline-First Bug Tracker Embedded in Git, with Bridges
I create issues for myself. I sometimes create branches with informal issue keys like `<three letters>-<incrementing number>`. I use `git bug webui` to type edit issues. That’s it.

All I would like as a nice-to-have is support for making those issue keys that I use automatic. For me it’s all local to me so it doesn’t need to be globally unique. But those N-letter prefix does make collision less likely on a single project. I’ve mentioned it here: https://github.com/git-bug/git-bug/issues/75#issuecomment-19...

chrishill89··on Git Bug: Distributed, Offline-First Bug Tracker Embedded in Git, with Bridges
Thanks. I can’t push it to the issue tracker since they don’t have one. Bug reports go to the mailing list.
chrishill89··on Git Bug: Distributed, Offline-First Bug Tracker Embedded in Git, with Bridges
I use this to note potential issues in my copy of the Git project. For outright bugs though I report them to the project.
chrishill89··on What I've learned from jj
I’ve just used it two times in the last few months. One was to track down a commit which triggered a bug I found in Git. I wouldn’t be able to troubleshoot it myself. And I couldn’t send the whole repository because it’s not OSS. But with a script to reproduce the bug and half an hour I was able to find the problematic change.

I also tried to make a minimal reproduction but wasn’t able to.

chrishill89··on What I've learned from jj
Here’s the most relevant (to me) difference:

- The real unit of change lives in Git

- The real unit of change lives on some forge

I want it to live in Git.

chrishill89··on fd: A simple, fast and user-friendly alternative to 'find'
This is an active fork: https://github.com/dathere/qsv
chrishill89··on Show HN: "Git who" – A new CLI tool for industrial-scale Git blaming
The Git project has a script for that: https://github.com/git/git/blob/master/contrib/contacts/git-...
chrishill89··on An epic treatise on error models for systems programming languages
> "They're just another kind of value" and "Just crash and restart" are both wrong, error handling is much more complex than that.

Regarding the first slogan: it is correct that they are just another kind of value. But they look more complex. The reason isn’t the error values as such. The reason is that they always (except for `None | Error`) appear together with a rich normal-value. In turn they end up looking more complex, simply by context-association.

And treating errors as “just another kind of value” necessitates general value-manipulation facilities. Because we expect to manipulate normal values in whatever ways we want. But for errors we tend to get stumped once we want to do something else than whatever the default is, which might be (using your example) to accumulate errors instead of bailing out after finding just one.

But we tend to get stuck trying to simplify errors and their processing too much. When really we should move towards generality; the “just” in “just another kind of value” should mean that we can use a lot of whatever we have (not invent new things). “Just” shouldn’t hint at “and so it’s easy/simple”.

chrishill89··on An epic treatise on error models for systems programming languages
What a great resource.

I’ve been in the errors-as-values (like Rust and Haskell) camp for over a decade. But I think you need to go a bit beyond that philosophy (i.e. not just as a slogan) in order to truly unlock what treating errors as values can be like.

Errors as values are just values. Not any more complex than other values. But except in cases like `None | Error` (no regular value or an error) you at least double the work that you need to do. The thing about the “happy path” is that it’s just one channel of information. With errors you layer on one more channel.

I think that has lead me to be careful about treating errors in the same way I treat other kinds of values. Because simplifying errors down to just one value, like a string, simplifies the whole error channel. But how is error-as-string useful if you are using the code as a library and not just erroring out and letting the end-user/operator read it with their primate brain? Well that lead me to add variants, sum types for all the different error possibilities. Along with associated data like how this int was larger than some exected bound.

Already here I run into a sort of sub-error problem: only advertising what variants of an error that I can return from a function. It seems that this isn’t supported directly in Rust, which is what I was using at that time. So if I just advertise that I can return all the error variants (which is a soft lie) then I might have to deal with them downstream.

More expressive error values seem to both (1) burden the implementation (sub-errors) and (2) the clients who have to consume them.

So do I use less expressive errors so that they become easier to handle for client code?

Well it seems here that I’m still stuck treating errors as special values. Because I still seem hesitant to treat errors in their full generality as values. Think of a value:

1. It’s not just a single “atom”/scalar like an integer

2. It can be in a list or a dictionary/map or on a stack or some other data structure

3. It can be transformed with a map or reduce operation since you either want to transform the error or leave out some details

Maybe the library should provide a list of errors. But the client only cares about the last one. So it takes the head of the list. Maybe the library should provide a detailed error containing associated data about the bounds, expected and actual. But the client just cares about the error type and elides the rest.

Maybe this doesn’t seem useful for errors. But take the `Hurdle` type in the article. These are effectively a mandatory (product type) list of warnings. Maybe the client code only cares about five out of eight of them so he filters out the rest. Or he doesn’t care about them at all so he just ignores them.

And this seems to lead into iterators and lazy evaluation. In high level languages at least we’ve come to a point where we value pretty functional, declarative, and coarse transformation of data. “Coarse” in the sense that we use a few datastructures and operations to process the data without worrying about special datastructures and micromanaging control flow. But crucially these constructs don’t have to allocate N intermediary collections. They can be lazy. So maybe errors (and warnings (hurdles)) should be as well. Maybe you can have a very rich error/warning declaration while only generating them when the client code wants it. And what happens when you make a nod (at least) that all the error computations can be elided if only, say, you just kill the application on the first sight of any problems (like in a script)? People are incentivized to treat errors as any other value without fearing that it is wasted effort (computation).

Now on the other hand I do program in Java. And personally, philosophically, I am fine with exceptions for the type of application advocated in the article: report errors way up the stack. But there I’ve been running into the same problem: treating exceptions as not-quite-values. For any other custom class I would create however many fields and methods I need. But for custom exceptions I just give them a name. And maybe pass a string to the constructor. Which I could do for a built-in or third-party error. (So effectively I just give them a new name.)

But why? There must have been a block in my mind. Because Java exceptions are just classes. And they can have fields. And ultimately we log errors that bubble all the way up. So why not record all the associated data as fields on the exception and use the error handler (with `instanceof`) to gather up all the data into a structured log? That should be a ton better than just creating format strings everywhere and sticking the associated data into them. And of course you can do whatever else with the exception objects in the error handler; once it’s structured you can use the structure directly, without any parsing or indirect access.

chrishill89··on Colliding with the SHA prefix of Linux's initial Git commit
I know. I don’t understand why the Linux Kernel uses a hard-coded abbreviation instead of just letting Git figure out how long it should be (edited my comment now).
chrishill89··on Colliding with the SHA prefix of Linux's initial Git commit
> Presumably the problem is that these tools only take the abbreviated hash into account. Not also the subject:

Well the first mentioned script:

> > Tools like linux-next's “Fixes tag checker”,

has `get_full_hash`[1] which uses the subject to search through the abbreviated matches.

Edit: And that check was added two weeks ago by Kees [2].

[1]: https://github.com/kees/kernel-tools/blob/trunk/helpers/chec...

[2]: https://github.com/kees/kernel-tools/commit/5bf6a1e71df59a23...

chrishill89··on Colliding with the SHA prefix of Linux's initial Git commit
Presumably the problem is that these tools only take the abbreviated hash into account. Not also the subject:

    <abbrev. hash> ("<subject>")
You also have another data point. You only need to search in the history from the commit that you are reading. Assuming that the "Fixes" commit is an ancestor of the commit whose commit footer you are reading.

I always just assumed that tools would take all the data into account. Which means that you both need to collide with the abbreviated hash as well as the subject. Now I don't do that since I just copy-paste the hash, but I would quickly notice in case the subject is different (and likely the commit message and the diff just look irrelevant).

I don't understand why the Linux Kernel has this hard-coded rule[2] -- again, you were going to get collisions eventually, so the tools should have just taken all the data into account (at least the subject) from the start. The recommendation in the Git project is to use `git show -s --pretty=reference`, without any fiddling with the abbreviation:

    <abbrev. hash> (subject, ISO date)
Although the Git maintainer uses `--abbrev=8` since git-show will just use a longer abbreviation in case the output would be ambiguous[1].

They could have used this instead if they wanted simpler, future-proof tooling:

    Fixes: <full hash>
Just like tools like git-revert and git-cherry-pick do.

[1]: https://lore.kernel.org/git/xmqq34j5h7v9.fsf@gitster.g/

[2]: Edit: hard-coded as opposed to Git just figuring out how long the abbreviation should be based on how many objects there are.

chrishill89··on I'm daily driving Jujutsu, and maybe you should too
You want to rewrite each commit snapshot in isolation? Like fix the formatting for each?

Sounds like `git test fix` from git-branchless.

https://blog.waleedkhan.name/formatting-a-commit-stack/

chrishill89··on I'm daily driving Jujutsu, and maybe you should too
> Or maybe we need something more than plaintext to store code and the meta data around it before we can get to this point. I'm sure it has been already done multiple times in different ecosystems but there must be a way to standardize the information about code and use it in IDE and versioning tools.

This idea of storing more structured information (in the VCS or on disk) keeps coming up but I don't think it's necessary at all. The disk and the VCS can stay dumb. You can reconstruct the rich information after the fact.

It boils down to storing the right information and extracting it. Maybe you want to guard some checkpoints from bad/unstructured data. But that's just a check[1]. The serialization can be simple.

Well say you are blocked from storing bad data (according to the structure). Say you can also detect refactors like code moves (better than Git does, I guess just line moves). Again the storage can be dumb so that's fine. But you still need intentional commits; you won't get helped by a 500-line diff where all kinds of back and forth to solve three and half separate concerns are buried.

First of all you need good data in order to have something to extract. Which requires commit discipline.

[1]: And you don't get much from storing in a more rich format by itself. Just that you committed/stored syntactically correct/compilable artifacts. That's not nearly as rich as what you want. Because there are thousands of different intents (like you allude to) like refactorings -- all are valid in their own way so the structured representation can't have an opinion on it.

chrishill89··on Rust's Two Kinds of 'Assert' Make for Better Code
Thanks. Yeah, I've also been having similar thoughts for logging systems. In particular right now we (at my job) configure the log level through compile-time switches. It would be nicer to be able to toggle things dynamically. Yes, including per module instead of just "debug" or "info"

This is Spring Boot so it seems possible.

chrishill89··on Rust's Two Kinds of 'Assert' Make for Better Code
I agree with the others here that classic assert statements are in a weird spot. You enable it when testing, in other words when the data flowing through your program is likely to be constrained and lazy. And then turn it off when your program is about to meet real data (production). There's some wrinkle to this where you might want to log certain things in production instead of aborting with an assertion failure. But let's assume failing hard and fast is preferred for now.

Basically there are three cases.

1. The performance hit of the assertion checks are okay in production

2. It has a cost but can be lived with

3. Cannot be tolerated

For number two there is some space to play with.

What I've wanted is something more fine-grained. Not a global off/on switch but a way to turn on asserts for certain modules/classes/subsection, certain kinds of checks, and so on. It would also be nice to get some sort of data about the code coverage for assertions which have lived in a deployed program for months. You could then use that data to change the program (hopefully you can toggle these dynamically); maybe there is some often-hit assertion that has a noticeable impact which checks a code path that has been stable for months. Turning it off would not mean that you lose that experience wholesale if you have some way to store that data. I mean: imagining that there is some history tool that you can load the data about this code path into. You see the data travels through it and what range it uses. Then you have empirical data across however many runs (in various places) that indeed this assertion is never triggered. Which you can use to make a case about the likelihood of regression if you turn off that assertion.

That assumes that the code path is stable. You might need to enable the assertion again if the code path changes.

Just as a concrete example. You are working with some generated code in a language which doesn't support macros or some other, better way than code generation. You need to assert that the code is true to whatever it is modelling and doesn't fall out of sync. You could then enable some kind of reflection code that checks that every time the generated code interacts with the rest of the system. Then you eventually end up with twenty classes like that one and performance degrades. Well you could run these assertions until you have enough data to argue that the generated code is correct with a high level of certainty.

Then every time you need to regenerate and edit the code you would turn the assertions back on.

chrishill89··on ASCII Delimited Text – Not CSV or Tab Delimited Text
> For machine-readable formats, I'd argue length-before-text is simpler to parse than splitting by separators

Yes.

> even before having to add extra application logic to handle this method being fallible (e.g: if you want to use it to store filenames, you now need a check and pop-up about some OS-valid filenames not being supported).

I don't see why you'd need that. CSV does not have anything like that.

That's a higher-level concern. This ascii-delimited format (like CSV) is supposed to be a stupid row/column format. And also simpler to implement than CSV.

chrishill89··on ASCII Delimited Text – Not CSV or Tab Delimited Text
Then I don't understand. Like the sibling comment said the really "problematic" character in TSV is the line feed. But a tab can occur as well.

The format that I described does not already exist in the form of TSV. And further based on your original comment I would have thought that both TSV and this format would be discarded as not-useful.

chrishill89··on ASCII Delimited Text – Not CSV or Tab Delimited Text
Thanks for that. Handy. I think I’ll have use for that myself.
chrishill89··on ASCII Delimited Text – Not CSV or Tab Delimited Text
Whoops, meant to link to qsv https://github.com/jqnatividad/qsv
chrishill89··on ASCII Delimited Text – Not CSV or Tab Delimited Text
You can disallow those metacharacters in the data proper. Then you have a format that can store any utf8 or whatever except the non-whitespace control codes without any escaping. That solves a problem in an opinionated way. Just like how json is opinionated (utf8 only).

You can convert to another format if you need something crazier than rows and columns consisting of normal text.

chrishill89··on ASCII Delimited Text – Not CSV or Tab Delimited Text
You can use ASCII-separated values in qsv.[1]

For the unlikely event that you are dealing with data with the metacharacters: qsv will use some other control character as the “quote” character to deal with that.

← PreviousPage 2 of 3Next →