87 karma · joined May 27, 2022
I'll take that bet.
I did <insert research notes> and found no other places in the code base where this needs to be fixed.
And as the cover letter for a single patch (if needed/not cowered by the commit message).And also like a commit message on the iterations on the patches. So for a patch series that go over three versions the note may say what updates where done in versions 2 and 3.
And other than that I use notes for:
- Private notes on how I’ve manually tested the commit
- Link to CI
- A localized changelog for customers (who are not technical)
All I would like as a nice-to-have is support for making those issue keys that I use automatic. For me it’s all local to me so it doesn’t need to be globally unique. But those N-letter prefix does make collision less likely on a single project. I’ve mentioned it here: https://github.com/git-bug/git-bug/issues/75#issuecomment-19...
I also tried to make a minimal reproduction but wasn’t able to.
- The real unit of change lives in Git
- The real unit of change lives on some forge
I want it to live in Git.
Regarding the first slogan: it is correct that they are just another kind of value. But they look more complex. The reason isn’t the error values as such. The reason is that they always (except for `None | Error`) appear together with a rich normal-value. In turn they end up looking more complex, simply by context-association.
And treating errors as “just another kind of value” necessitates general value-manipulation facilities. Because we expect to manipulate normal values in whatever ways we want. But for errors we tend to get stumped once we want to do something else than whatever the default is, which might be (using your example) to accumulate errors instead of bailing out after finding just one.
But we tend to get stuck trying to simplify errors and their processing too much. When really we should move towards generality; the “just” in “just another kind of value” should mean that we can use a lot of whatever we have (not invent new things). “Just” shouldn’t hint at “and so it’s easy/simple”.
I’ve been in the errors-as-values (like Rust and Haskell) camp for over a decade. But I think you need to go a bit beyond that philosophy (i.e. not just as a slogan) in order to truly unlock what treating errors as values can be like.
Errors as values are just values. Not any more complex than other values. But except in cases like `None | Error` (no regular value or an error) you at least double the work that you need to do. The thing about the “happy path” is that it’s just one channel of information. With errors you layer on one more channel.
I think that has lead me to be careful about treating errors in the same way I treat other kinds of values. Because simplifying errors down to just one value, like a string, simplifies the whole error channel. But how is error-as-string useful if you are using the code as a library and not just erroring out and letting the end-user/operator read it with their primate brain? Well that lead me to add variants, sum types for all the different error possibilities. Along with associated data like how this int was larger than some exected bound.
Already here I run into a sort of sub-error problem: only advertising what variants of an error that I can return from a function. It seems that this isn’t supported directly in Rust, which is what I was using at that time. So if I just advertise that I can return all the error variants (which is a soft lie) then I might have to deal with them downstream.
More expressive error values seem to both (1) burden the implementation (sub-errors) and (2) the clients who have to consume them.
So do I use less expressive errors so that they become easier to handle for client code?
Well it seems here that I’m still stuck treating errors as special values. Because I still seem hesitant to treat errors in their full generality as values. Think of a value:
1. It’s not just a single “atom”/scalar like an integer
2. It can be in a list or a dictionary/map or on a stack or some other data structure
3. It can be transformed with a map or reduce operation since you either want to transform the error or leave out some details
Maybe the library should provide a list of errors. But the client only cares about the last one. So it takes the head of the list. Maybe the library should provide a detailed error containing associated data about the bounds, expected and actual. But the client just cares about the error type and elides the rest.
Maybe this doesn’t seem useful for errors. But take the `Hurdle` type in the article. These are effectively a mandatory (product type) list of warnings. Maybe the client code only cares about five out of eight of them so he filters out the rest. Or he doesn’t care about them at all so he just ignores them.
And this seems to lead into iterators and lazy evaluation. In high level languages at least we’ve come to a point where we value pretty functional, declarative, and coarse transformation of data. “Coarse” in the sense that we use a few datastructures and operations to process the data without worrying about special datastructures and micromanaging control flow. But crucially these constructs don’t have to allocate N intermediary collections. They can be lazy. So maybe errors (and warnings (hurdles)) should be as well. Maybe you can have a very rich error/warning declaration while only generating them when the client code wants it. And what happens when you make a nod (at least) that all the error computations can be elided if only, say, you just kill the application on the first sight of any problems (like in a script)? People are incentivized to treat errors as any other value without fearing that it is wasted effort (computation).
Now on the other hand I do program in Java. And personally, philosophically, I am fine with exceptions for the type of application advocated in the article: report errors way up the stack. But there I’ve been running into the same problem: treating exceptions as not-quite-values. For any other custom class I would create however many fields and methods I need. But for custom exceptions I just give them a name. And maybe pass a string to the constructor. Which I could do for a built-in or third-party error. (So effectively I just give them a new name.)
But why? There must have been a block in my mind. Because Java exceptions are just classes. And they can have fields. And ultimately we log errors that bubble all the way up. So why not record all the associated data as fields on the exception and use the error handler (with `instanceof`) to gather up all the data into a structured log? That should be a ton better than just creating format strings everywhere and sticking the associated data into them. And of course you can do whatever else with the exception objects in the error handler; once it’s structured you can use the structure directly, without any parsing or indirect access.
Well the first mentioned script:
> > Tools like linux-next's “Fixes tag checker”,
has `get_full_hash`[1] which uses the subject to search through the abbreviated matches.
Edit: And that check was added two weeks ago by Kees [2].
[1]: https://github.com/kees/kernel-tools/blob/trunk/helpers/chec...
[2]: https://github.com/kees/kernel-tools/commit/5bf6a1e71df59a23...
<abbrev. hash> ("<subject>")
You also have another data point. You only need to search in the history from the commit that you are reading. Assuming that the "Fixes" commit is an ancestor of the commit whose commit footer you are reading.I always just assumed that tools would take all the data into account. Which means that you both need to collide with the abbreviated hash as well as the subject. Now I don't do that since I just copy-paste the hash, but I would quickly notice in case the subject is different (and likely the commit message and the diff just look irrelevant).
I don't understand why the Linux Kernel has this hard-coded rule[2] -- again, you were going to get collisions eventually, so the tools should have just taken all the data into account (at least the subject) from the start. The recommendation in the Git project is to use `git show -s --pretty=reference`, without any fiddling with the abbreviation:
<abbrev. hash> (subject, ISO date)
Although the Git maintainer uses `--abbrev=8` since git-show will just use a longer abbreviation in case the output would be ambiguous[1].They could have used this instead if they wanted simpler, future-proof tooling:
Fixes: <full hash>
Just like tools like git-revert and git-cherry-pick do.[1]: https://lore.kernel.org/git/xmqq34j5h7v9.fsf@gitster.g/
[2]: Edit: hard-coded as opposed to Git just figuring out how long the abbreviation should be based on how many objects there are.
Sounds like `git test fix` from git-branchless.
This idea of storing more structured information (in the VCS or on disk) keeps coming up but I don't think it's necessary at all. The disk and the VCS can stay dumb. You can reconstruct the rich information after the fact.
It boils down to storing the right information and extracting it. Maybe you want to guard some checkpoints from bad/unstructured data. But that's just a check[1]. The serialization can be simple.
Well say you are blocked from storing bad data (according to the structure). Say you can also detect refactors like code moves (better than Git does, I guess just line moves). Again the storage can be dumb so that's fine. But you still need intentional commits; you won't get helped by a 500-line diff where all kinds of back and forth to solve three and half separate concerns are buried.
First of all you need good data in order to have something to extract. Which requires commit discipline.
[1]: And you don't get much from storing in a more rich format by itself. Just that you committed/stored syntactically correct/compilable artifacts. That's not nearly as rich as what you want. Because there are thousands of different intents (like you allude to) like refactorings -- all are valid in their own way so the structured representation can't have an opinion on it.
This is Spring Boot so it seems possible.
Basically there are three cases.
1. The performance hit of the assertion checks are okay in production
2. It has a cost but can be lived with
3. Cannot be tolerated
For number two there is some space to play with.
What I've wanted is something more fine-grained. Not a global off/on switch but a way to turn on asserts for certain modules/classes/subsection, certain kinds of checks, and so on. It would also be nice to get some sort of data about the code coverage for assertions which have lived in a deployed program for months. You could then use that data to change the program (hopefully you can toggle these dynamically); maybe there is some often-hit assertion that has a noticeable impact which checks a code path that has been stable for months. Turning it off would not mean that you lose that experience wholesale if you have some way to store that data. I mean: imagining that there is some history tool that you can load the data about this code path into. You see the data travels through it and what range it uses. Then you have empirical data across however many runs (in various places) that indeed this assertion is never triggered. Which you can use to make a case about the likelihood of regression if you turn off that assertion.
That assumes that the code path is stable. You might need to enable the assertion again if the code path changes.
Just as a concrete example. You are working with some generated code in a language which doesn't support macros or some other, better way than code generation. You need to assert that the code is true to whatever it is modelling and doesn't fall out of sync. You could then enable some kind of reflection code that checks that every time the generated code interacts with the rest of the system. Then you eventually end up with twenty classes like that one and performance degrades. Well you could run these assertions until you have enough data to argue that the generated code is correct with a high level of certainty.
Then every time you need to regenerate and edit the code you would turn the assertions back on.
Yes.
> even before having to add extra application logic to handle this method being fallible (e.g: if you want to use it to store filenames, you now need a check and pop-up about some OS-valid filenames not being supported).
I don't see why you'd need that. CSV does not have anything like that.
That's a higher-level concern. This ascii-delimited format (like CSV) is supposed to be a stupid row/column format. And also simpler to implement than CSV.
The format that I described does not already exist in the form of TSV. And further based on your original comment I would have thought that both TSV and this format would be discarded as not-useful.
You can convert to another format if you need something crazier than rows and columns consisting of normal text.
For the unlikely event that you are dealing with data with the metacharacters: qsv will use some other control character as the “quote” character to deal with that.