Ncdu 2: Less hungry and more Ziggy
dev.yorhel.nl
dev.yorhel.nl
Keep in mind that "maintainence" according to them means "pretty much what I've been doing for the past few years: not particularly active in terms of development, just occasional improvements and fixes here and there." If you're trying to apply this to situations such as "py-cryptography" -- they are not very comparable.
Pretty much every license in existence has a line like the following: "THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED". This is especially true when the upstream never declared "support" for a given use case to begin with, such as different Libc implementations, big-endian platforms, etc...
Don't complain about free beer.
Had py-cryptography done the same thing (i.e. made the same announcement and promises, and followed through on them), I'd think we could say the exact same thing about py-cryptography as above too.
I can (and have) actually written quite a bit about this, but I’m not impressed by arguments where people quote the terms of a license to justify them being a jerk. I maintain several open source projects, and I can tell you that the license tells you nothing but what you must do to avoid being sued. It is a legal contract, not a guideline for how you should behave, unless you’re one of those people who goes around providing the bare minimum courtesy to others as is necessary to not get in legal trouble (we have a word for those people: jerks).
When I work on software other people use, I am a considerate person who understands that my code will often be used in ways I did not expect it to be. I think people that do is are really awesome! As a software maintainer, I believe it is my duty as a nice person to extent some amount of courtesy to them. Of course, there are limits to everything, but that limit is absolutely not at the point where my legal liability ends.
So, coming back to this topic: I have seen projects that take the stance you have in response to people reporting things like “I can’t compile this software for my platform anymore :(“. I think that is a legal, but jerk move. In this case the maintainer not only understood the concerns that people may have from the changes, but they also committed to providing basic support for it even though they obviously didn’t have to do anything. A good maintainer recognizes that a minimal bit of effort can pay off handsomely in goodwill, and I just wanted to call attention to that fact.
https://github.com/CyberShadow/btdu/blob/master/ARCHITECTURE...
https://github.com/bootandy/dust
> There is, after all, nothing more annoying than having to get re-acquainted with every piece of software you've installed every time you decide to update...
Looking at you $EVERY_MODERN_GUI_APP
> Ncdu 2.0 doesn't work well with non-UTF-8 locales anymore, but I don't expect this to be a problem nowadays.
I can believe this, but has there been any investigation into how common non-UTF-8 locales are nowadays? GNOME has been broken with non-UTF-8 locales for years (I can find a 6-year-old bug report that still has not been resolved) so that may have pushed users that previously were using non-UTF-8 locales to UTF-8, but that is only GNOME and does not apply to users of any other environment, who may have continued using non-UTF-8 locales without issues.
For like me who are curious:
- C version source: https://code.blicky.net/yorhel/ncdu/src/branch/master/
- Zig version source: https://code.blicky.net/yorhel/ncdu/src/branch/zig
The code is nice and very well commented.
> The code is nice and very well commented.
I opened the first file in src/: browser.c. Let's take the function browse_draw_mtime() and start picking nits :-)
It has a buffer of 26 chars, which will (for a regular C string) mean 25 characters + 1 NUL terminator.
char mbuf[26];
But in the end, it prints with the printf()-like function:printw("%26s", mbuf);
the count there is supposed to exclude the NUL terminator. So it should be 25 and not 26. Note that it cannot cause a problem, but it may indicate that the author didn't carefully grasp the exact definition of the format.
Before that, in a branch the mbuf buffer is filled with the strftime() function. This function is stupidly defined by the C standard and POSIX didn't make it better. It is not the author's fault, but one has to account for its flaws.
strftime(mbuf, sizeof(mbuf), "%Y-%m-%d %H:%M:%S %z", localtime(&t));
One would assume that the result string is truncated if it happens that the result would be too long, so that code would be fine. Well, not quite, it just protects from writing after the buffer. The standards say that in cases when the buffer is not long enough, the content of the result string is indeterminate :-/ Actually, it could even be not a string (not a properly terminated string).So one should check the value returned by strftime(), check if it is 0, and act accordingly.
Again, it is not dangerous, since the printw("%26s", mbuf) won't read after the buffer. But it may write garbage, for example after year 9999, when the expected result string is too long for the buffer.
-------------
Then in the last file, util.c, there are those macros:
#define oom_msg "\nOut of memory, press enter to try again or Ctrl-C to give up.\n"
and #define wrap_oom(f) \
which at some point does: write(2, oom_msg, sizeof(oom_msg));
This is going to emit a NUL character after the newline, because sizeof() of a string literal accounts for it. So, to stop after the newline is emitted, it should be sizeof(oom_msg)-1.Let’s…not? I don’t see why you would read “this code is nice and well commented” and immediately comb the code to try to find dubious “bugs” with it? It might have been mildly relevant if you responded with “no, I don’t think the code is actually that readable, for example look at this part” but to jump in when nobody made any assertions of correctness or safety and then bring up minor problems is just strange. If someone says “I think she is very pretty” do you respond with “let’s look at her parallel parking skills, shall we? Eh, I’ve seen better”?
So should I have said, "baaaah, this sucks!"? As it nevertheless far from being the worst code ever, I'd rather warn that what I am going to do amounts to nit-picking, and put a smiley to show that I don't want to disparage the work presented to me.
Well, you didn't really talk about those things the first time, you tried to find bugs instead…
To sum up, I'd say I quickly looked at 2 of the most common sources of mistakes:
- off-by-one errors: here, in the sub-species which concern how the NUL terminator is counted (or not counted), whether it is by the code itself, or by calling standard functions which behave differently in that respect (sometimes for good reasons, sometimes not, but in all cases one should be careful and check their spec twice);
- error management: what happens, what could we get as a result when an input is not in a typical range, how can we deal with it; standard functions are not especially regular in that respect, one should check that aspect of their spec too, for it can be unintuitive (as in the case of strftime() here).
Then I explained how the consequences were not serious in that case. In a well commented code, I'd expect a small comment to explain why this check is not done, why one can emit that without much care, and so on. Otherwise, one cannot know if it works by design or by chance; also if future modifications should happen with those pieces of code, problems could arise.
Emitting a NUL character to a (pseudo) terminal is not serious. In the worst case it could mess the display for 1 or 2 lines before it recovers, but generally it won't even have a visible effect, the terminal will simply ignore it.
However, if that gets written or redirected to a file (here, a log file for example), then the NUL character is quite present. Since a civilised text file oughtn't to contain a NUL character, a program which expects well-formed text files may not be ready to handle it properly. For example, reading such a file with the standard C function fgets() or similar will land you in a world of troubles. fgets() only stops when it encounters a newline character, so the NUL character will be included inside the string returned by fgets() (here, coming right after a newline, it would be at the start of the next line); but when you'll ask for the string length, it will be 0 (NUL being the first character), and printing the string will be the same as printing an empty string. Basically, you have lost the line content.
So it is better to be as correct as possible from the start. It is also often easier than thinking exhaustively of all possible implications.
> Exporting an in-memory tree to a file. Ncdu already has export/import options, but exporting requires a separate command invocation - it's not currently possible to write an export while you're in the directory browser. The new data model could support this feature, but I'm still unsure how to make it available in the UI.
1. breadth-first search? sometimes most of the directories are tiny and when one with a lot of files is found, I'd like it to be measured as the last one. I think that right now directory tree is explored using BFS, 2. add a way to display currently calculated directory tree state. Not sure about ncdu2, but ncdu doesn't show much apart from the current item, number of files and total size. I'd like to be able to use the approximate information it has gathered so far
And once again, thanks for the great work! I really appreciate it.
By this I mean: why are both Windows Explorer and all major Linux file managers incapable of showing total recursive directory size in bytes the same way OS X's Finder can? ncdu traverses home folders very quickly, why can't this be built in to a GUI file manager as an option, as Finder's Calculate All Sizes is?
I cannot be the only user for whom this is an absolutely indispensable feature --- one that actually keeps me from migrating entirely to Linux.
Btw, you forgot Rust :)
(Dust is "du + rust").