Curl doesn’t spew binary anymore
daniel.haxx.se
daniel.haxx.se
if(isatty && (outs->bytes < 2000) && terminal_binary_ok) {
if(memchr(buffer, 0, bytes)) {
memchr() returns the first occurrence of the byte 0 (your second argument), or NULL.So a few things:
* what if your output is more than 2000 bytes?
* what if your output is binary but doesn’t contain a byte 0?
* what if your output is a normal UTF-8 string but contains a byte 0? ( see https://stackoverflow.com/questions/6907297/can-utf-8-contai... )
This is interesting to me as I am developing a tool that parses files in search for bugs and I need to ignore binary files. What I am doing right now: when checking a line, if it is not a valid UTF-8 string I skip the file. It's not really nice as I am doing this verification for every file's lines...
2. Then the check yields a false negative, which is not a problem.
3. Then your UTF-8 string is unprintable, and the check will yield a true positive. The UTF-8/ASCII NUL character is not printable, despite being valid.
If one only assumes ASCII/UTF-8/Shift-JIS/similar, then a blob containing a null byte is guaranteed to be unprintable, while a blob not containing a null byte may be printable. That's good enough for a warning, telling you that you're doing something that you might not have intended.
Given that UTF-8 has become standard, it means that you will never realistically get a false warning, but may still get bonkers output. You can always overrule if you have a fetish for UTF-32.
You could just check the first X bytes. Also, I'm guessing curl doesn't print out to the terminal if the data is more than 2000 bytes anyway?
> 2. Then the check yields a false negative, which is not a problem.
the binary will be printed on the screen, that's a problem
> 3. Then your UTF-8 string is unprintable, and the check will yield a true positive.
How is it that the zero byte is part of UTF-8 then?
> You could just check the first X bytes. Also, I'm guessing curl doesn't print out to the terminal if the data is more than 2000 bytes anyway?
Why wouldn't it? 2000 bytes is just 25 lines by 80 characters.
Of course it does. You can curl the concatenated content of the library of congress to your terminal if you want to.
> the binary will be printed on the screen, that's a problem
No. Because printing the binary to screen is the current behaviour in all cases, the goal of this change is to reduce the incidence of it for quality of life.
> How is it that the zero byte is part of UTF-8 then?
Flash News: unprintable characters are part of unicode. NUL is one of them.
It is a valid code point, just not a printable character. Unicode encodes every character that is or was in common use, not just the printable characters; this includes the control characters at the beginning of the ASCII table.
Most programs detect binary files like this. Here's git's version of the same function: https://github.com/git/git/blob/master/xdiff-interface.c#L19...
The announcement says that "curl will inspect the beginning of each download", and I think that comparison just turns off the check after at least 2000 bytes have already been output (see a few lines below the change you quoted, where outs->bytes is incremented by the amount of bytes that were output).
what if your output is binary but doesn’t contain a byte 0?
I guess curl will incorrectly recognize the binary as text.
what if your output is a normal UTF-8 string but contains a byte 0?
I guess curl will incorrectly recognize the text as binary, and you can use `-o -` to override that and output to the terminal anyway.
This doesn't sound like a really good test.
bool isatty = config->global->isatty;
[...]
if(!config)
return failure;...What if the first byte is 0? That would cause the if condition to fail, and the output would be treated as plain text.
From what I understand, this means that no wild UTF-8 string would have a NUL byte anywhere in it, no?
I ran into this once with an Objective-C program that created filenames from strings found in files. When presented with a string containing NUL, the code ran, but didn't really work. I'd get "foo" and then ask it to append ".txt" and the result would come back as "foo" still! And it depends on the context in which you use it. Using it as a filename truncated, because you ultimately go through the POSIX-level calls that use C strings. Printing it in the debugger truncated, as evidently it used C strings at some point in that process. Displaying it in a text field in the UI worked perfectly fine, though, as apparently that path never uses C strings and the NUL character is just an invisible zero-space character in the middle.
It's not really a valid character to print to a terminal and most terminals ignore it nor is it particular valid in a "text file".
As for the rest of your questions if the bytes in the file are randomly distributed which is common with compressed binary files the chance that there is not a zero byte in the first 2000 bytes is 0.03% which seems low enough.
This is rarely true I believe, but binary files should often use the 0 byte because it sounds "practical" to use (going with the instinct here). So I'm guessing this test is "good enough"
It will make curl fail if a 0 byte is outputted within the text. There are many reasons why this will happen. Software errors, UTF-Encoding, special use cases...
In fact, I run into this problem with grep from time to time and it is super annoying:
> grep somefile.txt stuff
Binary file somefile.txt matches
So this change makes the code base more complex and the behavior fuzzy.I prefer elegant software that behaves predictably.
I assume that you also inserted a NUL byte into your somefile.txt to make grep treat it as binary.
cURL behaves predictably now as well as before. It will warn you if you try to print unprintable data to your terminal, but can be told to ignore it, and won't affect output redirects.
I often end up doing things like: watch --color juju status --color or "ls -c | less -r", etc.
Which could mess up your terminal, so curl is doing the right thing.
> UTF-Encoding
On Unix, terminals would be completely broken if they used UTF-16. Other Unicode Transmission Formats don't use null bytes.
> special use cases
Such as?
> In fact, I run into this problem with grep from time to time and it is super annoying:
That's not applicable here, because grep only prints parts of the file, while curl prints out the whole file.
Changing behavior of something that is as popular as curl to protect edge cases is strange, to say the least.
Which was the right change. It prevents data loss for negligible added complexity.
Edit: It would appear that the answer is `isatty()`
cURL already use isatty checks elsewhere, such as to determine if progress info should be printed (will only show if you're not outputting to tty).
So that's not so much tricking isatty as much as using the system in the way it is intended. If you do that without the required background information then you have only yourself to blame.
Its purpose is not just to make isatty return true. If you spin up that giant subsystem to trick an application into a different behavior on a simple pipe, then I believe it is appropriate to call it a hack, trick, workaround or similar.
You can also LD_PRELOAD an isatty shim that returns true. This is occasionally useful when piping the output of tools which disable ANSI colorization with no switch to re-enable that.
Some programs will check if stdout is a tty prior to writing to stderr, for example, which means that if you're doing,
foo | bar
foo's stderr will go to the tty, but would lose coloring, for example.
For example, `babel`, a JS transpiler, loses the color in its error messages when you pipe its output.And if you're the kind of person that needs this, the solution is a one-line alias.
I guess what really bad things can happen (and no I'm not talking about blindly piping to a shell to eval)?
Sure it messes up the terminal of which I usually type clear and/or reset. If it's really bad I usually just kill the terminal.
echo -e "\ec\n\e]55;/tmp/ouch.php\a"
I believe they've all removed that functionality by now, but there might be some older copies around here and there.
Why would anyone do this? If I see some unknown, interesting file, I might run cat, head, tail, less or vim on it. If it's binary then maybe I'll use xxd. But it wouldn't even occur for me to pipe it to a shell.
You could argue this is one of the reasons less doesn't interpret control codes by default. As it would let applications hide stuff like that to redraw the screen.
In general, it's a bad idea to have a hidden state such as a configuration file. Imagine, for examples, scripts that expect the standard curl behaviour, but your configuration file changes that.
(Of course setting shell aliases has the same problem...)
“please don’t make the behavior of a command-line program depend on the type of output device it gets as standard output or standard input.
[…]
There is an exception for programs whose output in certain cases is binary data. Sending such output to a terminal is useless and can cause trouble. If such a program normally sends its output to stdout, it should detect, in these cases, when the output is a terminal and give an error message instead. The -f option should override this exception, thus permitting the output to go to the terminal. ”¹
1. https://www.gnu.org/prep/standards/standards.html#User-Inter...
$ curl news.ycombinator.com > /dev/null
% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
100 178 0 178 0 0 369 0 --:--:-- --:--:-- --:--:-- 370
You have to add an extra arg to get rid of that progress meter. $ curl news.ycombinator.com -o /dev/null
% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
100 178 0 178 0 0 948 0 --:--:-- --:--:-- --:--:-- 951
I mean, it's effectively doing tty detection, but I think it's fair to argue a semantic difference, and it's very similar to how ls strips colors and disables listing format when piped, which is mentioned by the next paragraph in the coding standards.Compatibility requires certain programs to depend on the type of output device. It would be disastrous if ls or sh did not do so in the way all users expect. In some of these cases, we supplement the program with a preferred alternate version that does not depend on the output device type. For example, we provide a dir program much like ls except that its default output format is always multi-column format.