Unexpected places you can and can’t use null bytes
eklitzke.org
eklitzke.org
If you actually NEED strings with nulls in them (although I couldn't think of a reason why), you'll need to use/find/create APIs with length fields.
Well, it might. If it ever uses that string internally with some function that expects a null termination then it'll probably still get truncated.
One I've seen in the wild is an encoded data structure, a blob of null terminated strings passed around together.
Conceptually it was dealt with as one string, which happened to contain nulls.
I remembered seeing nul in strings in git, not sure if the same place (maybe?) but a quick search shows up some real world usages https://github.com/git/git/search?q=nul
PHP supports NUL, and the result is a HUGE amount of extra support code to handle it, spread all through the PHP codebase (introducing TONS of bugs over the years), all for the 0.00001% use case. Not a good trade.
Complaining about lack of NUL support would be like complaining that most text editors don't display the DC1 control character. It's just not useful in 99.999% of use cases, so there's no point in supporting it. You'd use a specialized tool instead.
So rather than, say, simply using the binary array data type already present in C, they decided that they needed to subvert strings to support passing non-string data, resulting in tens of thousands of lines of support code to parallel the C runtime library, thousands more lines of bridge code for any outside libraries they might want to make use of, and almost a million lines of PHP implementation code (not including extensions!) carrying this extra cognitive burden in order to support the 1% use case, resulting in countless bugs and security vulnerabilities whenever someone used the C library functions, or stopped on NUL (or not) by mistake.
Sounds like a great plan.
But only one of these is possible. Why? because there's an api for writing contents to a handle that takes a length, but not one for opening a file that does.
And you can't “create” api's for what the kernel doesn't support. These are limits on what is possible without forking Linux or some other kernel.
[edit, to clarify the question]
Like for example, the ASCII code '\0' is still a valid "thing" (it's a byte sequence, kind of). But in other contexts (other languages maybe), NULL is not a "thing" per se; it's more of a non-thing. How does a C programmer see the difference between NUL and NULL?
More specifically, "NUL byte" is the character that has value 0, whereas "null pointer" is the pointer that has value 0.
OK, I'll try to stop being the pedant. I'll let the real pedants poke holes in what I said...
Where is this enforced? You can assign between notional types without a problem. Whatever the context, providing 0 will have the same effect no matter how you labeled it.
char c = '\0';
void* p = (void*)0;
int a = (int)c;
int b = (int)p;
then a == b
resolves to true. In that sense, they are the same - which I believe is your point. In that point, you are completely correct.The difference is that I can't do
*c
In that sense, they are not the same - not in the sense of numberic value, but in the sense of type. Also, c is 8 bits, and p (at least these days) is 32 or 64. (I pity anyone who ever had to work in an environment where p was 8 bits!) So they are the same numerically, but they are different both in type and in memory footprint. $ cat zero.c
#include "stdio.h"
int main(int argc, char* argv[]) {
char nul = 0;
void* null = 0;
if( nul == null ) {
printf("compared char to pointer; they are the same\n");
} else {
printf("found a difference between char and pointer\n");
}
return 0;
}
$ gcc -o zero zero.c
zero.c: In function ‘main’:
zero.c:7:10: warning: comparison between pointer and integer
if( nul == null ) {
^~
$ ./zero
compared char to pointer; they are the same
$
You get a warning, but not an error, for making the comparison. By contrast, assigning the integer zero to a void* isn't even a warning -- it's just a natural thing to do. There isn't another way to set a pointer to NULL. There is another way to set a character to 0, the '\0' syntax, but that's not a warning either.C will think nothing of adding '!' to 'P' and getting 'q'. That's not strange because addition is a pretty normal thing to do with integers. You're right that a char variable should only occupy 8 bits of memory, but that's an implementation artifact, not a theory of what the value '\0' means. That value is unambiguously the integer zero with infinite precision. The reason it only occupies 8 bits is that you can't let it have infinite bits.
For that matter, what does (uncasted) p = c do? How about c = p?
[Edit: You updated while I was writing; my first question you already answered. Warning, not error.]
$ cat pointers.c
#include "stdio.h"
int main(int argc, char* argv[]) {
char Z = 'Z';
char q = 'q';
void* null = 0;
printf("Z is \\x%02x\n", Z);
printf("But if it were a pointer, it would be %08x\n", Z);
printf("Watch this:\n\n");
null = Z;
/* %p to print a pointer value */
printf("Our void* is now: %p\n", null);
q = null;
printf("And q is: %c\n", q);
return 0;
}
$ gcc -o pointers pointers.c
pointers.c: In function ‘main’:
pointers.c:12:7: warning: assignment makes pointer from integer without a cast [-Wint-conversion]
null = Z;
^
pointers.c:17:4: warning: assignment makes integer from pointer without a cast [-Wint-conversion]
q = null;
^
$ ./pointers
Z is \x5a
But if it were a pointer, it would be 0000005a
Watch this:
Our void* is now: 0x5a
And q is: Z
$
The assignments are warnings. They work just like you'd expect them to work.Notice all the different printf flags? This is why you need them.
char *p = 0;
is fine, but intptr_t a = 0;
char *p = (char *)a;
might not do what you expect (set p to the NULL pointer).It's true that "NUL" is the usual abbreviation for this value in character code charts/standards.
https://en.wikipedia.org/wiki/%00
But what the hell does it do??? In Safari and Firefox, I get an nginx 400 Bad Request page from en.wikipedia.org. But in Chrome, it seems to be redirecting to a google search for the same url, when I type it into the address bar. Well, that's meta. Chrome won't even let me drag-and-drop that icky %00 terminated url from one page into another page to navigate there -- it angrily rejects it and sadly animates the evil url back to where it came from (though dragging it into an existing or new tab mysteriously works). But actually clicking on that link immediately goes to a blank purgatory page with the url "about:blank#blocked". Those are Chrome's stories, and it's sticking with them.
At least this works:
https://en.wikipedia.org/wiki/%01
>Special Page
>Bad title
>The requested page title contains an invalid UTF-8 sequence.
>Return to Main Page.
Oh, yeah -- PHP:
https://webmasters.stackexchange.com/questions/84008/url-enc...
If not the actual NUL character, then at least the Unicode Symbol for NUL redirects to the page on the Null character.
https://en.wikipedia.org/wiki/␀
>Null character
>From Wikipedia, the free encyclopedia
>(Redirected from ␀)
>For other uses, see Null symbol.
...But then again, shouldn't the Null symbol ␀ redirect to the page on the Null symbol, which it actually is, not the page on the Null character, which it only symbolizes?
Firefox makes a GET request to "https://en.wikipedia.org/wiki/%00", the server returns an HTTP/2 400 Bad Request. Presumably because the web-server considered the URL invalid.
Chrome decides the string isn't a valid URL up front. So it does what it normally does when you enter random junk in the address bar; it searches for it.
The dirty secret of URLs is that no one can quite agree on which ones are valid or how they should be canonicalized.
We can take WHATWG's spec as a modern way to handle URLs [1]. If I'm reading it right (50/50 chance!) the URL would be considered valid by that spec.
See also this article from the developer of curl: https://daniel.haxx.se/blog/2016/05/11/my-url-isnt-your-url/
Fun story. An engineer was working to migrate an old system from Python 2 to 3 before the Python 2 EOL deadline. The engineer decided to use the str type to represent URLs. Chaos ensued when suddenly non-UTF-8 URLs don't work any more. Turns out back when that system was designed, people were directly URL-encoding binary data into URLs.
Perhaps it should. That page actually did not exist when the redirect was created:
https://en.wikipedia.org/w/index.php?title=␀&action=history
https://en.wikipedia.org/w/index.php?title=Null_symbol&actio...
See the table here https://en.wikipedia.org/wiki/ASCII#Control_characters
However, the standard does talk about null characters throughout as being a character code of value zero, as opposed to NULL the macro.
gets and scanf("%s") are also horrifically unsafe. gets is well-known to be unsafe (to the point where you'll almost certainly get a compiler warning for using it). However, scanf("%s") is unsafe for exactly the same reason (no bound on the buffer length) yet will not produce a compiler warning. Add to the fact that these functions will accept null bytes, and you have a very dangerous buffer overflow waiting to happen.
if (*s && *s != '\n' ...)
and never: if (*s != '\n' ...)I've seen this byte people when junk gets written to a filename (either accidentally or maliciously). Especially in shells but also in other programming languages. Issues that aren't always handled well include file names that:
* include a newline or some other control characters
* start with a `-`
* aren't valid UTF-8
The other trick that makes dangerous operations always more safe is prepending "./" to the path.
It was the only directory starting that way, so I typed "rm -rf ~<TAB><ENTER>"
Hit control-C a half second later, but the damage was done. Fortunately most of my important files are backed up.
Lesson: when deleting files with tricky names, write the command without flags first, then add "-rf" after the path is confirmed correct.
Windows has many more restrictions. [1]
[1] https://kizu514.com/blog/forbidden-file-names-on-windows-10/
But yes, Windows reserves (16 bit encoded) ASCII control characters and some special symbols like question mark and colon.
As for . and .. I believe that on UNIX they are just symlinks, though they may be treated specially by applications? On Windows they don't really exist but some APIs emulate them when resolving paths.
No, they are treated specially by the kernel and/or the filesystem. They behave similar to hardlinks to the corresponding directory (hardlinks to directories aren't allowed nowadays, these two are an exception). The special treatment by applications is that many applications hide all files and directories starting with a dot, which happens to also apply to these two.
Quite a few file systems require names to be valid text in a specific encoding, not arbitrary byte sequences.
Also, note the use of a plural in “the file systems in the path”. A file systems mounted in another one can change the rules halfway-through. There is not single fixed set of restrictions to file names.
Macs let you use "/" in file names, but instead used ":" as the path separation character.
It seemed to work fine and dandy at the time, until you ran "restore" and discovered your backups were corrupted!
https://en.wikipedia.org/wiki/GatorBox
https://news.ycombinator.com/item?id=20007875
Just for fun: guess what happens when you create a Mac file name with a slash in it today?
It works!
Or it seems to work. But behind the scenes, the Finder and Mac user interface libraries actually convert the "/" to a ":", which you can see with "ls". But at least it doesn't corrupt your backups!
Is there an (perhaps obvious, or not) common usage of NUL byte literals being passed around, not for the purpose of terminating strings? Just terrible ye-olde file formats?
echo -n $'\0' > nul
This doesn't work for the reason stated in the article. The argument is instead interpreted as a string ended at the NUL byte and the file will be empty. BTW you can get around this with printf since it processes escape sequences internally. printf '\0' > nulprintf("%.*s", (int) nbytes, str);