*(char*)0 = 0; – What does the C++ programmer intend with this code? [video]
youtube.com
youtube.com
I would probably use something like this when hacking on an Apple II with a zero page.
I have also written programs that ran on VMS which dumped memory starting at '0' without getting any segfaults.
Yes, I said if you had registers--if you were cheap you could buy your PDP-10 without registers and all the instructions that referenced registers would use core instead.
Since the registers were in effect just 16 words of fast semiconductor memory overlaying the first 16 words of slow core memory you could do anything in them that you could do in regular memory, including running code out of them.
If you had a small loop that did a lot of iterations you could sometimes get a significant speed boost by copying the loop into the registers and running it there.
Turbo-C used to check the zero location to see if it had changed and would issue a warning if so (in the real mode days).
[0] My other computer is a PDP-11,
It's an old language. Some of the fundamental mistakes were just things that were common at the time and only look so bad in hindsight (e.g. strings, locales, nullable pointers, half the k&r standard library). That doesn't mean we can't and shouldn't do better where we can though.
C and C++ targeting WASM, despite all the security message of how great it is, memory locations inside of the same linear memory segment can still be corrupted, thus providing a way to influence the overall execution logic inside of the sandbox.
At time C was written, swapping in/out of memory measured in literal minutes & when fraction of second of execution time, including swapping in/out of memory, was more than several times the average yearly salary of the day (excluding sneaker net intervention)
Consider how a surgeon would respond if told not to use a scalpel because of the risk of accidental injury when using a sharp tool.
We learn from our mistakes - to which I'd add, sometimes we can afford to make mistakes (home programs) and other times we can't (safety-critical code).
You see the same arguments against static analysis, unit tests, strong type systems, etc. The evidence seems to favor systems which prevent errors over artisinal expertise.
This casts flexibility vs safety as a tradeoff, which it is.
Surgeons are constantly killing people[1], and every time, it turns out it was because the surgeon disabled a known, recommended safety rail, and whenever anyone points out that they should stop disabling the safety rails, they insist that they know how to operate without them, it's those other people that don't. Plus it's sooooo inconvenient for an operation[2] to take five more minutes, they're too good to have to deal with that.
[1] introducing security vulnerabilities
[2] code changeset submission
Honestly, checklists and procedure matter more than most people want to admit.
1] introducing security vulnerabilities
[2] code changeset submission
This is my favorite one
I would agree the same for developers, if similar practices would be enforced everywhere instead of having people calling themselves enginners just because they like how the word sounds.
More quality and process validation, less cowboy programming.
A brick layer doesn't turn into a Construction Engineer, only because they think they know everything about building houses, and have somehow built their own during long weekends.
Likewise a coder out of a bootcamp, isn't someone versed in what actually entails to be a Software Engineer, regarless how cool the title might sound like.
Originaly using anonymous unions & placing largest bit count variable as first item in union was the context needed to align the "char" correctly -- vs. casting.
Technically, can point at anything in an unaligned manner, just need to align access to machines address boundary to use a "fetch the value of a complete valid memory alignment address". One can hand roll the appropriate shifts/masking to get at the set of bits loaded (which usually let compiler do). Awk and settable end of line marker good way to safely visualize this.
Today you can still work with (say) TI DSP chips that spit complex FFT pipelines results once per cycle and have absolutely no 8-bit hardware addressing or masking abilities as they're lean mean RISC optimised machines for pure 32 or 64 bit floating point operations.
They have C compilers and CHAR_BIT==32 (or 64).
* C preprocessor - it's a text substition operation that can be non standard
* C "the language" - just the syntax folks, no library functions here.
* C "the standard library" - for many of us the K&R stdlib was just a proof of concept example of how to code a library, feel free to ditch it and write up your own string handling, for example.
Which brings us to, say, a multi channel marine seismic processing system that has an IBM PC type design with a custom motherboard that has six TI DSP boards slotted in and you're writing code for the user interface to a real time signal aquisition and processing system and writing onboard code for the numerical processing on the DSP boards.
Now, each board handles a streamer cable, each cable has a number of microphones, an external boomer in the water is triggered and the reflective soundwaves from the ocean floor, and the soundwaves that penetrate the seafloor and later also partially reflected by density change layers, are all captured by the microphones.
You have keyboard+mouse I/O between the real time window manager user interface, shared memory I/O between the PC memory and memory on the DSP cards, analogue soundwaves going to sample ports on the DSP cards, block memory I/O going from the PC to a SEG-9 tape recorder, prepared lines of memory going to a plotter to build an image ...
There's a lot going on.
But, as far as the code compiled for the DSP boards, that mostly handles I/O by linking reaction code to interrupt triggers - when the sampling hardware interrupts to signal another bit of soundwave from microphone[i] is ready, that's stashed in a FiFo queue to be pushed into the DSP handling pipe and to be saved in a raw sample buffer.
When a raw sample buffer is full an interrupt is triggered to take the entire buffer via DMA transfer to be handled PC side by sending it to the SEG-9 tape "of (raw) record", when a processed sample buffer is full an interrupt is triggered for a different type of DMA transfer that takes processed data to display, printer, and to a different SEG-9 track for processed data.
Not much of this is the standard C library, so you can see why you might write your own low level handling.
This is significant because the following two snippets of code are not required to be equal in C.
char* c = (char*)(0);
char* d;
memset(&d, 0, sizeof(d));
c == d; // this does not have to be true
That's what the spec says, but are there any computers around anymore where nullptr is not zero? The spec should really go with the times IMHO, old hardware can be supported via platform-specific language extensions.
Yeah that's absolutely fair. I actually can't think of a single platform, embedded or otherwise, old or new, where the null pointer is not address 0.
> Depending on the ``memory model'' in use, 8086-family processors (PC compatibles) may use 16-bit data pointers and 32-bit function pointers, or vice versa.
And it isn't just zero anyway because there's segmentation in play.
[]
*()++
Technically, can only do address "1" with first bit set as a valid, usable address on hardware allowing addressing smaller than 8 bits (standard hardware epsilon factor).So really, address one is literally address 0 + epsilon, where epsilon is minimal addressable bit group, typically 8 bit clean without seg fault.
Although, for standard x86, epsilon size would depend on which ring/boot method level & asm addressing method available (8,16,32,64).
That's also totally valid on WebAssembly btw, address 0 is a regular read/write location.
#include <stdio.h>
int main() {
for (int x = 0;;x++) {
putc(*(char*)x);
}
return 0;
}
I was mezmerized by the strings from the uncompressed BIOS ROM being dumped back at me, "American Megatrends".. etc. Eventually this process crashes of course, when it runs past the end of mapped memory locations.Then came the realization you could alter any memory location, and further, you could write a tiny TSR to do things...
What were you trying to do and how did this work? I am interested in understanding the interplay between -funwind-tables, -fasyncronous-unwind-tables and -fexceptions.
I had an inline asm statement that could throw exceptions (by indirectly calling exception throwing functions). GCC assumes that asm cannot throw, and the function the statement was in did not have any throw statement nor call any other potentially throwing functions nor any non-optimized out memory access, so the compiler omitted generating unwind tables for the function even with -fasync-unwind-tables.
The result was that throwing from the asm would at best not call destructors, at worse, crash with a corrupted stack.
By adding a non-optimizable dummy memory access right after the asm (and jumping over it from the asm, so it actually doesn't get executed), I forced GCC to generate the unwind info . I also had to make sure that no compensation code was needed for the asm statements themselves, but after that things worked out fine.
I wouldn't really use this in a finished product, it is very fragile and just happened to work on the version of GCC I was using. The right solution would be for GCC to add an attribute to mark asm statements as potentially throwing.
I had long wanted to understand the interplay between -funwind-tables, -fasynchronous-unwind-tables and -fexceptions since the last is language specific but the first two are not. GCC docs are not of much help in understanding what exactly is going on in each case so i guess i need to do some research and experimentation.
Where along the determinate in the matrix?
do you mean the browsing plugin, or using base GPT-4?
The point is the journey and the discussion around what the line of code does, and how that discussion is valuable. The point isn't to tell us what the code itself does.
Anyhow it's just barely possible that this assignment was a hack to make this sort of code work on a machine where address 0 is (uninitialized) nonzero and can be harmlessly made zero.
Certain higher ups determined that behavior was a business requirement when I refused to reimplement it on a new platform.
https://docstore.mik.ua/manuals/hp-ux/en/B2355-60130/chatr_p...
Wonderful fun when they shipped a Kerberos library that unconditionally dereferenced a potentially null optional pointer-to-struct, relying on the executable to be set in the “null tolerant” mode!
You can relocate the vector table and to block it off with the MPU if you have one and memory to spare.
This is writing sizeof(char) (== 1 almost everywhere) zero to address zero. It is not using a NULL macro or other predefined symbol.
In the real world, this would generally write a byte to address 0000:0000, leading to UB because it would fuck up the divide-by-zero IV.
PS: I used Borland C++ 3.1, Microsoft C++ 3.x and 4.5x, Watcom, and early GNU.
https://c-faq.com/null/null2.html
https://c-faq.com/null/machexamp.html
Actual ways to do what you want to do are described in
https://c-faq.com/null/accessloc0.html
but technically speaking the pointer with a constant zero assigned to it _is_ a null pointer (which can be implemented as whatever bit pattern), independent of the preprocessor macro.
sizeof char is 1 by definition everywhere.
/pedantic
Parentheses are required around char because it's a type.
/pedantic
sizeof is an operator in C, and does not need parenthesis any more than pointer operator *. It is true that programmers frequently think of it as a function and use parenthesis.
To begin with, sizeof has two syntaxes: the first, which is the one you seem to refer to, is simply
sizeof expression
where expression involves variables and constants, not types. The second is sizeof (type)
where the parentheses are mandatory.Then, even in the first syntax, even if sizeof is listed among the operators, even if it doesn't look any different from "pointer operator ", nonetheless it has strange priority rules. For example
sizeof (T) *x
If it was a regular prefix operator obeying priority and right-to-left evaluation, this would mean: dereference x, cast it to T, and return its size. Instead the C standard forces the compiler to interpret it as: take the size of type T and multiply it by x.-----
unless initial property is start of dynamic operation, in which case, holding almost anywhere begins at the first operation after the start of the dynamic operation. process / lambda / epsilon calculi is just symbolic math. address 0 static, everything else dynamic.
per math, dimension N is static, to be able to "change things up" in dimension N, need to to be almost everywhere higher than dimension n. Edge cases are weird in any dimension. Guess why logicians just do the equivalent of C's !0
(cast classic logic) A=1 (cast boolean logic) B=0
C statement !(!B == A) hold everywhere and almost everywhere depends on how read C spec to interpret A & B.
The literal 0 is treated specially, so this could indeed be one of those 'turns into a weird bit pattern NULL pointers', if such a thing existed in the wild anymore.
But you're correct in that there probably haven't been any since the turn of the century or whenever the last Univac mainframes got turned off.
execl takes a variable-length, null-pointer-terminated list of character pointer arguments, and is correctly called like this:
execl("/bin/sh", "sh", "-c", "date", (char *)0);
Due to ececl being a variadic function it can not take advantage of a prototype to instruct the compiler that one of its arguments needs to be treated as a pointer context.Here in godbolt, clang compiling C simply deletes the code in the function past and including the null pointer dereference.
https://godbolt.org/z/9aqWPazsP
> This is writing sizeof(char) (== 1 almost everywhere)
1 everywhere. sizeof's unit is "how many chars". For instance there was a cray machine that could only access 64bit words. sizeof(char) is still 1, with 64bit chars.
> zero to address zero. It is not using a NULL macro or other predefined symbol.
NULL is defined as literal 0.
My first kernel was 1.0.9 released alongside Slackware 2.0, offering initial support for IDE CD-ROM drives and experimental support for ELF files, by the way.
Modern CPUs with virtual memory means the question is a lot more complicated. Every process in a modern OS gets it's own address space so you can write to 0 but it could go anywhere (even virtualized to disk) and all the actual hardware is not directly accessible (must go through the OS).
I'm not sure I'd call this ability "useful" except if you're writing an operating system. This is vast simplification but when your computer boots it's effectively in a mode that allows reading/writing to anywhere. The OS kernel has direct access to all the hardware and then it limits access when running user processes.
The address can be changed with the LIDT instruction and operating systems nowadays will just put it wherever, but for backward compatibility it is expected to still be at 0000:0000 (not sure how this is handled nowadays in UEFI, but it should still be possible t o set it up that way).
And yes, some addresses are special. (AFAIK, on all current mainstream architectures.) This is the expected way to set those signal handlers, output (and input) data, configure devices, etc.
That said, there are some gotchas on using specific addresses in C. AFAIK none apply to x86, but it's something you usually do in assembly.
As for why it's address 0, well, it has to go somewhere, every machine has a CPU so everyone needs an interrupt table even if they don't have much memory. And when memory was precious there was no sense wasting even one byte of it; 0 was a real address on your physical memory chip, so why not use it just like any other?
(The fact that it's "address 0" for "division by 0" is just coincidence as far as I can see; division by 0 just happens to be the first kind of possible CPU interrupt. Perhaps it was the most common one?)
From the numerous responses here, it's clear that people interpret my question as about how the hardware itself works, which isn't at all what I was asking about; I'm aware of how stuff like this works at the assembly level, but my understanding was that in C and C++, trying to write arbitrarily to "special" addresses like that would be considered undefined behavior (often resulting in segfaults). When I read the comment I responded to above, it surprised me, so I wanted to check whether I understood what was said correctly. It's honestly kind of confusing to me that so many people seem very upset by the idea that a stranger on the internet might have a misconception about how hardware abstractions are exposed via compiled code to the point that they feel the need to explain in detail how hardware works but not actually answer the question I asked.
They're not saying this is, like, a portable standard way to handle division by zero in C++. You're right that it would be undefined behaviour under the standard (but a C++ compiler for real-mode x86 would be expected to support it, at least implicitly; obviously this specific case is not a particularly useful, but C++ is used in embedded settings and setting a custom interrupt handler is something its users want and expect).
A decent, well-behaved language would do some kind of structured error handling on divide by zero, like throwing an exception. IMO that includes any C++ compiler worth bothering with (though again the standard makes it undefined behaviour so it's possible that some compilers don't). But, the way the runtime of such a decent C++ compiler would actually implement that would be by setting up an interrupt handler for the divide by zero interrupt (that would contain code to construct the exception etc.), and by performing this write to address 0 you're overwriting (the pointer to) that interrupt handler. So, this line of code would cause your program to behave (almost certainly) badly on the next division by zero, even if you were using a well-behaved C++ compiler that normally handled division by zero gracefully.
(OTOH with a maliciously pedantic C++ compiler that division by zero would already be undefined behaviour, so in practice, since most C++ compilers tend to be maliciously pedantic, you might be no worse off than you were before that line).
The original post you replied to was just talking about the somewhat interesting details of what would actually happen because of the quirks of what these addresses are used for on that hardware (e.g. the fact that address 0 is supposed to contain a pointer to the handler, so by setting it to 0 you cause the CPU to start executing the interrupt handler table as code, is kind of interesting - not as a point about C++, but as a point about funny emergent behaviour of hardware), not about what this is specified as doing or the normal way of doing things in C++. I don't know why you were downvoted.
What got missed though, is ther has to be an "unused"/"reserve" bit(s) space in order for things to run without requiring additional specific hardware operations.
DEC provided the necessary hardware MMU to do actual real time multi-processing/multi-user access in feasibile/practical manner.
But yes, the interrupt table was my first thought when reading the headline.
A byte is CHAR_BIT bits, where CHAR_BIT >= 8. (It's exactly 8 on most implementations; DSPs are the most common exception).
short and int are both required to be at least 16 bits wide. It's possible for int to be 1 byte (sizeof (int) == 1), but only if CHAR_BIT >= 16.
If I'm being pedantic, I might add something like
#if CHAR_BIT != 8
#error "This code assumes 8-bit char"
#endif
But realistically, if I'm using headers defined by either POSIX or Windows, that's probably enough of a guarantee. (Though I'd still use CHAR_BIT rather than 8 to refer to the number of bits in a byte.)There are hundreds of instances of (char)0=0; in github, and there are none of the bit_cast variant, so if your goal is to inform people how to read code, starting with the one humans might actually encounter in the wild makes sense.
Not true. It's illegal to execute that statement because of undefined behavior. But it's legal to have: if (false) { *(char*)0 = 0; }
(And of course it's legal to #ifdef it out or comment it out, but that's too much cheating.)
Reference if you don’t get the joke: https://www.reddit.com/r/shittyprogramming/comments/3bmszo/t...
Until accessed, a memory location & variable location can be expressed as an indefinate integral.
Memory location access request implies definate integral.
1) evaluation value indefinate of integrand variables: "A" less than/equal to "B" less than/equal to "C" less than/equal to "D". Variable a is start of sequence. Variable D is end of sequence.
2) integrand variables D & A define the continuous range of PC group of bits
3) integrand variables C & B define the cohntinous range of 8 bit char
4) Accessing bits outside the integrand limits implies math equivalent of computer segfault.
In order to not generate a segfault, only positive delta, and epsilon is 0 or 1. aka The "alignment request" implies that A MUST be equal to B; C less than or equal to D.
K&R C short hand for above is anonymous union.[0]
Physics / Mechanical Engineering center of gravity approach rotated 90 & implimented via punch card reader that doesn't skip cards much more interesting physical list processing visual than calculus number line (IMHO). Physical form bit to bulky to carry to/from class room. Fortunately, electrial engineering split the difference via 1/2 way point (45 degree angle). Computer scientests get the autonoma / semantic analysis difference. Software Engineers, by completing the circle with & with out a sigh of "e e e e e e ..."
Wait for C++ coder to get disney sponsorship for a stack frame version of/translation to George Lucas expressions / Steven Speilburg expressions.
aka * wars vs. * trek crossover comparisons, perhaps with less () than lisp/scheme take.
Ideally, in UTF format so no "7 samuri/8 bit kleen", AND/OR, lower API, protocol droids distractions.
In interm, learn python, show a python version of beetles 'let it be' and 'eigen a feeling'[1] parody.
Although, perhaps 'irq feeling' more on topic with an appropriate reference to irq hooks & NULL dumps.
Perhaps get DMCA'd by Wierd Al for using wrong character color scheme.
".plan 0[]", long[sigh] different discussion?
[0] : sigh[long] : http://news.ycombinator.com/item?id=3367392
"Eigen Maze by this pointer" may be theme song for this post; but no letting things be in order to let this stack down to get this pointer back to complete 'init & kill -9' autonoma sequence.
What's passed in there? Some googling let me know it's used on macOS, but every result was too generic to be helpful.
(obligatory in C and OS theory, other languages might differ)
Note: think there's a way to convert this to some shorter, massively recursive C pointer, to pointer .... to pointer declaration explaination.
Ideally after the appropriate cast of characters has had a no-fault OS page(s) performance.
That, of course, would be part of the talk: that this is perhaps an OS-specific feature that the original coder was maybe trying to trigger, and that it is still OS-dependent.
[1]: https://en.wikipedia.org/wiki/Zero_page states this to be the case, but mmap(2) seems to disagree / suggests 0 is mappable.
My objective was to test the crash reporter tool.
Pointers and type safety still causing issues decades later.
I know on some microcontrollers (e.g. Arm) addresses 0x0 and 0x4 are usually used to define the initial stack pointer and entry point and then are never needed again.
Later, you probably want to detect the presence of null pointers by checking if &maybeAStruct is either 0 or valid. If you accidentally test the value at that address, you'll get a non-zero and then run code you didn't intend to run. By defining address 0x0 as value 0, you'll avoid this issue. That has two problems.
Alternatively, you define address 0x0 to erase the initial stack pointer so that a nefarious actor can't somehow memory dump and then find out where the program starts its execution. I don't see the benefit for several reasons.
-----
Scenario 1.
The first problem is that you shouldn't write code poorly. If you meant to check a variable's address and instead check its value, that's on you, your test suite, your compiler, and your debugger (and by extension, your colleagues and educators).
Second, writing to address 0x0 on a Cortex usually involves writing to flash... which is not a temporary/nonpersistent change, more involved than simply writing to an address (you usually need to enable the flash controller and set some peripheral registers), and would usually not be allowed on a production device because that's where the bootloader is or that section of memory has been write protected following good practices.
Scenario 2.
The more I think about this one, the less it makes sense. Anyone with debug port access or a memory dump is already inside the castle, well past any defenses e.g. zeroing the main stack pointer. Plus, the same caveats from Scenario 1 apply: if you COULD write to 0x0, it is almost always a terrible idea.
-----
Grand takeaway? This is probably a talk relevant to PC and not embedded. A Cortex-M would definitely complain about this code. Even on an AVR (atmega) changing 0x0 would be changing R0. I don't even know if that register is directly writeable during execution---probably because it's how you set some bootloader bits---but it's definitely dangerous.
With that in mind, I searched for "virtual address space windows" to figure out what goes in an executable at address 0x0. After reading Microsoft's first two articles, it's still unclear what goes in 0x0. Wikipedia? No answer. Next result? Nothing.
Finally, a page about Windows ME/98/95 states address 0x0 is "available to the process" but is "not writable".
After ten more minutes of fruitless searching, I'm willing to just test what happens, because I expect that's not even writable without side effects during execution.
#include <iostream>
int main() {
*(char*)0 = 0;
std::cout << "test" << std::endl;
}
and then, "g++ test.cpp -o test", and ...[1] 869 segmentation fault ./test
Linux? Fail.
-----
"cl test.cpp /link /out:test.exe"...
This compiles and runs but does not display the test string. (If you remove the 0x0 assignment, the string displays; no surprise)
Interestingly, I can't delete test.exe right away. It has stayed open:
ERROR: The process with PID 12776 (child process of PID 19728) could not be terminated.
Not only does PID 19728 not exist, I have to use an administrator console to "taskkill test.exe", so this would probably never belong inside user space code on Windows.
-----
Thus, this short code snippet seems like trouble, and the answer I would hope for as an interviewer is, "yeah, I don't write code like that, sorry."
(Coincidentally, I would say the same in response to the professors giving microcontroller exam questions that try to stump students with multiple casts and dereferences in a single line *ahem Moreno*. Stop teaching students methods they will never, ever use in a shop that cares about its code!!)
Now I watch and learn.
For embedded, this might do something nonawful on a device with memory management e.g. Cortex-A. This exercise is left to the reader.
Moral of the talk? Most of computing is extremely complicated, and you can hand-wave and abstract much of it away unless it's directly in your field. In this case, that would be a compiler engineer or low-level kernel developer.
Humorous 'memory' explaination of various levels of 'memory'[1]
Without reviewing the video, looks like a coding artifact related to segment 64k memory of 8/16/32/64 bit dos/MS windows (vs. arm / sparc with flat address space).
far/near pointer delclaration would clarify things a bit.
This is a c language artifact. Using in C++, one would need to impliment/override the default C++ supplied memory management to avoid memory management gotchas.
The code example was a way to make sure that pointer variable on a non-unified memory architecture was initialized to point within a given memory segment.
aka Reference zero's out the non-unique bits and leaves the "memory reference bits for given memory segment" alone / segmented NULL.
Zeroing out the non-unique memory bits still allows one to make use of the upper address bits to find out where in memory the given
memory segment is (useful for implimentation of setjump(), figuring out machine byte order, 1st time in segment, etc.).
[0] : Expert C Programming, Deep C Secrets by Peter Weinberger
The code can do something shifty at run time, without a power surge, to the address reference before the address is de-referenced
For all the gory details see Unwinding the Stack: Exploring How C++ Exceptions Work on Windows - https://www.youtube.com/watch?v=COEv2kq_Ht8