A primer on some C obfuscation tricks
github.com
github.com
https://stackoverflow.com/questions/15393441/obfuscated-c-co...
It's a digital clock (which has to be compiled and run once per second to be accurate).
Of course, if you're interested in obfuscated C, you can't miss the International Obfuscated C Code Contest, which is where most of these evil tricks show up. Submissions for this year's IOCCC are still open: https://www.ioccc.org/
The IOCCC has been running since 1984, and there are some absolutely marvelous gems: https://www.ioccc.org/years.html. A great rabbit-hole to dive down if you're stuck at home ;)
({_:&&_;});
({});
({;});
All of these rely on a gcc extension ('statement expressions').________________________________
a = '-'-'-'
You can also do: a = '/'/'/'
To generate a 1 instead of a 0.________________________________
printf("%d %d\n", 0 == sizeof(count = 2, count++), count);
This works because:1. There are two forms of sizeof; sizeof(T) where T is a type, and sizeof x where x is an expression.
2. There are 'comma expressions'; if you have (x, y) where x and y are expressions, then x is executed and the expression evaluates to y.
3. Parameters to sizeof are not evaluated (which is important because otherwise the value of the 3rd argument would be undefined, since there's no sequence point between evaluation of function parameters).
Oh, but they are ;) Try running sizeof on a VLA sometime.
(If it's a parenthesized type name, it's unclear what it means to "evaluate" it.)
$ cat sizeof.c
#include <stdio.h>
int main(int argc, char *argv[])
{
printf("%zu\n", sizeof(int[argc - 2]));
}
$ make -B sizeof
cc sizeof.c -o sizeof
$ ./sizeof foo
0
$ ./sizeof
17179869180 ASM generation compiler returned: 0
Execution build compiler returned: 0
Program returned: 44
./output.s
PS: I did not examine the assembly code *(c ? &x : &y) = v;
return (char *[]){"No", "Yes"}[!!x];
(Note: for the second one, I use the "selecting from a temporary array" and !! separately; if I only had two options then I'd use x ? "Yes" : "No" of course.)Is !!x equivalent to x here?
It fits semantically with "I'm using x as a boolean, but it may not be 1".
I use !=0 in only one context, if(strcmp(a,b) != 0) to check if two strings differ. Normally if an if() expression I don't check explicitly for 0 or not 0, except for strcmp class of functions, because of the ordered comparison of it (<0, 0, >0) , its "boolean" value is inverted, so somewhat counter intuitive if used with implicit boolean values.
Cast to bool has me extremely nervous because ... It works in C99's _Bool type, but I have lots of bad memories of security bugs when using this approach with pre-standard custom bool types, which are still in common use including but not limited to Win32 and COM, or Objective-C.
The bug I am thinking of looks like this:
long long ll = 1LL << 32;
bool b = ll;
puts(b ? "yes" : "no");
I just tested to confirm. This prints "yes" when using C99 stdbool, presumably because they standardized it to work. If you change bool to a common choice for pre-standard bool typedefs (int or char, say), it prints "no" because the large nonzero value doesn't fit in that type. "Just use the standard type" you might say. But if we were working somewhere that required legacy nonstandard bools that is not our choice, and further, a later refactor on may start using those other types some day even if we make the correct choice today. So there is an argument to be consistent about avoiding this type of bug.Any nonzero expression is true, and 0 is false. But the result of !, or ==, ||, or any of those other boolean operators is 1 for true.
Recall relatedly that C didn't have a bool type until the 1999 standard, which introduced <stdbool.h>. C++ had it earlier, but for a long time the C way to store a boolean expression into a variable might have been int, or some nonstandard typedef. This is why so many libraries and frameworks in the C world define their own boolean type.
yes
yourCode = bad;
temp = bad;
bad = good;
good = temp;
conclusion = (yourCode == bad)
As we can see from evaluating this code, in this context, it isn’t bad.In general, if something is bad you can change that thing or change the context. Your job, your living arrangements, etc. People often are slow to change the context, even when a small change can make all the difference.
You might want to create an include file if you have a lot of code to check.
mycodeisalwaysgood.h seems self-explanatory.
Combining multiple booleans into an integer and using that as a switch or index is another related technique.
To modify an old saying slightly, perhaps "one person's obfuscation is another person's simplicity and elegance."
I understood the whole list of these obfuscation tricks upon first reading but its still obfuscation.
Creating an array instead of an if-else-switch like the second line is fine but for just two elements and in combination with !!x (and !!x in general) its just nonsense and doesn't help at all.
I'm all for code-density and avoiding repetition so I too use ?: whenever I can but not on the left side of an assignment. That's just an obfuscated if-else-statement. If normal code results in very long and "branchy" code you can rearrange your code in a better way.
int countNonZero = 0;
for (int i = 0; i < values.size(); ++i) {
countNonZero += !!values[i];
}
Also c++ has explicit operator bool();edit: not a great example as values[1] >0 would be better.
The reason I'm pointing this out is that often enough I had to work with code where this distinction was unknown to the author and you come across code like if (!!x) which is, again, just nonsense.
Only if that's an array of unsigned ints, as otherwise the code would count only positive nonzeros.
Screen space is cheap. Say what you mean. Clever today is usually a headache tomorrow.
I switch between many languages often enough so anything even remotely esoteric gets forgotten instantly and causes my brain to pause at suspicious line and making me think too much about the trees instead of the forest.
If however I was only using C and nothing else maybe I'd be catching this bug myself, one never knows ;)
Do you know why need to use pointer magic here?
Tried (c ? x : y) = v; and wondering why that doesn't work.
if (2^3 == 8)
puts("two cubed is eight");
if (5^2 == 25)
puts("five squared is twenty-five");2^3 => 1
5^2 => 7
if (2^3 != 1)
puts("two cubed is not one");
if (5^2 != 7)
puts("five squared is not seven");
it prints those messages too.The trick is operator precedence.
This is not true. If a decimal integer constant value cannot be represented in type "int", the next candidate type is "long int". If the value cannot fit in "long int" either, the next type to try is "long long int" in C99 and "unsigned long int" only in C89.
#include <cstdio>
void sw(int s)
{
switch (s) while (0) {
case 0:
printf("zero\n"); continue;
case 1:
printf("one\n"); continue;
case 2:
printf("two\n"); continue;
}
}
[0] https://gcc.godbolt.org/z/Q26LWG void sw(int s) noexcept
{
switch (s) {
case 0:
printf("zero\n"); break;
case 1:
printf("one\n"); break;
case 2:
printf("two\n"); break;
}
}Not a C expert, so just curious.
> return (char *[]){"No", "Yes"}[!!x];
I prefer return (!x<<2)+"Yes\0No";&"Yes\0No"[!x>>2]; // same number of characters
to shut it up.
Can't be unseen, I can't believe I never thought of that.
It could just as easily throw a compilation error to index a constant with an array rather than the other way around. I don't think this works in Rust even if the resulting machine code for array indexing is the same.
Now, of course an array type isn't a pointer type, but as "indexing" isn't one of the very few cases where an expression that has an array type isn't converted to an expression with a pointer type, you aren't really indexing an array, but a pointer to its first element.
So index[array] isn't an historical accident. It might not have been deliberate, but it follows naturally from the nature of the language.
Go very much follows the same discipline. Speed and simplicity of compilation constrain the syntax, most notably the lack of generics. Goroutines, channels, etc, only require minimal syntactic and compiler support. Contrast that with Rust--Rust front loads everything into the parsing phase--lifetimes, async, etc. Deep AST analysis and transformation is everything for Rust. Of course these days people abhor even the possibility of allowing something like index[array], so even a compiler like Go goes out of its way to disallow it.
Not since the 1999 ISO standard.
> int main(){ return linux > unix; }
A conforming compiler must diagnose "linux" and "unix" as undeclared identifiers. Many C compilers are not conforming by default.
Or maybe not.
fn main() {
return return return return return return return
}
You can do silly things in any language.Likewise, in your example, I doubt that this could bring anything. Is there a practical obfuscation method based on that quirk in Rust, or a reason for keeping it in the syntax? I cannot tell, but maybe you can?
Overall, the problem that I have with this is not that it is silly, it is that it makes it harder to understand and maintain. In some cases, it requires active engineering to fix the issue, but the language should be designed so that most of these problems are taken care of by design.
Also, many people seem to enjoy the fact the C can be bent that way. I don't mean to remove that from them, I just think that for system programming, it should be less permissive. Perhaps a 'strict mode' could be devised, not at the syntax level like in javascript (which I suppose couldn't be avoided), but as compiler flag (like the c++ people did it).
That's a very good point you are making here.
However, as I said elsewhere, I don't see the use of being able to either write
typedef int ...
or
int typedef ...
and I don't think that proper code generation would require it.
inline with_all_ColinIanKing_standards()
5) Surprising math:
int x = 0xfffe+0x0001;
It's trying to parse the e as scientific notation I think.
surprising_math.c: In function ‘main’:
surprising_math.c:4:10: error: invalid suffix "+0x0001" on integer constant
int x = 0xfffe+0x0001;
^~~~~~~~~~~~~error: invalid suffix '+0x0001' on integer constant
Edit: Clang also gives an error. Mscv seems to compile. I wonder who follows spec. I assume not mscv ...
I'm sure it could have been made stricter, allowing 0xfffe+0x0001 to be treated as 0xfffe + 0x0001 -- but the solution is simply to write 0xfffe + 0x0001 in the first place. The language grammar has to be consistently defined even if it leads to surprises now and then.
int main() {
fork();
printf("choo");
}int main() { printf("choo"); fork(); }
Here, "choo" can be printed twice, even though we fork after printing. This is a result from line buffering when the flush happens after the process is forked. Essentially, the output buffer is copied when forking and therefore duplicated.