What does the ??!??! operator do in C?
stackoverflow.com
stackoverflow.com
Appendix A (Reference Manual) of the book broadened my outlook on programming languages by providing me a glimpse of what goes into formally specifying a programming language. Section A.12 (Preprocessing) of this appendix specifies trigraph sequences. Quoting from the section:
> Preprocessing itself takes place in several logically successive phases that may, in a particular implementation, be condensed.
> 1. First, trigraph sequences as described in Par.A.12.1 are replaced by their equivalents. Should the operating system environment require it, newline characters are introduced between the lines of the source file.
Then section A.12.1 (Trigraph Sequences) further elaborates trigraph sequences in more detail. Quoting this section below:
> The character set of C source programs is contained within seven-bit ASCII, but is a superset of the ISO 646-1983 Invariant Code Set. In order to enable programs to be represented in the reduced set, all occurrences of the following trigraph sequences are replaced by the corresponding single character. This replacement occurs before any other processing.
??= #
??/ \
??' ^
??( [
??) ]
??! |
??< {
??> }
??- ~
> No other such replacements occur.> Trigraph sequences are new with the ANSI standard.
[edit]
For example ISO/IEC 646 is ascii with punctuation replaced by other characters.
On the Compis II computer (a CP/M machine built on the 80186 CPU), there were places for {|}[\] in the character set, but they were in the top half of the 8-bit characters and not generally useful for programming.
As opposed to say, “Learn You a Haskell for Great Good! A Beginner's Guide” which is 881 pages and doesn’t even moderately cover the prelude.
Anyway, C is an amazing language and I keep a K&R on my phone as a pdf
Page 1 starts: "Chapter 0: Introduction".
That says something about the C programming language design, that I as a deeply stack based FORTH programmer and explicitly parenthetical LISP programmer find horrible.
Well that's what I do. If you're having to look it up to write it, you're going to have to look it up to read it again down the line.
I've been programming in C for years and never had an issue with operator precedence.
For example: I really like HP calculators, so I'm in several facebook groups for fans of RPN/RPL and HP specifically. Sometimes a few of them go way too far out of their way to try to demonstrate how inferior algebraic systems must be.
For the record, my copy of K&R wants to open into either section 7.6 or appendix B. No idea what this says about me, though.
https://www.youtube.com/watch?v=fKHaNIEa6kA
It's about the poor people who read your code that relies on both you and them having perfectly memorized every single little detail of operator precedence and associativity, instead of simply and consistently using parenthesis.
Quick without looking: can you tell me what the precedence and associativity of the ternary ?: operator is?
The designer of PHP got it wrong (which isn't surprising given his proudly self proclaimed contempt towards computer science and incompetence at parser writing), but then millions of PHP programmers also learned it the wrong way.
https://en.wikiquote.org/wiki/Rasmus_Lerdorf
Do you really want any of those people who were corrupted by PHP messing around with your code, if you relied on it being one way, and they assume it works the other way?
It's not that you can't tell what it actually does, it's that you can't tell what the person who wrote it actually meant, which is more important than what it actually does, especially when it has bugs.
Don't do many operations on one line, AND do use parenthesis, AND do use indentation, with no exceptions except for very simple expressions. Take every opportunity to use line breaks and vertical alignment to make symmetry and repetition and nesting visually obvious, like:
float distance =
sqrt(
(x * x) +
(y * y))
Redundant parens, plus breaking expressions into multiple lines and indenting according to depth, unambiguously express programmer INTENT, so the reader doesn't need to wonder if the person who wrote it had a clue or was just showboating.Just use parenthesis, and put a comment on it, sailor.
https://wellcomecollection.org/works/m33njwx3/items
My copy of The Little Schemer won't open to page 13 because of the jelly stains.
https://vpb.smallyu.net/[Type]%20books/The%20Little%20Scheme...
I'm glad you enjoy Forth so much, I guess. I'm sure postfix will catch on any day now.
> Just use parentheses
Not sure what that has to do with memory usage..?
mrguyorama just implied that the only reason you would check the operator precedence chart would be to shave a few bytes off the size of your source code, which has not been a reasonable reason to do anything for many decades, and yet C programmers seem to like to do it anyway.
I suppose you could compare it to a table saw. C is one without a guard or any other safety measures, so you need to be careful not to cut your fingers off. More modern languages have the guard and break etc.
For general use you probably do want all the safety bits, but occasionally it is useful to be able to take it off to do a weird cut on a weird bit of wood.
None of that necessarily means you will cut your fingers off though.
It wasn't, so any C application that is more than a toy hello world with stdio, pings back into POSIX for any kind of meaningful work, that wants to stay cross platform.
Basically it the the C runtime library, that wasn't part of ISO.
I use JVM, .NET, Web and C++, not caring if the runtimes are bare metal or running on top of an OS, type 1 hypervisor, or whatever.
If you're downloading a JVM binary, you're missing out on the build step. It's C dependent, friend. How do you think that VM interfaces with the OS? Go on. Try it. ldd the java executable.
It's libc all the way down. C itself is a sort of "VM" specification utilized to create the tools to run the tools to build the tools that make other high level languages possible.
Unless you create something entirely custom in platform specific assembly, you're running on C at some level.
POSIX is an IEEE standard (example [1]). POSIX defines the Operating System API. You can see the C implementation of this API here[2].
> so any C application that is more than a toy hello world with stdio, pings back into POSIX for any kind of meaningful work
Simply calling printf relies on writing to a file descriptor. A "Hello world" application on linux uses posix. ANY hello world application uses posix. Even your Java Hello world App will call into the posix APIs. `System.out.println` isn't magic. It calls into the C posix implementation.
If you want to do anything in any language (write to files, create threads, allocate memory, network communication), you need to go through the OS. POSIX is what defines that OS interface.
> I use JVM, .NET, Web and C++, not caring if the runtimes are bare metal or running on top of an OS, type 1 hypervisor, or whatever.
So you use POSIX, you just don't think about it.
I'm not sure what you're trying to say. The Portable Operating System Interface (POSIX) is specified in an ISO standard, and basically specifies what a UNIX operating system's programmable interfaces are.
https://en.wikipedia.org/wiki/POSIX
POSIX also specifies stuff like "awk must be made available". Is that what you think the C programming language specifies?
Number of times I've seen trigraphs in "real code": still zero. I hope it's the same for you.
All of this was pretty much fine in the context in which it was written, but these days bullet-proofing things is pretty much mandatory and K&R’s elegance disappears in the face of such challenges.
Perhaps that mindset is part of what made C survive for so long and in such diverse roles.
Great memories!
I don't remember whether trigraphs were not supported by the compiler at the time or whether we just wanted to avoid completely unreadable code. Not experienced in VM/370 administration we spent weeks to modify the system to use some international EBCDIC codepage.
The system never saw much use, everybody preferred Unix workstations where programming in C was a natural thing.
I've pasted it here for convenience (formatting fixed, thanks child comment!):
// Are you there god??/
??=define _(please, help)
??=define _____(i,m, v,e,r,y) r%:%:m
??=define ____ _____(a,f,r,a,i,d)
main(__)<%____(!_(-~-??-((-~-??-!__<<-
??-!!__)<<-??-(!!__<<!!__))+-~-~-??--~-~
-~-~-~-~-??-(-~-~-~-~-??-!!__<<-~!!__),-
??-!__))<%??>%>_(__,___)??<____
(printf("please let me die??/r%d bottle%s"
" of bee%s""""??/n",(!(___
%-~-~!!___))?--__+!___++:__+!___++,!(__-!!___)
&&___%-~-~!!___??!??!!(___%-~-~!!___??!??!__
-(-~!!___))?"":"s",___%-~-??-!!___<-??-!!___?
"r on the wall":"eeeeeeer! Take one down,pass ??/
it around")&&__&&_(__,___),"mercy I'm in pain")??<??>??> For code blocks, prefix each line with two or more spaces.Small nitpick, however I am happy you linked the page.
But think the cpp has to go away first, after enough sed.
https://grayson.sh/blogs/using-piphilology-to-hide-strings
https://www.gnu.org/software/gawk/manual/gawk.html#Signature...
But trigraphs have gotten old even for IOCCC. In the guidelines for recent years, they specifically mention "We tend to dislike programs that ... obfuscate by excessive use of ANSI tri-graphs": https://www.ioccc.org/2020/guidelines.txt
Of course they also had to limit the number of characters per symbol to 1 in order to unambiguously support the "ab" syntax for multiplying a and b (or however you wanted to overload the "absence of white space" operator), but fortunately they mitigated that little problem by making C++ fully supports Unicode, so you had thousands of single character Unicode variable names to choose from. His prophetic intuition was spot-on, now that there are so many expressive and inclusive Emoji characters to use for single character variable names!
https://www.stroustrup.com/whitespace98.pdf
I really appreciate Bjarne Stroustrup's clean simple design and coherent long term vision for C++2000, and I'm looking forward to using three dimensional white space overloading in C++3D.
In the meantime, you can use multidimensional analog literals: http://www.eelis.net/C++/analogliterals.xhtml
+:= in Algol68
=+ in B
+= in C
while (x --\
\
\
\
> 0)
printf("%d ", x); for i in range(10): #{
print(i)
#}
or: for i in range(10): #BEGIN
print(i)
#END
You can even mix-and-match them, like: for i in range(10): #BEGIN
print(i)
#}
or: for i in range(10): #{
print(i)
#END
or turn them inside-out, like: for i in range(10): #}
print(i)
#{
or: for i in range(10): #END
print(i)
#BEGIN
Python is extremely flexible that way, and can easily strangle and eat all other languages.> Almost every country needed an adapted version of ASCII, since ASCII suited the needs of only the US and a few other countries. For example, Canada had its own version that supported French characters.
> Many other countries developed variants of ASCII to include non-English letters (e.g. é, ñ, ß, Ł), currency symbols (e.g. £, ¥), etc. See also YUSCII (Yugoslavia).
> It would share most characters in common, but assign other locally useful characters to several code points reserved for "national use". […]
> Because the bracket and brace characters of ASCII were assigned to "national use" code points that were used for accented letters in other national variants of ISO/IEC 646, a German, French, or Swedish, etc. programmer using their national variant of ISO/IEC 646, rather than ASCII, had to write, and, thus, read, something such as
ä aÄiÜ = 'Ön'; ü
instead of { a[i] = '\n'; }
> C trigraphs were created to solve this problem for ANSI C, although their late introduction and inconsistent implementation in compilers limited their use. Many programmers kept their computers on US-ASCII, so plain-text in Swedish, German etc. (for example, in e-mail or Usenet) contained "{, }" and similar variants in the middle of words, something those programmers got used to. For example, a Swedish programmer mailing another programmer asking if they should go for lunch, could get "N{ jag har sm|rg}sar" as the answer, which should be "Nä jag har smörgåsar" meaning "No I've got sandwiches".http://www.righto.com/2019/11/ibm-sonic-delay-lines-and-hist...
The next-gen was far more common.. The IBM 3270 terminal hooked to a local controller that talked to the mainframe. Could also hook a printer to the controller, you could print screen and simple forms independently from the mainframe.
You know all this, but I've always thought it was cool, and try to refresh my understanding of the setup. I no doubt have many details wrong.
#include <iso646.h>
#include <stdbool.h>
#include <stdio.h>
#define is ==
bool is_whitespace(int c) {
if (c is ' ' or c is '\n' or c is '\t') {
return true;
}
return false;
}
int main() {
int current, previous;
bool in_word;
while ((current = getchar()) not_eq EOF) {
if (is_whitespace(current) and not is_whitespace(previous)) {
putchar('\n');
} else {
putchar(current);
}
previous = current;
}
return 0;
}I quite like them, but then again, I have been writing way too much python lately.
edited to add: I really like "Modern C" and just re-checked -- no mention of the preprocessor feature!
I love it when a candidate blows through my easy, medium and hard questions and leaves me scrambling.
<: and :> are [ and ]
<% and %> are { and }
%: is #
(since C99, and expanded a bit later than trigraphs)‘Unfortunately’, none of the characters used here can be coded using trigraphs, so you can’t use trigraphs to generate digraphs in source.
"Are question marks fine?"
"Yes."
"I'll come up with something."
"Whether it's computer languages or human ones, as soon as you get into a discussion about the correct parsing of a statement, you've lost and need to rewrite in a way that's unambiguous. Too many people pride themselves on knowing more or less obscure rules and, honestly, no one else cares."
I've used uppercase-only terminals, and I've used ancient C, but not at the same time.
Come to think of it, didn't they remove trigraphs in one of the more recent iterations of the standard?
Unicode is a worthy successor to trigraphs -- no need for pre-processing!
Guess with tri-graph elimination & awk getting unicode support will have to gawk C with cpp using pipology theory.
But think the cpp has to go away first, after enough sed.
https://grayson.sh/blogs/using-piphilology-to-hide-strings
https://www.gnu.org/software/gawk/manual/gawk.html#Signature...
So if you started programming anywhere after the point in time when you needed to hand off your code to a punch card operator, you're unlikely to have seen them.
#include <stdlib.h>
int main()
{
int *foo = malloc(sizeof(int));
return 0;
}
That works in C, but not in C++.There's actually another subtle different in there that main() means "unspecified arguments" in C, and "no arguments" in C++. ("No arguments" in C would be main(void).) However, it's no longer commonly used that way in C, but casts from void * to other types is very common in C.
The modern C++ way to do this ~safely isn't legal C, and yet the type pun isn't safe in C++. I believe using memcpy() to launder the bits is legal in both languages and in some cases your compiler can figure out what you're doing and not actually emit the unnecessary copy.
You have to explicitly pass -std=c17 (or whatever) to get standard-conforming behavior including trigraphs.
Took me a while to figure out that "trigraph" was referring to some part of "??!?!!?!????" and not "WTF".
Valid Invalid ??? (Exercise for the reader to decide if this is a trigraph or not)
The latter is a very common idiom in Julia code, which I found obscure and puerile at first (“look how smart I am”), but have come to appreciate as concise and natural by now.
For example:
function fact(n::Int)
n >= 0 || error("n must be non-negative")
n == 0 && return 1
n * fact(n-1)
end
https://docs.julialang.org/en/v1/manual/control-flow/#Short-... #define and &&
#define and_eq &=
#define bitand &
#define bitor |
#define compl ~
#define not !
#define not_eq !=
#define or ||
#define or_eq |=
#define xor ^
#define xor_eq ^=
I suppose that allows for code like this: if (x or not y or not z) {
return 1;
}
https://en.wikipedia.org/wiki/C_alternative_tokens template <typename T>
void print(T const bitand foo) {
std::cout << foo << std::endl;
} void print(auto const bitand foo) {
std::cout << foo << std::endl;
}
Since C++20.It made for some amusing group projects when I got to university, when classmates had never seen those operators and were trying to figure out where they were coming from and why I would write such silly things. I trolled them by replacing all my brackets with `begin` and `end` in the next assignment before moving to the standard use of C operators for the rest of the class.
[1] https://stackoverflow.com/questions/1642028/what-is-the-oper...
https://en.wikipedia.org/wiki/BCPL
This is the earliest example of this sort of thing i'm aware of - is there an earlier example?
Also, BCPL supported // for comments, again, probably the first use of this sequence.
This comment on the SO post made my day. :D
1.c:1:11: warning: trigraph ??< ignored, use -trigraphs to enable [-Wtrigraphs]
Is there a preprocessor directive to enable support out of curiosity? int main() {
[](){}()
}
is still wierd.Wonder if there will be a request for an emacs macro to handle the replaced cpp trigraphs? [2]
[1] https://zygoloid.github.io/cppcontest2018.html [2] https://www.emacswiki.org/emacs/CppTemplate
Anyway, probably obscure enough:
int main() {
[]<class=void>(){}();
} int main() {
[]{}();
}