White space does matter in C23
gustedt.wordpress.com
gustedt.wordpress.com
When people say "whitespace matters", what it means is whether the meaning changes when you have more than 1 not more than 0.
inta;
vs
int a;
To which I can just say "duh!"
I agree. After reading the blog post, it's clear that the author should have called it "syntax errors do matter in C23"
DO 10 I=1.10
Which got interpreted as: DO10I = 1.10
Whereas the programmer wanted: DO 10 I=1,10
For a DO loop. Conversely, with SSW a language will evaluate these two expressions differently: inta = 10;
int a = 10;And not everything that is possible is worth doing; e.g. I once designed a language ("Leazy") where keywords don't have to be declared, and can be used as variable names, just to show I could still write an LL(1) recursive descent parser for it. You don't want that in anything for daily use, as it introduces confusion.
final (C++11)
override (C++11)
import (C++20)
module (C++20)> A further ten character sequences are restricted keywords: open, module, requires, transitive, exports, opens, to, uses, provides, and with. These character sequences are tokenized as keywords solely where they appear as terminals in the ModuleDeclaration, ModuleDirective, and RequiresModifier productions (§7.7).
https://docs.oracle.com/javase/specs/jls/se12/html/jls-3.htm...
Pff. You can have your cake and eat it too: disallow whitespace in variable names except no-break space. ;)
Writing shell script under the assumption that filenames do not contain spaces is a liberating experience. I want more of that! And it is nearly possible, by tr ' ' 0x00A0'ing every call to fopen, (probably as an option for mount).
CREATE TABLE []([] []);
Which will create a table with zero length name containing one column with a zero length name and zero length type. And yes you can do all the regular SQL against them providing you quote the zero length name.PL/I has that, too, because the designers thought you couldn’t expect programmers to know all keywords. For PL/I, that’s a correct assumption. Implementations can have hundreds of keywords, and some of them are single-letter (http://bitsavers.trailing-edge.com/pdf/ibm/series1/GC34-0084... pages 19-25 mentions A, B, E, F, P, R, S, V and X)
In many contexts yes but as a blanket statement no. Calca allows whitespace in names and it is a delight to write as a result
e.g:
[<Property>]
let ``Reverse of reverse of a list is the original list`` (xs:list<int>) =
List.rev(List.rev xs) = xs[1] https://fsharp.org/specs/language-spec/4.1/FSharpSpec-4.1-la...
The main use case here seems like self-describing properties that will be hardly referenced elsewhere. In principle this doesn't really need any new syntax, as you can put the description into the attribute (`[<Property>]` here) and the name itself can remain arbitrary or even be made anonymous. But we've got textual identifiers instead. So I guess that some code does refer to those textual identifiers, and if that's the case, being able to ignore invisible differences when comparing textual identifiers looks like a good idea as well.
https://web.stanford.edu/class/me200c/tutorial_77/03_basics....
This feature makes some tokenization ambiguous without context -- is MODULEPROCEDUREFOO to be interpreted as "MODULE PROCEDUREFOO" or "MODULE PROCEDURE FOO"? But tokenization without any reserved words is a tricky problem anyway.
#include<stdio.h>
#define r(R) R"()"
int main()
{
puts(r()[0] ? "C99" /* r() evaluates to "()" */
: "C23" /* r() evaluates to "" */);
}
Output: https://gcc.godbolt.org/z/Wj3s6KEGKI have used that trick here:
https://www.ioccc.org/years.html#2015_yang
(C23 wasn't a thing back then, but the same trick can be used to differentiate C++11 from C++98).
x*=02//* */2
...which was used to differentiate C89 and anything later.(More detailed explanation, if you can speak Japanese or tolerate machine translation, is available on: https://mame.github.io/ioccc-ja-spoilers/2015/yang.html )
Are you sure about that? I only see u, u8, U, and L defined as encoding-prefixes.
https://en.cppreference.com/w/c/23
On the plus side, I now have a GNU extension detector.
Can't you just check `__STDC_VERSION__` ?
(Or, is your way a "But, where's the fun in that?" exercise? Actually, if you're an ioccc entrant, "where's the fun in that?" does suddenly become a leading hypothesis ;-)
I can think of only one exception. In function-like macro definitions, the opening parenthesis `(` must directly follow the identifier. Though I guess the newline is significant in macro definitions in general, too.
Are there other places where white space matters?
Most pervasively, between any two letters and numbers. `unsignedint` is different from `unsigned int`, as is `inta` from `int a`. Similarly, `1 1` is not the same as `11`, `0x 1` is not `0x1`, and `0. 0` is not `0.0`.
`&&`, `||`, `<<`, `>>`, `++`, `--`, `//`, `/*`, `*/`, and all the `sign=` operators also don't allow whitespace.
New lines are significant for single-line comments (and macro definitions as you said).
Spaces are significant in escape sequences as well - "\n" is different from "\ n". Also related to string or char literals, newlines are also not allowed at all.
Spaces are also significant for an obscure feature which was officially removed in C23: digraphs and trigraphs. Before C23, sequences like `??/` (but not `? ? /`) were alternative ways of spelling many other characters. For example, this was a valid program:
int main() ??<
return 0;
%>
This also could be detected, as `??/` represents the `\` character, which, if it appears at the end of a single line comment, makes the next line part of the comment.So, the following function will tell you at runtime if the compiler supported trigraphs:
int supportsTrigraphs() {
// detector ??/
return 0;
return 1;
}• In the stringification via `#`, whitespaces are significant but any run will be converted to a single canonical space character.
• Macros can only be redefined when conflicting definitions are identical, where all whitespace separations are considered identical and trailing whitespaces are ignored. So you can put multiple copies of `#define FOO 3 + 4` with differing number of whitespaces, but `#define FOO 3+4` is not considered identical.
> Similarly, `1 1` is not the same as `11`, `0x 1` is not `0x1`, and `0. 0` is not `0.0`.
There is some interesting surprise here because the tokenization specifically looks for the preprocessing number (the `pp-number` non-terminal), which is a superset of both integers and floating points. There exist a set of preprocessing-only numerals, like `0xe+4`, which would be a syntax error once got past the preprocessor.
> Spaces are significant in escape sequences as well - "\n" is different from "\ n".
This is most visible when the space is between `\` and a newline. GCC and Clang accepts this as an extension but will issue the "backslash and newline separated by space" warning.
> Before C23, sequences like `??/` (but not `? ? /`) were alternative ways of spelling many other characters.
Note that digraphs (`%>` here) are still allowed. Trigraphs were problematic as they were very early textual replacements (even earlier than the tokenization!), while digraphs are just separate tokens with the same semantics.
The JSON appears to mentions that this is a regression affecting `U"string"` where U is a macro (that expands to a string literal).
Obviously there are numerous examples of where whitespace always mattered even in prior versions.