> Idiom #120 Read integer from stdin
> Read an integer value from the standard input into variable n
int n[15];
fgets(n, 15, stdin);
Really?> Idiom #120 Read integer from stdin
> Read an integer value from the standard input into variable n
int n[15];
fgets(n, 15, stdin);
Really?> Idiom #137 Check if string contains only digits
> Set boolean b to true if string s contains only characters in range '0'..'9', false otherwise.
char b = 0;
for (int i = 0; i < strlen(s); i++) {
if (! (b = (s[i] >= '0' && s[i] <= '9'))) break;
}
I appreciate the funny assignment-and-test-and-early-break in one (although I'd hardly say it's idiomatic), but I could do without the quadratic strlen(). int n = strspn(s,"0123456789");
BOOL b = (s[n] == 0);You raise an interesting point. It got me thinking about how big-O notation has failed us in some ways: it teaches us to ignore constant factors.
In big-O, an algorithm that makes 1000 comparisons per element is no different from one that makes a single comparison per element. They are both linear time. But you can't deny that one of these will likely take 1000 times as long as the other.
Of course, like you, I favor simple and readable code over grotesque code that is hard to understand and mentally verify.
When you have some big constant factor that's not good.
But when you have a big exponent that's not workable anymore, even for smallest input.
"In the abstract machine, all expressions are evaluated as specified by the semantics. An actual implementation need not evaluate part of an expression if it can deduce that its value is not used and that no needed side effects are produced (including any caused by calling a function or accessing a volatile object)."
Object files are an implementation detail not known by the C standard.
The relevant question is "does the standard say that it does not have side-effects?" (is a pure function). @skissane's sibling comment to yours provides the explanation of how the compiler can deduce that it's a pure function.
I think the standard would fall apart if you read it under the assumption that anything not explicitly forbidden can happen. Instead, you should read and find out what the side effects are (and same for undefined behavior, unspecified behavior, implementation defined behavior, etcetra).
"Accessing a volatile object, modifying an object, modifying a file, or calling a function that does any of those operations are all side effects, which are changes in the state of the execution environment." (There are more details if you care to dig in)
Of course nothing stops you or me from making extensions to the standard, but analysing things from the perspective that some implementation might extend strlen to have visible side effects goes too far into whataboutism for my taste, unless there are real world examples to make it a relevant point.
Writing it for human readers can be used to argue for the original implementation if you think of for(i=0;i<expr;i++) as the C idiom for "iterate expr times" and here you want expr to equal the length of the string, which is what the (obviously pure) function yields. No unnecessary variables and assignments -- no clutter. It perfectly describes the intent.
I try to write code that is as plain and simple as possible, and make it obvious what it does and that it has no gotchas. I want my code to be understandable both for new developers, and for my future self, who will surely be less smart than I think I am today.
Here is the original code, with the loop body elided:
for( int i = 0; i < strlen(s); i++ ) {
}
It is trivial to rewrite this as: for( int i = 0, n = strlen(s); i < n; i++ ) {
}
This is a very common idiom, and now it is perfectly clear what the code actually does. Obviously, it only calls strlen() once.Of course, as I pointed out in another comment, if you're writing a loop that iterates over a C string, you never have to call strlen() at all! You can just use the canonical C string loop:
for( int i = 0; s[i] != '\0'; i++ ) {
}
Or if you like brevity (which I like too): for( int i = 0; s[i]; i++ ) {
}It shows what happens if you take away the compilers' ability to reason about 's': it could be a global variable, and 'foo' could be modifying it, so now the compiler has to call strlen on every iteration.
So even though 'strlen' gets optimized out in the original version, it's quite a maintenance hazard since any minor change could inadvertently change the complexity from linear to quadratic. It also doesn't get optimized in debug builds, making those far slower than necessary.
It depends on the compiler, it's version, it's flags, and likely "the position of the moon".
Of course the compiler is only allowed to do transformations that the spec permits. But it's impossible for a human being to anticipate the exact outcome. It's more like: "Compiler, do something that has the same outcome as this code I show you here". The output can be than something that doesn't resemble the input even slightly!
There's obviously nothing wrong when the compiler is so smart that it sees some patterns and transforms your code into something much more efficient. Only that there's not much difference to what happens when you use a high level language. In both cases you in fact don't control the exact code that gets executed, and in both cases you rely on the smartness of your compiler to produce some efficient code, "whatever" you've written.
That's why I think it's mostly a function of the code-style how performant or efficient some language can be (to some extend of course). When you write low-level style code (even in a high level language) a smart compiler will (hopefully) create something like what you would get form writing your code in C/C++.
When the compiler is optimizing it might very well realize that the parameter to strlen doesn't change and the output can be saved first.
If the intent is to compute the length only once, then we can save the length in a variable.
Of course; as others have pointed out; there is no need to call strlen at all in this example.
It is difficult to optimise because you need the compiler to evaluate and prove at compile time that both the loop cannot affect the result of the function call and the function call will not affect the loop.
It's a common and easy optimisation to simply move function calls like this out of the loop.
For the cost of 1 line of code I've regularly seen 10%, 100%, 1000% speed ups.
It's actually one of the most common optimisations to do in non-compiled / "slow" languages if you know how functions are evaluated, you see that the cost of a simple getter function call can be the most expensive part of a loop.
It's great and sometimes even astonishing what GCC and LLVM can do.
Also it's clear that even the smartest compiler can't magically optimize any code.
My point was more about the fact that compilers for lower level languages like C/C++, exactly the two named, use the most "magic" possible and that it's therefore almost impossible to anticipate upfront how their generated code will look like. But C/C++ claim that you have the most possible control over the code. My point was that this is only true to some extend, and that you can get almost equally good generated code using a less low level language just by writing code in a style matching the usual low level languages. (Especially than you need to think about loop invariants and such like you said)!
So my point was more: The claim that you have "total control" over what happens at runtime when using a language like C/C++ is false.
Seeing this example and at the same time people discussing pages long (while using even de-compilers) given that source snippet how the generated code may or may not look like reminded me of that, like I said "funny", fact about the "total control" C/C++ gives you.
It's not an issue, of course. It's just an observation and I was reminded of it.
If you have a good mental model of modern CPUs, in a simple loop, you can estimate what you think the bottlneck of the function will be, either by counting the micro ops or the number of stack / heap memory reads or memory allocations, etc, to estimate what is really happening to work out how you can optimise it, otherwise you're just shooting in the dark trying random combinations of flags or code not understanding why something worked or didn't work.
At least in C/C++ that model works.
In slow languages like python/javascript/etc, doing simple operations doesn't translate down to the very low levels at ALL. Generally if you imagine the worst possible way you can think of for how something will execute in a simple loop and multiply it by 10, it might be close.
Also, code like this should always be put inside a function that returns a value, not just written inline. Making it a function allows simpler and more understandable code too.
The funniest part is that it is not necessary to call strlen() at all! The whole thing can be written in a single pass over the string. Here is how I would code it in C:
int OnlyDigits( char str[] ) {
for( int i = 0; str[i] != '\0'; ++i ) {
if( str[i] < '0' || str[i] > '9' ) return 0;
}
return 1;
}
Try it here:What about adding a check of str[0] == 0 -> return 0 Also, giving char str[] will make it char* str. Which can be null. This may cause reading a random memory location (possibly segfault or use-after-free)
edit: I get the comments but empty string still contains no digits. Given the regex would be ^[0-9]+ (+ instead of *) What I want to say is string has a numerical value or not. Empty string is NaN.
It is synonymous with “none of the characters is a non-digit”, which is true for the empty string.
Yes, and that was deliberate on my part, as it meets my expectation of what such a function should do in this edge case.
The problem statement was "Set boolean b to true if string s contains only characters in range '0'..'9', false otherwise."
To my mind, the question "is every character in the string a digit" should be equivalent to "are there any non-digits in the string" (with the answer inverted, of course).
Returning 0 (false) for the empty string makes those questions not equivalent. It makes the empty string a special case.
Of course the real problem is that the problem is under-specified. It should call out specifically what should happen for an empty string, because as illustrated here, this is something where reasonable people may disagree.
> Also giving char str[] will make it char* str. Which can be null. This may cause reading a random memory location (possibly segfault or use-after-free)
Well yes, of course. The point of my comment wasn't to write bullet-proof library-ready code, it was only to illustrate two things: code like this should always go in a function, and the entire task can be accomplished in a single pass through the string.
Thanks for keeping me on my toes!
But that's the future of programing! Just ask Microsoft or JetBrains.
Soon, with the help of AI, any random dude will be empowered to write software!
Big layoffs are to be expected as AI will take over most of the high-paying jobs in the software industry.
Belief me. /s
- There are no standard integer types that take 15 (decimal) digits to represent.
- The array contains ints instead of chars
- Why would you use fgets() instead of just gets()? (Though I don't touch C very often so perhaps that is considered proper style)
- Obviously no conversion of the digits into else, let alone specifying a base or handling a `0x` prefix for hexadecimal or a minus sign for negative numbers.
I assume it's because gets() ranks as "-10: It's impossible to get right" on Rusty's API Design Manifesto? (http://sweng.the-davies.net/Home/rustys-api-design-manifesto)
man gets
...
SECURITY CONSIDERATIONS
The gets() function cannot be used securely. Because of its lack of
bounds checking, and the inability for the calling program to reliably
determine the length of the next incoming line, the use of this function
enables malicious users to arbitrarily change a running program's func-
tionality through a buffer overflow attack. It is strongly suggested
that the fgets() function be used in all cases. (See the FSA.)Nitpick: that is irrelevant. The code reads in at most 14 characters.
> Why would you use fgets() instead of just gets()?
You don’t use gets because it doesn’t exist anymore. It got removed in C11 (it rightfully was deemed so bad that backwards compatibility was sacrificed). You can use
char *gets_s( char *str, rsize_t n )
, though.