Fixing C
embedded.com
embedded.com
What about buffer overflows (it happens to the best of us)? Mounds of undefined behavior exploited by optimizing compilers (even in seemingly correct code)? A type system that is strict enough to get in one's way, but not strict/powerful enough for strong guarantees?
Some years ago I read a post on a forum for electronics/embedded programming, which basically said "The system of C header files is actually very good, because without it you would have to manually declare all functions you use." This article reminded me of that.
I'm sure there are some good technical reasons that make this difficult/expensive, but there aren't even any half measures available. Not even something that requires the pointer be at the start of the originally malloc()ed memory and can return an error code if the pointer is somewhere in the middle of the heap or off in lala land.
A function like that would at least make defensive programming possible and not require people to manually pass in the size of their buffer all of the time.
EDIT: That might sound snarky, it's not meant to be. I just mean that you can do something like this:
https://gist.github.com/joeld42/74f67cde97a789467a2f681961b8...
people complain about C's memory manager, but really it's one of the most flexible out there because you can do stuff like this. Well written C libraries will provide hooks to use a custom allocator if needed.
So it's great for your own internal code, but of pretty limited value once you start pulling in libraries.
Most decent 3rd party libraries will allow you to plug in your own allocators, and if they don't, they should. Typically you do something like #define LIBFOO_ALLOC my_alloc before including things, or they have some kind of function like setAllocator().. For example, here's how to do it for Freetype (a pretty large 3rd party library). http://www.freetype.org/freetype2/docs/design/design-4.html
If a 3rd party library didn't let me do this, I would probably refuse to use it (or patch it). And you can always search/replace malloc and free if you're using some 3rd party code that doesn't do that. It really doesn't take that long.
Also, malloc and free are specified to be weakly linked, so you can implement your own and replace the cstdlib ones even if you don't have the source available. This isn't a great option, but it works, especially if you're just using it for a diagnostic build.
If you're talking about C++, that's yet another thing the STL screwed up. You can usually find ways to work around it but yeah, it's a bigger problem.
http://stackoverflow.com/questions/1281686/determine-size-of...
With Ada, it is straightforward to create the kind of tight, fast code you'd write in C. You can control your own garbage collection and it has a fabulous type system. And - unless you need templates - it can do pretty much everything C++ can do as well.
I don't use Ada due to its verbosity. That's makes refactoring hard, and I refactor constantly. But only yesterday I picked up my copy of _Concurrency in Ada_ and flicked through to look at some ideas. There's lots of good stuff in there. I guess at the back story - a team who have spent years with difficult enterprise problems, who have thought really hard about them, and then rolled their conclusions back into the platform.
In the near future, a language called "flash" or "rock" or "bourne" will hit the hn frontpage. People will be amazed by its bootstrap website, its features, and will compare it well to Go. It'll get a whole lot of momentum. Weeks later a greybeard will publish an article, "Rock vs Brand X". The author of Rock will come clean in the comments saying that Rock is just a syntax transform that feeds to gnatmake, and that the whole thing was an April Fools joke.
http://www.embeddedconf.com/boston/scheduler/session/mars-at...
The full conference pass is $1300 but you can watch his session and many other sponsor sessions with the free pass.
> The answer to that is that if you need
> more than 3 levels of indentation, you're screwed anyway,
> and should fix your program.
Breaking the code into manageable chunks is a huge part in writing clear, maintainable software. A high amount of curly braces implies deep nested loops and/or branches, which also implies a high complexity (measured by cyclomatic complexity).Also, IDEs can auto insert braces and indent the code. Unless you're using notepad, I don't see the problem.
What I would like would be a way to portably create a stack frame and pass it around as an object and then call a function with it.
Everything should be made as simple as possible, but not simpler.
Sounds like node and its moronic packaging ecosystem.
git clone --depth=1 git@github.com:git/git.git
find git -name '*.c' -type f | xargs grep -P '^( {4}|\t){4}' | wc -l
11375
git clone --depth=1 git@github.com:torvalds/linux.git
find linux -name '*.c' -type f | xargs grep -P '^( {4}|\t){4}' | wc -l
732937
Stupid shitty broken software. </s>"$ find git -name .c -type f | xargs fgrep '{' | wc -l"
19288
"$ find linux -name .c -type f | xargs fgrep '{' | wc -l"
1408368
It's not as if Torvalds wrote the entire codebase by himself, anyway.
Git and linux were both written (originally) by Torvalds.
And both use a fair amount of deep indentation.
Whether >3 level of indentation is "good", IDK. But it happens a lot. And I'm a practical man.
to start:
https://gcc.gnu.org/onlinedocs/gcc-4.7.0/gcc/C-Extensions.ht...
we see that bare C has:
- no vectorization/SIMD support
- no hinting at likely branches
- no way to prefetch memory
- no way to block inlining
- no actual inlining!
whoever decided that the compiler should be allowed to ignore the 'inline' keyword ....
additional issues off the top of my head:
- const != immutable so 'const' is relegated to being a keyword for generating compiler warning
- RVO is implicit.. so just pray it happens!
The language is frankly just too old. Half of those features probably just simply didn't exist in hardware when the language was designed
I'm desperate for a better language
I'm not sure what you mean here. On embedded targets, static data declared 'const' will be put in with the program memory, and so will be definitely read-only. Casting the pointer and writing to it will cause a hard fault (or segfault, or whatever the equivalent is on your platform).
int main(int argc, char argv) {
const int *b = 9;
b = 10;
printf("%d",b);
}
what does (should) this print? (It compiles with warnings on llvm 7.3 / clang703 OSX)
const int \*b
Means a pointer to a thing that is const. The pointer itself (which is on the stack in this case) is NOT const. const int *b = 9;
*b = 10;
^^^ This will NOT compile. const int b = 9;
int main() {
int *a = (int *)b;
*a = 10;
}
^^^ This will compile (possibly with warnings). If you run it on an embedded target, this will crash: b will have been put in flash, so trying to write to it is a hard-fault.Yes, C allows you to do 'unsafe' things with pointers. You aren't going to fix that without throwing away C and starting again.
Using a pointer to iterate through a string/array/other data structure, that is stored in const is a completely standard thing to do.
To the compiler, this:
const int *a;
a = NULL; // perfectly OK
*a = 0; // does not compile
Is a pointer to a constant int, not a constant pointer to an int.You're probably confusing it with this:
int *const a;
a = NULL; // does not compile
*a = 0; // perfectly OK
Which is a constant pointer to a regular int.Of course you can combine both as follows:
const int *const a;
a = NULL; // does not compile
*a = 0; // does not compile
Which is a constant pointer to a constant int (which, of course, makes no sense at all in this case since 'a' is uninitialized and cannot be initialized without an unsafe cast, but it's a perfectly valid statement in C).> It's all about the underhanded trick(s) with the pointers, yours will never compile. I'm forcing clang to do horrible things.
Indicates, to me, that you see this as a strange behaviour that you're forcing the compiler into when in fact this is exactly the intended (and expected) semantics of const.
I switched the code to this:
#include <stdio.h>
int main(int argc, char **argv) {
const int b = 9;
b = 10;
printf("%d",b);
}
And I get a very clear "error" and nothing compiled with gcc, icc, and clang. Did you maybe have a left over binary from a previous compilation? I didn't try OSX, but it's possible the real moral would be to always pay attention to warnings. #define begin {
#define end }
and many other Pascalisms. As a result, his code did not look like C at all. IIRC, the code had gone unmaintained since the late 80's, which is not surprising in embedded systems. The code I had to fix was truly cringe-worthy.I vehemently disagree with Ganssle's article. Curly braces are the way to go. I am quite comfortable with pythonic indentation now, but remove one 'if' in complex code, and we have to change all the code below it manually. Cut-paste some code, and we have to manually take care of indentations. It's a pain.
Instead of his proposal, I'd want cleaner syntax for bit-wise addressing, which is currently handled via cumbersome unions or mask macros.
One of the most beautiful and elegant ideas in syntax is that of pattern matching. You want your syntax to match your semantics. That means, for instance, that your arrays (being a fixed, ordered sequence of items) should be presented syntactically as a fixed, ordered sequence of symbols. Your functions (being means of converting inputs to outputs) should syntactically separate inputs from outputs. (This is why I also despise INOUT and it's brethren.) And the most elegant languages will even use pattern matching for control flow, by syntactically differentiating different branches in parallel.
Part of C's brilliance was using pattern matching to declare variable types. So when you write this:
int *a;
If you solve for `a`, you get an `int * `, and if you solve for `* a`, you get `int`. That kind of general purpose symmetry is the goal.However, BEGIN/END as delimiters of blocks of code are far to imperative. The words refer to positions in the code, not to the structure of the code itself. This allows it to create noisy ambiguities:
FOR i=0; i < 10; i++
WHILE j > 0
....
END FOR
END WHILE
This is nonsense, because the terms are too granular and don't work in relative position. It may be more explicit, but explicit is only good when you're actually making decisions. Nobody should be deciding to overlap loops or blocks of code. Explicit is not good when there's one best way to do something, and `{}` always implicitly denotes a block. Nesting is easy, and doesn't need to be explicit. It just needs to be readable.Maybe Rust or something new will come along sometime soon to fix all the issues that exist with C.
Oh, and for those who haven't seen it, there was a cool guide to writing neat, modern C on here a few months ago. https://matt.sh/howto-c
Worth a read in my opinion.
I feel much the same way, and I think it's the general principle of freedom over security that makes it fun. You can do lots of things that other languages wouldn't allow, and despite not needing to most of the time, the fact that such power is there if you want to use it is what I like. It's a bit of a refreshing environment compared to all the other "safe" languages.
That's okay, just please don't have this attitude while writing production code. If my system is compromised because of yet another buffer overflow exploit in some C library, I don't care if you felt the wind in your hair while programming it.
I think Keith Thompson's (not related to Ken Thompson) critique of the link you posted is an even better read: https://github.com/Keith-S-Thompson/how-to-c-response
I am using Nim for years now, and it is much more convenient than C while retaining the same performance. C functions can be imported and used seemlessly. Nim offers several optional garbage collectors which can be turned off for embedded systems. Nim has a package manager (Babel), and Nim's macros which are powerful like Lisp macros (way ahead of C++ macros) make it possible to create special DSLs for test cases and webservers, or to use Perl's awesome regular expressions with native Perl syntax.
Homepage: http://nim-lang.org
Nim on Arduino: http://disconnected.systems/nim-on-arduino/
Embedded Stack Trace Profiler: http://nim-lang.org/docs/estp.html
Nim for scientific computing: http://rnduja.github.io/2015/10/21/scientific-nim/
A sinatra-like web framework for Nim: https://github.com/dom96/jester
"We just switched from Rust to Nim for a very large proprietary project": https://news.ycombinator.com/item?id=9050114
I know many programming languages and styles. Nim is one of the most amazing, powerful, practical and effective programming languages ever - a real gem. I am still wondering why so few people know about it. Probably just because of lack of promotion.
It's called Nimble now: https://github.com/nim-lang/nimble
Edit for the nitpickers: Yes, arrays decay to pointers and that's the real problem. Sorry.
This kind of thing keeps coming up in vulnerability reports.
This really isn't complicated. It's hardly obscure. There's no excuse.
Worse, it's an easy way to get burned on the sizeof() function, especially if you at some point refactor the code and put that chunk in a function separate from the original declaration.
This is why C programmers get gunshy about relying on that information and instead just treat strings like pointers all of the time.
Good programmers do nothing of the sort.
If you're talking about the old stdlib.h calls, don't use them with blind pointers where the size isn't known. I'd say use the "n" versions but you don't even have to do that.
'C' gets easier when you do faux RAII and make "objects" with internal state, at least think about not using dynamic allocation for everything and think a little bit about what might make the API usable by the extremely lazy.
This article is contentless... I mean there's absolutely nothing informative in it. And anyone who things that curly brackets is the thing to fix in C must have been never using C in the first place.
For projects where you are free to choose your language, Nim is very modern and powerful. It compiles to C so you can use it as easy replacement.
That's definitely the best idea I've seen.
It can get much worse. Consider a function that takes two strings as arguments, then returns the longer one. Constness of the result ideally depends on the input. At the very least, when both inputs have the same constness you'd expect the result to also have that constness.
I cannot remember the last time this happened to me.
> I’d prefer requiring matching begin and end blocks, with the end statement indicating which block is being closed.
So you can close an outer block before closing inner blocks? Or you can omit closing of a block altogether? Why - doesn't this create loads of nasty ambiguity?
> Sure, careful indentation tells us the same thing, but so much code today has been modified by so many people that the original engineer’s careful indenting often becomes hopelessly mangled.
"Often"? Just fix that problem if, indeed, it IS a problem.
> I often put an indication of which brace is closing which block in the comments.
As others have pointed out, if you need to do this, you have bigger problems.
I do actually like Python's approach, although I don't have enough experience with it to judge if it's a big improvement.
struct much_improved_c { ... };
struct much_improved_c an_api_that_doesnt_suck() { ... }; for i in {1..20}
do
echo $i
done
What is the purpose of having the `do` and `done`? And with if statements, what's the purpose of having `then` and `fi`? Curly braces fix that issue and are much easier to read. Not to mention that literally every single real editor has a plugin for making sure your braces always match up.I think `else if` to `elif` would be a much better change to fix people who write `else <newline> if` and other things that make it hard to read code. However, it should be noted that you couldn't do some pretty horrible macro hacks that way (but maybe if you made the no-brackets format of blocks no longer valid, then that problem would be fixed).
Not to mention that `end <block>` is syntactic clutter. We all know what block it must terminate -- that's why brackets work.
Why don't we get rid of the semi-colon and delimit the line of code via the symbol EOL, or maybe use the carriage return.
Here is a much better start: http://blog.regehr.org/archives/1180
"Proposal for a Friendly Dialect of C"
This is the only thing that I miss in C. I always want to declare local functions, and return pointers to them.
I would definitely accept GCC lock-in, but so far GCC does not have this feature. Clang does, hoever, via _Blocks, but the syntax is not very C-like.
I'm not saying every closure can be done without allocation. I'm saying many useful ones can be though.
> I’d prefer requiring matching begin and end blocks, with the end statement indicating which block is being closed.
The author's magic wand sounds like C macro to me.
The great problem of C is that almost everything is a pointer and they potentially can all be invalid. On top of this there's no easy way to check wether they are valid. This is the reason, why memory safety is so hard to get right with C.
Of course it would break a lot of existing C code, which is why it will never happen in any language named "C".
(Braces though? Really?!?)
while (1)
{
< many lines of nested code >
} // end of whileHow is this a problem with editors/IDEs doing it all for you? If your editor/IDE does not do this, then.. why don't you have this in your setup?
To make matters worse, you actually need to compile the program in order to determine which macro gets expanded into what amount of braces...