I love coding in C
lord-left.github.io
lord-left.github.io
C is basically it for embedded development though. I've gotten so tired of recompiling and waiting to flash a chip with C, that I've started learning ANTLR and making my own language. The idea is to have a language which runs in a VM written in C, and allows you to easily access code which has to be in C. Sort of like uPython, except it can easily be used on any board with a C compiler with a small amount of setup, and C FFI is a first class citizen. Also coroutines masquerading as first class actors with message passing, since I always end up building that in C anyway.
The only issue I have with C is that there are no good reasons for some of those missing features to be missing.
Take, for example, namespaces. Would it be a problem to implement them as, say, implicit prefixes?
Namespaces don't need to be mangled at all if they are interpreted as symbol prefixes.
I'd prefer that, say, namespace foo::bar resolved into foo_bar for all symbols (even that meant risking namespace naming collisions) to not having any support for namespaces in C.
In fact, this approach is already used to implement pseudo-namespaces, so that wouldn't be much of a stretch.
It seems like you want the language to be more complicated for no benefit. Why not use C++ at that point?
You're missing the whole point of namespaces. The goal is not to replace foo_bar with foo::bar. The whole point is that within a scope you can type bar instead of foo::bar, or bar instead of foo::baz::qux::bar, because you might have multiple identifiers that might share a name albeit they are expected to be distinct symbols.
using namespace kind::ofa::long_name;
or
using kind::ofa::long_name::bar;
is the point.
#include <string.h>
__prefix__ str; /* in scope: "","str" */
strlen("Hi!"); /* try "strlen",done,ignore "strstrlen" */
len("Bye"); /* no "len",try "strlen",done */
__prefix__(foo_) { void bar(void); } /* "foo_bar" */Me I want range types like Ada. Real array types. I think I want blocks/coroutines.
aka CDC
> a language which runs in a VM written in C [...] easily be used on any board with a C compiler with a small amount of setup, and C FFI is a first class citizen. Also coroutines
Use Lua. It has all of that already.
-Static types that match C for seamless interop.
-No GC
-First class actors
-Ahead of time declaration/allocation of actors and messages, with automatic async-like passing on a coroutine/actor waiting for an allocation, and automatic disposal of the actor after prolonged allocation failure.
-Hot swapping actors from a running system in the field
-Over the wire debugging and inspection of a running system in the field
AtomVM, an erlang VM for small devices is closer than LUA to what I'm looking for. However, I really want to create an environment made for embedded development, not try to tack a higher level runtime onto something like an ESP32 and have devices drop every 2 months from a memory leak.You can implement that for embedded systems too, but as there is no obvious best way to implement it (Von Neumann machine is so different) you'll constantly get annoyed if you are the type that likes to use the weakest hardware possible as different implementations have different costs and so you'll always be second-guessing the abstraction.
Rust using one of the actor libraries based on async, might suit embedded really well if cross platform support matures a bit more. Rust cross-compiling requires a C linker & compiler for the target platform but Rust doesn't respect the standard environment variables for setting them using LD/CC/CXX environment variables. Though it'd be really interesting to see if anyone can setup Rust to run on an embedded WASM to allow introspection / real co-routines. IMHO, that'd be awesome.
Currently I've settled on a Forth (in Forth style, an implementation I wrote https://github.com/elcritch/forthwith/) to give me an interactive environment without GC and stable timing (important!) with ability to add C function calls readily. Interactivity is big for me in the area I'm currently working in. Forth can easily be extended to have tasks / co-routines too, though I never tried implementing them. And of course, Forth macros are like exercises in puzzle solving.
AtomVM does look interesting, but early stage. Erlang VM actually matches embedded really well. The actor model in theory lets each process run it's own GC which makes the GC model much simpler than other dynamic languages. Modern OTP is large and not suited for embedded as you mention, but pairing it down to basic actors and processes would fit well on many modern embedded MCU's. Elixir and Nerves is great for SBC/Linux embedded computing and some boards like the Omega2/Licheepi Nano. There's also GRiSP (https://www.grisp.org/) that runs BEAM on top of RTEMS.
Also bitwise operators, but using functions instead isn't a huge problem (and supposedly 5.3 or so fixes that).
I do recall that in 1991, I was the only student in my "Introduction to Programming Using C" 101 class (as a Freshman in college) who could understand pointer arithmetic. That class ended the careers of many aspiring Computer Science majors.
I eagerly learned C++ a few years later. I did not understand at the time what shitpile OOP is. I was totally in OO because it was the "industry trend." Someday we'll fully recover from that entire folly.
I personally only got in trouble in C when trying to be too clever.
The truth is that decades of projects written in C have clearly demonstrated that people can't write secure programs in C.
The whole idea that many eyes or 'safe' languages or some other magic bullet are going to solve this is so far off-base that it makes productive discussion impossible because everybody starts to focus on the particular tool at hand whereas the real problem is an instance of of the class 'architecture', not of the class 'tool'.
however i disagree with both of you that writing safe C is hard. most of the problems i see relate to incorrect pointer usage or memory management. use stack allocated containers from libs very well tested and pass references. this removes entire classes of bugs from popping up. one thing you can’t do with C is be lazy.
https://en.wikipedia.org/wiki/Forth_(programming_language)
also, system level REPL is some hardcore nerdity.
Incidentally, while Forth comes from the era of ALLCAPS language names, even its creator Chuck Moore calls it Forth: https://web.archive.org/web/20040131054056/http://www.colorf... ; and Lua never was spelled with all caps.
You may want to check out Moddable's XS runtime[1], which is exactly that. It's written in portable C, and "XS in C"[2] makes it easy to interoperate with C code you may write. And even if you decide to roll your own runtime, there's tons to learn from the source code.
[1] https://www.moddable.com/faq.php#what-is-xs [2] https://github.com/Moddable-OpenSource/moddable/blob/public/...
I haven't gotten to building the VM yet, and I'm hoping to just use someone else's and then contribute the tools I need to their ecosystem.
This is the entire reason for me. With the way IoT is blowing up C actually might be gaining market share. I've done my research and things like TinyGo exist, and you can sort-of compile Rust for microcontrollers if you use all of the right crates, but it feels pretty risky to do so.
"Any sufficiently complicated C or Fortran program contains an ad-hoc, informally-specified, bug-ridden, slow implementation of half of Common Lisp."
In modern world, any sufficiently complicated C or C++ program embeds a runtime of some higher level language. More often than not, that language is formally specified, well-supported, and fast.
Sometimes it’s indeed LISP like in AutoCAD, but more popular choices are JavaScript (all web browsers), LUA (many video games and much more, see Wireshark), Python, VBScript or VBA (older Windows software), .NET runtime (current Windows software and some cross-platform as well, like Unity3D game engine). These days, it’s rarely a custom language runtime, but it still happens, like Matlab.
I'm running VSCode right now with many tabs open and extensions running, and it's using less than 100mb of memory. Doesn't seem so bad to me
You know... any embedded programmer reading that comment is probably laughing hysterically. In most cheap or energy efficient microcontrolers, more than 4 Mb is considered a luxury.
There is no virtue in using less memory for the sake of it. If a desktop application is taking 100mb, you can have a lot of those running before modern boxes start to struggle.
* 32-bit vs 64-bit application(64 bit will bloat all pointers)
* Methods of loading and editing documents(modern editing attempts a lot of introspection based on highlighting syntax or project environment)
* Character encoding support(Supporting current Unicode rendering would almost immediately take you beyond 4Mb)
Interactive editing is a pretty memory-hungry task compared to a simple viewer or batch processor. There's more reason to keep things cached. If you actually go back to old text editor versions from the days of 4Mb desktops, you'll find yourself missing stuff. Not so much that you can't get by, but enough to give pause and consider taking the hit.
uh... is it ? no tabs open and less than 12 extensions, the code process itself takes 120 megabytes, and there's also 1.2 gigabyte of electron processes running along with it
Many of them are international standards, see ANSI INCITS 226-1994 for LISP, ECMA-262 and ISO/IEC 16262 for JavaScript, ECMA-334 and ISO/IEC 23270:2018 for C#. Lua and Python aren’t standardized, but even so, they both have a comprehensive formal spec.
> calling Python or JavaScript any efficient sounds a bit unwarranted
Both are much slower at number crunching compared to C or C++ but they aren’t that bad, either. Most JavaScript VMs feature a good JIT compiler. Some Python runtimes have JIT too.
> Consider any JavaScript desktop application, its power and memory usage.
There’re good ones, like VSCode. On average they’re indeed not great, but I think it’s just an unfortunate consequence of low entry barrier. It’s easy for inexperience people to get started with the technology. And the ecosystem is misjudged based on the output of these inexperienced developers.
Maybe having most of this stuff in c libs with scripts wrapped around them will make it easier to migrate. Keep the same py but swap out the lib from C to some rust that's still useful in a crate for pure rust projects.
More often than not? Almost none of the languages mentioned are "formally specified", and most of them are hardly fast either...
They all have written specs. Many of them even have these specs standardized, see ANSI INCITS 226-1994, ISO/IEC 16262 and ISO/IEC 23270:2018.
> most of them are hardly fast either
It’s borderline impossible to be fast compared to C or C++. By that statement I meant 2 things.
(1) They’re likely to be much faster compared to whatever ad-hoc equivalent is doable within reasonable budget. People spent tons of resources improving these runtimes and their JITs, it’s very expensive to do something comparable.
(2) On modern hardware, their performance is now adequate for many practical uses. This is true even for resource-constrained applications like videogames or embedded/mobile software, which were overwhelmingly dominated by C or C++ couple decades ago.
Bookmarking that comment. Will totally quote that whenever the need arises.
On Windows any language able to speak COM can use it, regardless what language was used to create the COM component.
Same applies to mainframes and their language environments.
Well there's your problem. Instead of thinking about your problem and devising a C solution, you resort to pounding the square peg into the round hole. Your problem isn't C, it's your proficiency at it. What's worse, instead of correcting said ignorance your solution is to produce more code using a new language. This only adds to the gross pile of trash that is every language/framework that tries to solve programmer ignorance and laziness with what amounts to wishful thinking. This is why software really sucks nowadays. Everyone wants results without reading the manual or putting in real effort. They want languages with built in hand holding and magic inference. That's not how any of this works.
I get it, programming and C sucks because it's so damn tedious. I suck at C too. But I actively study books and similar problems people already solved using the language. Programming is very hard. The best programmers I know are also very good with math and puzzles (e.g. Ken Thompson was a chess geek). They're puzzle solvers. They don't cheat and try to smash or duck tape the pieces together. They thrive on studying strategy meaning reading books, papers, and other people's code. Learn how to solve the puzzles strategically instead of fighting them.
That's actually how _all_ of this works. If we required everyone to handle the complexity of everything, we wouldn't be able to make technological progress. Yes, there's the tension (there's always a tension) between having a deep understanding of an underlying technology, and therefore be able to use it better, and trying to get to a point where you can get away with not understanding. But it's undeniably considered a success story in technology if the technology gets to a point where people can use it to basically its full potential, without needing to understand it.
When writing interrupt service routines for serial communications on an obscure architecture with multiple heaps (one 16 bit addressable and one 32 bit addressable), I could have written 2000 lines of assembly. Instead, I used a system of assembly macros and wrote about 300 lines of that.
When writing firmware in a C which lacked decent coroutines, I could have just made a giant while loop and put all of my logic in it. Instead, I wrote a quick implementation of coroutines, a scheduler, and a basic async/await implementation using the C preprocessor. This also saved thousands of lines of boilerplate code.
The fact that C has no good method of implementing a generic hashmap has nothing to do with my proficiency in the language. My proficiency in C also has nothing to do with my love of sweet, glorious type inference, or my deep felt appreciation for the conveniences modern IDEs and languages provide. Writing software has never been better!
Maybe I'm misinterpreting, but what I'm hearing is: "You don't need to use Vim to get more done, you just need to get better at typing. Not that I know anything about how fast you type." Or... you know, I could just use Vim.
Implementing cooroutines in C is pretty old hat: https://www.chiark.greenend.org.uk/~sgtatham/coroutines.html
If you just put together a linked list of structs which represent the routines, the time they last slept, and what they're last waiting on, you can just iterate through it and you've got most of the functionality of async/await. I never bothered to implement separate stacks and register restoration, opting to just use global variables for state that needed to be saved between awaits. It would have been neat to do that, but I think it would have turned an afternoon spent saving time into a multiday project, and left me permanently paranoid about corrupting the stack or registers. Would have been fun though.
That is not how your original post read to me. It sounded like you came from the dynamic/scripting language world and are trying to hammer those features into C. You sound quite proficient, I'm not trying to shit on you or anything. I'm tired, it's late, and no one cares but it sounds like you want something like an mbed that runs python or FreeRTOS? Why reinvent the wheel if someone already did this? Or is it simply that you're stuck using C on your "weird" platform? In that case, I'd say have a look at Nim, there's an embedded branch. It's a dynamic language that compiles to C89 so perhaps it's buildable on your platform. If it's a personal project then have fun.
It will explode and kill people eventually. That's a sure thing, no matter how good you are.
The only relevant question than is: Have you been still around or did you manage to find some safer workplace far away?
It will? Some of the stuff in this thread is really over the top.
I think, it even does a better job than C++ which share some of the C syntax.
The problem i guess is complexity and productivity. Zig here is a hit while Rust is a miss as much as C++.
The question Zig was trying to answer is: Can we have a simple, free and powerful language as C but with a more modern approch?
While Rust in its cruzade to be seen as a C++ competitor were not aiming at simplicity and user ergonomics.
So my feeling is, Zig is much more compeling as a C substitute than any other language out there, and probably thats the reason why your friend like it more than others.
That isn't to say that the article doesn't have a point though, far from it.
Without cache coherency and ILP, programming in any language would be insane. It's not like programming in x86_64 assembly suddenly opens a world of possibilities because its "truly low-level". You gain very little extra control over a given platform by switching to assembly over C, that's what we mean by "low-level".
C maps cleanly onto the instruction sets provided by chip manufacturers. It provides the option to the programmer to optimize structures for use in vectorization if they so choose, or to optimize for some other objective like size for a binary wire protocol or limited memory space.
The very nature of having a choice about memory layout of structures and the ability to cleanly link with the platform ABI is what makes C low-level. Obviously the inner-workings of a modern CPU don't map cleanly to the C virtual machine. However, there's no convincing evidence that greater control over cache invalidation is what's holding back performance of those CPUs.
So C is too high level for GPU programming, and needed to be extended.
It certainly opens a world of possibilities regarding the vector units. That's one place where inner loops hand-written in assembly still have an edge. On the other hand, with contemporary CPUs, high quality assembly coding is a highly specialized job skill by itself, so for general purpose engineers, learning it is most likely to yield a bad cost/benefit.
The chips move heaven and earth to maintain the fiction that those instructions actually direct what they do.
The number of fundamentally different kinds of cache, and specialized state machines not directly accessible by instructions, in a modern chip would boggle your mind.
e.g.
>Why is there no pointer arithmetic?
>Safety. Without pointer arithmetic it's possible to create a language that can never derive an illegal address that succeeds incorrectly. Compiler and hardware technology have advanced to the point where a loop using array indices can be as efficient as a loop using pointer arithmetic.[1]
I wish this meme about Go and generics would die already. C allows you to define an array of int and and an array of char and have those be two different types. Go allows you to do this with slices and maps as well, because slices and maps are also built in types in Go. This is not "generics".
[1] With the largely irrelevant exception of the C11 _Generic stuff.
If anything, more modern languages (eg Java) have to contort their data representations to build, say, “structs of arrays” that are highly efficient on modern CPUs (and trivial to express in C!)
I’m very curious to hear of any recent languages that are designed expressly with “mechanical sympathy” in mind.
[2] https://ispc.github.io/ispc.html#the-ispc-parallel-execution...
[3] https://ispc.github.io/ispc.html#structure-of-array-types
They jump through hoops so that C programs and the programs of all the other languages designed to run in the same execution model run as fast as possible. It's not to indulge C programmers but to support the vast body of existing software.
Hardware engineers need optimization targets just liker anyone else, and it's an eminently reasonable one.
Also, it's not like loads of improvements not related to the C execution model haven't been made. Demonstrably, the C model is not holding us back.
Anyway, if you want to move on from the C execution model, great. Now you need to introduce a practical transition plan, which should include such details as how and why we should rewrite all of the existing performance sensitive software designed to work in the old model.
We need to move beyond memory-unsafe paradigms already. Assembly intrinsics aren't great but they're certainly better than C.
Never use it, myself, except configuring the date format in my menu bar.
Because the speed advantage is so huge, people went through the pain of learning this new model and redesigning algorithms to better fit it.
So even with GPUs you have a similar situation..
The current microarchitectures for general purpose computing, that is OoO execution, cache coherent shared memory with a flat memory model, and mostly transparent and coherent caches seems to be optimal. In fact, far from being designed around C, C and siblings had to evolve to support the model well (a memory model, explicit SIMD builtins, support for vectorization and offloading etc).
It is entirely possible this is only a local optimum, but I have yet to see a plausible model for a better architecture.
It is more likely that, now that silicon is cheap and CPU designers are struggling to find ways to use it, extensions to support higher level languages might be added: more fine grained cache control for message passing, extra tags to help GCs, and more stuff that I can think of.
However, when I'm carefully using C++, I don't have to fiddle around with memory management and still get fast performance and better type checking.
Yesterday I wrote a Python script to generate C code test cases. String manipulation in Python is very easy.
The right tool for the right job, I suppose.
For example: I understand why the standard was written the way it was but syntactic closures in C++ are way worse than GNU C.
[1] https://gcc.gnu.org/onlinedocs/gcc/Nested-Functions.html
Although I’m not sure I agree about the “ergonomic” thing.
I remember GNU GRUB going through and getting rid of all of them.
I'm not familiar enough with ARM - what does ARM Cortex give you here?
I saw a talk about weird machines. The upshot of it is if you can clobber the return address on the stack you can usually create a weird machine that you can then program[1]. Leads me to believe that if you are putting entrusted data and return addresses on the same stack you're in a state of sin security wise.
https://en.wikipedia.org/wiki/Weird_machine
[1] Also saw an exploit on an embedded system where they leveraged that even though the memory was locked you could set the program counter and read/write registers via jtag. Using that they were able in short order to recover the devices security keys and re-flash it with their own code.
Eg, https://news.ycombinator.com/item?id=21635551 with alternating ro+x and rw-x pages holding thunks and fptr+data pairs.
I have a list of the structure members. I generate an assert for each member being a certain value. The python script builds/runs the C code. On failure, I parse the assertion failure, get the actual value, change the C assert string, rebuild-rerun. Continue until the program succeeds.
I've visually verified the decode is correct once. I want to keep the decoder tested if I change the code again so I'm generating complete test coverage.
Note that Python's official type checker (mypy) is pretty good (within the limits of type erasure); I started the project I've been leading up at work this past year in statically-typed Python, and it's been a great experience.
https://docs.python.org/3/library/typing.html#typing.TYPE_CH...
JSON = Union[Dict[str, JSON], List[JSON], str, bool, int, NoneType]
TypedDict isn’t for arbitrary JSON, but it is really cool.Looks like there is a work in progress though[1] and people in the ticket provided some workarounds[2][3] for now.
[1] https://github.com/python/mypy/issues/731
[2] https://github.com/python/mypy/issues/731#issuecomment-53990...
[3] https://gist.github.com/catb0t/bd82f7815b7e95b5dd3c3ad294f3c...
I do this quite a bit. Basically, anytime there is something very boiler-platy (like HTTP APIs, configuration management, test cases, etc.) I write metadata in JSON that describes what I'm adding, and then at build time a python script will parse the metadata and generate all of the C or C++ which is then compiled. I even have the python generating comments in the generated C/C++!
I used to use Jinja2 for generating the C/C++ files, but now I just use f-strings in the latest version of vanilla Python3.
{
"name": "backoff_mode",
"group": "telemetry",
"description": "We don't want to clog up networks with tons of failing attempts to send home telemetry. This parameter configures the device to back off attempts according to the specified mode",
"cpptype": "std::string",
"enum": ["none", "linear", "exponential"],
"default_value": "exponential"
}
At build time, a few things are generated, including a .h file with the following: struct telemetry_config {
...
std::string backoff_mode = "exponential"; //Possible values: ["none", "linear", "exponential"]
};
The boilerplate for getting and setting this parameter is also generated. Thus, just by adding that simple piece of metadata, all the boilerplate and documentation is generated and you as the developer can focus on the actual logic that needs to be implemented regarding this parameter.There is nothing wrong with generating code in a pragmatic way. It is just strange to see all the time people using workarounds for things that a tiny DSL could probably solve better.
It's just equivalent to a preprocessor.
A small one-off run under your control is an exception. But people learn to fear the thing, and apply that fear everywhere.
> A small one-off run under your control is an exception. But people learn to fear the thing, and apply that fear everywhere.
Why does it usually suck? Writing code that writes code seems like an optimization that would bolster productivity.
One large factor is that when somebody really dislike the result of a code generator, they most commonly cope by editing the generated code by hand. With other metaprogramming techniques people can't do the same, so they fix the issues.
The end result is that code generators are usually a set of mostly-functioning tools that create bad code that is almost, but not exactly impossible to understand but that you will be required to change at some point.
Few years ago it fully rewritten from C++ to C.
Can't answer for now, but check next awesome lists:
• AWESOME C — https://github.com/Bfgeshka/awesome-c
• AWESOME C — https://github.com/kozross/awesome-c | mirror — https://notabug.org/koz.ross/awesome-c
• AWESOME C — https://github.com/aleksandar-todorovic/awesome-c
∗ AWESOME C++ — https://github.com/fffaraz/awesome-cpp
Modern C apps with a GUI typically build on GTK but that toolkit has become massive. Modern GTK3 apps are built on top of dbus, pango, atk, cairo. This is not necessarily a bad thing though and libraries like ATK provide accessibility which is important. At the time GTK+ was created the competing toolkit at the time, under Unix, was Motif and it had the nickname "bloatif". Now we've reached a point where GTK3 is much much larger and "bloated" than a 2019 Motif application.
If we include C++ the Qt and WxWidgets toolkits are also massive. FLTK is still a lightweight option though.
Yes, I was also meaning the ui, which in this one is amazingly fast. I tried a few filters and they seem fast too, but being not a gfx expert I have no way to judge. But the ui is incredible; same speed on my home PC which has a mechanical disk. Making an external library of its gui primitives and functions could be an interesting project to be used on small embedded boards.
Think, AzPainter itself already could be used on small embedded boards ;)
For source code of AzPainter's `mlib` toolkit look here:
- https://github.com/Symbian9/azpainter/tree/master/mlib
Since AzPainter v2.1.3, `mlib` sources shipped with AzPainter now licensed under GPLv3 terms, but there is AzPainter theme editor (older `mlib` minimal demo app) which sources licensed under BSD terms:
- http://azsky2.html.xdomain.jp/linux/mthemeeditor.html
Also, there are few other apps based on `mlib` toolkit
- http://azsky2.html.xdomain.jp/linux/azpainterb.html
This way you can have the speed and memory tightness of C++ where it matters, but do the other stuff, like initial data parsing and munging or overall control and sequencing in Python, thus avoiding the general clumsiness and unproductivity of C++
Static typing is a superior approach that makes maintainability and refactoring orders of magnitude easier.
Much better type system than C, syntax is very similar to Python, compile times are excellent and speed analogous to C++.
Python is a great tool for a certain class of jobs. I saw a company doing systems programming in Python and originally thought it was a great idea, and then I learned of the hidden(or not) dangers of Python.
Yes, you definitely need to pick the right tool.
The code takes one file and outputs another, it's a classical 'unix filter' and it relies on a couple of very high performance libraries that others provided, battle tested and ready to be plugged in. As the project matures I come across the occasional wart in the language, something that I know I could have done more elegantly in a language with richer constructs (lists, list comprehensions, that sort of thing).
What surprised me most after working on this project for a couple of months though is how well the language fits the problem, that's something that I did not really think would be the case when I started out with it. But as time passes and things 'find their place' it is indeed just like the author writes, I - still - love coding in C. Even if I know there are better alternatives out there and I should probably get invested in one of them one of these days at the 75% mark or so of my career I'm still very happy that I picked C at the start and never once did I imagine that it would still serve me this well 37 years later.
Most of the other 'hot' languages of the days that went by in the meantime have all been lost, and yet, C is still here, and likely will still be here for years to come. To me it's like an old knife. Not pretty, known to hurt you if you abuse it but still plenty sharp and well capable of doing the job if you treat it well.
It is therefore irresponsible to use C for a new project like this, unless you know your code will never be exposed to malicious input. (Note that if you share the code with anyone else, it's hard to have such confidence.)
So, given all this doom that hangs above my head, what should I have used?
The powerful libraries growing up around it make dropping down to the dangerous level mentioned largely unnecessary, without giving up performance.
If you want to avoid objects containing references (or raw pointers), then you have to avoid using lots of standard library stuff --- `string_view`, C++20 `span`, or even STL iterators.
It is very likely that Rust would be a good choice, especially if the libraries you used have Rust equivalents.
As for libraries, the most important ones that I rely on deal with the loading of .wav files, writing of midi files, decoding of mp3s and fftw ( http://www.fftw.org/ ).
There are Rust libraries for all those things (including Rust bindings for FFTW itself), though I can't speak to their quality. I can confidently say that using Rust libraries in a new Rust project is a lot easier than using C libraries in a new C project, especially if you're not on Linux.
I've yet to come across anybody outside of the HN crowd that knows about or uses Rust. Not a single company that I looked at in the last year used Rust, that's 41 of them and I asked every one of them which programming languages are in use.
> Compilation speed is a problem for some projects, but not for small projects like yours
ECT cycle (edit, compile, test) is a pretty big factor during early stage development. Running a test in a second versus running a test in ten seconds would ruin my day. I'm not sure what the current speed of compilation is for Rust but the only time that I looked at it (very early on) it was so slow as to be unusable from my perspective. I'd hope that has substantially improved. Think of what I'm doing as explorative programming, I'm both trying to understand the problem and trying to solve it at once.
> Understandable, but writing exploitable code because you don't know a suitable alternative language very well is not a good long-term situation.
That may be true. At the same time, Rust may not be a good long term solution either, languages come and languages go, and before I invest a year or so into a new eco-system I'd like to see it has staying power. It is interesting that you recommended Rust and not say Java which has pretty good performance, is memory safe and is in production for well over a decade, for this particular use case I would choose that over Rust.
> Define "immature". It's not one of the world's most mature languages, but it is definitely mature enough to write production code for many contexts.
There isn't a month or we see an announcement of the next point release of Rust on HN. I don't have the luxury of tracking a moving target next to working a full time job and having a family, then doing this hobby project besides. If Rust wants to see mainstream adoption for stuff like this because you seriously believe that writing a new project in C is irresponsible (which smacks of FUD, and is one of the reasons why I think Rust is one of the more toxic language communities on the internet, I have never seen a Java proponent use that kind of language) then I hope you agree that getting Rust to some kind of long-term stable form in the very near future is a must. This whole 'the sky is falling' tactic is rather off-putting, at least, it is to me. The same happened to perl, which was supposed to be - and still is according to some - the best thing since sliced bread.
> There are Rust libraries for all those things (including Rust bindings for FFTW itself), though I can't speak to their quality.
That's good to hear.
> I can confidently say that using Rust libraries in a new Rust project is a lot easier than using C libraries in a new C project
For you, as a Rust user, yes. But you are forgetting that for me that is not at all easier, and that using C libraries in a new project is a lot easier than using Rust libraries for me because I happen to know how that works.
Vantage point is a bit of an issue here, I take it that you are fluent in both Rust and C and that you have decided to hitch your wagon to Rust. I'm fine with that and no doubt in the long run you will be proven right but to me it smacks of 'you should do as I do' rather than that it is the best for me.
Learning a new language just for some project is a very high degree of friction to add, it would slow me down tremendously and it might lead to the project being abandoned rather than progressing at a quite acceptable speed given my constraints in time.
> especially if you're not on Linux.
I wouldn't dream of developing software on anything else, but I totally respect other people's choices in their platforms.
Interesting. High school students and university students that my kids know are using Rust (in New Zealand, not some high-tech mecca).
> Think of what I'm doing as explorative programming, I'm both trying to understand the problem and trying to solve it at once.
OK. I might recommend Julia then.
> It is interesting that you recommended Rust and not say Java which has pretty good performance, is memory safe and is in production for well over a decade, for this particular use case I would choose that over Rust.
Sure, Java sounds like a fine choice. I didn't know much about your requirements when I suggested Rust. People often choose C because of strict performance or deployment constraints, which Java often can't meet, which makes Rust often a safer recommendation.
> There isn't a month or we see an announcement of the next point release of Rust on HN.
That's because they release every six weeks. That doesn't mean the language is unstable or you will keep having to make changes to your code. It means there's a steady flow of incremental improvements that preserve backwards compatibility.
> you seriously believe that writing a new project in C is irresponsible (which smacks of FUD, and is one of the reasons why I think Rust is one of the more toxic language communities on the internet, I have never seen a Java proponent use that kind of language)
In hindsight I shouldn't have mentioned Rust at all. I do sincerely believe that propagating C code is irresponsible and the software industry needs to recognize this. It sounds harsh, but I think that's partly because software developers have historically taken too lightly the consequences of their choices (even for hobby projects).
People like me who are zealous about the industry moving away from unsafe code tend to be big Rust fans, because Rust finally makes that possible for the systems/embedded domain where C/C++ were for such a long time the only viable option. I guess we have skewed the Rust community to some extent. Mea culpa.
> it smacks of 'you should do as I do' rather than that it is the best for me.
When we write code and distribute it, we have an effect on the world, and then I think it behoves us to consider more than just "what is the best for me".
For hobby projects where the code is never distributed, these issues are mostly moot ... though it's amazing how often such projects escape sooner or later.
I believe that you are sincere in this. At the same time I ask you to recognize that the way you - and many other proponents of 'safe' (safe between quotes because computer programming will never be safe) languages approach this is combative and therefore ultimately un-productive.
The old proverb says that you will catch more flies with honey then with vinegar and headbutting with people and calling them irresponsible is not going to get you where you want to be. It will have the exact opposite effect, people will dig in and their resolve will strengthen rather than weaken. This is because people tend to be invested in their tools and their creations. Going for a full on frontal confrontation about this is counter productive, that's just human psychology, which is in cases like these as important - if not more important - than having a technological edge or being right.
Zealotry is the last thing you want to take with you in your toolbox if you want to make change.
To be fair other languages where there first, they just lost due to UNIX uptake across the industry.
Had any of those language been more sucessful regarding adoption across mainstream OSes, and a language like Rust wouldn't be required.
GC is a problem for OS/embedded programming. And before Rust, safety required GC (or an unpalatably restrictive programming model like MISRA). Even if Symbolics or the C# Windows Vista work had been more successful, Rust would still be needed.
You also wrote: "Not pretty, known to hurt you if you abuse it but still plenty sharp and well capable of doing the job if you treat it well."
To me it seems you know it can be dangerous, but like it too much not to use it.
I could have chosen Java, Python, Go or C++ instead. Instead I chose the tool I'm most familiar with because that will get me focused on the problem rather than on the tool. It's a disadvantage of getting older: there is less time left to waste.
> To me it seems you know it can be dangerous, but like it too much not to use it.
It is mostly a familiarity thing, not a like or dislike.
The ones who like C and C++ are like Toretto from the Fast and Furious franchise choosing a classic muscle car for a race. They are super fast, give you much more options, but because of that, they are not the safest.
A muscle car requires more skill to drive and win than a modern fast-shiny-and-safe car. But because you know they are more dangerous to drive, is likely that you will try to improve your skills to make it less likely to have an accident with it.
Thats one of the reasons why im still skeptic to jump into the Rust bandwagon, once you turn into a good muscle car driver, its hard to choose for safer but more limiting options, but you totally get it why someone without experience in fasters cars would choose for the safer version.
(If you want to write C full time professionally too, contact me!)
I've gotten into the habit of pushing memory allocations as far back into the user program as possible (where the user asks for the size of the internal struct, allocates and casts, and deals with ownership himself) to allow more flexibility to the users of my libraries, and also to remove entire classes of memory issues from the libraries themselves: https://github.com/kstenerud/c-cbe/blob/master/tests/src/rea...
As a bonus, it avoids the headaches of mixing multiple allocators (malloc, new, [NSObject alloc], JNI, etc) when used in cross-language codebases.
It does, unfortunately, complicate the API a little bit, but I find the tradeoff to be worth it in terms of safety.
The most common issue I see in C codebases is excessive memory allocation. In most other languages, allocating a chunk per instance is the only way, though the implementation likely does some kind of pooling/slab allocation behind the scenes. In C you have the freedom to arrange your allocations any way you feel like, and treat any segment of N bytes as a value of size N. As an example, most collections (vectors, sets, maps etc); tend to only work with pointers; which is a huge missed opportunity.
My two favorite kinds of collections in C are embedded linked lists [0] and deques. One reason is they both provide stable pointers (which is the main disadvantage of vectors), the other that they minimize the number of memory allocations.
I recently implemented a pool allocated deque library [1] in C that uses embedded linked lists to keep track of memory blocks and provides value semantics, which I hope will make my interpreters easier to implement as well as increase performance. Time will tell.
[0] https://github.com/codr7/libceque/blob/master/source/libcequ... [1] https://github.com/codr7/libceque
But I think "getting close to the lower levels" is still great. The only time I write C-like code is for an Arduino. Turns out that even if your electronics is rusty, you can do some very useful things in the real world by automating a simple motor and a few LEDs. And whenever you have to interface with individually numbered output pins, it feels pretty low level.
tldr: if you want to get close to the metal, it's really useful to start by switching to hardware that facilitates this.
It's not like playing with an Arduino is getting me closer to understanding the underlying chemistry and physics. Which are also interesting to learn about...
But personally I think it's for the best that reality is inexhaustibly complex (from the perspective of human individuals).
Since you mention the Arduino, when you start out with it, you're most likely turning on digital pin e.g. D7 with digitalWrite(7, true), it's easy to understand and quick to implement.
But spend enough time on the Arduino (or C or Assembly or embedded in general) and you'll soon realize that all that function call is doing is basically the same as doing PORTD |= 0x80. That's a bit tougher to understand, but even quicker to implement, and once you've understood how that all works together, a lot of everything about computers and digital electronics becomes so much easier to understand.
Here's a more detailed guide for the Arduino example: https://www.instructables.com/id/Fast-digitalRead-digitalWri...
If you want to be reliably close to the metal, I think the best options are either to actually write assembly, or use a language with a simpler abstract machine (preferably, no undefined behavior), that statically compiles to code that maps closely to what you wrote, and lets you generate "idiomatic machine code" (e.g. unboxed integers). Safe Rust fits the bill, but there are probably other more obscure options.
I decided to work on this year's Advent of Code in C and it's been a blast so far. I finally "get" pointers, malloc/free, etc. (and I have a newfound appreciation for Ruby things like CSV.each).
Since most of my coding work is at a much much higher level (web backend & frontend stuff), I probably won't spend too much time in C in the regular course of business, but it's massively helped me better have an intuitive understanding of previously vague things like "allocate on the stack vs. heap" and "pass by reference vs. value" -- I knew what those words all meant but didn't really have a feel for what they actually did. C helps with that.
Go is nice, but tends to be very verbose and also is garbage collected. It's for a different usecase.
Edit: Rather than responding to all the comments individually: There's obviously a lot of code already written in C. That doesn't mean it makes sense to write new code in C. Sometimes C is your only option, which I already said in my original comment.
None of that means there is any reason to start a new project in C if it's going to run on Windows or Linux, and it's certainly not the case that you have to write your program in C if you want to call it from another language. Sorry if this disappoints anyone, but it's true. For the traditional uses of C, such as writing a text editor or a library for numerical computing that you're going to call from another language, there's no longer a sensible argument for using C.
Linux (like the BSDs) being C is an organizational phenomenon. There is no plausible technical reason for it not to be, or include parts in, C++. History is path-dependent; how we got here matters more than what is around today.
There is really no excuse for systemd being C. It is really a drag on its development.
On the OSes that aren't just plain UNIX clones, C has a much less relevant role, and in some of them it is even being phased out, even if it might take a couple of decades yet.
Pascal UCSD also used a mix of interpreter and later AOT compiler.
Most mainframes don't use C as their original implementation language, and the oldest still being sold (Unisys ClearPath MCP) was designed about 10 years before C was born, in an Algol derived language, that uses instrics instead of Assembly.
If one broadens their horizons beyond UNIX, and looks into the history of computing there are plenty of interesting languages and OS architectures to dive into.
The Internet is full of digitalized version of papers and computer manuals from those days.
What happened instead was that everybody had their own Pascal dialect with the extensions it needed to be actually useful. Meanwhile, C had them all, already, and was about the same everywhere. Soon after, C++ came along, and there was again only one (although Microsoft delivered workable template support very late).
Lesson in Network Effect there. Also, shooting for usefulness out the gate.
Meanwhile, the Pascal crowd went off with Modula-2, Modula-3, Oberon, Apollo Pascal, UCSD Pascal, Object Pascal, Delphi, and countless others, none compelling. Apparently Delphi (still) has economic importance. The rest, not so much.
This part isn't quite true. All that's needed is to be able to define an exported function with a specific name and calling convention. Libraries can be written in any language that supports these features.
Even C++ is a pain in the behind, because everything that doesn't translate needs to be encapsulated.
As a result, it's very rare to find code written in different languages with exported C APIs from my experience.
Add name mangling, ABI issues and crazy compile times to that.
It was potentially safer until they added Exceptions, which tipped the scales so far the other way I can't believe we're still having this discussion.
But you won't when writing C++ unless you absolutely have to, because it's a major pita to encapsulate everything.
Entire books [0] have been written about exception safety in C++. Getting it right in combination with RAII, (copy)-constructors and assignment operators is very tricky. Which touches the main issue with C++ from my experience, mixing foot guns with higher level convenience.
Embedded linked lists [1] is a good example.
Seriously, live and let live. I'm fine with you enjoying and preferring C++, go for it. But consider taking a good look in the mirror and asking yourself why you feel the need to spread FUD about C to feel good about your choice.
[0] http://www.gotw.ca/publications/xc++.htm [1] https://github.com/codr7/libceque/blob/master/source/libcequ...
The key is to make sure all your cleanup code is in destructors where it will be exercised frequently as part of the normal operation of your program. In the event an exception is thrown, a bunch of destructors run, as usual, and your program ends up in a known clean state, with no extra effort.
Rust is the best candidate, but it operates on a completely different level. (That said, I don't think one ever need that much control on a PC or a server. Not even on kernel programing.)
There are other modern alternatives nowadays as well, including for MCUs.
C has not evolved because C++ has, and because the computers got a million times faster, so even Python is usually good enough. If you need something that does what C does, but better, C++ is right there waiting. Change your makefile, or rename some files, and suddenly you have a C++ program, that is then easy to improve as much as you like.
Well, it has, in a way: see the latest C standard and compare it with the Kerninghan&Ritchie edition.
That's what you think. Transistors have evolved a lot, as have programming languages.
It's the only language your chip vendor will give you, but there is very little reason ever to use their compiler, anymore. And they have no say in what you use to generate your machine code.
Anyone that is open mindend can get Pascal, Basic, Java, Oberon, C++ and Ada compilers, although I admit that the Ada option is only available for deep pockets.
Several mistakes in c:
Using string interpolation preprocessor macros makes it so that you have to 1) learn a whole second language to read other people's code (most portable code requires using #defines) and 2) it's very difficult to trace through code where you don't know where some values are coming from.
Header/code segregation seems like a huge mistake. It's far more effective to have a single file with well-defined exports. Even the #include system is broken, because it's not explicit or obvious which directories one must go to to find header files (they might not be local to your project)
Sigil order. The fact that there is a guide saying you should "spiral out" to understand how to read a type should be evidence enough that this is a huge mistake.
I get it that there are professional c coders that will say that's part of knowing the language. I'm not a professional c coders, I program in another language and very rarely need to dip into something lower, say for performance or mutability. Thus the c code that I should write is dangerous, I need safety, and simplicity. Luckily, I found another language that hits those spots.
Got as far as playing around with a web site that used Rustler NIFs on a Phoenix website. Why? It was just fun to figure out how to get something from the web frontend through Elixir and into some Rust numerical code.
I don't think it had much practical use and knowing me, would be ridiculous to go back and touch it again (likely). I've been thinking about doing Rust to WASM but I'm still not sure if it's just throwing a bunch of complexity at something just because. It holds my interest a lot more than just doing everything in JS, though.
C was also more enjoyable to me than other languages. Something about compiled languages just feels really nice. But, it's also great to have a REPL available right there. Like with Elixir.
And I like C (although I don't code in it), because Delphi/Lazarus can interface with C very easily, you know, there is a lot of C libraries around the world :)
C is fantastic for small and very optimized components. If you want it portable, from microcontrolers to GPUs, it is the best choice too.
But if you need rapid prototyping, large scale projects within an organization or some libraries and tools to make the perfect, animated, templated and themed user interface you should think twice before choosing C. Java and .Net are freaking good at database integration and memory allocation/dealocation. Erlang is fantastic for multitasking and multithreading. Ruby on Rails or Python on Django are much better to prototype a website nannying a database.
You already know the saying "if your only tool is an hammer..."
The bottom line is to echo your first statement, C is a tool, good for some tasks, not so good for others. Choose accordingly.
#include <stdio.h>
int main() {
char a = 3;
char b = 8;
char result = a * b;
printf("%u %u %u %u", sizeof(a), sizeof(b), sizeof(result), sizeof(a * b));
return 0;
}
PRINTOUT/x86-64: 1 1 1 4Surprising results on x86, probably not on Arduino or embedded platforms.
The phenomenon is integer promotion in C.
EDIT: "Char isn't an arithmetic type!" The same is true for short, too.
EDIT2: The consequences are: Might need to manually %= 0x100 instead of relying on data type's inherent modulation in multi-operational arithmetic, making rotation a necessarily pre-scan operation for LE/BE conversion instead of a process-transparent operation and I had a third one... hm. Also, maybe speed loss?
If you want to write a new programming language that does the same from the get-go, make C as one of the compilation targets.
As other commenters have pointed out, use the right language for your use-case. Then, I agree that exporting a C interface is best.
"Make sure you have homebrew installed and then `curl http://goodluckbro.io/doomed | sh` then go read the Medium blog post to write your first hello world"
And don't get me wrong, I love and use whatever-shiny-lang-v2 too, I just really like the simplicity of C (and yes I acknowledge that comes with the ambiguity and undefined behavior too).So... what is the equivalent of doing this in C?
* install C compiler (no, there is no install script, go google and install yourself and yes, the process is different on different platforms)
* read a blog post (possibly on medium) and write hello.c
* read another blog post to figure out compiler flags
* ./a.out
A bit less convenient, but not that different
I think you can argue for both sides of the argument, but what is the obsession for "the real"? Is the JavaScript running in your browser "less real" than the x86 Assembly that is compiled once again to Intel Microcode?[1]
> When I saw a memory address outputted to stdout, I felt that I was as close to the hardware that I could be without literally opening up my machine and stroking my ram modules like a crazy person.
Sorry to shatter your enthusiasm but the memory address you are seeing is very likely to be "virtual" and not physical thanks to paging: yet another damn abstraction![2]
-----
> "If I have seen further, it is by standing on the shoulders of giants." --Isaac Newton
The Magic of Software: How I Learned to stop Worrying and Love Abstractions
http://www.thecaucus.net/#/content/caucus/tech_blog/373
----
Because when the shiny abstractions and bloated frameworks layered like a Guinness World Record stack of pancakes† that attempt to hide all the underlying complexity of the hardware break down and leave you with a problem, be it CPU, memory, or I/O, then anyone without an understanding of what's going on at least at the assembly and architecture level will be up the proverbial creek without a paddle. I've made good money bailing people out of these kind of problems.
† About 60 cm or 2.5 feet apparently. https://www.eater.com/2012/2/21/6611971/this-is-officially-t...
I know you can find bolt ons, but they aren't terribly useful since no 3rd party libraries use them. You have to resort to stuff like antirez's sds.
That's not to say that the rest of C is bad, but there's nothing to be gained from defending C strings; they're bad and you should probably use an alternative and that alternative should probably have a length and a buffer, even better if you can find this out of the box in your language that otherwise works a lot like C.
typedef struct {
size_t size;
char *buffer;
} string_t;
“un-C”.That then implies at minimum a sanitation of user- and data-input of char array, and maybe an absolute maximum overrun char count protection at execution, but once your default is to keep track of every string length and not the opposite of keeping track of specific strings, I'd kinda agree with above that it's not really "C"... more like your own implementation of a(n overrun-safe) string library?
Inherent to the concept of the char array is that a size value would be redundant, for it is the position of the null-terminator that already contains/is this.
The analogy for the string concept of a char array's missing null-terminator is actually a wrong size value.
I ask myself, why are strings safer?
Nope, plenty of alternatives have existed throughout the history of computing
https://en.m.wikipedia.org/wiki/System_programming_language
> C as the Ur-language
Not really, ESPOL came up in 1961, while C only became a thing in 1969.
"Physicality" was setting CPU register bits directly via the I/O panel of my UYK-7 back in the day.
Happy not to be cruising at that altitude these days (to use an equally inappropriate metaphor).
The article clearly states this is a subjective statement and isn't meant to be taken literally.
Zig made some moves to that direction but not yet enough and is not really ready for production anyways. Also I did some experiments and initial impression was good but it failed to compile used libs written in C even though it claims to be a C compiler as well.
There are lots of small towns still emptying out. The people who remain are more likely to say they like them.
When I was learning C++, a popular idiom to employ was to use it as "a better C". Don't make your own classes, but leverage good "string" impls, templates, namespaces, declaration semantics, RAII, etc, basically stuff that just your C... better.
I can appreciate good nostalgia, but I can't help but think i'd be pining for some of those betternesses that came along later.
Felt like I finished reading the intro here, only to realise I scrolled to the bottom with only 2 paragraphs to read more. Guess I was expecting a bigger read, perhaps with some code references, as it was interesting :)
A: “C tbh”[0]
For actually getting things done there’s go and python of course.
Doing it in C obviously. I can’t really consider Go embedded because I’m maxing the chip out now in C without including a Go VM or whatever you call it.
This isn’t all as rare as you might think.
If every feature everyone wanted was added to C, it would just turn into a bloated mess like C++, or like PHP, with tons of bad decisions baked into the language that can't really be avoided. Features like this belong in libraries - the stdlib should remain as minimalistic as possible.
The C++ evolution is essentially an iterative cycle where people hackily create new language features with insane preprocessor and template hackery, and then the Standards Committee comes in and takes the most popular hacks and tries to turn them into real language features. And then people create new hacks on top of that.
C says, you want object oriented polymorphism, you've got structs and function pointers, what more do you need?
No, C says your first mistake was wanting object oriented polymorphism. Your second mistake will be implementing it.
These programs, by the time I came aboard and started sharing code with those to make the games I was involved in, had also "succumbed" to this force.
The shared bedrock of all these games was code that used certain Z80 registers as pointers to take advantage of instructions that read or wrote fixed offsets from those pointers. The first byte of the record pointed at by the pointer was the type and the next 4 bytes were the position. The rest of the record depended on the type.
That was us doing the best we could with 8-bit processors when our competition was using 68000.
But the main issues that I've had with Go is when the Go runtime is causing your code to very much not act like a C program. To be fair, I work on container runtimes and similarly "lower level" tooling so this is probably not an experience that everyone else has.
"Oh but it's powerful", yes, a chainsaw is powerful but even that has an emergency brake. C has none of that.
(Not to mention compilers that take "undefined behaviour" as a code word for "bamboozle time")
So I have to hope for the best when some piece of hardware, coded in C, gets plugged into the network.
Morris Worm is 30+ year old, and yet best we can do is having platforms like CHERI, Solaris SPARC ADI, ARM MTE, which also happen to be quite specific.
So you'll be hoping for the best for some time to come, no matter what the stuff you install was written in.
Σ UB + Σ memory_corruption + Σ logical_errors ≥ Σ logical_errors
But every language that can reach the hardware will have the ability to wreck things in spectacular and hard to predict ways. Every new language ever was touted as the one that would finally solve all our problems. For Java and COBOL the historical record borders on the comical. I have no doubt that the same will go for every other language that we just haven't found the warts in yet. Two steps forward, one step back, that seems to be the way of the world in the programming kingdom.
Sending digital signals over analog medium (read: always) can fail, and may do so regularly depending on hardware itself and environment conditions.
Simple examples: Digital 3.3V signals sent over 10m, high-capacity lines (lose a bit here and then), or an overclocked CPU undervolting just at the right time, etc...
EDIT: I'd also argue that UB is a type of logical_error. Not the fact that UB exists within the language, but in that the logic failed to account for the scenario?
UB is its own kind of error, specially since ISO C documents over 200 use cases, and unless you are using static analysers, most likely won't be able to find out that are failing into such scenarios as no human is able to know all of them by heart, whereas logical errors are relatively easy to track down, even by code review.
Unfortunately liability is not yet a thing across the industry.
https://cve.mitre.org/cgi-bin/cvekey.cgi?keyword=stack 3496 entries https://cve.mitre.org/cgi-bin/cvekey.cgi?keyword=pointer 2389 entries
https://cve.mitre.org/cgi-bin/cvekey.cgi?keyword=java 2034 entries
A possible outcome is that you would trade pointer bugs etc. for "Java bugs" if Java embedded were used everywhere. Embedding a complete runtime increases the attack surface alot.
Many of those Java exploits are on C and C++ written code layer, yet another reason to get rid of them in security critical code.
According to Microsoft Security Research Center and Google's driven Linux Kernel Self Protection project, the industry losses due to memory corruptions in C written software goes up to billions of dollars per year.
Several reports are available.
Would you like to tell us how it compares to e.g. Java's runtime as well?
Then, unless we are speaking about a non-conformant ISO C implementation for bare metal deployments, it provides the initialization before main() starts, floating point emulation, handling of signals on non-UNIX OSes, VLAs.
I have been at this since the BBS days, I don't care about Internet brownie points.
The only thing left is actually having some liabily in place for business damages caused by exploits, I am fairly confident that it will eventually happen, even if it takes a couple of more years or decades to arrive there.
I would never think of "floating point emulation, handling of signals on non-UNIX OSes, VLAs" as anything resembling a "runtime". These are mostly irrelevant anyway, but apart from that they are just little library nuggets or a few assembly instructions that get inserted as part of the regular compilation.
By "runtime", I believe most people mean a runtime system (like the JRE), and that is an entirely different world. It runs your "compiled" byte code because that can't run on its own.
Dishonest is selling C to write any kind of quality software, specially anything conected to the Internet, unless one's metrics about quality are very low.
Assume whatever you feel like.
> 70% of the vulnerabilities Microsoft assigns a CVE each year continue to be memory safety issues
https://msrc-blog.microsoft.com/2019/07/18/we-need-a-safer-s...
If you are going to argue that is 'cause Windows, there are similar reports from Google regarding Linux.
Punting all of this extremely important reasoning to "the programmer" is ridiculous when it's been proven over and over again that even incredibly smart C programmers get this wrong.
Why not disallow undefined behavior and instead have the compiler emit: "Hey, this branch is never taken, if you want us not to compile it remove the dead code". Seems like that's exactly the same sort of putting the onus on the programmer that andrepd mentioned, but without the side effect of millions of device drivers are zero day remote code execution vulnerabilities waiting to happen?
Even Rust made a mistake with restricting undefined behaviour -- they accidentally allowed creating references to fields in packed structures (resulting in possibly unaligned references which is undefined behaviour in Rust) in safe code. These days you get a warning telling you to use unsafe or copy the value, but they have yet to make it an error. For comparison, C also gives you a warning for the same problem (or an error with -Werror).
I also want to point out that modern compilers have different sanitisers (-fsanitize=...) which cause your program to abort() on memory unsafety or other violations. It's effectively a better valgrind which is compiled into your program.
I don't know how anyone who has studied the history of vulnerabilities and security issues and code would think this is a hard tradeoff:
> If yes, your code is now slower but there isn't an easy way to disable this "feature" (if x is heap-allocated or passed as a pointers, how will the compiler know what kind of bounds check is sufficient?).
> If no, you have undefined behaviour and crashes.
The latter choice has resulted in how many millions or billions in losses, perhaps even lives lost?
Okay, maybe you can add a small overhead to array accesses (which is what Rust did), but you can't stop undefined behaviour entirely. Rust has plenty of undefined behaviour (most can only be triggered through unsafe -- which is the whole argument behind the language -- but as I mentioned there are some cases where even they made a mistake and made an operation safe when it should've have been).
In C/C++, undefined behavior and unsafe operations abound and it is nigh impossible to isolate what is "important" to review from a security standpoint.
You can't "disallow" all undefined behavior in C without fundamentally changing the language to something different (Which would also likely make it useless for it's intended use-case). Some of the details related to undefined behavior are simply not possible for the compiler to detect without extra information, hence why they are undefined. `Rust` has the same thing, you can invoke undefined behavior anywhere in the program through use of `unsafe` anywhere else. And Ex. How is the compiler supposed to determine if a particular function is ever called with a `NULL` parameter? How does it know if a particular signed addition may overflow? How does it know if two parameters to a function alias in an invalid way? How does it know if a particular pointer lacks a `const` and is actually backed by read-only memory? And most importantly, how does it determine all of this at compile time?
With that not all undefined behavior-based optimizations are unexpected or warrant a warning or error (In-fact, I would say at the very least the majority of them don't). Undefined behavior is much more complicated then just dead-code removal. Even then though, you don't always want dead-code removal warned about - Imagine getting a warning about every inline function or macro that results in some dead-code.
To be clear, I'm not saying there aren't obvious problems that have come from this approach, and I think some of the choices for undefined-behavior are clear mistakes now even though they made some sense in the 80s or 90s (And some of those can usually be turned off, thankfully). And I think there are also some additions that could be made to the language that could solve a lot of the existing problems (Like non-null pointers, or fat pointers, though I'm not exactly holding my breath...), but the existence of undefined-behavior is not in-and-of-itself the actual problem with the language, and "disallowing" it outright is simply not possible. I would actually wager undefined-behavior exists in some form in a lot more languages than you'd expect, just less documented. Rust certainly still has it, despite it's numerous safety-focused features. And even if you add things like fat pointers to C that would clearly reduce possible out-of-bounds mistakes, the language would still need to provide some way to create one from a size and a pointer, and that still allows for undefined-behavior cases to happen - my point being, while adding such a thing is still clearly better, it does not actually remove the undefined-behavior possibility.
It's not very expensive because the branch will never be taken except when it results in an out of bounds access. When that out of bound access happens you care more about the security benefits than performance.
I'm curious if this still stands in general today, given today's advanced branch predictors and the fact that the CPU tends to be memory-bandwidth-bound, thus you have more "free computation" while waiting for memory (what GPU programmers call compute to memory ratio).