“C is how the computer works” is a dangerous mindset for C programmers
words.steveklabnik.com
words.steveklabnik.com
> Tools like [Valgrind] are exceptionally useful and they have helped us progress from a world where almost every nontrivial C and C++ program executed a continuous stream of UB to a world where quite a few important programs seem to be largely UB-free in their most common configurations and use cases...Be knowledgeable about what’s actually in the C and C++ standards since these are what compiler writers are going by. Avoid repeating tired maxims like “C is a portable assembly language” and “trust the programmer.” Unfortunately, C and C++ are mostly taught the old way, as if programming in them isn’t like walking in a minefield.
I’ve mentioned before here, but when I found a reasonable number of bugs in critical library used throughout one of the FAANGs written by their most senior and brilliant engineers by running a fizzier on it for a few hours, my opinion of those languages changed.
Integrating Valgrind and Clang’s checker, and running AFL for many CPU years was critical in unearthing some really funky crashes, hangs, overflows. Without those tools it’s not clear that those bugs would have been fixed. In fact, one of them affected me a few months prior and AFL helped uncover the codepath that was failing.
Out of curiosity, how many bugs you found by using these tools could have been avoided by using a “watertight” memory management system [0], with strong decoupling of pointer and object lifetimes?
[0] https://floooh.github.io/2018/06/17/handles-vs-pointers.html
I think almost any JITed memory safe language will be faster than using handles for all object access. At least Java, .NET, JS, etc under the hood can avoid "double dispatch" of memory access. And you can use things like arenas to ensure same objects are allocated adjacently, etc.
[1] With great power comes great responsibility
Who cares what your tool for solving the problem is. Use Java, use JS, use python, go, rust... whatever suits for the task. Why do you want to compare yourself to C. Just stop it. No one with enough expirience in coding will ever care what language you have used to acomplish the task. Unless the task is going to be purely written, slow, buggy, CPU consuming, will have security issues and so on. This can be acomplished with half-baked (I hate the expression, but sorry) coding monkey in any language. Why do you even care about C. It doesnt matter, all that are complaining about it in whatever way possible or compare to it, will probably never be involved in a task where C excels, so why bother?
---
jeffdavis: the problem is that people would like to use wrong tool for the job. And the fact that they cant frustrates them, blaming everything else except the fact that they picked the wrong tool. As you need expirience for picking the right tool. A lot of expirience or rather beeing fluent in all the tools. Which makes it a problem if you dont dedicate a lot of time to this.
---
jeffdavis #2: > choosing the wrong tool really is a great way to understand your tools and the problem more deeply.
YES! Exactly! I did mistakes, a lot of them, like pragma packing the structure and send it to system on other side of network in other endianness. Overwritting EIP number of times before I understood it. Or best one, translating command.com to my language using edit :D Used wrong call convention. But my goal was always to understand assembler. To understand C. To understand C++. And I have made quite a lot of code in all 3. Then learnt any language that crossed my way and it looked like interesting (Rust is still on my list) and used each in at least one project, hobby or not. Now everyone learns one or two languages, without any background knowledge and then starts to preach it, prove that language written in X is faster than language that X is written in and so on. Who cares! It makes me puke.
I don't think that's a fair assumption. You may be underestimating the number of developers who have a need for a systems programming language.
System language? Make it. I dont care, if it will work better than C for the task, I will take it. If it is just a newage nonsense wasting resources for the sake of deallocating and taking care about buffer boundaries as it is so hard to take care and track your memory I will skip it. I dont care if I take a hammer or brick for using a nail, I will use what works best and my requirement is that I dont need 20 people to hammer in one nail for the sake of no one getting hurt due to putting its finger between hammer and nail. Someone would rather have a safety inspection. Also a fine decision as far as I care. Not my choice but I couldnt care less.
---
Anyway, kids, downvoting me to censor what you cant handle, I am over 40 years old, developing for my whole my life. I have seen hypes, new "revolutionary" ideas, repacking old technologies as new, stealing ideas (I see this lately 24/7) at least 20 new "revolutionary" languages, cpus, mainframes, clouds, any possible way of another human trying to rip you off (google, fb,...), lies, deceit, preaching, evangelism,... Do you really think I will care about it? It is just sorry truth that you will have fun of enjoying kernels written in JS on a browser, on a linux kernel that no one develops any more. Have you seen movie Idiocracy? You really should. I cant wait for it to happen. And then will I ask you if you are sorry, that 20 years back you didnt rather decide to learn instead of copy/pasting and spitting over everything you are unable to handle. (your parent should told you that, but they didnt, I am sorry)
----
saagarjha: sure, ignore the part about censoring, that is not part of debate. About everything else. You get used to it in a same manner as using floats in js. But this is my expirience. It is completely ok, to have different one. Some people are memorizing decks of cards, i consider that something difficult. They probably dont. They have learned to do it. And I am so sorry if dont feel like I need to prove it to you. Anyway, thank you for the cycript.org link, the only problem is that i cant use userspace solution, but it might come handy some day.
> I am just hooking linux kernel calls, I will never take JS for this. Pure C.
Why not? Hooking native code in other languages is already fairly common: http://www.cycript.org
> Valgrind? I have a header with #define MALLOC/FREE/REALLOC/... which returns larger buffer with guards at beginning and the end while after 25 years of C I hardly do anything more problematic that this.
Valgrind does significantly more checks than your solution.
> System language? Make it.
"Systems language" is a very ill-defined term, often used to gatekeep people.
> If it is just a newage nonsense wasting resources for the sake of deallocating and taking care about buffer boundaries as it is so hard to take care and track your memory I will skip it.
Keeping track of your memory is hard. If you say "no, it's easy for me", I'd like to see a non-trivial sample of your code that doesn't have memory safety issues.
Engineers make bridges that don't collapse (unless not maintained for decades) and planes that don't fall from the sky (unless overruled by management). By the evidence, it is not easy for people who just can't be bothered to take the time to make anything correctly.
So, there are languages for engineers to use to make things where it matters if they are right, and languages for everybody else, where apparently it doesn't. Now, all we need is for engineers not to need to call into libraries not written by engineers.
Laziness may make the problem more likely, but even a disciplined team working on security-critical software can make mistakes[0].
[0] https://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2016-0777
A primary factor that determines whether a particular language is suitable for a task is its runtime performance.
> constantly trying to compare everything to C.
C has little runtime overhead and incredibly mature compilers, so it is an excellent target to use when measuring another language's runtime performance. "Close to C" is simply a synonym for "the language adds little runtime overhead".
Have a gnarly problem in a Python program that Haskell would be perfect for? Too bad. You'll spend more time figuring out how to transform the data and get the Python code to call into Haskell than you'll save by using the right tool. And in the process, the overhead, complexity, and bugs introduced in this process turn the right tool into the wrong tool.
And that's the reason why C is so important. It actually does work with pretty much any language, because it has a defined ABI, simple data types, minimal runtime, can work on any data formats unmodified, no GC. Very few languages have this superpower.
C++ mostly has that superpower, but some features don't entirely work with unmodified structures (RTTI and virtual functions both require adding magical fields to the class structure). Rust seems to have this superpower (does dynamic dispatch a different way than C++, so can always work on unmodified structures).
The JVM is an interesting case and the ecosystem of languages built on it do work better together because they have the exact same runtime (GC, etc.).
To summarize: we'd all like to use the right tool for the job. But we can't, except for the languages that go out of their way to make this effective, and that list of languages is very short.
I consider it significant code smell when a project has multiple languages, particularly when each language seems tied to a particular developer.
This isn't really true.
Trait objects are a double pointer: one to the vtable, and one to the data. C++'s dynamic dispatch uses a single pointer to a structure that has the vtable and the data. The memory layouts are very different.
Clarification: the memory layouts between C++ dynamic dispatch and rust dynamic dispatch are very different.
Rust imposes no requirements on the structure layout, whereas C++ adds some magic data. The magic data means it's awkward to make C++ do dynamic dispatch on a C struct pointer; but rust can do dynamic dispatch on a C struct pointer with no problem.
But if you want to be able to optimize both globally and locally, well that's just an engineering problem!
Two examples spring to mind:
1. Lisp. Lisp's metaprogramming facilities allow you to build the language into the specialized tools that you need, while still having the common runtime and base language.
2. Unix. There a hundreds (possibly even thousands) of little programs running on my Linux machine written in probably dozens of languages. For the most part, I have no idea (nor do I care) what language they are written in. That's because the Unix model puts a strong emphasis on the protocols for communicating between programs.
Unix is a reasonable example of a polyglot environment, but at a high cost. Lots of serialization/deserialization through text. That has a high cost in terms of bugs, complexity, lines of code, and inefficiency.
Yup, that's more or less the argument, though the argument need not only apply to lisp.
> Unix is a reasonable example of a polyglot environment, but at a high cost.
I'd rather say that Unix is an example of how to effectively drive down the cost of a polyglot environment to a point where it is outweighed by the benefits.
Now, of course the cost is still there, but I would argue that beyond a certain scale you cannot optimize globally on a single runtime (perhaps not even with something like lisp) and so the cost of global consistency is outweighed by cost of being unable to optimize locally.
To put it more succintly, you could not build a system as complex as Unix in a single language[1]
[1] Although Ala Kay's work at VPRI suggests you can, if you choose a sufficiently expressive base language. But even that may have it's limits.
Unix doesn't put a strong emphasis on protocols. It just says "everything is text, except when it's not". It's not very helpful.
OK, I probably misspoke by claiming it put an emphasis in protocols, but I think it's fair to say that Unix does emphasize component integration by sharing data instead of trying to integrate through a common runtime.
I do think that having everything be (mostly) text is helpful though. It's true that the data formats are often poorly specified, but that is compensated for by the fact text is a format that has a very rich set of tooling. text editors, regular expressions, parser-generators, etc. all make it possible to capture, analyze, and manipulate the text data exchanged.
Perhaps a better example of protocol-based integration would be the internet. It has it's flaws as well, but it's also enabled collective engineering projects on a vast scale.
Those mechanisms aren't textual, they are byte-sequence based.
It turns out that many text-based protocols are more popular than binary/object alternatives. The penalty for textual redundancy is negligible/acceptable for many of them. Where not, Unix allows people to use unix/create new mechanisms to implement specialized protocols (rpc, protobuffers,...).
The reason C is so successful is named ABI.
C (and C++ to some extend) can be used to produce library that can be used from any higher level language in existence. And this is strictly required in the heterogeneous world we live in.
If you have to write a system components, long lifetime, implementing a protocol / database / format / any-low-level-staff: You do not have choice, safety or not. It is going to be C or C++, because it is the only thing reusable in most other environment.
Python has pybind11, Java has JNI, Lua operate beautifully with C/C++, Node is itself C++, go has cgo, Ruby is in C, any proper programming language can interface with C or C++.
They are the only languages allowing that, and that is why they are still alive, very well alive, even if they are unsafe.
As long as new language authors will be more obsessed about supporting new fancy feature instead of providing a well defined ABI (C compatible). The situation will not change.
Rust might be the new comer in this area that has its chance. But for the time being, it is still too young: you have to map every Rust API to a C one manually to make it exportable.
This might seem like nitpicking but it's an important distinction because you can't assume that C data structures can be passed between different platforms (e.g. over a network or even stored in a filesystem).
Symbian did not had a C ABI for example, nor do mainframes, in ChromeOS or Android you also don't have a C ABI exposed to userspace, in Windows a C ABI also won't help you much if you need to speak COM or now UWP.
That's actually a great way to learn. I do that intentionally sometimes to force myself to understand a problem better. Sometimes it even works[1]!
I'm half serious here. Clearly when something is important and needs to be timely, choosing the right tool is important. But my comment is not sarcastic, either: choosing the wrong tool really is a great way to understand your tools and the problem more deeply.
... turns out to be an extremely tall order; the 2017-11-17 working draft of the C++ standard is 1,448 pages in PDF format.
At some point, I wonder when programmers who are required to care about correctness throw up their hands and say "This language isn't reasonably human-sized for a user to know they're using it correctly." At which point the immediate next question should be "Why don't you use something else?"
(In my experience, the only sane counter-answer to that last question is "I'm programming on an architecture where the {C/C++} compiler is the only one with any real expressive power that anyone has ported to this physical system." This is an answer I understand; most other answers will raise an eyebrow from me).
I did that over 20 years ago and have never looked back.
To be fair, that contains both language documentation and library documentation. The language specification portion itself is only ~400 pages, which is smaller than the Java Language Specification and JS's specification (both about 700-800 pages, although JS does include its [meager] standard library in there).
I manage a c++ team btw. Our app has to be fast.
Still, have you guys looked at alternatives -- and seriously evaluate them? Rust, Zig, Nim, D, others?
If you tell me "we can't afford to, we have too much work" then that's also a valid answer (for a while at least).
Good lord, if that's the kind of human capital that's involved in making next-generation languages, then I have no doubt what will still be used to write serious software in 2080.
(Hint: not Rust.)
For any particular combination of compiler, OS and computer architecture there are no 'undefined behaviors'; you only care about the concept if you want portability to a different compiler/OS/architecture.
So the idea that C++ sucks because of 'undefined behavior', so let's use some random language without an ISO standard or any portability guarantees at all is completely and utterly insane.
C++ alternatives still need to grow up to this kind of mixed language development, where I can have .NET code with C++ libraries and easily debug across them on the same Visual Studio session.
Same applies to Java and C++ development experience.
F.ex. when coding in Erlang/Elixir I can use an excellent bridging library between their VM and Rust. There's also an excellent support for working with C libraries. Not sure what else can be expected.
Maybe I am not reading you correctly and I apologise if so. It just kind of sounded like "the newer languages must be compatible with 20+ different ABI standards if they want us the older programmers working in C/C++ to adopt them"?
Back on the original topic, I completely understand if a project has too much baggage and sunk cost so as to make even a partial migration (and hardening via using memory-safe languages and/or professional paid code analysis tools) unrealistic.
It is like telling someone to delivering games in Unity and Unreal that they should use Amethyst instead, without understanding what it means to actually having a team working in those ecosystems.
Naturally Rust will get there, however it took C++ 30 years to be where it is today, and that is what any replacement attempt should take into consideration.
That being said, I agree -- existing teams can't be expected to just run away to the new thing, that much is true.
CLI tooling is the best it ever was in history (from where I stand at least). But IDE tooling isn't a first class citizen for many languages these days, is what I was saying.
Just for reference, my first UNIX was Xenix.
Don't see the point why people buy expensive computers to just use them like we did at the university lab with IBM X terminals.
The main difference was that instead of Slack we had xterms dedicated to run talk sessions.
Java's main issue is the barrier to entry is lower because it has good tooling, which means really bad, inexperienced programmers can write bad code, and have more tooling to bail them out, allowing them to continue down that path until they have a huge ball of mud.
There ought to be some kind of phrase for this "the tooling curse".
GDB isn't bad, IMHO, but the inability of IDEs to deterministically parse templates / index code without huge problems is problematic.
More like the inability of text editors to parse templates; most IDEs rely on a full compiler backend that can usually figure things out.
The ultimate result, given the slow compile times of c++, is that you cannot know in your editor if your code is correct.
My experience is frankly more that the usual IDEs I know of can't even handle syntax coloring 100% correct when it comes to C++…
I'm surprised you haven't seen this point made many times. It would be a full-time job on the internet asking people why they have code completion issues with C++ IDEs.
template<typename T>
void foo(T bar) { bar.buzz(); }
Other than that, finding foo<T> is usually easy, i've had good experience with qtcreator or clangd, you just have to be very very sure that your include paths are correct, qtcreator is good in that it uses CMakeLists as project-files so all this is done automaticaly.It turns out if you use enough layers of templates, the compiler has to carry around an awful lot of state to figure out what program it should actually output.
This is another downside to a 1400 page language specification; whether the user can hold most of it in their head or use a reliable, safe subset to avoid sharp edges, the compiler and tools still have to be aware of all of the layers of complexity in both the code a developer is writing and often whatever tricks and quirks developers chose to usein the libraries the developers code is depending on.
template<class _Tp, class ..._Args>
inline _LIBCPP_INLINE_VISIBILITY
typename enable_if
<
!is_array<_Tp>::value,
shared_ptr<_Tp>
>::type
make_shared(_Args&& ...__args)
{
return shared_ptr<_Tp>::make_shared(_VSTD::forward<_Args>(__args)...);
}For that matter, most of the so-so devs resumes that come across my desk will list "HTML and CSS" as skills. Even among experienced devs, being legitimately good with HTML or CSS is extremely rare. I'd give my eye teeth to have a dev on staff who was an actual HTML expert (knows the w3c docs, can write clear, correct, semantic html using the right tags, organizes and expresses their intent, etc.), but since the barrier to entry is barely above basic literacy, so finding one is tough.
Portability. A C library can be trivially linked with any other language. But one will be hard pressed using a python library form Ruby, for example.
I realized recently that Python could have saved a decade by having a means to "interlink" modules allowing the use of python 2 modules in python 3. Much easier in the C world; there have been incompatible syntax changes and incompatible link changes BUT not both at the same time.
Python has a very rich runtime, so there are some tricky problems to solve if you want this to be perfect. e.g., circular references between the GCs. However, there are simplifying assumptions that can dramatically reduce the difficulty, and these assumptions might not significantly hinder the language-upgrade use-case.
We've known about this approach to embedded interpreters for a couple of years, but we've not found a sufficiently compelling use-case for anyone to develop it beyond a simple proof-of-concept.
In general, it seems like there most compelling use-case for inter-language interfaces is writing core libraries in something fast, (mostly) runtime-free, and portable, especially for codes that no one wants to write twice.
Though I spend most of my time writing codes in slow, rich runtime languages, I've been on the lookout for a better technology to replace C. There are quite a few interesting options, and some have even less of a runtime than (dynamically-linked) C!
Where can I pip install this from?
Or is there a reason that none of the huge numbers of python users delaying their transition as much as possible until all their libraries were updated built it?
A couple of signatures have changed, but the general approach looks like: https://gist.github.com/dutc/eba9b2f7980f400f6287 or https://gist.github.com/dutc/2866d969d5e9209d501a
The above will launch ("embed") a Python 2 or Python 1.5 interpreter from within a Python 3 interpreter. If you use `dlmopen` and `LM_ID_NEWLM`, the guest interpreter will have its own linker namespace. In other words, the guest interpreter (and any DSOs it opens) will be totally isolated from the host interpreter.
The above shows the use of `PyRun_SimpleString`. Since `PyRun_String` accepts a `PyObject* globals` and `PyObject* locals`, you could build a very basic bridge in <30 lines of code by using serialisation to copy and convert. (For user-defined classes, you would need something better.)
I've given a many talks about this at various Python and PyData conferences. (I've mentioned it a number of times on HN, both in response to complaints about the Python 2→3 transition and in response to comments like yours suggesting this solution to that problem.)
Though presented as a joke, the approach could be made to work with some effort. I can't speculate on why no one has ever followed-up on it, and I can't speculate on why the Python 2→3 transition has been so difficult for some users. Perhaps in some places, financial or organisational arguments are more influential than technical arguments.
It's fairly rare that you actually need to be in the same process. (The thing I wrote was a wrapper for unshare(), so it did strictly need to be in-process, but that's an unusual use case.) Even if you want to share large amounts of data between your process and the library, you can often use things like memory-mapped files.
And there are some benefits from splitting up the languages, including the ability to let each language's runtime handle parallelism (finding out that the C library you want to use isn't actually thread-safe is no fun!), letting each language do exception-handling on its own, increased testability, etc.
If you're a Ruby application, sure, you can use a Python library by forking off a subprocess.
Just make sure to document the installation requirements - in addition to having Ruby and the correct gems installed, you also need Python (which version?) and the correct pip libraries installed.
If you're a Ruby library, instead of an application, forking off a Python process is a nonstarter, unless you want to propagate those requirements out to every single application using your library.
I'm mostly thinking that if you already have a Python library you really like, chances are you already have a Python environment capable of running it. (The environment doesn't have to be related to your Ruby environment at all! If you want to get Ruby through your OS and Python through a Docker container, or whatever, that should work fine.)
That's a bullshit argument very close to FUD.
You don't need to know the 1400 pages of the C++ standard to use safely a subset of it. Specially with proper tooling to help you.
Do you really think that every web frontend developer knows every W3C/ECMA standard every time they create a website ? Or that they have even an idea of feature the complexity of the Web browser they use. One tips: no they don't.
Most importantly, if someone does find a way, it's an error in either the spec or the browser implementation. It's not flagged as "undefined behavior" and we go on with the assumption some web pages just break your browser or the OS hosting it.
I can tolerate a thousand-page spec for a system that can't crash unsafely; it's a much bigger risk for one with unchecked pointer indirection as a feature.
This is a narrow view where you focus only on memory safety and buffer overflow. Undefined Behaviours goes way wilder than that.
Many problems coming by UBs do not goe into any crash, just wrong results and wrong behaviour which is something even worst.
And currently, JavaScript is full of that. Almost everything that JS can not handle (or do not know how to handle is undefined) or purely implementation specific.
There are circumstances where you want the features C++ brings to solve a problem. But solving a problem with dynamite is still solving a problem with dynamite; there's a lot more ways it can go wrong than solving the problem with a steam drill, to torture an analogy a bit.
Excepted that a language is not absolute, is not dynamite or not dynamite. It is what you do of it depending of the subset you use. C++ is not exception.
... which goes back to the topic of the HN post; that's only one way to look at the state of a running program, and it's a way that has strengths and weaknesses. The utility of it comes at the cost of the program failing in ways programs in other languages structurally cannot fail (unless you do something truly exotic, like implement a subset of a C++ compiler or interpreter in that language).
JS is ... well, JS
This is at least for me the first time I've seen them discussed in the same breath, which is revolutionary: C++ the language bringing a blue screen to a desktop near you; JS the language bringing your startup to its knees.
I guess I never realized that while they are polar opposites, there are probably some subtle cultural similarities that would make for some hilarious unearthing
Why don't you use something else?
Because, most (all?) of those other languages fall into one of three traps, they also have a lot of undefined behaviors, they define the behaviors in ways that don't reduce actual programming bugs (javascript!), or they have so rigorously defined the language with underlying assumptions (say a really strong memory model, signaling NaN, overflow exceptions, etc) that the performance is sub-optimal on any platform that doesn't exactly fit the expected machine description.Really, C isn't hard if you burn your copy of K&R and stop trying to be so damn clever. It also turns out to be a lot more readable, if a bit more verbose, if one ignores a lot of its "features" and pretends its pascal with a single statement, without side effects per line. That includes using pointers in any kind of arithmetic (or type casting), instead using them only as though they were C++ references. (and a few other basic rules).
So, while I don't think rust is a particularly good language, I'm also starting to think that everyone should be forced to use it early in their career so they are forced to consider object ownership and lifetime in a rigorous way. Then when they move to C/C++/etc they wont be foot-gunning themselves at every turn.
* There's no distinction between registers and memory in C. A function parameter that is "register volatile _Atomic int" is completely legal and also makes absolutely no sense if you want C to be a "thin" abstraction over assembly.
* There's no "bitcast" operator that changes the type of a value without affecting the value. The most common workaround involves going through memory and praying the compiler will optimize that away (while violating strict aliasing semantics to boot).
* No support for multiple return values (in registers). There's structs, but that means returning via memory because, well, see point 1.
* Missing machine state in the form of vector support (which is where the lack of the bitcast becomes particularly annoying). The floating-point environment is also commonly poorly supported.
* Missing operations such as popcount, cttz, simultaneous div/rem, or rotates.
* Traps are UB instead of having more predictable semantics.
* No support for various kinds of coroutines. setjmp/longjmp is the limit. You can't write zero-cost exception handling in C code, for example. Even computed-goto (a source-level equivalent of jump to a location held in a register) has no C representation, though it is present in some compiler extensions.
Access through a union allow the same region of memory to be interpreted as different types.
No support for multiple return values (in registers). There's structs, but that means returning via memory because, well, see point 1.
No mainstream compiler will ever not optimise that.
…only in C.
> No mainstream compiler will ever not optimise that.
Actually, it's mandated by the System-V ABI for structures of an appropriate size.
RE: Return values, you'd be surprised. You can't assume you'll get properly optimized code for something like that in WebAssembly for example, despite the fact that you're using an industry-standard compiler (clang) and runtime (v8 or spidermonkey).
Not with modern compilers: https://godbolt.org/z/e6sRqh
I'd be really interested to hear experiences of people who have done it though, even as a toy.
There is a small, very restrictive class of LLVM IR that is going to be effectively stable, and even portable. Stripping debug information goes a long way to making your IR readable by newer versions of LLVM, and staying away from C ABI compatibility (especially structs) can make your code somewhat portable.
Technically, you are supposed to use unions for that, though it's a pain. The PostgreSQL codebase is compile with -fno-strict-aliasing so that a simple cast will get the job done, but obviously that's technically out of spec.
1: http://www.open-std.org/jtc1/sc22/wg14/www/docs/dr_257.htm
2: http://www.open-std.org/jtc1/sc22/wg14/www/docs/dr_283.htm
gcc/clang offers builtins for most of these. And if you give the compiler the right march parameter it'll quite likely turn your hand-rolled code for these operations into the x64 instruction you're targeting. Try it on Compiler Explorer.
Doesn't the div function do this?
> * There's no "bitcast" operator that changes the type of a value without affecting the value. The most common workaround involves going through memory and praying the compiler will optimize that away (while violating strict aliasing semantics to boot).
> * Missing operations such as popcount, cttz, … or rotates.
Thankfully, these are coming in C++20.
Writing something like:
sum = x + y;
if(sum < x || sum < y) {
...
}
And then hoping the compiler will optimize your if statment to a single overflow CF check is a bit silly.What you would generally do is assume that x and y can't be overflow. I.e. you need to have a rough idea what quantities you will process. And put a few checks and assertions in strategic locations.
Return values in registers are a machine level optimization. Whether something is in a "register" is up to the CPU today on most CPUs. If you set a value in the last few instruction cycles, and are now using it, it probably hasn't even made it to the L1 cache yet. Even if it was pushed on the stack. That optimization made sense on SPARC, with all those registers. Maybe. Returning a small struct is fine.
There's an argument for multiple return values being unpacked into multiple variables, like Go and Python, but that has little to do with how it works at the machine level.
Missing operations such as popcount, cttz, simultaneous div/rem, or rotates.
Now that most CPUs have hardware for those functions, perhaps more of those functions should be visible at the language level. Here's all the things current x86 descendants can do.[1] Sometimes the compiler can notice that you need both dividend and remainder, or sine and cosine, but some of those are complicated enough it's not going to recognize it.
Traps are UB instead of having more predictable semantics.
That's very CPU-dependent. X86 and successors have precise exceptions, but most other architectures do not. It complicates the CPU considerably and can slow it down, because you can't commit anything to memory until you're sure there's no exception which could require backing out a completed store.
[1] https://software.intel.com/sites/landingpage/IntrinsicsGuide...
I think what's fair to say is: the "abstract C machine" tends to have a minimal amount of concepts on top of the machines instruction set, compared to other languages, and generally provides the least friction if you need to poke a memory address or layout a structure in a very specific way. It's just easier to say it's closer to the metal because that's a mouthful.
what do you mean by these two? not sure about the former. as for the latter, of course you can read uninitialized memory, unless you mean something else?
Edit: Some excerpts from the C standard:
6.7.8 Initialization
10 If an object that has automatic storage duration is not initialized explicitly, its value is indeterminate.
3.17.2 indeterminate value
either an unspecified value or a trap representation
3.17.3 unspecified value
valid value of the relevant type where this International Standard imposes no requirements on which value is chosen in any instance
6.2.6 Representation of types
6.2.6.1 General
5 Certain object representations need not represent a value of the object type. … Such a representation is called a trap representation.
This page says value is indeterminate, which is either unspecified or a trap, as you say: https://wiki.sei.cmu.edu/confluence/display/c/EXP33-C.+Do+no...
But.
This part says that reading an indeterminate value is, in fact, undefined behavior (line 11 in the table): https://wiki.sei.cmu.edu/confluence/display/c/CC.+Undefined+...
From 6.2.6.1¶5:
>Certain object representations need not represent a value of the object type. If the stored value of an object has such a representation and is read by an lvalue expression that does not have character type, the behavior is undefined. ... Such a representation is called a trap representation.
A read from uninitialized memory is not always UB.
Per http://www.open-std.org/jtc1/sc22/wg14/www/docs/dr_451.htm , the current standard is unclear in some respects, but the latest committee view is that that under the current standard any library function (including memcpy) may exhibit undefined behaviour when called with uninitialized memory, even when the uninitialized memory is of character type or is propagated through values of character type.
int x;
if(x == 0) foo();
if(x != 0) foo();
That is reading the uninitialized value twice. Since it is unspecified, it does not have to be consistent, so you could get the same behavior as non-zero for the first reading, and zero for the second reading. Changing your code to: int x;
if(x == 0) foo();
else foo();
will give different output (same if you use !=).In the compiler discovers UB, the Standard places no requirements of any kind on the program or compiler. It is free to launch missiles, or (more likely) assume this code cannot be reached, and omit it from the program, along with any code that reaches it unconditionally, and any check that would send control that way. Such elision is the basis for many important optimizations.
Implementations are free to define things left undefined by Standards. For example, "#include <unistd.h>" is UB by the ISO Standard, but defined by Posix, which implementations also adhere to.
https://stackoverflow.com/questions/11962457/why-is-using-an...
What is the point of not making it UB? You can't AFAICS possibly rely on any useful way on it behaving 'precisely and consistently' so just make it UB anyway?
int a; printf("%d", (&a + sizeof(int)));
> The value of an object with automatic storage duration is used while it is indeterminate.
But reading the value without using it seems fine?
For the second, https://www.ralfj.de/blog/2019/07/14/uninit.html
(the second one is a bit more Rust focused but the core idea is the same)
A modern CPU willy nilly reorders instructions, stores, and creates new "virtual" registers out of thin air to dissolve stall inducing dependency chains. It can also split instructions, fuse them together and even entirely remove some instructions -- as long as it doesn't affect the end result.
Generally (within limits of CPU memory model) the only thing you're guaranteed is that eventually the result is what you'd expect from sequential execution. That of course applies on the core you're running on, otherwise you naturally need to synchronize with other cores.
Luckily there are tools to find out, which work to some extent, IF you can run it on same CPU model (and really, whole system, incl. memory subsystem!) as what you're interested in.
a = 10;
b = 20;
c = b + 30;
bar = a + b;
x = 1;
y = 2;
z = y + 3;
foo = x + y;
... in order to remove dependency stalls might be executed as: b = 20;
y = 2;
a = 10;
x = 1;
c = b + 30;
z = y + 3;
bar = a + b;
foo = x + y;
This would still be true, even if 'x' and 'a', 'y' and 'b' etc. variables (well, registers) had same name. In that case CPU would just make up new "variables" as required and do the same transformation!Note that several decades ago, there were processors that weren't capable of actually keeping the state correct after a processor exception, so if you got a division-by-0 error, your program counter had advanced by 30 or so. (This is, I believe, part of the reason why traps are undefined behavior in C: if your processor can't guarantee any state after a machine trap, it's impossible to implement a programming language that has even the loosest guarantees).
https://en.wikipedia.org/wiki/Tomasulo_algorithm.
So no, the consistent serialized ISA-conforming state might be reconstructed by replaying after the exception (interrupt) arrives.
Alternatively, interrupt might be served only after execution reaches the next checkpoint.
On x86, a lot of instructions are close matches, yet some are removed entirely (say "xor eax, eax") or fused into one micro-op, (like "cmp #123, eax / je <address>").
Future CPUs might even do some data flow analysis to optimize code even further.
Say speculative "constant" folding based on runtime profile to remove chunks of code from hot inner loops.
Or to replace longer instruction patterns with HW optimized implementation, if that's what it takes to get some extra performance for the next year's model.
See the diagram here https://software.intel.com/sites/default/files/managed/9e/bc...
One of those only worth of time to learn as hardware engineers, but understanding both of them will give you an introduction to Spectre vulnerability.
printf("%x", *(int*)0x12345678);In short, how much does the concept of a "well-defined C program" differ from the concept of a "C program" as implemented in practice? Can we say that most C programs, or even a medium-sized percentage of C programs, are well-defined?
Implicit in the Standard is the notion of invalid pointers. In discussing pointers, the Standard typically refers to “a pointer to an object” or “a pointer to a function” or “a null pointer.” A special case in address arithmetic allows for a pointer to just past the end of an array. Any other pointer is invalid.
An invalid pointer might be created in several ways. An arbitrary value can be assigned (via a cast) to a pointer variable. (This could even create a valid pointer, depending on the value.) A pointer to an object becomes invalid if the memory containing the object is deallocated or moved by realloc. Pointer arithmetic can produce pointers outside the range of an array.
Regardless how an invalid pointer is created, any use of it yields undefined behavior. Even assignment, comparison with a null pointer constant, or comparison with itself, might on some systems result in an exception.
I'm not a language lawyer, but I suspect this means that even initialization to the wrong literal might well be undefined behavior. What makes you confident that it's not, and instead is merely implementation defined?
What if it's not an "invalid pointer", but a pointer to a memory-mapped IO address, ROM, etc? I grew up learning C on 16-bit machines in the early 90's. Hard coded pointer values were very, very common.
> What if it's not an "invalid pointer", but a pointer to a memory-mapped IO address, ROM, etc?
Yes, this is central to the question. And how is the compiler to know? Is it safe to presume that the compiler can't know, and thus can't presume undefined behavior? I think the answer is in the comments you linked where 'supercat' replies to 'Peter Cordes':
The Standard makes no attempt to mandate that all implementations be suitable for low-level systems programming, nor does it in any way imply that it's possible to have a quality implementation that is suitable for low-level or systems programming without it supporting behaviors beyond those mandated by the Standard (and which might not be processed predictably by implementations that aren't suitable for systems programming).
Which is to say, yes, for a compiler implementation to actually be useful for low-level programming, it must behave in a predictable manner when given literal addresses. Unfortunately, it may be possible for a C compiler to be "standards conforming" without actually being useful for this purpose. One can only hope that at least some compilers will continue to "do the right thing" despite that lack of explicit requirements.
It is possible for a C implementation to be conforming witout supporting low-level programming. Sometimes it even makes sense, like if you're running C code on GraalVM's LLVM bitcode interpreter. But there is no trend of "standard" C implementations following this route. While modern C compilers like to aggressively exploit undefined behavior, they generally make reasonable decisions for implementation-defined behavior. In this case, all major compilers will compile
*(int*)0x12345678
to the obvious assembly, and what happens then depends on what your memory map looks like.(Caveat: the compiler will still perform normal optimizations like removing unused loads or redundant stores. If the address actually points to memory-mapped I/O, you need `volatile` to prevent that.)
Let me quote wikipedia here:
> The Pentium 4 can have 126 micro-operations in flight at the same time. Micro-operations are decoded and stored in an Execution Trace Cache with 12,000 entries, to avoid repeated decoding of the same x86 instructions. Groups of six micro-operations are packed into a trace line. Complex instructions, such as exception handling, result in jumping to the microcode ROM. During development of the Pentium 4, microcode accounted for 14% of processor bugs versus 30% of processor bugs during development of the Pentium Pro.
Like, that's one paragraph, there are others I can cherry pick to show just how complicated the execution of microcode is. Unless you're an actual engineer on processor development and have insider info, I find it highly suspect that you could look at a decently sized chunk of assembly code and really know what happens with the microcode.
Sure you can have a table that says, it takes ballpark about this long. Ok fine, but that's not the same thing as assembly mapping cleanly to what the processor is actually doing.
Micro-ops are part of the micro-architecture of the processor, and are in hardware. They are not patchable and are not software.
But yeah, microcoded instructions are relatively rarely executed.
People have this common misconception that the programmable micro-code is what your CPU is actually executing, and x86 somehow translates into instructions for it, and this was really just because of a conflation of the terms "micro-code" and "micro-op".
Admittedly, Intel isn't the best at this term either. They have several places in the Architecture Manuals where they refer to the "micro-code synthesizer" when they mean "micro-op synthesizer"; this really has nothing to do with the micro-code ROM.
Also AIUI the microcode controls the issuing of the micro-ops.
> It would be way too slow to make that programmable
Then what is the "processor microcode updates" updating? I think this may just be a terminology mixup.
Dunno if this helps, FYI from https://stackoverflow.com/questions/17395557/observing-stale...
...and I can't copy/paste it. In the above link, look for 'embarrassing' by Krazy Glew.
https://twitter.com/uops_info/status/1202950247900684290 https://github.com/llvm-project/llvm/blob/master/lib/Target/...
I'd say "C is not how the computer works ... but it's much, much closer than nearly every other language."
M..............................................Java
Look. C is much closer to how the machine (M) works.
I think you're talking about the JVM. Anyway, some processors have native support:
That's... disturbing. Hope they're not building medical devices.
As for former colleagues, oh well, we all learn and grow. I chose to bail out because something as non-deterministic as a C's code stability when cross-compiled for anything more than two systems turned me off. Different strokes, different people.
We tend of think of everything as an ideal model in a vacuum, forgetting that in real life that the implementation details are often gory with caches and microcode and fault tolerance and all sorts of hidden details put in there to make things easier to work with or faster.
It’s how the hardware is modeled.
Hardware works as physics allows, at energies we can’t comprehend.
A computer is designed to a spec we can (sort of) comprehend. Spectre and such being evidence we don’t fully comprehend it.
It’s a recursive back and forth of this is a structure, this is the favored algorithm for that structure.
The hardware is the structure. Physics provides the algorithms for operating on that structure.
Thinking like one can see inside the machine always struck me as absolutely ridiculous.
It exists in the world defined by physics. Crawling into our imagination, linking abstract pictures from textures together arbitrarily, doesn’t mean we’ve discovered something new.
We know all kinds of stuff can be modeled on a compute because reality already gave us the math model.
We’re just brute forcing implementation looking for simple models for things we already have
Now if you care about performance, you might not want to ignore the real machine, but it really bis the hardware's job to work like the machine code tells it to. And when it doesn't, that's considered a bug^H^H^H exploit.
I've noticed in, say, an interview context, that people without a lot of C experience have a lot of trouble reasoning about how an allocator might work, or where memory comes from, or even that some constructs in their preferred language result in memory allocation. With C you can't hide these details.
Yes there are abstractions beneath you, both in software (your OS has set up the MMU) and hardware (cache, out of order execution) and also in the C standard itself (assumptions about pointer types, the precise meaning of the stack). This doesn't take away from the point that reasoning about allocations is more of a thing.
Those abstractions are there mostly for the purpose of speeding up code execution, but also for enabling hardware designers to think about the machine in terms of real-world concepts and not just as an incomprehensible mind-blowing collection of a few billion transistors.
It's abstractions all the way down. I don't know why they're so maligned by programmers, they're the only way any work gets done.
This is why saying sth is "not how the machine works" is a fallacy of black and white thinking and not useful unless one is ready to be very specific abt how it resembles the machine and doesn't.
wrt C, the smartest thing to do is use it wisely and avoid dragons, the same way we e.g. use chemistry to make useful products without performing operations in unstable environments. And surround it with an appropriate layer of cultural expectations - keeping humans in the loop. You don't just do what the computer tells you, you use it to help you make judgments, but ultimately take responsibility yourself.
One can see this in Rust's efforts to not only develop a language, but a conscious culture around it. This will be successful until, like a 51% attack, it is outmaneuvered.
Anyway...
Dictating that all architectures implement UB the same way would have significant overhead, since it would force programs that don’t rely on UB to run in emulation on all but one architecture (at best).
Implementation-defined behavior works as you say, where each compiler can define how it responds to certain circumstances.
[1] well, you can today by swapping malloc for a GC-enabled implementation, but the fact is that almost nobody does.
You don't need GC for this (at least, in the sense of "we'll call free for you"), you just need to verify that things passed to free were handed out via malloc.
Of all the potential issues that might be attributed to C, undefined behavior (UB) is the one that creates few to no problems at all.
To me, complaining about UB is like complaining about the quality of a highway once they intentionally break through the guard rails and start racing through the middle of the woods.
The reason why some behavior is intentionally left undefined is to enable implementations to do whatever they find best in a given platform. This is not a problem, it's a solution to whole classes of problems.
C is not unsafe if you only write safe code. But humans are fallible, and "just don't write unsafe code" is not a solution.
If you want to go with that analogy, UB is a kin of you intentionally pointing your gun to your foot, taking your time to aim it precisely on your foot, and in spite of all the possible warnings and error messages that are shouted at you... You still decide that yes, what you want to do is to shoot yourself in the foot because that is exactly what you want to achieve.
And then complain about the consequences.
Yeaaaah not really:
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
static char a[128];
static char b[128];
int main(int argc, char *argv[]) {
(void)argc;
strcpy(a + atoi(argv[1]), "owned.");
printf("%s\n", b);
return 0;
}
compiles with no warning using clang (8), even with -WeverythingAnd that’s a super easy case, it’s not a use after a dynamic callback somewhere freed a pointer it did not own, or a function 5 levels deep (or worse in a dll) didn’t check its input pointer which only became occasionally null years later.
> UB is a kin of you intentionally pointing your gun to your foot, taking your time to aim it precisely on your foot, and in spite of all the possible warnings and error messages that are shouted at you…
Because they could hardly be more wrong, the average UB is completely invisible to static analysers because it's a dynamic condition which the type system is not smart enough to move to static. That's why it's an UB and not, say, a compile error.
Now, I think that particular argument is poor. The more defensible version would be to look at the proportion of errors that occurred in C that would not occur in your language of choice, versus the number of errors that occur in this other language that would be unlikely to occur in C. And for completeness it would be good to do some sort of evaluation as to whether they were other gotchas like performance of the resulting code and programmer productivity. I expect that in the grand scheme of things, C is not the best option, but 100% perfection is an absurd goal post and makes it a lot easier for people to dismiss your argument.
I looked through the documentation and cannot find what hardware it runs on and how to set it up. The Getting Started for real hardware returns a 404 page: https://doc.redox-os.org/book/real_hardware.html
It is still in use, because when governments want safety above anything else, they buy such OSes, not UNIX clones.
https://www.unisys.com/offerings/clearpath-forward/clearpath...
The numerous and regular terrifying security bugs that are being found in ancient tools (some 30+ years old) are a living testament that the philosophy you quoted just doesn't translate that well into the reality of humans programming the current breed of computers.
Hence the existence of languages like Rust and the recent general strive towards more correctness and more compile-time catching of potential problems (especially in light of GCC 10's new `-fanalyzer` flag).
But even more than that, I think a lot of people would like to overlook that programming is fundamentally unsafe. Even if you're working in a safe language, that language's runtime is almost certainly in an unsafe language, making syscalls through a library written in an unsafe language, interacting with an OS written in an unsafe language, interacting with hardware running on firmware written in an unsafe language.
Sure there should be a boundary somewhere. But whether that boundary should be a type system or valgrind/cppcheck/etc. differs case-by-case.
y = x;
x += 1;
x -= 1;
true == (x != y);
But I don't know of any actual integer machine that would want integer overflow to actually be UB.Your metaphor brings to mind an image of someone who can see the guard rail, and then deliberately chooses to expend quite non-trivial effort to break them. This is not an accurate metaphor.
As I'm not in the metaphor fix-up business, because I believe they tend to mess us up in precisely this sort of way, let me simply directly point out this gets it backwards. The default is that you will write a program shot through with undefined behavior, and you must make quite substantial, intellectually-challenging efforts above and beyond the intellectual challenge of writing something in C at all that at least sometimes works, to write your program in a way that doesn't have UB in it. It has no resemblance to driving down a road in a normal fashion and then bashing guard rails down.
Sure, if those guardrails are 1mm tall and so barely visible, then that's a valid analogy to C. I hope you can see why this is not an ergonomic design that will lead to safe programs (or safe highways).
C did not have a standard for 18 years, it was first created in 1972, and ANSI C didn't appear until the spring of 1990. I'm not suggesting that it took them 18 years to write a spec, it's that it wasn't really needed for a long time. Expecting Rust to have one after five years is aggressive. Some languages have done this! I think we'll end up somewhere between what C did and what C# did, in this regard.
However, you do also bring up something that's a good point when it comes to this discussion: a lot of people treat "a spec" as a binary thing: a language has it, or does not have it. But specs are written by humans, and have their own mistakes, bugs, errata, etc.
One of the reasons that we haven't wanted to declare "this is the spec for Rust" is that we want to have a good, possibly even formally verifiable spec. The reference is basically at a fairly reasonable informal spec level.
[1] https://fzn.fr/readings/c11comp.pdf
Right now I see more efforts coming out Microsoft and static analysis tool vendors than any WG 14 mailing.
This has required a fair bit of assembly language for old machines with a much more minimal profit motive - in fact, non-existent - just so that I can keep my machines alive - which is the reason I'm digging into things as a bit of a challenge to myself.
There is so much great code to read, when you can disassemble. And, even when you can't.
I have a veritable fleet of old systems, many of which are still getting new titles released at astonishing levels of proficiency, yearly now .. CPC6128, ZX (Next!), C64, Atmos, BBC, etc., all have new tools being written for them, including C compilers.
One thing I note is that the tools are everything, because all tools get 'degraded' by new hardware releases. Hardware isn't driven by the compiler designer - the compiler designer is nothing until the hardware guy prints something.
Returning to older computers has meant understanding their architecture, and surprisingly enough after 30 years of mainstream programming, going back to this world is rewarding. There is a lot of joy in writing in C for the 6502, or even Z80, in real mode.
But, always disassemble.
And, in each case, its quite easy to disassemble - get an understanding of registers and memory usage and code inefficiencies, and so on.
In a way, GPU shader languages "fit" parallel GPU architecture (though not cache-aware).
Maybe... limited loop code-length to fit in cache; No pointer chasing (though you can workaround anything in a TM).
Some java subsets for very limited hardware might be instances.
Cell SPUs had explicit user managed cache and host DMA, and Intel ISPC and Unity Burst compiler/ECS (variants of C and C#) are good examples of programming environments that try to more explicitly encourage parallelism and/or optimal cache utilization.
just a nitpick, but explicitly user managed caches aren't. They are really better known as scratchpads. A cache is supposed to be mostly transparent except for the performance implications.
What you mean is that the popular Instruction Set Architectures (x86, ARM) don’t expose these pragmatics. (Though the microcode architectures of the processors executing these ISAs probably do.)
There are other ISAs (e.g. Itanium, most Very Long Instruction Word ISAs) that do expose these pragmatics, constraining what instructions can appear when in an instruction stream. Which in turn means that assembly code for these ISAs needs to be written with awareness of these pragmatics.
https://queue.acm.org/detail.cfm?id=3212479
HN discussion:
A lot has changed since then, but computers put up a lot of effort to pretend that they still work similarly, and most other programming languages have put up a lot of effort to figure out how to make things work for a "C computer".
The big iron of the time (and a few years later) had some things that did not work like C, too, such as single-level storage addressing, tagged architectures, and capability-based addressing.
Thinking that the C language is a representative model for computer history, or indeed for computer architecture, is a bad way of thinking of this. During the 1980s and 1990s there was a lot of shoehorning all of that architectural stuff into C compilers, with extra keywords and the like. (e.g. __interrupt, _Far, _Huge, __based, _Decimal, _Packed, digitsof, precisionof, and so forth)
So there were computer chips sold in the 1970s that worked a lot like C thinks. You are right that there were also ones that didn't.
Going forward from close to 1980 you added more and more features to all architectures that diverge.
I think that the hegemony of C for low-level programming is harming us significantly because it puts the other models to obscurity. Try to write in C for something like F21 chip or J–Machine...
Undefined behavior or use of anything not strictly included in POSIX.1c may result in the launch of nuclear weapons.
The self-quote for the thesis for this article comes from a conversation with tptacek in the HN thread from the first post.
I mean, I guess you could argue that it's produced an interesting discussion. There's actually a pretty interesting question here: is the purpose of a site like this to distribute interesting news, or foster interesting discussion? I personally feel like it's more the former, but maybe the latter matters too.
I remember we started our University course of low-level programming from writing in machine codes. No jokes, took binary representation of assembly instructions from that manual and wrote in HEX editor. Then we disassembled simple C applications, understanding unfolding our high-level instructions into x86 assembly. Passing parameters through registres, through stack, two models of stack restoration (yeah, cdecl and stdcall, that moment I realized what does it mean), and other stuff you won't be ever bothered to think about writing in C/C++. And only after that started writing in actual Assembly. THIS is knowledge how computer works. If you know only C, you know nothing.
-- Micha Niskin
Not so much these days. How C works and how x86-64 assembly works is very similar. Of course, that is not the whole story and you could argue that how x86-64 assembly works is not how the "computer" works because it is a high-level abstraction of the CPU... But then what does it mean to talk about how the computer works?
What high language could I use that gives me most performance because it accurately matches the underlying platform?
C is close to how a computer works.
Assembly is the next step down, and is not nearly as hard as the internet would lead you to believe.
Step down from that; read a book like "Computer Architecture: A Quantitative Approach", and learn about cpu-s and memory (also Agner Fog's site, official amd and intel doc, many other).
Down from that is electronics (logic gates, 8bit adder, etc., down to transistors and capacitors).
And beyond is physics.
Everything else is either computing theory or buzzwordy bullshit.
To repeat: C is a high level language, no matter what anyone says.
But there is an interesting catch to it: because C is so important, and programmers at large know so little "real C", vendors do their damndest to make it work for them.
A well-known microcontroller forum in Germany is constantly pushing that you need to make your code safe in the case of concurrent access to a variable by a main loop and an interrupt. By making it volatile.
Strictly speaking that's plain wrong. You need some kind of mutex.
Practically speaking, no embedded C compiler vendor would dare use an "optimizing opportunity" and surprise his customers there.
That other languages don't qualify, either, is beside the point. Ruby or Java aren't mistaken to be close-to-metal.
Point in case: you don't even see anything about Harvard vs. von Neumann architecture in portable C sources. Not to speak of micro code or anything like that.
I'm disputing that the C execution model is adequately and exhaustively modelling the idiosyncracies of real existing modern hardware, where "modern" means "almost anything in my lifetime".
It is, but what happened though, is that assembly/machine language no longer is as close to the real machine as it was in the past, it is an abstraction on top of micro code.
The high level languages that you mentioned are further away, and competition abstracted away from the assembly.
main thread -> interrupt processing -> resume main thread.
If you're modifying a variable in the interrupt you should declare it volatile so as to avoid the compiler optimizing it and introducing unexpected behavior. Adding a mutex would just be worthless overhead.
Of course, it didn't do that. Because vendors imparted their compilers with additional knowledge about common programming patterns and uses that aren't standard C.
Well, TI doesn't do that thankfully. At least not for the MSP430.
But assembly doesn't have the following concepts built into it, you have to manage those yourself if you want them:
- The notion of functions and return values. CALL/RET implements subroutines, not functions.
- Linkage that is more complex than jumps or call/returns. The ABI between units of code is something you have to define and do youself.
- The notion of associating types with variables. Variables in assembly are simply labels for memory locations, but don't really constrict action or tell you how many memory locations are in the variable. Types, including valid actions, data widths, etc. are completely up to you and must be done manually in assembly.
I like this article on this topic: https://raphlinus.github.io/programming/rust/2018/08/17/unde...
That was true for the PDP-11 that the article series mentions and it was true for the early home computers of the '70s and '80s but it became increasingly less true as processors became more and more complicated under the hood. That's kind of the point of the article series but it's not as big a deal as the author makes it out to be because it was already people writing in C were aware of.
1. You do lots of pointer arithmetic. But I don't think being practiced at pointer arithmetic is important for developing as a programmer. Just learn what it is.
2. Local variables in C exist on the stack, not the heap. But I don't think programming in C is a great way to learn about stacks. Except when you try to return a reference to a local variable from a function, everything feels the same as in other languages.
If you're making a list of languages you want to try out, I wouldn't put C above 5th place. I say this as someone who spent 12 years primarily coding in C.
Edit: Oh I get it--your manager, who doesn't believe in leaky abstraction or UB, will beat you if the program fails.
The computer is the thing that has the web browser running on it.
The computer is the thing that has the code editor running on it.
The computer is the thing that has the compiler running on it.
The computer is the thing that provides a virtual hardware interface for running programs.
The computer is the thing that the virtual hardware interface interacts with.
The computer is the thing inside the hard drive/graphics card/other component that the computer talks to.
The computer is the thing inside the processor that pre-processes, jits, or otherwise transforms commands from the computer for the computer so the computer ... ... ...
The c1 is the thing that has the web browser running on it.
The c1 is the thing that has the code editor running on it.
The c1 is the thing that has the compiler running on it.
The c2 is the thing that provides a virtual hardware interface for running programs.
The c3 is the thing that the virtual hardware interface interacts with.
The c3 is the thing inside the hard drive/graphics card/other component that the c2 talks to.
The c4 is the thing inside the processor that pre-processes, jits, or otherwise transforms commands from the c2 for the c3 so the c1 ... ... ...
Is this correct?
I wasn't being super-precise with my comment, and my point is that there is no one thing that "is the computer". There are many computers in your computer, and many definitions of what a computer is.
I've never had to go below the HAL, or inside the JIT, but to some people that's where "the" computer really is. If that's where the computer really is, then I'm not a computer programmer (HINT: I'm not, but I do write code sometimes and get paid for it).
So if I write code, but it doesn't tell the computer what to do, what am I even doing?
Well, I write code that gets consumed by pre-processors and lexers and parsers and compilers and linkers and build agents and I don't know how the soup works, and you probably don't either, you may know more or less of the of the alpha-bits in the soup, but it's still soup. The soup is turned into jello and fed to a virtual machine, which is a fake computer, but it works well enough.
---
If I were to be more precise, I would say that there's a von Neumann machine C0 that is represented by some hardware abstraction layer C1 by a operating system kernel. There may be a virtual machine C2 running in the C1 context.
But, inside that C0 von Neumann machine there are a bunch of microprocessors that may be called computers that may have Turing complete instruction sets.
---
Or you could say, anything that can be Turing complete is a computer, in which case your web browser is a computer (among many others).
You can tell the computer what to do by writing a program that passes through six levels of abstraction before the fate of any electrons is affected by what you wrote.
You can also tell the computer what to do by pushing a lot of keys on a keyboard. We're lucky; our keyboards have like a hundred keys on them. Some of our predecessors had eight switches and a ninth key or switch to "ACCEPT" the current switch-bank state.
You can also tell the computer what to do by causing patterns to be stored on magnetic or optical media and read back later.
You can also tell the computer what to do by plugging a wire into the back of it and varying voltages on that wire using another machine a hundred miles away.
Part of the art of programming is knowing that there are no bright, solid lines between these ways to tell a computer what to do (but for a given task, some are clearly more applicable than others ;) ).
If not, what is it?
I've read some intriguing things about D, though (which came along more recently, interestingly enough). Apparently it's in production at several large companies.
[1] https://andrewkelley.me/post/zig-cc-powerful-drop-in-replace...
More than that though the C and C++ competitors don't have the tools and libraries.
When all these new languages get made, they never ease back on the language part and prioritize the next rough area. The language is argued about debated and added to with features for years or decades. Meanwhile build systems, guis, text editing with suggestions and syntax checking etc. all go into the pile of being a mess of half supported tools pieced together as an exercise in frustration for the brand new curious user.
You have never worked on a C compiler, I take it. All production compilers will at some point turn even straight-line code into a DAG and relinearize it without care to the original structure of code. Any semantics beyond that which is given in the C specification (which are purely dynamic, mind you) are not considered for preservation. In particular, static intuitions about relative positions (or even the number!) of call statements to if statements or other control flow are completely and totally wrong and unreliable.
This is the main reason why people keep harping on the "C is an abstract machine." When people think of machine code, there is an assumption about the set of semantics that are relevant to the machine code, and that set is often much, much broader than the set of semantics actually preserved by the C specification.
Structured assembly semantics are what C users in a lot of domains expect and production compilers like clang at -O3 will obey the programmer in all of the cases that are necessary to make that code work:
- Aliasing rules even under strict aliasing give a must-alias escape hatch to allow arbitrary pointer casts.
- Pointer math is treated as integer math. Llvm internally considers typed pointer math to be equivalent to casting to int and doing math that way.
- Effects are treated with super high levels of conservatism. A C compiler is way more conservative about the possible outcomes of a function call or a memory store than a JS JIT for example.
- Lots of other stuff.
It doesn’t matter that the compiler turns the program into a DAG or any other representation. Structured assembly semantics are obeyed because of the combination of conservative constraints that the compiler places on itself.
But bottom line: lots of C code expects structured assembly and gets it. WebKit’s JavaScript engine, any malloc implementation, and probably most (if not all) kernels are examples. That code isn’t wrong. It works in optimizing C compilers. The only thing wrong is the spec.
…as I take it, for languages without undefined behavior.
> Structured assembly semantics are what C users in a lot of domains expect and production compilers like clang at -O3 will obey the programmer in all of the cases that are necessary to make that code work
No, they will not, and the people working on Clang will tell you that they won't.
> Llvm internally considers typed pointer math to be equivalent to casting to int and doing math that way.
LLVM IR is not C, but I am fairly sure that it is still illegal to access memory that is outside of the bounds of an object even in IR.
> A C compiler is way more conservative about the possible outcomes of a function call or a memory store than a JS JIT for example.
Well, yes, because JavaScript JITs own the entire ABI and have full insight into both sides of the function call.
> That code isn’t wrong. It works in optimizing C compilers. The only thing wrong is the spec.
The code is wrong, and that it works today doesn't mean it won't work tomorrow. If think that all numbers ending in "3" are prime because you've only looked at "3", "13" and "23", would you say that "math is wrong" because it tells you that this isn't true in general?
You’re kind of glossing over the fact that for the compiler to perform an optimization that is correct under C semantics but not under structured assembly semantics is rare because under both laws you have to assume that memory accessed have profound effects (stores have super weak may alias rules) and calls clobber everything. Put those things together and it’s unusual that a programmer expecting proper structured assembly behavior from their illegally (according to spec) computed pointer would get anything other than correct structured assembly behavior. Like, except for obvious cases, the C compiler has no clue what a pointer points at. That’s great news for professional programmers who don’t have time for bullshit about abstract machines and just need to write structured assembly. You can do it because either way C has semantics that are not very amenable to analysis of the state of the heap.
Partly it’s also because of the optimizations went further, they would break too much code. So it’s just not going to happen.