And why oh why in C with so many safer options available today ..
And why oh why in C with so many safer options available today ..
2. Because this is exactly the kind of task C was intended for: systems programming.
3. Because I want this to be a light-weight and portable program.
4. Because I have 30+ years of C-experience and happen to like the language.
(Preemptive response to downvoters: please have a sense of humor about yourselves. :P)
This is an accurate parody of my opinion. I'm unable to say "good C" with a straight face.
> I seriously think the C-hater crowd on HN is just disappointed that they can't reason well about pointers.
Let's be (hopefully) pessimistic and assume I'm a bottom 10th percentile programmer, and assume (I believe optimistically) that every programmer half an iota better than me is physically incapable of writing a single solitary buffer overflow, double free, use-after-free, or other pointer usage related bug. Perhaps they are robots!
I'm of the opinion that the remaining bottom 10th percentile remains a significant problem and security hazard. My continued employment with C++ seems to indicate we lack either the technology to detect, or the will to fire, this bottom 10th percentile. I feel that a hatred of C is a pragmatic step in helping reduce this problem, by reducing the usage of C by my fellow incompetents, by reducing the usage of C in general. Do you disagree? (I extend this stance to C++, despite it being my day job.)
http://www.cvedetails.com/vulnerabilities-by-types.php
Continuing the assumption - I'm disappointed that I can't reason well about pointers. I'm a bit more bummed about being in the bottom 10th percentile. I'm significantly more distressed that people can't reason well about people who can reason well about pointers - that, by virtue of being human, even they will make mistakes.
Being slightly less pessimistic: I'm fucking awesome. C isn't. What Dunning–Kruger effect?
And where would that lower 90th be if it weren't for the fact that these days people are primarily being taught in memory safe languages? A key difference here, and one of the reasons I would resist percentile ranking for this discussion, is simple exposure and practice.
Reasonable. Programming skill is not a 1 dimensional trait... and it's a bit bloody minded, even if it were. I threw myself in the fire with the hypotheticals because it made me uncomfortable, even as a hypothetical for the sake of argument, to be doing such a reduction.
> Should you therefore mistrust by default the software written by those folks, due to being burned by output from the lower 10th working in the same language?
Sure, absolutely. This manifests as a willingness to review their commits for mistakes, and to support their efforts in trying to reduce the chances of making certain types of mistakes, and an insistence they do the same for me.
And once I've gotten a better handle on their skill, assuming they're good, we'll both invariably regret it when I start to get complacent with my reviews and they check something in with a problem we've now both missed. It wouldn't surprise me if the best programmers I've worked with sneak more bugs past me than the worst as a direct result of this...
> And all this despite the fact that we have some mitigations these days, such as address randomization, fewer pages being executable, privilege separation, safer coding styles and practices being pretty well known in some circles, etc.?
Yes, absolutely. Why even indulge in any of these mitigations if you trust yourself and your fellow programmers with unsafe pointers? Each of these is vitally important specifically because we can't. And unfortunately, these mitigations are imperfect solutions. A horrifying amount of software disables them outright - and even when enabled, exploits still show up in the wild - they just get more convoluted and end up with lower success rates.
Better language choice is yet another potential mitigation. For me that usually means C#. This too is imperfect: I'm still more than able to write an unchecked buffer overflow with an unsafe block, or a marshaling API call, etc. ad infinitum - but it's an improvement.
> And where would that lower 90th be if it weren't for the fact that these days people are primarily being taught in memory safe languages?
Spending more time learning about and fighting memory bugs instead of shipping features, a sad waste of productivity. (I don't think that's the answer you were hinting at, but I'm not sure what is. An aside: My field still worships and is primarily taught memory unsafe languages.)
You should check out Rust and Go. Their users are most definitely among those who fail to understand C pointers.
The benefits range from overstated [bounds checking is a thing, but amongst experienced practitioners using modern, safe styles and conventions, lack of memory safety is not really the huge issue people say it is] to sometimes irrelevant [I read that the concurrency model aims to eliminate the possibility of race conditions entirely - who cares? In many domains races are a fact of life, benign, or totally manageable, so it appears to solve the wrong problem].
As for Go, I do acknowledge the background of its creators [venerable, as another commenter said], but most stuff I see written by "rank and file" members of that community is very clearly written by people who mostly deal in higher-level languages.
In any case, bounds checking is a tiny side street in how Rust ensures memory safety: ownership and lifetimes are the important novel things (lifetimes especially are novel, only research languages have had anything similar until this point, and I don't know if any have had exactly the same system).
In Rust the story is totally settled: in order from most to least preferred/recommended: use references, no pointers, owning pointers and then start thinking about reference counting/other shared ownership.
Rust doesn't solve race conditions: it solves data races, which are also problematic (they cause memory unsafety). It is impossible/very hard to avoid arbitrary race conditions, but data races are the only ones that can cause memory unsafety directly.
The first is very clearly a borrowed C++ feature (from RAII and smart pointers). I'm unfamiliar with the 2nd but it sounds like just making the compiler yell at you more when you break scope rules (like C++ will happily let you do with dangling pointers - I will give you that dangling pointers are bad).
> Rust doesn't solve race conditions: it solves data races, which are also problematic (they cause memory unsafety). It is impossible/very hard to avoid arbitrary race conditions, but data races are the only ones that can cause memory unsafety directly.
This sounds like gibberish to me. When I said "race conditions" please read it as "data races". Many of which will not "cause memory unsafety directly". Many of which low-level code such as a kernel will need to not be abstracted from. (Imagine writing a page fault handler in a language that tries to hide data races from you... Doesn't sound like a great idea.)
Even the Linux kernel has very, very few intentional data races. Most of them rely on undefined behavior and Linus has complained before about kernel code breaking when GCC actually tries to exploit that behavior. In those rare cases, it is perfectly reasonable to use unsafe code, along with copious comments explaining why it's safe.
You may want to read this article: https://software.intel.com/en-us/blogs/2013/01/06/benign-dat...
This, in my opinion, suggests that what I read about Rust makes it sound unsuitable for certain domains where concurrent writes and data races are an inherent fact of the universe, a reality that the language seems to deny and be working against.
(Didn't even mention lock free algorithms, for example I was reading one day about [I think it was] the dentry cache using lock free algorithms... You can bet there is part of that code that yes, expects data races to happen and handles them.)
Just to be completely clear, since we didn't define what we were talking about: a data race exists when there are two or more unsynchronized accesses to the same memory location, at least one of which is a write. What happens when that occurs is undefined behavior in C, whether or not you are writing a kernel. On some hardware the result is totally unusable (the write is corrupted). Additionally, the compiler and (on some architectures) the processor are free to reorder operations such that other threads will see your writes in the wrong order, or not at all.
This is all hardware level, the kernel has to deal with these problems too. The "relaxed" ordering in the C++11 memory model is essentially the lowest guarantee you can do anything useful with (other than have something like a racy counter, where you literally don't care at all what the value is, which is basically the only place the kernel has intentional data races at all). For some architectures and operations it resolves to a noop (relaxed stores and loads on x86, for example, never require additional synchronization beyond what the compiler provides).
The Linux kernel takes advantage of these semantics for lock-free or mostly-lock-free algorithms like RCU. There's a lot of code in those implementations specifically to ensure the correct barriers and orderings are in place so that there can't be data races. There's no reason they couldn't be implemented in Rust, and it would probably be easier since they could rely on LLVM's implementation of fences and so on rather than reimplementing it for each architecture. And if LLVM didn't have what they needed, they could just implement it in inline assembly (which is what kernel developers already do anyway).
If you define data races to not include race conditions that your code is handling (by placement of atomics, fences, locking etc.) then maybe your points make some sense but it is exactly the code handling these that I am saying does not fit with what I have seen rust advocates claim. You have to reason about these things at some level, and what I hear people advocate on hn is to sidestep it all by preventing sharing, which may be a fine idea but doesn't work for everything.
Lastly, in my mind, if your lock free algorithm works by letting the races happen but using the result of an atomic op and/or fences to ensure consistent behavior, I would still call that a (benign) race. That you needed atomics, fences, or compiler hints to get what you want seems orthogonal.
> If you define data races to not include race conditions that your code is handling (by placement of atomics, fences, locking etc.) then maybe your points make some sense but it is exactly the code handling these that I am saying does not fit with what I have seen rust advocates claim.
Its actual guarantees for atomics are not very strong and let you handle all of those things in safe code. What's interesting is that it can guarantee correct behavior for types that have more interesting semantics than "a bytestring," like mutexes or smart pointers. The implementations of those things are still unsafe, but you can generally provide a safe API to them, which you can't do in most other languages. This is because Rust has the notion of thread safety built into the language and enforced by the type system. For example, Rust's `shared_ptr` equivalent, `Arc<T>` is defined to be thread safe to share and send to other threads if and only if `T` is also thread safe to share and send to other threads. That doesn't mean you can't build up a safe type from the lower level atomics, or Rust isn't suited for this domain. It means you can write safe APIs in Rust that you can't in C++.
> You have to reason about these things at some level, and what I hear people advocate on hn is to sidestep it all by preventing sharing, which may be a fine idea but doesn't work for everything.
Rust just forces you to do just enough synchronization to not have undefined behavior. It's not a dogmatic language and doesn't believe there is one right way to do concurrency. Channels were just removed from the prelude for precisely that reason--they're not particularly favored over other concurrency mechanisms.
> Lastly, in my mind, if your lock free algorithm works by letting the races happen but using the result of an atomic op and/or fences to ensure consistent behavior, I would still call that a (benign) race. That you needed atomics, fences, or compiler hints to get what you want seems orthogonal.
Data races are well-defined, and what you're describing aren't data races, benign or otherwise. You basically never want data races, just like you basically never to dereference a dangling pointer. Rust prevents things you basically never want, it doesn't prevent writing useful things like lock-free data structures.
Lock free algorithms, or even a lock implementation, does not jive with this. You let the race happen, and you safely detect when you lost or won the race, and then you do stuff accordingly. To use your pointer analogy, it's OK to have a stale pointer in RAM or in a register at a moment in time. It's a violation of the invariants to dereference it.
As for the rest, perhaps I have been misinformed or led to some outdated info on rust when I saw things about no sharing. I will be sure to take a look.
A data race is essentially defined as a race condition between reads and writes where at least one of those is non-atomic. Rust's type system allows one to enforce atomicity by default (e.g. using atomic instructions, or wrapping data in a mutex).
I said "safely" detect, so yes, the final observation that you have won or lost the race must come from an atomic op. However, a "data race" susceptible read can totally be part of the process. The most common idiom I have seen for a lock free atomic read-modify-write has been to do an "unsafe" read, then act like it was OK, then issue compare-and-swap to determine if it really was OK and do the write if successful. (A failed compare and swap means you need to re-fetch and try again.) Yes you need to make sure fences are OK and the compiler is not caching/reordering the read and you need to account for the famous "ABA problem", blah blah blah, but it's what people do and I am not making it up for the purposes of an HN discussion.
The typical RISC approach to atomics, load-link/store-conditional, also encourages what I will call "my kind of thinking" on this issue, though it does so with specialized instructions. You do a special read, then you use ordinary non-atomic register operations, then you do the special store which fails when you lose the race. I will give you that these are specialized instructions and not ordinary C assignments but I would suggest looking into them if you have not already, the semantics are very educational. (And they may just convince you that disallowing data races is an overly restrictive thing to do in some scenarios.)
// let the "race" happen:
old = compare_and_swap(some_shared_memory, 1);
if old == 0 {
// we won!
}
This is perfectly valid data-race-free code if the CAS is atomic, but is invalid if the CAS is not (assuming no other memory fences).But sure, if you're manually fencing then you can have data races that are benign; however, just because something has some relatively rare uses doesn't mean it's wildly unsafe for the general case, and so disallowing by default it helps the correctness of the vast majority of concurrent code.
In any case, Rust allows one to opt-in to that sort of behaviour using `unsafe` locally, e.g. one would use `unsafe` deep in the internals of the implementation of the lock-free data structure and with careful vetting to ensure it's correct, and then all users can benefit with the compiler ensuring concurrency-safety by default.
Rust tries not to completely disallow behaviour, just make memory safety the default. The programmer can override the compiler via `unsafe` if they truly know better.
// Performs some modification on the input.
extern int f(int);
volatile int global_var = /* ... */;
int expected;
int new; // assuming C and not C++, where "new" is reserved
do
{
// Do a speculative read
expected = global_var;
// Make some (non-atomic) modification
new = f(expected);
// loop until the speculative read was OK. (won the race)
} while (compare_and_swap(&global_var, expected, new) != expected);My point was, Rust supporters have no problems understanding pointers.
Go isn't in the same ballpark nor do the authors (venerable as they are) seem to have considered that many people need consistent power over their run time execution when microseconds matter. stop the world is not a viable memory management trade off. Also managing memory in c++ these days is pretty "easy".
I've leived for years with people spending hundreds of man hours a year trying to get around issues caused by STW pauses in high throughput java applications. Makes no sense.
The portability of C is also huge because you want this running on a large number of hardware platforms. Everything from tiny home routers to the biggest server imaginable.
For something like an ntpd daemon you also want a language that many developers read and understand. C is used by more developers than pretty much any of the newer safer languages.
It might be old and weird, but C is still really hard to beat.
I think it's safe to say that he's forgotten more about C/kernel/low-level/... programming than HN'ers will ever know.
I don't think that's how you're supposed to software.
Overhead is important since Linux is designed to run on a whole spectrum of devices. Abstractions mean you are bringing in more code i.e. greater complexity, space, overhead. And frankly nobody could care less about programmer's time/effort/sanity in this case. It is a tightly focused, specific utility that shouldn't see too many changes after a certain point.
Like the difference between a ntpd.c vs ntpd.go is going to matter for anyone...
If you really find yourself stuck with an embedded device small enough that a few megabytes of RAM matter then you can still use the old C implementations, it's not like they are going away.
Abstractions mean you are bringing in more code i.e. greater complexity, space, overhead.
In this case abstractions mean you can probably write the whole thing in under 2 kloc, get memory safety for free, and base it on a network stack that other people actually use and debug independently.
It is a tightly focused, specific utility that shouldn't see too many changes after a certain point.
Umm. Yea. Right. Tell that to the ntpd guys with their 100kloc codebase...
Why not write that part in C and leave the rest to the more suitable language?
I would guess that's a rather small portion of the codebase, essentially the callouts to adjtime().
I don't buy into the low-latency FUD. C isn't a magical realtime language either; what happens when the system is under high load? What about delays in the kernel/network stack?
As a layman I would think incoming frames will have to be stamped with a hardware high resolution timestamp at arrival anyway, regardless of the language. And yes, this internal timestamp will have to be compared immediately before setting the hwclock. I would argue this is the only part that really needs to be written in a GC free language.
I might be wrong on this, but I'd like to see a stronger argument than "could well be significant" for writing yet another generation of a baseline system daemon in an unsafe language...
And speaking of latency: If this really is such an issue, why hasn't NTP long been moved into the kernel?
About 2 million servers were sold world-wide in 2014Q2
Assume 25% runs UNIX and NTPD -> half a million servers
Assume NTPD uses 0.1% of machine resources -> 500 fully loaded servers.
Assume 100W/server -> 50 kW
50kW for half a year -> 220,000 kWh
Assume 500g CO2/kWh -> 110 tons of CO2
QED: I think CPU overhead is a very relevant concern for time synching programs.
Bold assumption. It uses 0.0006% on my slow 6 years old box (over 13 days uptime). Let's be generous and say it's 0.001%, then you get 1.1 ton of CO2, which is the equivalent of the production of about 32.8 Kg of beef. I don't think such dramatic conclusions can be made from those facts.
Second, have you checked the memory footprint of NTPD ? "machine resources" is more than CPU cycles.
And don't forget: That number was one quarters purchase of servers, running for half a year. All the servers bought in 2010, 2011, 2012, 2013 and the other half of the year were not included.
Things add up.
284K, which is a whopping 0.04% of available memory. I don't see how this increases CO2 output though.
> Things add up
To a still very insignificant number, dwarfed by the beef consumption of the people reading this.
I'd say there are better arguments in favour of using dated languages than this.
I can change how their computers synchronize time.
Removing 0.1% load from the machine will not result in 0.1% of energy savings, very very far from it if in any at all. And that's not even what we want to calculate. We want the difference with hypothetical alternative in a safer language, which is hard to guess, but is unlikely to be in the same order or two with your calculations.
I get that sometimes you just want to hack in a language that's familiar to you, but then call your project a fun side project and not the "NTPD replacement" (and yes, I know who phk is).
To clarify, I'm not claiming that everyone should start using Rust. Now was I claiming that PHK is somehow unqualified for the task, far from it. I'm not even saying C is a bad language (I like C), or that people should stop using C altogether. But if you're building an security-critical application designed to be used in the future across millions of servers, and one that requires low-latency, I simply don't understand why you wouldn't strongly consider Rust, which offers strong guarantees about memory and type safety along with low latency and a very modern and complete standard library. Or Ada. Or OCaml.
The whole of the Unix philosophy is enforcing this policy of strong abstractions by making each part of a complex overall system its own binary, such that they can only communicate using the file I/O abstraction, possibly via pipes. C was therefore designed for a system which promotes abstractions.
PS: I would like to see a toolchain for Go that emits portable C instead of statically linking in a giant runtime... Lightweight camping vs. YAGNI hoarding. Seriously, the weight of runtimes would be the biggest barrier to serious systems deployment, outshining the slowdown, productivity and safety tradeoffs. (Eg Go would need to play nice with embedded constraints including memory usage as well.).
Deleted comment
Standards are good and useful things, but they're only fully written down in retrospect.
[1] http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1570.pdf
Haskell + stream fusion gives better results than well tuned C in some very typical situations http://research.microsoft.com/en-us/um/people/simonpj/papers...
Even just comparing C++/C, C++'s templating mechanisms mean that generic containers can be optimised on a per-type basis.
Abstraction and overhead are not as tightly coupled as they used to be, and not introducing Heartbleed-style bugs is worth a lot in my opinion
This is more a matter of "the right language for the job" and for timing-critical systems programming running as root, that language is C.
And there are plenty of common abstractions I would love see added to the C language: basic linked lists, byte-endianess and packing for struct members, validity intervals for integers and FP variables (like Ada!)
Unfortunately my taste seems to be the direct opposite of ISO-C which have instead wasted time giving us another thread-API.
In most cases it isn't a bad thing at all. But a use case like a timekeeping daemon, where a single unexpected garbage collection pause could cause a world of problems, it most definitely is more pain than gain.
System time is far, far too critical to take chances with.