Getting Past C
blog.ntpsec.org
blog.ntpsec.org
That said, there are actually drawbacks of Rust compared with Go, IMHO. When facing a moderately large project written by others, the ergonomics for diving into the project is not as smooth as Go. There is no good full-source code indexer like cscope/GNU Global/Guru for symbol navigation across multiple dependent projects. Full text searching with grep/ack does not fill the gap well either since many symbols, with their different scopes/paths, are allowed to have the same identifier without explicitly specifying the full path. That makes troubleshooting/tracing a large, unfamiliar codebase quite daunting compared with Go.
Github project: https://github.com/jonathandturner/rls
Announcement: https://internals.rust-lang.org/t/introducing-rust-language-...
There are also many other tools that provide indexing, eg. ide [plugins], kythe, and the rust language server.
You can also use https://github.com/nrc/rust-dxr to index rust code via DXR.
IIRC ctags also works with Rust.
RLS should cover this pretty well too once it happens.
However the point was of course to make a safe-land layer which contains the needed functionality. The point was never to have a function "safeioctl" in the first place.
And, honestly, one of the features is "vibrant community of developers." Even if Go and Rust were bad languages, which they're not, they'd still be better choices than C-with-custom-in-house-restrictions-and-libc-wrappers, simply because of the communities around them. If you're writing in Rust, someone else has already written the safe C wrappers, and if they haven't, there's a community of people who will code-review your wrappers for safety and merge them into a centrally-maintained project, which is extremely useful.
To answer your question though, the advantage of using C is certainly not the notorious bad-habits standard library. C is fun and productive exactly where there's just you and some bits and bytes to bang around. Coding in the small. Not platforms and architectures.
Just as you can statically link libc, without it showing up in the symbol table.
tangentially, ats has had some success in that area since it compiles via c. but I agree that c is less work to get set up with.
All the way through the 80's up to the early 90's, C toolchains only mattered to those that were working on UNIX workstations and servers. On home computers it was just, yet another language, with compilers generating slower code than junior Assembly programmers.
So of course Rust has to catch up a few decades of market use.
What you are talking about is value range analysis. Rust does a limited amount of it, and if it can prove safe indexing it will remove the runtime checks.
Moving to an abstraction would potentially solve some overflow issues and requires constant attention. Moving to a safe language likely solves all of them and you get checked by the compiler instead.
Rust and Go are not your only choices here. C++ with aggressive use of modern features and aggressive non-use of everything inherited from C is also a fine choice, but enforcing that discipline is hard, and the compiler won't help you a lot. Honestly, a dynamic language like Python or Ruby is also a fine choice in many cases, although NTP might be too latency-sensitive for that to work. (But it might not be! Premature optimization and all that.)
False. Buffer overflows in C can overwrite the program's memory, so it can be hijacked and supplanted with the attacker's code. This cannot happen in Rust (unless unsafe code has the vulnerability), or any memory safe language.
Sure you can implement a safe array/buffer abstraction and use it in your C programs that abort on invalid indexing. Now how many actually do this? Very few given the prevalence of C programs on vulnerability disclosure lists.
Idris.
Basically, in Idris I can have a function concat which takes a Vector<N, T>, a Vector<M,T>, and produces a Vector<N+M, T>. You can have arbitrary expressions there, and even things like a function which returns a different type based on its boolean argument.
So most signatures will just carry dependencies through, but some things are dependent on runtime input and that's when the "N" part of the type will be tracked at runtime. At least, that's how I think it works.
Dependent types help eliminate large swaths of bounds checks. Not all of them, and I'm not sure how much better it does than a good optimizing compiler.
It can be done by static analysis by compilers in other languages also. Knowing size at compile time is easy most of the time and we are not talking at all about this case. Your answer that Idris solves that is wrong in the context that this discussion started (dynamically allocated memory/not knowing the size compile time). You still need to trade time/space comparing to solution without bound checking, doesn't matter which language you use.
I am getting info that I am submitting too fast? so answer to your other comment will need to wait.
The nice thing about dependent types is that in all the cases where you don't need a bounds check -- where a program invariant enforces that you are within bounds -- your code will not contain a bounds check regardless of optimization. Only in the cases where there are runtime dependencies that can't be resolved will there be issues, and in thee cases you need a bounds check anyway.
It's not perfect/complete, and like I said optimizers can get most of the way there anyway. I just wanted to note that there are better solutions for bounds checking.
I think the history of exploits proves that it's always acceptable for server code. The only possibly justifiable place to omit bounds checking is isolated, high performance numerical code, as used in scientific simulations. But this code isn't exposed to the public, so its vulnerabilities aren't important.
You need unsafe to do some things in Rust, sure. But usually this code is isolated and auditable, and nowhere near the linecount of the rest of the application. Even the operating systems written in Rust have pretty conservative use of unsafe (redox, Phil's OS, etc). Most of your code won't be working with raw memory operations, most of it will work with zero-cost abstractions on top of raw memory.
And that's what I wrote.
> Most of your code won't be working with raw memory operations, most of it will work with zero-cost abstractions on top of raw memory.
For some web applications sure but industry is not constrained only to writing web applications especially if you want to get into C market space. In HPC or time/space constraint environments this matter.
You might want to define further what you mean by "raw memory here". Rust lets you work with arrays and vectors and the heap safely just fine. These are designed as zero-cost abstractions over raw memory (slices, Vec<T>, Box<T>). The equivalent C (with relevant bounds checks) wouldn't be any faster. Rust does not let you do things like call out directly to malloc/free safely. But that's okay. The existence of these abstractions means that you rarely have to do this.
> And that's what I wrote.
.... huh? that's exactly why I put "sure" there, that means I agree with that statement, what followed was why I disagree with the conclusions.
Trading efficiency for safety for unaudited code should be difficult. C does not make this difficult. Rust makes this difficult.
Auditing should be supported for this kind of code. C does not support auditing in any meaningful way. Auditing much easier in Rust since it explicitly identifies code that may be unsafe.
In conclusion, I disagree emphatically that C and Rust are somehow equivalent even when you are dropping down to unsafe code.
Check out https://en.wikipedia.org/wiki/Return-oriented_programming for more information.
Rust arrays/ vectors are safe-by-default. To use the unchecked, unsafe version requires using the 'unsafe' keyword.
let v = vec![0, 1, 2]; unsafe { let x = v.get_unchecked(5); }
This means you can basically grep audit for vulnerabilities, and the above code should be very rare.
Neither does Rust.
> The subscript checking variants of C and C++ have to use "fat pointers" which carry along size information.
So do Rust's Slices.
> The overhead for this is large and nobody uses that.
People use std::vector all the time for this purpose in C++. It has about the performance you'd expect, with very little overhead except where you want it in bounds-checking.
I don't think there's actually a performance difference here. Rust's default is safer because it requires dropping to unsafe code to do something dangerous, but the same optimizations are available in both.
Again, I haven't actually dug into this; maybe someone more knowledgeable about this can point me in the right direction here?
I imagine Rust does something similar, copying bytes if the underlying type has the `Copy` trait and calling some actual code if not, but I'm not familiar with the details.
[1]: http://en.cppreference.com/w/cpp/concept/TriviallyCopyable
> copying bytes if the underlying type has the `Copy` trait and calling some actual code if not,
It does not. Moves and copies are both "memcopy these bytes", the only difference is if you can use the previous copy or not. (This is also, of course, subject to the optimizer, which may elide the copy.)
> either way some constructor of the object must be called if it exists (though it might be inlined and optimized away).
Yeah, this is what I was getting at; this has to happen in C++, but not in Rust. You are right to point out that this only matters for things that aren't trivially copyable.
It would be cool to have those things, but it also means that there's less "magic" stuff going on, which is nice. And it makes the semantics of stuff like this a lot simpler.
> Does `drop()` get called on objects that have been copied from?
Nope. In fact, Copy types can't have a Drop at all, but types that move don't have their Drop impl called when they move.
struct Foo { s1: String, s2: &str, }
where s2 is always intended to point at s1's backing storage. What's unfortunate here is that Rust will disallow this, as it doesn't understand that s2 is pointing to some data on the heap, with a stable address, not the parts of the String struct in s1 that are part of the struct itself. So what this means is, in plain Rust, this type isn't movable, Rust is concerned about the invalidation.
However, you can get around this restriction with some unsafe code to teach Rust about it; this is the premise of the "owning-ref" crate.
This isn't the case. They only need to borrow it. A borrowed value isn't allowed to move (the borrow itself can be copied around and shared within the scope of the borrow), so that works out.
Most Rust iterators only borrow the container or iterator they operate on. It's only explicitly moving iterators like .into_iter() (which extracts elements by-move) that don't.
Rust's answer to copy ctors is Clone, which is always explicitly called. Variable use in rust is a move. Trivially copyable types (Copy) will be copied without compile-time invalidating the old type.
Lack of move and copy ctors in rust greatly simplifies things like this, and makes it very explicit when code is running, but the trade-off is that intrusive data structures are hard to do on the stack in rust.
The self-reference would point to the old object. See this for an example: http://ideone.com/sEFtbN
(Intrusive datastructures on the heap are doable with some tricks.)
They're not very essential so it works out.
I have rust installed. I would be curious to benchmark it and see if is indeed faster than the same thing in C, using a user-defined bounds checked array.
I'm not feeling very well unfortunately, not really in the mood to code.
Hope you feel better soon.
Of course, bounds checking is worth paying a cost! As are many other common foot-guns like remembering to free memory. Which starts to raise the issue of when it would actually be necessary to reach for Rust instead of (arbitrary example) Java. Because a language with enforced bounds checking is not really the same kind of thing as C, and we've already had languages that are safer than C for decades.
(Languages with dependent types can help move bounds checks to compile time, but Rust does not have those)
Rust (which I have thus far only tinkered with) looks pretty appealing to me, though. Do you suggest using C instead of Rust? If so, why?
I'd probably suggest a higher level language if you can get away with it (portability, performance, etc), but I am not really advocating for using one language over another here.. What I am advocating is writing a thin layer of abstraction to avoid buffer overflows in the already extant code base. I suspect it would be a better choice in terms of cost/benefit than re-writing the code.
Who's saying this? It doesn't really make sense, taken out of context as it is here.
> inform me that rust has "Zero-cost abstractions"
IMO at best the array bounds checking thing is an extremely poor example of a "zero-cost abstraction"; at worst it's not an example of it at all. As I understand it, the term refers to things like static dispatch on closures and trait methods, iterator fusion, and generally any place where the compiler can transform high level abstractions into really efficient low-level code where in other languages/implementations you might incur a performance penalty, for instance by dynamic dispatch to heap-allocated closures, or allocation of an intermediate vector at every "link" in an iterator method chain.
If you want a true speed comparison, you'd have to define more than just "read an arbitrary string". For example, encoding requirements. Your C is going to treat it as just a bag of bytes, and you can do that in Rust too, but it's not the most natural way; you'd convert to UTF-8, which isn't free. Stuff like that.
use std::io::{self, Read};
fn main() {
let stdin = io::stdin();
let mut data = vec![];
stdin.lock().read_to_end(&mut data)
.expect("Reading stdin failed");
}Also, "safe by grep audit" means "safe according to a human." The argument of course is that it lowers the surface area of what a human must be trusted to verify. I'm still not convinced by that argument, because human error is a thing. And for actual systems programming, "very rare" may not be true.
Certainly. I didn't intend to imply otherwise.
> Also, "safe by grep audit" means "safe according to a human."
Again, totally correct here.
> The argument of course is that it lowers the surface area of what a human must be trusted to verify. I'm still not convinced by that argument, because human error is a thing. And for actual systems programming, "very rare" may not be true.
Well, given a codebase where both safe and unsafe code exists, the amount of unsafe code is strictly less than the amount of both safe and unsafe code. So it does reduce the amount of code needed to audit, even in a very atypical case where a ton of the code is unsafe.
It's true that a project may use egregious amounts of unsafe. That would be unfortunate. Rust is still safer than C in that case, since it just defines more behavior (like arithmetic overflow), but I certainly wouldn't pretend that the rust code should be trusted.
When writing rust one should certainly strive to write less unsafe code, and to always document the invariants required for unsafe code to be safe.
Rust is not 100% safe 100% of the time, I'm only arguing that safe defaults are critical, and that grep auditing is a powerful tool.
In general when talking about safety in a language it's about the level of explicitness required to trigger unsafety.
I like the distinction made in the nomicon (https://doc.rust-lang.org/stable/nomicon/meet-safe-and-unsaf...) -- Rust comprises of two distinct languages. You have everyday Rust, which is completely memory safe, and "unsafe Rust", which looks similar to everyday Rust but is not safe. `unsafe {}` blocks are your FFI between the two. Looking at unsafe blocks as FFI is IMO a very useful mental model especially for understanding the changes to invariants involved.
AFAIK, cargo doesn't have any feature to point out when a crate contains unsafe code—so you pretty much need to grep the source of every crate you consume for "unsafe".
The issue you're talking about is related, but different.
Auditing FFI is a whole other challenge, however :(
Not in binary libraries, hence why it is important to have a culture to only use unsafe if it really must be used.
You do have C libraries which you access through FFI. This is inevitably unsafe. We should be auditing more there. Though IMO it's still manageable, for most crates.
For example, most OS code is _not_ interrupt hooks or malloc but the rest of the OS. Most of Postgres is not reading data quickly, but higher-level abstractions.
Large-scale systems programming will always be mostly higher-level abstractions, because that's the only way to write large programs. Name any "systems programming" OSS project, choose a random C file, and you'll see that most code does not require pointer arithmetic except because that's how you do things in C.
You don't bit-twiddle for 200k lines of code, so being able to limit dangerous stuff like that to the 5k lines that actually need it makes work an order of magnitude easier.
True. Or slightly rephrased: Most of the unsafe code is centralized to some core pieces.
From the POV of a PG dev: The big problem using something like rust for something like PostgreSQL is its it's portability, stability and uncertainty about where things are going. We do five years of back-branch releases (and for many that's not even enough!). Language and tooling around the new crop of languages simply aren't mature enough for that yet.
So if you are in a so performance critical section that you start to care about the "zero overhead" kind of stuff and the cost of your (predicted) bound checks, you might be impacted by this non-zero overhead... (that is falsely widely believed to be zero!)
And don't get me started about debug builds with all mainstream compilers. The performance is then complete utter shit. This is made worse by the fact such an unsafe language needs debug builds more...
C++ was an interesting experiment in its domain, and has been and still is a success in some aspects, but honestly given core language modifications are needed to significantly extend the behaviour of core types (like vectors, unique_ptr, strings, etc. -- e.g. with the introduction of rvalue references and all the associated machinery and default constructors and so over), I'm starting to believe there are little advantages of this approach over integrating such fundamental types/concepts in the core language (I'm not advocating to do that for C++, I'm thinking about other/future languages). Then you can have more true "zero overhead" stuffs, whatever that means.
If I don't have this kind of performance needs, I have no reason to use C++ to begin with. I mean, why use such a monster if I don't need the crazy performance it may offer when used properly? And even then, C, D, or Rust may be viable alternatives.
The idea of ANSI C++ was to provide safety behind library calls instead of compiler switches.
I think I'd rather not have it panic at all. They were making a language from scratch and still didn't solve the crash-at-runtime problem that C has. Instead they just made a common C compiler option default and called it 'safe'.
It's worse than the fact that the Go creators ignored years of research after the year 1970 and didn't implement generics, for no good reason.
Rust's safety is a joke, and so is the idea that a lot of it is anything 'new'. They had the chance to fix it and they didn't. Boo.
My point is that panics are memory safe, contrary to what my parent said.
Before you say Address Sanitizer you should be aware that its authors intended it as a debugging tool and recommend against using it in production as it introduces new attack surface, as described in the "Address Sanitizer local root" mail.
You can also catch this panic at the thread boundary, so other threads in your program can keep running, unless the program was compiled to call abort() on panic.
(The alternative instead of .unwrap() would be to propagate the error, polluting the whole program with code to handle errors which never will happen. And since they will never happen, many programmers would simply start ignoring them - and in the process, they would by accident end up ignoring errors that can happen. Not a good situation.)
Looks pretty safe to me.
A panic isn't safe, and if you are going to claim that it is, you can get the same safety in C using gcc and clang features.
There's one important difference. As far as I know, the bounds checking of gcc or clang can either print a warning and continue, or terminate the process.
A panic in Rust, however, will safely unwind the stack (similar to a C++ exception, in fact it's the same mechanism) and terminate only the thread. The rest of the program can continue running, and even start a new thread to replace the terminated one.
You can see this in action when running "cargo test". A panic in a test (the assert!() and assert_eq!() macros, often used in tests, do a panic in case of failure) will not terminate the whole process; the rest of the tests still run.
People do not use the sanitisers nearly as much as they should, and nor are they designed to be used in production. There's has been CVEs issued for them caused by the testing mindset in which they are written.
In any case, how much a programming language helps the programmer get a correct program (I.e. the most general form of safety) is a thing of degree, not all or nothing. Rust does a lot more than C, pushing many many errors to compile time, meaning they're fixed before the code even runs. Sanitisers only catch these when they actually happen, and will presumably result in the runtime crash you're so concerned about, if used in production. The fact that Rust isn't dependently typed and so can't catch OOB at compile time is unfortunate in this respect, but don't throw the baby out with the bath water.
Scenarios where I frequently end up fixing other people's memory errors:
1. No Error Handling: not checking an error condition on a function that allocates, then using the uninitialized pointer anyway
2. Sloppy Error Handling: jumping to abort from an error without freeing what has already been allocated
3. Faith in \0: still using the old string functions
I'm on the fence about the whole thing, so others may be able to field something more compelling.
2) I can't remember if a panic in Rust calls destructors, which would clean up that memory. Can someone answer that please?
3) is only relevant in FFI scenarios in Rust, and everywhere else is irrelevant because Rust does not require the use of such footguns for basic string manipulation.
More exotic panics (like an infinite loop panic on a microcontroller) would not call destructors though.
Most panic impls will either unwind (which will call destructors) or abort/exit, so this is usually not a problem.
Why allow programmers to make mistakes? That was fine in the 70's when resources for compiler execution were limited. I don't see any reason for it today.
I mean, just look at the underhanded C contestants and especially winners for ways in which your program can completely blow up for extremely subtle reasons.
Isn't this also true of Rust, with its unsafe keyword? None of these languages are completely safe against programmer mistakes.
It's used sparingly. Not as sparingly as I'd like, but sparingly enough.
I think that makes a very significant difference.
Yes, carelessness in either could amount to trouble. No, they are not close to the same amount of risk because the relative amount of code in each is orders of magnitude different, and unsafe blocks can be heavily reviewed.
> None of these languages are completely safe
No, but the first time I encountered an unexplained segfault in Rust was the first time I've ever found & fixed the cause of a weird segfault message in just a few minutes, because the mistake was in the 12 lines I had written inside a clearly marked unsafe block rather than the other ~3000.
Without something like unsafe blocks, finding known bugs and auditing code for other "weird" memory issues means looking much wider (and often, screaming "what the hell are you talking about?!" at the screen on a routine basis).
For a philosophical counterpoint: Why allow anyone to do anything that might possibly be incorrect, harmful, or otherwise perceived by some to be negative?
I've looked at a lot of the talk surrounding "safe/secure languages", "safe/secure programming", etc., and yet every time I've heard people preach about the benefits, I feel like I just vehemently disagree. At a very deep and fundamental level, I feel like somehow we are sacrificing something more important in the pursuit of this "safety", this seemingly overpowering desire to make everything completely safe, mindlessly constricted, and stifling. It's not just software; the whole "war on terrorism" irks me in the same way. I imagine a "completely safe" world, the ones these "safe software" proponents appear to be striving for, would be rather dystopian.
A quote that immediately comes to mind is: "Freedom is not worth having if it does not include the freedom to make mistakes."
My serious point is that in practice, the example of C shows that if it is available and people understand that it is "performant" then you will see it all the time, including in libraries you are forced to use.
This isn't an accurate analogy at all because I can compile thousands of lines of Rust code without even thinking about needing to use `unsafe`. How many lines of C code can you compile without `cc`?
> My serious point is that in practice, the example of C shows that if it is available and people understand that it is "performant" then you will see it all the time, including in libraries you are forced to use.
Except this is already demonstrably false in the Rust ecosystem.
And there is no shortage of mistakes available for you to make, so don't worry about that.
Are engineering standards for public buildings evil because they stifle architects' freedom to design whatever crazy structures tickle their fancy? Should power tools not include safety features like blade guards because their users' freedom to accidentally kill or maim themselves must be held sacrosanct? Probably to both questions, the reasonable answer is "no".
If you're building something for yourself, and offer it to others only with clear warnings, then go nuts and make all the mistakes you want. Nobody's saying you can't do that; that sounds rather dystopian, a society where pointer arithmetic is illegal! But you can't treat a project people are meant to trust and rely on as your personal art project.
The physical analogy is good because even there one can see that there are different standards --- and, unlike what the "safe software" community seems to promote, engineers are not doing the equivalent of making every building strong enough to withstand a nuclear war and calling anything less "unsafe".
This also brings up another difference with software: the "absoluteness". In the physical world, no security is perfect. Locksmiths exist, and with enough determination, essentially anything can be broken into. But in the software world, with good encryption, that can never occur. Provably correct software can be employed for effectively unbreakable DRM and un-rootable/jailbreakable user-hostile devices, of which there is (fortunately) no similar real-world analogy I can think of.
Nobody's saying you can't do that; that sounds rather dystopian, a society where pointer arithmetic is illegal!
Given that there are attempts at even prohibiting Turing-completeness[1][2], I would not be surprised if that eventually happens. As it is, I'm sure there are already people who would consider you suspicious if you write software in an "unsafe" language, and from there it is not far to complete prohibition.
"The road to hell is paved with good intentions."
Yes. Except the equivalent in software is using something like Coq, where your intention is proven to be correct. (You even mention provably correct software in the next paragraph...) And yes, it would be awesome if we could prove every program, because all we would be doing is proving that the programmer's intent is correctly coded into the software, which doesn't limit freedom of expression at all. (I would argue that such is actually an even stronger form of expression, as it's afterwards impossible to misconstrue your intent.)
> Given that there are attempts at even prohibiting Turing-completeness[1][2], I would not be surprised if that eventually happen...
Rust is not one of these, and no one but you seems to have this in mind. So it's a bit of a strawman to be throwing it out here, no? But yes, obviously such a language would limit your freedom of expression.
You can extend that absolutist argument to say that inventing anything that could be used as a tool of oppression is unethical. Like, inventing plumbing may have done great things for human society, but it was ultimately unethical because when the police came to take away your general-purpose computer they used a pipe to hit your kneecaps until you told them where it was. Or, how about this: developing any society beyond the level of the most primitive hunter-gatherer tribe is unethical, because what are governments if not the agents of oppression themselves?
This is the sort of abstract position that can't really be argued with in a vacuum, so I'm not going to try. But it's also not a useful ethic for building a modern society free of large-scale oppression, because it completely ignores the practical realities of doing so.
In the mean time, buggy, exploitable software is out there in the real world hurting real people every day.
There is no philosophical beauty in typos or off by one errors. You are not being stifled when your compiler points out to you that, no, "coutn" is not a variable currently in scope. If we have a static proof that what you just did will never work, why would you still want to do it?
And if you had a reason, you still have a way to do it.
If you're trying to be clever, you need to prove that you know what you're doing.
We're talking about eliminating a class of programmer error through static analysis. (If you're writing a large C program, you arguably should be using static and runtime analysis tools on it anyway.) This is no different than using a strong type system. Yes, it limits your "freedom", but sending a "banana" to a `sine` function is nonsense anyway. Well, it also ends up that reading 10 bytes into a five byte array is also nonsense, and we now have tools to detect that nonsense and tell you that it is, in fact, nonsense.
Y'all please, please note that dreta said "... an array implementation that prevents that from ever happening..."
When the figure of "ten minutes" was used - that's really what it should be as a mean or median figure, with some long-tail outliers for knotty cases.
While I am (somewhat) sympathetic, I don't miss what it is - it's make-work. There's not some body of material on language, O/S and software design available today that wasn't available 40 years ago. People new better back then, and didn't do better.
But the endless, circular, Utopian arguments just don't cut it. It ain't the crate, it's the pilot. I am sorry that oh so many are put upon to actually - gasp - test their code but that's what this is.
The economics of a Rust or Go or whatever make perfect sense - in the long run. Just recall what Keynes said about the long run.
Just to be clear - I don't suspect that the things that pretend to replace C are bad, or no good. I've just seen multiple "pretenders to the throne"[1] and what would appear to have happened is that we just moved the pathology around.
[1] please excuse the horrible metaphor.
After a few iterations of that, one begins to think that perhaps this is a distinctly human problem that is not necessarily addressable by improved language systems.
My greatest concern is that I keep seeing people learn the same things, over and over, on the job. There's no real repository of literature to actually address any of this - each engineer appears to have to learn it mostly from scratch.
There's a famous bon mot from pubic choice economics - "something must be done; this is something; this must be done." I just hope all the new crop of languages are not that.
Finally: If you can, and if you can learn the right patterns ( isn't that true of all languages, though?), the thing I have gone to, again and again that is better than test equipment in terms of reliability continues to be Tcl.
It is called university and engineering degree in software development.
So long as one can convert it in a straight-forward way that preserves the current meaning of the program's statements. If not, the port might introduce new problems.
Edit: Btw, send me an email (see profile). Got something you might be interested in. One or two things.
That (email thing) would be awesome, Nick. Thanks.
- be implemented with macros and token pasting, and result in a ton of mental overhead because you'll have a pile of types like array_foo for an array of `foo`s, and array_bar for an array of `bar`s, along with a pile of corresponding `foo * array_foo_get(array_foo, size_t)` and `bar * array_bar_get(array_bar, size_t)` functions.
- or, have a runtime cost and lose type safety by storing void* and casting when accessing.
The first case is even worse than it sounds: e.g. I don't know how you handle arrays of types with spaces in them (like `unsigned char`, or `struct bar`) with a macro. And, we haven't even thought about const correctness yet, which would probably require having const_array_foo, const_array_bar (etc.) types defined too.
(And, of course, these only solve one facet of the problems with C's pointers: there's no way to defend against use-after-free or dangling pointers.)
alloca-style variable arrays is a whole other can of worms of danger and complexity.
https://gcc.gnu.org/onlinedocs/gcc/Variable-Length.html
Legal in C99, available in C++ on gcc/clang.
Pointing into subsections works fine. You just have to create a type for it. This solution doesn't have the same problems as strings because you don't rely on a terminating entry, and it's what languages like Rust or Java do as well.
You can allocate dynamic arrays on the stack in C just fine with alloca(). The only performance cost is when checking bounds, but since it's a dynamic array, it's the same cost you'd pay in Rust.
> A beginner will often try something along the lines of size = sizeof( myarray ) (which is incorrect).
Additionally, all the functions should be inlined anyway because the function call overhead will likely be as much or more than the actual code, and, more importantly, inlining enables other optimisations (removing the branch, vectorising the memory access, etc.). Once inlined, the code will be the same as the manual/macro-based approach of writing `if` statements around each array[index] access.
They're not an ideal solution by any stretch, but it's not the nightmare scenario you envision wrt generic data structures in C.
typedef unsigned char uchar;
DECLARE_AND_IMPLEMENT_ARRAY_API(uchar)It does make debugging a chore. I ended up with a Makefile rule to run the test suite through the preprocessor (and did some hackery to exclude #include of system headers), format it with clang-format, and build that. Not exactly pretty or easy, but it got the job done.
1. https://github.com/alpha123/yu/blob/master/src/yu_splaytree....
2. https://github.com/alpha123/yu/blob/master/test/test_splaytr...
3. Where this sort of thing is much easier and I am happier and more productive.
The CVE database. Just because you 'can' write such an array implementation doesn't mean you will, doesn't mean your third party libs will, doesn't mean any of your legacy code uses it, and certainly doesn't mean you will properly test said array implementation correctly.
The number of mitigations added to C compilers and OSes dealing mostly with C and C++ code. ASLR, W^X, /GS, -fstack-protector-all, AddressSanitizer, ... - note the lack of similar tools, or demand for them, for, say, JavaScript - despite it enjoying a similar ubiquity.
I ask this in bad faith: I encourage you to share a single nontrivial codebase which actually creates the abstraction you've described and religiously adheres to using it throughout. As to why this is in bad faith: I'm definining "nontrivial" here to mean using 3rd party APIs - which will operate on C style arrays, not your project specific safe wrappers - and thus by definition won't be "religiously" sticking to said abstractions when using said APIs. By these definitions, the codebase I'm asking for doesn't exist - by definition. Even relaxing the "third party" rule, I haven't actually worked on a nontrivial C or C++ codebase without buffer overflow problems.
Now, e.g. Rust will have the same problems when interacting with C APIs - and nontrivial programs will end up doing so eventually. However, by virtue of the language itself embracing safe-by-default, you're less likely to run into the same problems when consuming Rust APIs.
You can also use third party static analysis tools to ensure you're using a "safe C subset" (such as MIRSA C), but "nobody" does that.
That one "if" is (by definition) not zero-cost.
> C++ implementations obey the zero-overhead principle: What you don’t use, you don’t pay for. And further: What you do use, you couldn’t hand code any better.
Two points:
What you don't use, you don't pay for: if you don't use array indexing, you won't get a bounds check. In addition, you can call an access method without a bounds check as well, so it truly is only if you use the checked version.
What you do use, you couldn't hand-code any better: that bounds check is written the exact same way you'd write it in C.
Therefore, this is a zero-cost abstraction.
The Haskell `newtype` example I gave was meant to illustrate this, as newtype's are respected by the type system and then are treated as the underlying type at runtime.
Most code doesn't use bounds checking, because the branch is a safety net you should never hit, even in theory. Any code that does hit it is already broken. Correct programs using bounds checked indexing will in general be slower than but equivalent to a program where indexing instead results in undefined behaviour.
"You couldn't hand-code any better", well, I won't argue on that point, as it sounds contentious. ;)
_Should_ never hit is very different than will never hit...
My point is that Bjarne Stroustrup wasn't comparing against writing the exact same program the exact same way. He was comparing against what you'd get if you dropped down to ye olde C or Assembly and wrote the same algorithm there, without redundant work or waste.
The comparison shouldn't be the language's GC versus SteveGC, it should be the language's GC versus an ideal, manually implemented allocator. Equally it shouldn't be built-in bounds checked indexing versus manual bounds checked indexing, it should be built-in bounds checked indexing versus an ideal, manually implemented indexing scheme. If you want safety against out-of-bounds, it seems to me the ideal method would be a proof, not runtime overhead.
I don't know of a single language that comes with a GC that does this, do you?
> He was comparing against what you'd get if you dropped down to ye olde C or Assembly and wrote the same algorithm there, without redundant work or waste.
Right. I agree with this.
But basically, we are arguing over an extremely fine semantic, which is "should you even want bounds checks in the first place." If you don't, then don't use a method that has bounds checks. The one that does will have them. They'll both cost the exact same as writing it in C or assembly.
To put it another way, let's say I was a C++ developer on the fence about Rust. If I read this conversation, I'd see that indexing gets called "zero-cost" despite the overhead. Since tons of things in Rust are "zero-cost", like traits, closures, borrowing, etc., all of those things now have doubt cast on them. How can I really trust that these things are actually getting compiled efficiently?
If instead the conversation pointed out that this was one of a few cases where safety took priority over truly being zero cost, but that there were tools in place to mitigate the cost (iterators, unsafe indexing, LLVM), I'd have a much more positive outlook that focussed on what Rust did right.
That said, I can appreciate focusing on other things when talking about the principle; I only brought them up here because we were literally discussing them. I think there's much better examples when actually attemping to convince someone.
Java, .NET, Go, ML and Lisp compilers.
Escape analysis allows to do that, even if just in certain special cases.
Plus the more one uses value types and less heap, the GC needs to work less, specially if we take languages like Modula-3 into this mix.
Escape analysis may not use the GC for those variables, but it is still a pervasive runtime cost in both senses unless it's totally gone.
Why no malloc? It was originally due to the small memory sizes for code and data. The less standard library the better, and dynamic allocation may lead to heap fragmentation and a subsequent crash when malloc fails.
The problem isn't the language it's the developers.
New languages here and there every day. Replace this replace that. When, in the end everyone is simply reinventing the "wheel" over-and-over.
All these languages end up as assembly.
I honestly don't know who I'd put my money on between "AI takes over the world" and "programmers stop writing buffer overflows"
The "heartbleed in rust" example is a great one, and it arises in real life in many high level language APIs for file I/O and sockets. You have an allocation, and you have a count of available bytes coming back from a read() function which may be lower than the allocation size. So you are creating a "virtual" array bound from nothingness. Fail to respect it (without bounds checks) and you will see bugs.
If you reject that this is a valid way to write code, maybe in your API every read() style function will always return the correct size enforced by your JVM or whatever, but you will do too many allocations and over-tax the GC.
If you accept that this makes sense, then you must embrace a more C style way of thinking, where array bounds are created and destroyed at will and must be enforced through your own actions... And suddenly you see the other side of this coin, which reflects valid and true things about the universe, that you may want to chop up a buffer into multiple pieces - and that's OK.
(Now, I wouldn't be surprised if Rust has mechanisms to chop up arrays in the way I describe and enforce the bounds you provide it... Which would be handy. But frankly does not completely destroy the validity of the C approach or substitute for a proper understanding of it. Without that understanding, you will code more heartbleeds.)
[T]::split_at is probably what you're looking for.
Almost all array handling in Rust is done through slice types which are tagged with sizes.
However, you can't bungle the creation of a slice in rust without using explicitly marked unsafe code.
And it's not limited to the VM itself: check out npm "native extensions" like `json`. Not to mention glibc, or the OSes themselves.
By your definition, nothing is safe. And you're right ;)
I note that Firefox is using some Rust code now - so perhaps that will change at some point, for at least one of the common JavaScript implementations, in the not too distant future. I don't imagine we'll see it for the majority within the decade - but who knows, maybe I'll be pleasantly surprised.
I have less hope for the widespread adoption of OS kernels written in safer languages - given the general unwillingness to even use C++ there (although plenty of toy/'research' kernels in safer languages do exist.) Although maybe we'll see one within the next century? Perhaps a microkernel for use in containers?
Of course, that still leaves bugs in the JITs, compilers, hardware, 'legacy' native interop, unsafe{} blocks, ...
Maybe
I think currently there's no plan for it. Maybe after they finish servo to the degree where it supports all modern html features
I work on a C codebase that does this, although in the slightly weaker sense that it does drop the abstractions at a few isolated interaction points with external APIs (think openssl, linux system calls, and not a whole lot else). Yes, there is quite a lot of NIH. With essentially-uniform use of checked data structures, and an extremely comprehensive suite of automated tests getting run under ASAN (originally Valgrind), memory safety errors almost never get so far as being committed to the main branch. This is a complex, >1M SLOC distributed system that has seen several years of production use at this point, and as far as I can recall we have not seen a single memory safety related issue in production (a few have managed to to get as far as certification testing). General resouce-leak class issues have struck a few times, but are also pretty rare.
Proprietary, naturally, so I can't actually show you (sorry), but it absolutely can be done in practice. It isn't even really all that difficult, it just needs to be done from the start, and then you just need a bit of discipline to keep it up.
And more power to you. Note the beginning of the parent comment, however:
> Just because you 'can' write such an array implementation doesn't mean you will
So yes, even if the codebase you work on does have these 'mythical', hard-to-achieve properties, that doesn't mean that most or even many C codebases will.
Good engineering entails observing what problems actually occur and working to fix those. Memory safety issues do commonly occur in C codebases. Regardless of whether the fix in C is simple or even trivial, programmers aren't doing it. So, Rust has some value because it forces the programmer to produce code that is largely free from this type of issue.
Enforcing norms like 'be more disciplined when writing C' or 'stop using external libraries' is much harder than simply using a different language.
>
> […] it just needs to be done from the start, and then you just need a bit of discipline to keep it up.
Here's a neat idea: wouldn't it be cool and save a lot of time if the compiler did this for you automatically, from the start?
Of course, you say it's easy to do it manually, but something tells you your company might have needed to pay less for development if the compiler did it automatically with no human intervention required.
Rust, at least in this regard and probably others too, is no better than C, and for me it isn't enough to justify the horrible and complex syntax.
Rust checks at runtime and panics if your program exceeds the bounds. You can opt-in to asking if the bounds are exceeded and fail gracefully if you like, or if you want to promise the compiler you know for sure your bounds are tight, you can use unsafe blocks and act like C. Opt-in to danger.
C lets you do it with no checks. You have to opt in to the safe path of checking and failing. You don't automatically segfault if you exceed the bounds; instead you read arbitrary memory. Welcome to the land of undefined behavior. You may crash, but more likely, you will read some value from an unexpected place, and carry on executing incorrectly for who knows how long. Opt-in to safety.
That's what people mean by Rust is safe by default. And that's just bounds checks. Carry that notion over to pointers, references, threads, lifetimes, ...
What ever made you think "Safe Rust" meant "compile time checks of runtime values are possible" or "C is just as safe because it lets you index outside an array"?
Optimizers already try to prove index bounds to eliminate unnecessary checks, and static analysis tools to demand necessary checks. Turning the latter into compile time errors is a reasonable approach if your language can provide sufficient information to deal with false positives - likely by forcing you to add your own bounds checking to explicitly handle out-of-bounds cases.
A language John Carmack was using or researching at one point comes to mind, which had this kind of thing going on IIRC. I'm afraid I can't find it off hand, so I might recall in error.
And the point of this thread is that Rust's default (check index bounds and fail at runtime unless the check can be proven unnecessary by the compiler or optimizer) is safer than C's default (don't check anything by default and hit UB if the index is out of bounds).
GCC and clang both have sanitizers either built in or available for them. Sure, it's not default, but let's not act like there is no choice in C but to account for every OOB access while programming or to read memory you don't want to.
Furthermore, I never said that compile time checks of variables are possible, but rather we could move to using dependent typing, or at least a way to judge whether a variable would work as a subscript based on the type of the array and variable.
The Rust designers didn't do that. Instead they put in a feature common in the two most popular C compilers to "panic" at runtime instead of accessing memory. That's nothing. It's rubbish. And if you know to "catch" the panic, why don't you check the value of what you're subscripting with? Saying you can catch the panic is missing the point of unintentional OOB accesses, which is that they're unintentional.
C is just as safe with regard to OOB accessing, and to be honest that's pretty poor in 2017.
If you are using gcc or clang, you have more options. True. But not all C compilers give you those options. However, the point is moot, since I never said you can't catch these things in C; I said it wasn't the default. Which you agree with.
> I never said that compile time checks of variables are possible, but rather we could move to using dependent typing
You didn't say anything about dependent typing. You said "Rust is no better than C". And I'm pointing out that it is. Dependent typing may be even better in some cases; I'm not arguing otherwise.
> Saying you can catch the panic is missing the point of unintentional OOB accesses, which is that they're unintentional.
No one said you should catch the panic. You can use Vec::get() for example if you are using runtime-derived indices and want bounds checking in an ergonomic fashion.
And saying a panic for unintentional OOB is the same as in C is not true, since you get a panic by default in Rust, and to get one in C, not only must you be using a specific compiler or two, you must have the sanitizers enabled for every source file in your program. Not "by default" by any stretch.
> C is just as safe with regard to OOB accessing, and to be honest that's pretty poor in 2017.
It is nowhere close, and saying it is is pretty poor in 2017 as well.
And you are still ignoring pointers, references, lifetimes, threads, ... you know, the other things that also help Rust make "Safe by default" and C "dangerous by default".
https://doc.rust-lang.org/std/primitive.slice.html#method.ge...
C will actually read arbitrary memory, Rust won't, that's the difference.
We're talking about a situation where the length of the array is not statically known and you access it out of bounds at runtime. Rust checks first if it's out of bounds, and if it is, DOES NOT blindly read the memory anyway (as C would) but exits.
I don't understand how you could think the two situations are at all equivalent.
Yes, but when the project started the only existing compiler that met all requirements was C (also C++, although that was not chosen, by reasoning I disagree with). We are in a domain where we derive material benefits from the low-level control C gives us (we have a bunch of highly specialized memory management and I/O), and are not willing to accept GC pauses. There's a common sentiment that we would have used Rust if it had existed when we started, but it didn't so we didn't and so it goes.
1. Buffer overflows aren't considered the most insidious issue in C nowadays. That award would probably go to use after free, which is not so easy to fix.
2. In C, it is easier and faster to do the wrong thing. Compare "char buf[256]; strcpy(buf, foo); ..." to "array_t buf = array_create(strlen(foo) + 1); strcpy(buf.ptr, foo); ... array_destroy(buf);"
3. Buffer overflows do not in fact come up routinely in Rust the way they do in C.
Also it is actually a joke, since it still separates the pointers and length in two separate variables, instead of using some kind of struct.
The only thing it does is have functions with better semantics on the terminating nulls.
That it's not a by-default and forced language-feature and that most developer aren't going to spend those 10 minutes when they need an array.
They'll just use the language-provided array-implementation instead. Which in C is very, very unsafe.
Really? This sounds like idiomatic rust to me (heavy with enums).
It's a surprisingly important feature. Every large codebase I've worked on has clunky workarounds for storing heterogeneous types in collections.
Haven't seen a silver bullet; dynamic languages are great at this until you have to scale (either in lines of code, number of types or dataset size) and then they become unmanageable. Static languages force you to type a lot and build in-memory ETL-style transformations but scale better. Using SQL is another solution, but that separates you from the language's type-checker and can be another source of error.
Type safety at the serialization boundary was the promise of CORBA and protobuf but we need better tools. I hope we see more focus on ser/des types in the next generation of industry languages.
c-style unions are in nightly behind a flag; they're not stable yet.
struct TaggedUnion {
enum {T1, T2} type_id;
union {type1 t1; type2 t2;} u;
}
If the goal is to translate automatically to a rust enum declaration, that will be (a) possible to do with an automatic tool and (b) will give enhanced safety because it will force the type to be checked everywhere these values are used. x = [] of Int32 | Char
x << 42
x << 'F'
=>
[42, 'F']More data points will help to inform discussion, or at the very least add structure to the flame wars.
(What irritated me though was the switch to first-person narrative at the end).
Not strictly required. You can certainly disable it at the cost of more work (and understanding). Or live with it and reduce its impact as needed.
From a purely technical standpoint, Rust's borrow checker is able to catch data races, while D has no such functionality (unless I'm not caught up), which is a huge advantage in today's world of multithreaded applications.
But as someone else said, Rust has had way better marketing (and is newer, which probably catches some people as well.)
Automotive has unfortunately settled mostly on C and build a whole ecosystem around it (Autosar). With all that talk around self driving cars safer programming languages and tools would make a lot of sense.
Or perhaps it's a good opportunity for a language which offers transpilation with ANSI C as the target?
Last I checked, some of the popular IOT architectures (atmega, etc) were not included without some 3rd party plugins.
Nobody is going to be running NTPsec on it in any case!
> Nobody is going to be running NTPsec on it
Never say never. :) There are wifi, ethernet, and clock shields for Arduino, all of which are running microcontrollers, and many applications which would benefit from using NTP.
It's not clear to me why you'd choose an AVR for a new product unless you really, really needed a specific feature. Current generation ARM Cortex M0 devices are available at a similar cost and with significant performance gains, with the benefit that if you realise you need a more powerful core you can scale up to hundreds of MHz with effectively the same code base.
Microchip / Atmel (same company now btw) are doing an "everything and anything" strategy. PIC, ATTiny, ATMega, UC3, AND ARM chips are available from them.
It seems like the smaller 8-bit microcontrollers use less power and have better features for embedded engineers (ie: ADC converters, PWM, Real-time clocks, deeper sleep modes).
While ARM has the general benefit of being much faster from a computational perspective. But if you want to read the voltage from a simple thermo-resistor and then output it on an I2C bus... I'm thinking a classic 8-bit AVR is going to be superior over any ARM.
--------
I'm still seeing 8051 stuff pop up everywhere, and that thing was supposedly dead years ago.
Another hint is when you actually get in touch with a lot of complete product designs. The amount of ARM cores that you will see in there seems steadily to incline.
Cortex M0+ devices are "larger" than AVRs. You have more RAM, CPU power and unfortunately... complexity and power usage.
I'm thinking a good example is the LPC811 from NXP: very good specs and it really shows how good ARM M0+ systems have become: http://www.nxp.com/documents/data_sheet/LPC81XM.pdf
But its still got an order of magnitude more power usage and complexity over AVR's AtMega328pb. The ATMega's "active" mode is comparable to the LPC811's "sleep" mode.
Most of the LPC811's pins only can sink / source ~4mA, while the "biggest" pins can only sink / source 20mA. In contrast, all of the AtMega328pb pins can sink/source 40mA. Easily driving LEDs with only a resistor (instead of having to hook up a transistor or external buffer of some kind)
The LPC811 also shows what happens to cheap ARM chips: they are missing ADC converters, PWM, Real-time Counters (or a deep-sleep mode that still keeps the 32kHz clock active). Sure, you can buy these externally... but the AVRs and PICs have superior integration.
> It's not clear to me why you'd choose an AVR for a new product unless you really, really needed a specific feature.
Running an LED with just a resistor on any pin is a pretty nifty feature IMO.
------------
The AtMega328pb is a "bigger" and more expensive part though. Perhaps its more "fair" to compare the LPC811 against the AtTiny44A (which is also $1.50ish). ATTiny44A has similar power specifications as the 328pb.
FYI: Microcontrollers really ought to just be interpreting the WWVB radio signal (aka: the 60,000 Hz Atomic Clock radio signal throughout the entire continental USA). Alternatively, Microcontrollers easily connect up to GPS modules for an alternative radio signal / alternative time source.
If anything, Microcontrollers are a great interface to the 60kHz Atomic Clock signal and are therefore would probably be the best NTP-server.
In any case, the "real" best architecture is probably a microcontroller doing Radio Logic / digital signal processing for WWVB, and that is connected through a simple connection (ex: I2C that says the last announced time from the WWVB signal) helping a "bigger" Raspberry Pi. And the Raspberry Pi can handle the ethernet / server stuff
Erm, like, all of them? ARM has a 24% market share in microcontrollers, and 70% in 32-bit microcontrollers.
Source (page 21): https://www.arm.com/zh/files/event/1_2015_ARM_Embedded_Semin...
Here are a ton of Arduino-style boards with ARM chips: https://developer.mbed.org/platforms/
I just counted on nightly and got 20 distinct architectures, including all the le variants; every distinct architecture value, iow.
$ rustc --print target-list | cut -d- -f1 | sort | uniq | wc -l
20Also, it has been suggested by some experiences that older architectures often end up getting supported past when they really should be stopped out of sheer inertia, though I'm not having luck digging up the articles that prompt me to say this. Are there that many systems out there running ntp that can't run Go and/or Rust, and if there are, are there enough to be worth bending the project around? If the people running those things care, perhaps they should fork the project and maintain it themselves. Which is less harsh than it may sound, because they can still pull from upstream, and they probably just need to tread water rather than stay up with the latest & greatest.
(I should emphasize that the operative question is are there enough to be worth bending the project around, rather than whether there are any. Because there certainly are non-zero numbers of systems running ntpd that can't run Rust or Go. But who's going to pay to maintain them? Especially if classic ntp is still available?)
Specific to IOT devices, I would personally love to see secure everyting. NTP, TLS, SSH, etc.
LLVM support would benefit the developers more, but they have their own budgets and constraints, which are probably not conducive to submitting and maintaining a LLVM plugin.
(Specifically addressing IoT devices, IoT security is so abysmal in general that there are much easier ways to compromise one than targeting its NTP daemon, if it's even running one, and if it's even "bothering" to run a secure one like NTPsec.)
Dropping support for things is certainly not a decision you make lightly, but sometimes the greater good demands it.
That should hit most of the chips most people would care about. I presume rust would gain from the upstream LLVM support?
It's not really something to resolve; it's just the usual situation with a big dependency.
But realistically speaking it might be easier to change tools for a significant number of teams, because the software development community is in general poor at leadership and process. A rewrite seems more approachable for the average (not necessarily average skills-wise) dev compared to a change in attitude, self-reflection and incremental improvement at a department or company level.
Then the rust compiler will pummel everyone into submission. This can be a succesful strategy. :)
Anything that is optional gets pushed aside.
Lifetime inference etc is a hard problem, and if it could be done right now, compilers would just have static checks for all that.
There is a C to UNSAFE Rust transpiler, though: https://github.com/jameysharp/corrode.
Realistically, the situation where you have source access with known errors you're somehow barred from fixing seems vanishingly rare.
[1] shameless plug: https://github.com/duneroadrunner/SaferCPlusPlus
[2] https://github.com/duneroadrunner/SaferCPlusPlus-BenchmarksG...
Replacing all the pointers and arrays in your C code with the safer substitutes will eliminate the possibility of invalid memory access, in your code. Of course this doesn't prevent them from occurring in any unsafe libraries you use, including the standard library.
Also, there are certain behaviors in C that may not translate well to SaferCPlusPlus. Like exotic pointer arithmetic. Or, for example, you could imagine some C code that compares two pointers that point to items that have already been deallocated to see if they were previously pointing to the same item. Is that valid in C? Anyway, that kind of thing is not supported by the safe pointers.
There is not yet an automatic translator from C to SaferCPlusPlus, but the translation for most code is straightforward and direct. No new paradigms or "Rust borrow checker" type restrictions. You can check out the benchmark code I linked to in my comment to see examples of C++ code before and after conversion to SaferCPlusPlus. A little code reorganization sometimes helps to achieve optimal performance. But isn't that always the case? :) And of course, SaferCPlusPlus requires a modern C++ compiler and has dependencies on the standard library.
They are usually used in debug builds though, since they have a 2-3x runtime penalty.
[1] http://clang.llvm.org/docs/UndefinedBehaviorSanitizer.html
Just throwing it out there as something to keep an eye on!
Doesn't sound like C++.
Rust and Go, in various different ways (and to various degrees), make it actually safe.
My experiences getting things to compile across gcc and visual c++, dealing with strings (especially Microsoft's WCHAR), reliable integer sizes (pre stdint.h), and debugging templates were not things I would wish on anyone.
Re-doing some of my side projects in Go and Rust was a lot more enjoyable. I could focus on what I was doing instead of trying to work around deficiencies in the language and its libraries.
Oh well, at least they're moving from C, which will be a big win either way.
It's been said over and over since at least the Java times that creating OS processes for individual service invocations is bad for performance, but I've never seen proof for this statement in the form of a benchmark.
Even the OpenBSD developers (who know a thing or two wrt. security of memory allocation schemes) diss process-per-service-invocation architectures in their httpd implementation (eg. calling their CGI bridge "slowcgi" and favouring fcgi over it).
Isn't that inconsequential? I mean if there's a performance problem with CGI-like process-per-service invocations, why not target these problems at the OS level (or via pooling of network connections or whatever the bottleneck is)?
It doesn't have to deal with 40 years of bad legacy code written by sloppy developers.
You can obtain similar quality in a C modern code-base, using tools like static and dynamic analyzers. In fact, today the hardest issues came from multi-threading. I won't even dare to write multi-threading apps without helgrind/TSAN.
And Rust doesn't help in this regard. From: https://doc.rust-lang.org/nomicon/races.html 'So it's perfectly "fine" for a Safe Rust program to get deadlocked or do something incredibly stupid with incorrect synchronization.'
The big question for me is what will happen when those same developers writing insecure C and C++ code start writing rust? Will they find ways to subvert the safety checks? Will they even want to subject themselves to them in the first place?
Which higher-level abstraction to use is a question and depends on the problem at hand; for throughput computing in C, Cilk is perhaps the best today (I guess you could call my own project checkedthreads a bit of a Cilk knock-off.) AFAIK, Rust developers are more interested in higher-level abstractions than Go developers (Go's built-in goroutines being on par with threads in terms of ability of automated debugging tools to flag bugs; perhaps there's work in Go on higher-level abstractions but the vibe I felt, perhaps mistakenly, is that these are non-problems, which I disagree with: http://yosefk.com/blog/parallelism-and-concurrency-need-diff...)
I just left a job in which C code was being written from scratch, in 2016. The code was awful. Coverity didn't prevent it from being awful.
From: https://doc.rust-lang.org/nomicon/races.html
Data races are mostly prevented through rust's ownership system: it's impossible to alias a mutable reference, so it's impossible to perform a data race. Interior mutability makes this more complicated, which is largely why we have the Send and Sync traits (see https://doc.rust-lang.org/nomicon/races.html).
as far as I can tell you can't really create a C
compatible shared library from [Go].
Sure you can! (as of Go 1.5, iirc)Change your build command to use the c-shared build mode:
go build -buildmode=c-shared
Then, to change the linkage of the functions you want to export, use the magical comment "//export" directive: //export MyFunction
func MyFunction(arg1, arg2 int, arg3 string) int64 {
...(And you're bringing that whole runtime with you. Which might be okay! As always, it depends.)
all 37641 (100.00%) in 135 files
c 37311 (99.12%) in 66 files
asm 232 (0.62%) in 1 files
makefile 98 (0.26%) in 2 filesas a starting point look here: http://www.ntp.org/ntpfaq/. you can go from there to the rfc's (5905 - ntpv4) to get a more indepth look.
Of course, C++ might be "safe enough" for your use cases.
What do you mean by that ? Are you saying the standard decided to make it non-destructive moves ? Then yes, it is a possible scenario for error but also a performance advantage in some cases.
You can still access a value after it has been moved out. use-after-move is allowed by the compiler. It places (stdlib) types in an unspecified but valid state (I've seen C++ code reusing these types and assuming that the state after move is something in particular -- it's not).
Most optimizations you can do by reusing a moved value could be done automatically by the compiler (by using the same stack space since it statically knows that it's been moved out).
Linear types are a pretty well known pattern and could have been used as a part of the design of moves. This would have the additional benefit of removing explicit move constructors from 99% of all types out there. This has not been done (and can't be now).
Edit: The pretty rare optimization potential of runtime moves is pretty much negated by the fact that common things like reallocing a vector can't be made into a memcpy for most types, just because they have nontrivial move ctors which wouldn't exist in the compile-time move scenario.
You don't destroy it, you treat it as an actual "move", where the compiler will not consider the original variable as accessible after this point, and not allow accesses after the move. The memory it took up is free for reuse, and no destruction code is run on the side of the code doing the moving (in the case of a conditional move this gets a tiny bit more complex, but not too much). "It's variable will be still accessible in local scope" is exactly what I'm getting at; you can enforce at compile time that this isn't the case by simply disallowing access.
You shouldn't have local references active, just like how you shouldn't have local references to the contents of a vector when you push to it. You can already invalidate local references to a part of a struct when something gets moved in modern C++. Being wary of invalidating local references is an established concept in C++; this doesn't exacerbate that problem.
It's even better if you track scopes in references which gets rid of the dangling reference problem entirely but at this point you've reinvented Rust :p
To be clear, I'm talking of a completely different model that could have been used in place of the rvalue reference and move model. Making C++ have linear types now would be tough, but it could have been done before. There are different tradeoffs there, but I suspect it would have been safer.
> all variables should have a "not-a-value" state and throw exceptions on use.
You can make this a compile time error. Like I said, linear typing is a pretty well established pattern.
Not if constructing the object is costly. It is very well possible for an object not to have a default constructor.
This is why they get clubed together, regardless how much I like C++, I am yet to see the use of C baggage being successfully forbidden in enterprise teams, let alone if there are third party dependencies (which is always the case).
So far I have only seen modern, safe C++ being used successfully on a big project I was part of at CERN, where everyone on the team actually cared to write proper C++.
The enterprise is not the target here.
Rust and Go also need to match the Java/.NET ecosystems if adoption at enterprise level is a target.
And yes, there are AOT compilation toolchains to native code in Java/.NET available.
I was under the impression that these specific things were actually quite hard to do in Go. I believe that both setuid/setgid and seccomp_load change the current OS thread (only), and since Go multiplexes across multiple threads and gives programmers very little control over which ones are used for what goroutines, I'm not sure how you would, for example, apply a seccomp context across all threads in a Go program. setuid/setgid are currently unsupported for this reason, with the best method being "start a subprocess and pass it file descriptors" (https://github.com/golang/go/issues/1435).
I'd be interested to hear if others have found ways to actually do this reliably for all OS threads underlying a running Go process.
Switched to Rust and there was only had one hidden system call left, getrandom used to initialize the hashmap
One of them (lack of leap second support) is fresh on our minds since we just had a leap second. Some thoughts from the OpenNTPD folks are at http://undeadly.org/cgi?action=article&sid=20150628132834
Chrony is another alternate NTP implementation, here's their comparison vs ntpd and openntpd. https://chrony.tuxfamily.org/comparison.html
You're right that the article I linked argues the leap second isn't a big problem to client machines. I guess I sort of believe it, if you don't care if your client is a second off from true time once every couple of years.
(And I have made terrible experiences with JVM's GC, but never tried tuning).
Some people don't believe that extremist functional languages are needed (depending on the domain).
Validation. Oh, well.
I think you propose Haskell, and it might disqualify for an NTP daemon, for example in terms of debuggability (stack traces?).
A Haskell DSL that outputs some safer low-level code would be a more likely choice (for example http://hackage.haskell.org/package/atom), but Rust is both more popular and has more commercial support.
I've heard the idea of DSLs over and over again, but who actually does that? I know of course of sed, awk, regex, etc but what part of NTPsec is narrow enough in scope and large enough in volume to justify creating a DSL? (just asking -- I'm not familiar with NTPsec).
Care to give an example of such a language which allows for a guaranteed GC pause\resume configuration\declaration\etc in a critical section?
https://msdn.microsoft.com/en-us/library/system.gc.trystartn...
But surprisingly, this doesn't solve all code execution security problems from buffer overflows. Sometimes exploiters can find non-executable memory to change that allows them to change the program behavior to do what they want, such as changing the command string that gets passed to a normal execve call later in the program.
But even if you solve the security issues, buffer overflows are still a huge headache. A one byte buffer overflow can cause your program to crash an hour later, with almost no hope of figuring out why it happened. A developer can easily spend weeks tracking down a single buffer overflow crash.
I don't know about rust, but you can have an ArrayOutOfBoundsException in Java, and equivalent in Go and other safe languages
Changing function pointers or the return value on the stack are the normal things to do. In particular if you corrupt the stack you can simply return into system() with whatever arguments you want.
Note that doing that doesn't fix all of the problems that out of bounds array writes can cause though.
>> But NTPsec is a lot smaller and cleaner now at 62KLOC of C (that’s just 27% of the original size). It’s been brought up to pretty tight C99/ANSI standards conformance, and the few remaining platform dependencies are either already well isolated or can easily be made so.
Then they have a section about future plans and a short comparison of two possible languages. I'm surprised that you're surprised Go is virtually non-existent in this thread.
Hackernews comment sections are strictly for endlessly typing about Rust's memory safety features to seemingly no end. You will find this behavior under any article that mentions C or C++.
At this point I don't care about downvotes because the whole finding-valuable-comments-to-learn-from experience I've had in the past is getting harder and harder to attain to. So instead of bitching about that, I will just go upvote someone who is downvoted.
Of course we also can talk about D, Go, Swift, .NET Native, Haskell, OCaml as possible modern languages to use instead of C.
Hint, many uses of C actually only care about AOT compilation to native code, and never have spent even 1 second measuring performance.
Thanks, Eric. Now, FFS, learn how the Cathedral model WORKS, get OpenBSD, learn pledge(2) and stop spreading old FUD against C.
For a small one time fee of 1000 USD I can copy paste you 50-100 lines of C that provides an implementation of an array without buffer oveflows. Cheaper than switching to rust or you know, learning C properly.
I'd argue that for programs such as NTP, avoiding dynamic allocations on heap is not such a terrible goal.
It then follows that reading things safely is therefore a completely insurmountable problem in C. There is no possible way for me thin wrapper around read that works on bounds checked arrays, and not use "read" cavalierly in my code without thinking. Right?
It only got worse.
I think I recognise your name from some rather aggressive rust advocacy in another thread, so I'll try and break this down in a way that won't trigger you:
- Buffer overflows are trivial to avoid in Rust, AFAIK. I acknowledge this
- Buffer overflows are very easy to code in C, and have occurred many times in the wild. Again, I acknowledge this.
- My argument is - buffer overflows are easily preventable in C if you provide your own thin abstractions. The fact that people don't take these steps doesn't mean that these steps cannot be taken. Take the rust compiler itself: AFAIK it's now implemented in C++, an 'unsafe' language, and YET it manages to provide in abstraction in the form of a language that makes buffer overflows all but impossible. Do you understand now how it's possible create a safe abstractions in an unsafe language?
Who cares? What's "possible" is completely besides the point. What matters what is done. And so, yes, you can create safe abstractions in C. But if you need to interoperate with someone else's C code, you have to deal with how your abstractions interact with someone else's, or someone else's lack of them. The point of using a language like Rust or Go or... pretty much any non-C language with lots of traction is that you don't have to provide "your own thin abstractions." You use the same one as everyone else. The point of Rust or Go isn't to make things easier on you. It's to protect you from every other piece of code you interact with. The standard library. Libraries written in the language. Other people working on the same project as you. Yourself five weeks ago. You all use the SAME safe abstraction. Nobody has to roll their own. There aren't 25 safe abstractions for 20 people working on the project. You don't have to reason out or guess whether or not some library (or even the standard library) is doing the right thing to avoid buffer overrun. You don't need to defend your entire codebase against someone else's code not doing the right thing.
Because I am old enough to remember when C was only relevant to UNIX users and those of us that care about quality code were able to enjoy much better options.
It is not possible, because one seldom works alone, so regardless of whatever laboratory attempts to write safe C code, they all fall down when the team reaches a size of two, or dependencies to binary libraries are required.
C++ is also not free from this. Regardless of the language features we are able to use, there are always a couple of guys that code it like C, thus making an hole under the castle walls.
As for Rust, as any compiler writer knows the implementation language has very little to do with the language being compiled. If LLVM was done in Java, as Chris initially thought of doing, Rust would be using a Java compiler instead.
Regardless of this being wrong, it also doesn't prove your point. The compiler and the built code are two separate executables. Your "C plus abstractions" language still allows "C without abstractions" within the same executable. And that's the dangerous part, because now you're talking about audits to ensure that only your "C plus abstractions" language is the one being used.
Without generics these abstractions are hard to work with and hard to make work across libraries. I am reminded of https://thefeedbackloop.xyz/i-bet-that-almost-works/
> Take the rust compiler itself: AFAIK it's now implemented in C++
The compiler is written in Rust. It uses LLVM, which is in C++, for codegen, but in the future we might be able to turn that off too (https://internals.rust-lang.org/t/possible-alternative-compi...)
Besides, the implementation language of a compiler does not make the compiler an "abstraction" over the language. Sort of, but not really.
The only fair thing to say at present is that Rust's main/most useful/most performant/most whatever implementation is a mixture of Rust and C++.
I really enjoyed these posts by strncat:
https://www.reddit.com/r/programming/comments/5krztf/rust_vs...
> It's scary how unfamiliar C programmers are with the rules of the language they're using... it's very difficult to write correct / secure C code without undefined behavior even when you know the rules.
[...]
> It's not feasible to avoid undefined behavior at scale in C or C++ projects. It's simply infeasible. They are not usable as safe tools without using a very constrained dialect of the languages where nearly all real world code would be treated as invalid, with annotations required to prove things to the compiler and communicate information about APIs to it.
and of course, if lots of people fail at something, it will get proportionally harder with a larger team / larger codebase / different stakeholders demanding xyz, etc etc etc.
not chiming in about the basic disagreements in this thread, but the repeated claim that something that has caused trouble for countless professionals is "easy" just can't be right. if all the qualified people are "unqualified" because they all fail at something "everyone can do," the speaker is the one who's confused, and it's a lexical problem.
I mean people manage to do a lot of dangerous things due to sloppiness - cause car accidents, for example. I don't think it's hard to avoid that, and I don't think those people should be driving. I suppose others think we need wider roads and bumper cars.
This isn't an argument against safer languages, or higher level languages, or languages with built in bounds checking or anything else. This is more an argument against programming in a "close to the metal" language like C so thoughtlessly that buffer overflows are actually a serious issue.
Maybe you just need a pedantic mind. When first learning C, as soon as I figured out that writing "char buffer[BUFFER_LEN];" could cause your program to crash, I immediately set out to write a safe dynamic array implementation I could push chars onto it and I didn't have to worry. And I am by no means an expert C programmer.
Can you give me a similar set of primitives to manipulate memory in a temporally safe manner as well? What is the cost of that abstraction? How does it compare to a runtime's GC?
It will cost a single correctly predicted branch, which is effectively free on modern architectures. Any "safe" language will have to make this conditional branch too (Rust's File::read method will check the size of the slice).
As for the cost of the abstraction of bound checked arrays of arbitrary length, I can't imagine it being any slower than rust. I could be wrong. If you'd care to provide an example program in whatever language you are advocating, I can give you an implementation in C using a bounds checked array ADT, and we can compare notes.
- index checks aren't intrinsic for all cases where it is done, e.g. https://github.com/rust-lang/rust/blob/8f62c2920077eb5cb8132...
- being intrinsic is absolutely not required for the checks to be removed. Bounds checks are just a normal branch with a fairly straight-forward condition, and the always-true nature of those conditions is inferred in the same way for both the built-in indexing and non-built-in indexing. This is true in languages other than Rust.
An iterator is designed to just not do any indexing at all (neither using the built-in [] operator or one of the functions that implements manual bounds checks), because it instead just manually (unsafely) walks a pointer along the array.
It obviously still has to bound-check that it's not about to walk right out of the array. If you want to call this something other than a "bound-check", I think that's being overly pedantic.
But yes, strictly speaking you're right, an iterator does do check when it's reached the end of the array, it just doesn't do any indexing nor does it use the indexing checks built into the compiler (which is what you implied the iterators benefit from).
This is false. The index checks are written in Rust code as an impl of the Index trait on [T].
The checks being removed have nothing to do with this -- LLVM can prove that certain checks are unnecessary. C does the same, if you used a library that provided checked indexing.
What's different is that in Rust indexing is overall used much less often, because iterators are the dominant pattern, which completely sidesteps this pattern.
> I could push chars onto it and I didn't have to worry
And how did you implement this? Are you stitching together chunks of memory, or are you reallocating and copying? If you are reallocating and copying, what do you do when you have a point to the old address space that still exists? Now instead of manually updating memory and knowing you have to take care of pointers, your dynamic implementation might change stuff from underneath you without you noticing.
What's your point? My only argument has been that it's very easy and achievable to avoid buffer overruns in C. The fact that you assert most people won't do it is completely orthogonal to that.
And how did you implement this? Are you stitching together chunks of memory, or are you reallocating and copying? If you are reallocating and copying, what do you do when you have a point to the old address space that still exists? Now instead of manually updating memory and knowing you have to take care of pointers, your dynamic implementation might change stuff from underneath you without you noticing.
I use the standard library function realloc. The dynamic array is written in a fairly standard way, IE, I am not returning pointers to the allocated memory of the internal array. I access values inside the dynamic array by value (eg they are copied), not as a raw pointer to a slice of memory that could be realloc'd. It would never even occur to me to do that, so your example seems strange and far fetched.
X is possible!
X is not probable!
Its not a winnable debate. You are both different kinds of people. The prior an idealist, the latter a pragmatist. Both philosophies are good, both are correct. Before both of you continue your debate, you should both recognize this difference.
Realloc may automatically copy the range to a new memory block if the old block cannot be expanded, and when that happens the old range is freed. Any other pointers you had to items in that original array may become invalid every time realloc is called, and if it's automated by your dynamic array code, that could conceivably be any time you push an item onto that array.
This is the waterbed theory of complexity. You can push complexity down in one part, but that just causes it to pop up somewhere else. You can make array size management easier, but the cost is that when it needs to deal with array sizing it's abstracted away to the point where you can't be sure when it happens and when you need to fix problems it might cause, or you can deal with it up front and manually when needed and then when you are manually dealing with the size changes you should remember that the memory might be freed and you may need to deal with that. GC languages deal with this by having all the information to know exactly how to fix all the references needed, and Rust deals with it by requiring you to not have two references to the memory in that circumstance.
Note that it still prints 42 despite the allocated memory being freed. Line 7 copies the data at a specified index into x - x has no references to the malloc'd memory.
Because I'm not limiting the case to just core types. You don't only ever need arrays of chars, ints, doubles, floats, etc. Sometimes you need arrays of structs.
As a simplistic example, perhaps you have a large array of structs, and you want to iterate through them and add all the items that match your criteria to a shorter array of matches, which will be pointers to the real data. Adding a single item to the original array could cause realloc to invalidate every pointer in the array of matches. Of course there are ways around this, but someone starting work on the code might not necessarily expect that adding an item to an array would cause pointers elsewhere in the code to become invalid, unless they look at the implementation of your array code to understand what it's doing.
Here's how it works - if your array is of malloc'd structs, then when you access the element, you get back a malloc'd struct. At no point can you access the actual memory, only copies of the values contained at a given index.
So yes, I suppose if your bounds checked dynamic array was implemented by a complete novice or someone that doesn't know C, your hypothetical scenario could happen.
None of this changes the fact that it's easy for a competent C programmer (do I really need to specify this?) to completely avoid buffer overflows in his own code, which is the only thing I have argued.
The basic idea is that sometimes you might want a pointer to an array item, if that item is complex, not just a copy of it, as there's no need to be wasteful if it's a fairly large struct. Any pointers to that array might be invalidated if realloc is called on it. Knowing exactly when that happens means you can note that it might be invalid, and do something about it, but if it can happen any time you add items to the array, that means you need to check for whether it was reallocated every time, or assume it's always invalidated whenever you push an item on the array.
In the example here, I'm determining the struct with the smallest num field. As I keep allocating space to the array (which I'm doing explicitly here), I'm doing the incorrect thing, which is assuming I can continue to use the pointer, which may no longer be valid and just checking if any of the new items are smaller than the existing smallest. I should be recomputing from scratch. It's obvious when I'm calling realloc, but if I was just pushing new items to the array, it would not be obvious at what point it reallocated to a new location in memory unless I specifically checked.
What's happened in that case is we've traded the complexity of explicitly controlling memory allocation of arrays for the complexity of either not allowing pointers to array items or having to keep track of the array location with a separate pointer and checking that they are still the same prior to using any pointers to array items we've stored.
https://gist.github.com/anonymous/32df4da4949a6c1067c2be8f20...
so TL;DR, if you don't want to deal with copying structs, malloc them then push them to the dynamic array, and deal with their pointers. Else if you push a struct value on the ray, return struct values.
So, manually manage their memory allocation, but allow dynamic allocation of the array of pointers? Sure, there are some cases where that's useful, but if you're already managing memory for the structs themselves, you can probably just manage the memory for the array at the same time.
> Else if you push a struct value on the ray, return struct values.
So, like I said, "not allowing pointers to array items".
You can do this, but you aren't just making array access a little safer, you're also restricting quite a bit of what you can do for efficiency. If I'm going to throw away the ability to use pointers for efficiency, why am I even using C in the first place? I should just write it in some other language from the start. Presumably I used C because there was a need for that efficiency.
Yes.
Sure, there are some cases where that's useful, but if you're already managing memory for the structs themselves, you can probably just manage the memory for the array at the same time.
... then you have memory bugs. As your example code clearly shows. What you're suggesting (exposing the internal backing array of a dynamic array) is completely unorthodox and fraught with potential bugs, and I doubt if rust even does this internally.
All I can suggest is that you look up how dynamic arrays are typically implemented in C. The technique I describe is almost universally followed. This is also what happens with std::vector in C++ - std::vector doesn't manage the memory of the elements themselves, just its internal backing array.
I think this gets at the crux of people's problem with the idea that you can just work around the problem of manually allocating memory. It's a bolt-on to the language, and the behavior is dependent on the implementation chosen, and it makes the behavior fundamentally different than "native" C arrays, to the point that it might cause problems.
> This is also what happens with std::vector in C++ - std::vector doesn't manage the memory of the elements themselves, just its internal backing array.
Yes, but there's also usage directions for std::vector that specifically state and make very clear what iterators/pointers are invalidated on what actions. Encountering someone's home-rolled array routines may or may not allow you to easily make the same deductions. Are the routines for dynamic arrays, or are they for doing system cleanup at the same time, or have they been combined? Are there comments noting the reason for what's being done, and that certain operations may invalidate pointers, or are you left to intuit that yourself?
These are the problems with having a non-core (and not even a popular implementation to fall back on) way to extend the language. C++ is a step up in that it at least standardizes a bunch of core types so you can learn those and carry your knowledge of how they work around to different projects in the language. C's lack of this means that every project may implement something like this - or not - in their own way, with subtle usage differences.
The main benefit you would get from Rust in a situation like this (ignoring that it would likely either be built in or readily available through a crate), is that on encountering some home-rolled system, you can look for where it uses unsafe to find any problematic behavior you need to be aware of, because otherwise you are fairly protected. Worst case, the whole home-rolled chunk of code is riddled with unsafe blocks, and you know it's definitely something you need to hunker down with to figure out what's going on (assuming you need to use it).
Rust's unsafe is effectively an enforced comment around dangerous code. Put that way, I'm not sure many C programmers would really object.
There is some memory overhead per array with this method though.