The CVE database. Just because you 'can' write such an array implementation doesn't mean you will, doesn't mean your third party libs will, doesn't mean any of your legacy code uses it, and certainly doesn't mean you will properly test said array implementation correctly.
The number of mitigations added to C compilers and OSes dealing mostly with C and C++ code. ASLR, W^X, /GS, -fstack-protector-all, AddressSanitizer, ... - note the lack of similar tools, or demand for them, for, say, JavaScript - despite it enjoying a similar ubiquity.
I ask this in bad faith: I encourage you to share a single nontrivial codebase which actually creates the abstraction you've described and religiously adheres to using it throughout. As to why this is in bad faith: I'm definining "nontrivial" here to mean using 3rd party APIs - which will operate on C style arrays, not your project specific safe wrappers - and thus by definition won't be "religiously" sticking to said abstractions when using said APIs. By these definitions, the codebase I'm asking for doesn't exist - by definition. Even relaxing the "third party" rule, I haven't actually worked on a nontrivial C or C++ codebase without buffer overflow problems.
Now, e.g. Rust will have the same problems when interacting with C APIs - and nontrivial programs will end up doing so eventually. However, by virtue of the language itself embracing safe-by-default, you're less likely to run into the same problems when consuming Rust APIs.
You can also use third party static analysis tools to ensure you're using a "safe C subset" (such as MIRSA C), but "nobody" does that.
That one "if" is (by definition) not zero-cost.
> C++ implementations obey the zero-overhead principle: What you don’t use, you don’t pay for. And further: What you do use, you couldn’t hand code any better.
Two points:
What you don't use, you don't pay for: if you don't use array indexing, you won't get a bounds check. In addition, you can call an access method without a bounds check as well, so it truly is only if you use the checked version.
What you do use, you couldn't hand-code any better: that bounds check is written the exact same way you'd write it in C.
Therefore, this is a zero-cost abstraction.
The Haskell `newtype` example I gave was meant to illustrate this, as newtype's are respected by the type system and then are treated as the underlying type at runtime.
Most code doesn't use bounds checking, because the branch is a safety net you should never hit, even in theory. Any code that does hit it is already broken. Correct programs using bounds checked indexing will in general be slower than but equivalent to a program where indexing instead results in undefined behaviour.
"You couldn't hand-code any better", well, I won't argue on that point, as it sounds contentious. ;)
_Should_ never hit is very different than will never hit...
My point is that Bjarne Stroustrup wasn't comparing against writing the exact same program the exact same way. He was comparing against what you'd get if you dropped down to ye olde C or Assembly and wrote the same algorithm there, without redundant work or waste.
The comparison shouldn't be the language's GC versus SteveGC, it should be the language's GC versus an ideal, manually implemented allocator. Equally it shouldn't be built-in bounds checked indexing versus manual bounds checked indexing, it should be built-in bounds checked indexing versus an ideal, manually implemented indexing scheme. If you want safety against out-of-bounds, it seems to me the ideal method would be a proof, not runtime overhead.
I don't know of a single language that comes with a GC that does this, do you?
> He was comparing against what you'd get if you dropped down to ye olde C or Assembly and wrote the same algorithm there, without redundant work or waste.
Right. I agree with this.
But basically, we are arguing over an extremely fine semantic, which is "should you even want bounds checks in the first place." If you don't, then don't use a method that has bounds checks. The one that does will have them. They'll both cost the exact same as writing it in C or assembly.
To put it another way, let's say I was a C++ developer on the fence about Rust. If I read this conversation, I'd see that indexing gets called "zero-cost" despite the overhead. Since tons of things in Rust are "zero-cost", like traits, closures, borrowing, etc., all of those things now have doubt cast on them. How can I really trust that these things are actually getting compiled efficiently?
If instead the conversation pointed out that this was one of a few cases where safety took priority over truly being zero cost, but that there were tools in place to mitigate the cost (iterators, unsafe indexing, LLVM), I'd have a much more positive outlook that focussed on what Rust did right.
That said, I can appreciate focusing on other things when talking about the principle; I only brought them up here because we were literally discussing them. I think there's much better examples when actually attemping to convince someone.
Java, .NET, Go, ML and Lisp compilers.
Escape analysis allows to do that, even if just in certain special cases.
Plus the more one uses value types and less heap, the GC needs to work less, specially if we take languages like Modula-3 into this mix.
Escape analysis may not use the GC for those variables, but it is still a pervasive runtime cost in both senses unless it's totally gone.
Why no malloc? It was originally due to the small memory sizes for code and data. The less standard library the better, and dynamic allocation may lead to heap fragmentation and a subsequent crash when malloc fails.
The problem isn't the language it's the developers.
New languages here and there every day. Replace this replace that. When, in the end everyone is simply reinventing the "wheel" over-and-over.
All these languages end up as assembly.
I honestly don't know who I'd put my money on between "AI takes over the world" and "programmers stop writing buffer overflows"
The "heartbleed in rust" example is a great one, and it arises in real life in many high level language APIs for file I/O and sockets. You have an allocation, and you have a count of available bytes coming back from a read() function which may be lower than the allocation size. So you are creating a "virtual" array bound from nothingness. Fail to respect it (without bounds checks) and you will see bugs.
If you reject that this is a valid way to write code, maybe in your API every read() style function will always return the correct size enforced by your JVM or whatever, but you will do too many allocations and over-tax the GC.
If you accept that this makes sense, then you must embrace a more C style way of thinking, where array bounds are created and destroyed at will and must be enforced through your own actions... And suddenly you see the other side of this coin, which reflects valid and true things about the universe, that you may want to chop up a buffer into multiple pieces - and that's OK.
(Now, I wouldn't be surprised if Rust has mechanisms to chop up arrays in the way I describe and enforce the bounds you provide it... Which would be handy. But frankly does not completely destroy the validity of the C approach or substitute for a proper understanding of it. Without that understanding, you will code more heartbleeds.)
[T]::split_at is probably what you're looking for.
Almost all array handling in Rust is done through slice types which are tagged with sizes.
However, you can't bungle the creation of a slice in rust without using explicitly marked unsafe code.
And it's not limited to the VM itself: check out npm "native extensions" like `json`. Not to mention glibc, or the OSes themselves.
By your definition, nothing is safe. And you're right ;)
I note that Firefox is using some Rust code now - so perhaps that will change at some point, for at least one of the common JavaScript implementations, in the not too distant future. I don't imagine we'll see it for the majority within the decade - but who knows, maybe I'll be pleasantly surprised.
I have less hope for the widespread adoption of OS kernels written in safer languages - given the general unwillingness to even use C++ there (although plenty of toy/'research' kernels in safer languages do exist.) Although maybe we'll see one within the next century? Perhaps a microkernel for use in containers?
Of course, that still leaves bugs in the JITs, compilers, hardware, 'legacy' native interop, unsafe{} blocks, ...
Maybe
I think currently there's no plan for it. Maybe after they finish servo to the degree where it supports all modern html features
I work on a C codebase that does this, although in the slightly weaker sense that it does drop the abstractions at a few isolated interaction points with external APIs (think openssl, linux system calls, and not a whole lot else). Yes, there is quite a lot of NIH. With essentially-uniform use of checked data structures, and an extremely comprehensive suite of automated tests getting run under ASAN (originally Valgrind), memory safety errors almost never get so far as being committed to the main branch. This is a complex, >1M SLOC distributed system that has seen several years of production use at this point, and as far as I can recall we have not seen a single memory safety related issue in production (a few have managed to to get as far as certification testing). General resouce-leak class issues have struck a few times, but are also pretty rare.
Proprietary, naturally, so I can't actually show you (sorry), but it absolutely can be done in practice. It isn't even really all that difficult, it just needs to be done from the start, and then you just need a bit of discipline to keep it up.
And more power to you. Note the beginning of the parent comment, however:
> Just because you 'can' write such an array implementation doesn't mean you will
So yes, even if the codebase you work on does have these 'mythical', hard-to-achieve properties, that doesn't mean that most or even many C codebases will.
Good engineering entails observing what problems actually occur and working to fix those. Memory safety issues do commonly occur in C codebases. Regardless of whether the fix in C is simple or even trivial, programmers aren't doing it. So, Rust has some value because it forces the programmer to produce code that is largely free from this type of issue.
Enforcing norms like 'be more disciplined when writing C' or 'stop using external libraries' is much harder than simply using a different language.
>
> […] it just needs to be done from the start, and then you just need a bit of discipline to keep it up.
Here's a neat idea: wouldn't it be cool and save a lot of time if the compiler did this for you automatically, from the start?
Of course, you say it's easy to do it manually, but something tells you your company might have needed to pay less for development if the compiler did it automatically with no human intervention required.
Rust, at least in this regard and probably others too, is no better than C, and for me it isn't enough to justify the horrible and complex syntax.
Rust checks at runtime and panics if your program exceeds the bounds. You can opt-in to asking if the bounds are exceeded and fail gracefully if you like, or if you want to promise the compiler you know for sure your bounds are tight, you can use unsafe blocks and act like C. Opt-in to danger.
C lets you do it with no checks. You have to opt in to the safe path of checking and failing. You don't automatically segfault if you exceed the bounds; instead you read arbitrary memory. Welcome to the land of undefined behavior. You may crash, but more likely, you will read some value from an unexpected place, and carry on executing incorrectly for who knows how long. Opt-in to safety.
That's what people mean by Rust is safe by default. And that's just bounds checks. Carry that notion over to pointers, references, threads, lifetimes, ...
What ever made you think "Safe Rust" meant "compile time checks of runtime values are possible" or "C is just as safe because it lets you index outside an array"?
Optimizers already try to prove index bounds to eliminate unnecessary checks, and static analysis tools to demand necessary checks. Turning the latter into compile time errors is a reasonable approach if your language can provide sufficient information to deal with false positives - likely by forcing you to add your own bounds checking to explicitly handle out-of-bounds cases.
A language John Carmack was using or researching at one point comes to mind, which had this kind of thing going on IIRC. I'm afraid I can't find it off hand, so I might recall in error.
And the point of this thread is that Rust's default (check index bounds and fail at runtime unless the check can be proven unnecessary by the compiler or optimizer) is safer than C's default (don't check anything by default and hit UB if the index is out of bounds).
GCC and clang both have sanitizers either built in or available for them. Sure, it's not default, but let's not act like there is no choice in C but to account for every OOB access while programming or to read memory you don't want to.
Furthermore, I never said that compile time checks of variables are possible, but rather we could move to using dependent typing, or at least a way to judge whether a variable would work as a subscript based on the type of the array and variable.
The Rust designers didn't do that. Instead they put in a feature common in the two most popular C compilers to "panic" at runtime instead of accessing memory. That's nothing. It's rubbish. And if you know to "catch" the panic, why don't you check the value of what you're subscripting with? Saying you can catch the panic is missing the point of unintentional OOB accesses, which is that they're unintentional.
C is just as safe with regard to OOB accessing, and to be honest that's pretty poor in 2017.
If you are using gcc or clang, you have more options. True. But not all C compilers give you those options. However, the point is moot, since I never said you can't catch these things in C; I said it wasn't the default. Which you agree with.
> I never said that compile time checks of variables are possible, but rather we could move to using dependent typing
You didn't say anything about dependent typing. You said "Rust is no better than C". And I'm pointing out that it is. Dependent typing may be even better in some cases; I'm not arguing otherwise.
> Saying you can catch the panic is missing the point of unintentional OOB accesses, which is that they're unintentional.
No one said you should catch the panic. You can use Vec::get() for example if you are using runtime-derived indices and want bounds checking in an ergonomic fashion.
And saying a panic for unintentional OOB is the same as in C is not true, since you get a panic by default in Rust, and to get one in C, not only must you be using a specific compiler or two, you must have the sanitizers enabled for every source file in your program. Not "by default" by any stretch.
> C is just as safe with regard to OOB accessing, and to be honest that's pretty poor in 2017.
It is nowhere close, and saying it is is pretty poor in 2017 as well.
And you are still ignoring pointers, references, lifetimes, threads, ... you know, the other things that also help Rust make "Safe by default" and C "dangerous by default".
https://doc.rust-lang.org/std/primitive.slice.html#method.ge...
C will actually read arbitrary memory, Rust won't, that's the difference.
We're talking about a situation where the length of the array is not statically known and you access it out of bounds at runtime. Rust checks first if it's out of bounds, and if it is, DOES NOT blindly read the memory anyway (as C would) but exits.
I don't understand how you could think the two situations are at all equivalent.
Yes, but when the project started the only existing compiler that met all requirements was C (also C++, although that was not chosen, by reasoning I disagree with). We are in a domain where we derive material benefits from the low-level control C gives us (we have a bunch of highly specialized memory management and I/O), and are not willing to accept GC pauses. There's a common sentiment that we would have used Rust if it had existed when we started, but it didn't so we didn't and so it goes.
- be implemented with macros and token pasting, and result in a ton of mental overhead because you'll have a pile of types like array_foo for an array of `foo`s, and array_bar for an array of `bar`s, along with a pile of corresponding `foo * array_foo_get(array_foo, size_t)` and `bar * array_bar_get(array_bar, size_t)` functions.
- or, have a runtime cost and lose type safety by storing void* and casting when accessing.
The first case is even worse than it sounds: e.g. I don't know how you handle arrays of types with spaces in them (like `unsigned char`, or `struct bar`) with a macro. And, we haven't even thought about const correctness yet, which would probably require having const_array_foo, const_array_bar (etc.) types defined too.
(And, of course, these only solve one facet of the problems with C's pointers: there's no way to defend against use-after-free or dangling pointers.)
alloca-style variable arrays is a whole other can of worms of danger and complexity.
https://gcc.gnu.org/onlinedocs/gcc/Variable-Length.html
Legal in C99, available in C++ on gcc/clang.
Pointing into subsections works fine. You just have to create a type for it. This solution doesn't have the same problems as strings because you don't rely on a terminating entry, and it's what languages like Rust or Java do as well.
You can allocate dynamic arrays on the stack in C just fine with alloca(). The only performance cost is when checking bounds, but since it's a dynamic array, it's the same cost you'd pay in Rust.
> A beginner will often try something along the lines of size = sizeof( myarray ) (which is incorrect).
Additionally, all the functions should be inlined anyway because the function call overhead will likely be as much or more than the actual code, and, more importantly, inlining enables other optimisations (removing the branch, vectorising the memory access, etc.). Once inlined, the code will be the same as the manual/macro-based approach of writing `if` statements around each array[index] access.
They're not an ideal solution by any stretch, but it's not the nightmare scenario you envision wrt generic data structures in C.
typedef unsigned char uchar;
DECLARE_AND_IMPLEMENT_ARRAY_API(uchar)It does make debugging a chore. I ended up with a Makefile rule to run the test suite through the preprocessor (and did some hackery to exclude #include of system headers), format it with clang-format, and build that. Not exactly pretty or easy, but it got the job done.
1. https://github.com/alpha123/yu/blob/master/src/yu_splaytree....
2. https://github.com/alpha123/yu/blob/master/test/test_splaytr...
3. Where this sort of thing is much easier and I am happier and more productive.
Rust arrays/ vectors are safe-by-default. To use the unchecked, unsafe version requires using the 'unsafe' keyword.
let v = vec![0, 1, 2]; unsafe { let x = v.get_unchecked(5); }
This means you can basically grep audit for vulnerabilities, and the above code should be very rare.
Neither does Rust.
> The subscript checking variants of C and C++ have to use "fat pointers" which carry along size information.
So do Rust's Slices.
> The overhead for this is large and nobody uses that.
People use std::vector all the time for this purpose in C++. It has about the performance you'd expect, with very little overhead except where you want it in bounds-checking.
I don't think there's actually a performance difference here. Rust's default is safer because it requires dropping to unsafe code to do something dangerous, but the same optimizations are available in both.
Again, I haven't actually dug into this; maybe someone more knowledgeable about this can point me in the right direction here?
I imagine Rust does something similar, copying bytes if the underlying type has the `Copy` trait and calling some actual code if not, but I'm not familiar with the details.
[1]: http://en.cppreference.com/w/cpp/concept/TriviallyCopyable
> copying bytes if the underlying type has the `Copy` trait and calling some actual code if not,
It does not. Moves and copies are both "memcopy these bytes", the only difference is if you can use the previous copy or not. (This is also, of course, subject to the optimizer, which may elide the copy.)
> either way some constructor of the object must be called if it exists (though it might be inlined and optimized away).
Yeah, this is what I was getting at; this has to happen in C++, but not in Rust. You are right to point out that this only matters for things that aren't trivially copyable.
It would be cool to have those things, but it also means that there's less "magic" stuff going on, which is nice. And it makes the semantics of stuff like this a lot simpler.
> Does `drop()` get called on objects that have been copied from?
Nope. In fact, Copy types can't have a Drop at all, but types that move don't have their Drop impl called when they move.
struct Foo { s1: String, s2: &str, }
where s2 is always intended to point at s1's backing storage. What's unfortunate here is that Rust will disallow this, as it doesn't understand that s2 is pointing to some data on the heap, with a stable address, not the parts of the String struct in s1 that are part of the struct itself. So what this means is, in plain Rust, this type isn't movable, Rust is concerned about the invalidation.
However, you can get around this restriction with some unsafe code to teach Rust about it; this is the premise of the "owning-ref" crate.
This isn't the case. They only need to borrow it. A borrowed value isn't allowed to move (the borrow itself can be copied around and shared within the scope of the borrow), so that works out.
Most Rust iterators only borrow the container or iterator they operate on. It's only explicitly moving iterators like .into_iter() (which extracts elements by-move) that don't.
Rust's answer to copy ctors is Clone, which is always explicitly called. Variable use in rust is a move. Trivially copyable types (Copy) will be copied without compile-time invalidating the old type.
Lack of move and copy ctors in rust greatly simplifies things like this, and makes it very explicit when code is running, but the trade-off is that intrusive data structures are hard to do on the stack in rust.
The self-reference would point to the old object. See this for an example: http://ideone.com/sEFtbN
(Intrusive datastructures on the heap are doable with some tricks.)
They're not very essential so it works out.
I have rust installed. I would be curious to benchmark it and see if is indeed faster than the same thing in C, using a user-defined bounds checked array.
I'm not feeling very well unfortunately, not really in the mood to code.
Hope you feel better soon.
Of course, bounds checking is worth paying a cost! As are many other common foot-guns like remembering to free memory. Which starts to raise the issue of when it would actually be necessary to reach for Rust instead of (arbitrary example) Java. Because a language with enforced bounds checking is not really the same kind of thing as C, and we've already had languages that are safer than C for decades.
(Languages with dependent types can help move bounds checks to compile time, but Rust does not have those)
Rust (which I have thus far only tinkered with) looks pretty appealing to me, though. Do you suggest using C instead of Rust? If so, why?
I'd probably suggest a higher level language if you can get away with it (portability, performance, etc), but I am not really advocating for using one language over another here.. What I am advocating is writing a thin layer of abstraction to avoid buffer overflows in the already extant code base. I suspect it would be a better choice in terms of cost/benefit than re-writing the code.
Who's saying this? It doesn't really make sense, taken out of context as it is here.
> inform me that rust has "Zero-cost abstractions"
IMO at best the array bounds checking thing is an extremely poor example of a "zero-cost abstraction"; at worst it's not an example of it at all. As I understand it, the term refers to things like static dispatch on closures and trait methods, iterator fusion, and generally any place where the compiler can transform high level abstractions into really efficient low-level code where in other languages/implementations you might incur a performance penalty, for instance by dynamic dispatch to heap-allocated closures, or allocation of an intermediate vector at every "link" in an iterator method chain.
If you want a true speed comparison, you'd have to define more than just "read an arbitrary string". For example, encoding requirements. Your C is going to treat it as just a bag of bytes, and you can do that in Rust too, but it's not the most natural way; you'd convert to UTF-8, which isn't free. Stuff like that.
use std::io::{self, Read};
fn main() {
let stdin = io::stdin();
let mut data = vec![];
stdin.lock().read_to_end(&mut data)
.expect("Reading stdin failed");
}Also, "safe by grep audit" means "safe according to a human." The argument of course is that it lowers the surface area of what a human must be trusted to verify. I'm still not convinced by that argument, because human error is a thing. And for actual systems programming, "very rare" may not be true.
Certainly. I didn't intend to imply otherwise.
> Also, "safe by grep audit" means "safe according to a human."
Again, totally correct here.
> The argument of course is that it lowers the surface area of what a human must be trusted to verify. I'm still not convinced by that argument, because human error is a thing. And for actual systems programming, "very rare" may not be true.
Well, given a codebase where both safe and unsafe code exists, the amount of unsafe code is strictly less than the amount of both safe and unsafe code. So it does reduce the amount of code needed to audit, even in a very atypical case where a ton of the code is unsafe.
It's true that a project may use egregious amounts of unsafe. That would be unfortunate. Rust is still safer than C in that case, since it just defines more behavior (like arithmetic overflow), but I certainly wouldn't pretend that the rust code should be trusted.
When writing rust one should certainly strive to write less unsafe code, and to always document the invariants required for unsafe code to be safe.
Rust is not 100% safe 100% of the time, I'm only arguing that safe defaults are critical, and that grep auditing is a powerful tool.
In general when talking about safety in a language it's about the level of explicitness required to trigger unsafety.
I like the distinction made in the nomicon (https://doc.rust-lang.org/stable/nomicon/meet-safe-and-unsaf...) -- Rust comprises of two distinct languages. You have everyday Rust, which is completely memory safe, and "unsafe Rust", which looks similar to everyday Rust but is not safe. `unsafe {}` blocks are your FFI between the two. Looking at unsafe blocks as FFI is IMO a very useful mental model especially for understanding the changes to invariants involved.
AFAIK, cargo doesn't have any feature to point out when a crate contains unsafe code—so you pretty much need to grep the source of every crate you consume for "unsafe".
The issue you're talking about is related, but different.
Auditing FFI is a whole other challenge, however :(
Not in binary libraries, hence why it is important to have a culture to only use unsafe if it really must be used.
You do have C libraries which you access through FFI. This is inevitably unsafe. We should be auditing more there. Though IMO it's still manageable, for most crates.
For example, most OS code is _not_ interrupt hooks or malloc but the rest of the OS. Most of Postgres is not reading data quickly, but higher-level abstractions.
Large-scale systems programming will always be mostly higher-level abstractions, because that's the only way to write large programs. Name any "systems programming" OSS project, choose a random C file, and you'll see that most code does not require pointer arithmetic except because that's how you do things in C.
You don't bit-twiddle for 200k lines of code, so being able to limit dangerous stuff like that to the 5k lines that actually need it makes work an order of magnitude easier.
True. Or slightly rephrased: Most of the unsafe code is centralized to some core pieces.
From the POV of a PG dev: The big problem using something like rust for something like PostgreSQL is its it's portability, stability and uncertainty about where things are going. We do five years of back-branch releases (and for many that's not even enough!). Language and tooling around the new crop of languages simply aren't mature enough for that yet.
So if you are in a so performance critical section that you start to care about the "zero overhead" kind of stuff and the cost of your (predicted) bound checks, you might be impacted by this non-zero overhead... (that is falsely widely believed to be zero!)
And don't get me started about debug builds with all mainstream compilers. The performance is then complete utter shit. This is made worse by the fact such an unsafe language needs debug builds more...
C++ was an interesting experiment in its domain, and has been and still is a success in some aspects, but honestly given core language modifications are needed to significantly extend the behaviour of core types (like vectors, unique_ptr, strings, etc. -- e.g. with the introduction of rvalue references and all the associated machinery and default constructors and so over), I'm starting to believe there are little advantages of this approach over integrating such fundamental types/concepts in the core language (I'm not advocating to do that for C++, I'm thinking about other/future languages). Then you can have more true "zero overhead" stuffs, whatever that means.
If I don't have this kind of performance needs, I have no reason to use C++ to begin with. I mean, why use such a monster if I don't need the crazy performance it may offer when used properly? And even then, C, D, or Rust may be viable alternatives.
The idea of ANSI C++ was to provide safety behind library calls instead of compiler switches.
I think I'd rather not have it panic at all. They were making a language from scratch and still didn't solve the crash-at-runtime problem that C has. Instead they just made a common C compiler option default and called it 'safe'.
It's worse than the fact that the Go creators ignored years of research after the year 1970 and didn't implement generics, for no good reason.
Rust's safety is a joke, and so is the idea that a lot of it is anything 'new'. They had the chance to fix it and they didn't. Boo.
My point is that panics are memory safe, contrary to what my parent said.
Before you say Address Sanitizer you should be aware that its authors intended it as a debugging tool and recommend against using it in production as it introduces new attack surface, as described in the "Address Sanitizer local root" mail.
You can also catch this panic at the thread boundary, so other threads in your program can keep running, unless the program was compiled to call abort() on panic.
(The alternative instead of .unwrap() would be to propagate the error, polluting the whole program with code to handle errors which never will happen. And since they will never happen, many programmers would simply start ignoring them - and in the process, they would by accident end up ignoring errors that can happen. Not a good situation.)
Looks pretty safe to me.
A panic isn't safe, and if you are going to claim that it is, you can get the same safety in C using gcc and clang features.
There's one important difference. As far as I know, the bounds checking of gcc or clang can either print a warning and continue, or terminate the process.
A panic in Rust, however, will safely unwind the stack (similar to a C++ exception, in fact it's the same mechanism) and terminate only the thread. The rest of the program can continue running, and even start a new thread to replace the terminated one.
You can see this in action when running "cargo test". A panic in a test (the assert!() and assert_eq!() macros, often used in tests, do a panic in case of failure) will not terminate the whole process; the rest of the tests still run.
People do not use the sanitisers nearly as much as they should, and nor are they designed to be used in production. There's has been CVEs issued for them caused by the testing mindset in which they are written.
In any case, how much a programming language helps the programmer get a correct program (I.e. the most general form of safety) is a thing of degree, not all or nothing. Rust does a lot more than C, pushing many many errors to compile time, meaning they're fixed before the code even runs. Sanitisers only catch these when they actually happen, and will presumably result in the runtime crash you're so concerned about, if used in production. The fact that Rust isn't dependently typed and so can't catch OOB at compile time is unfortunate in this respect, but don't throw the baby out with the bath water.
False. Buffer overflows in C can overwrite the program's memory, so it can be hijacked and supplanted with the attacker's code. This cannot happen in Rust (unless unsafe code has the vulnerability), or any memory safe language.
Sure you can implement a safe array/buffer abstraction and use it in your C programs that abort on invalid indexing. Now how many actually do this? Very few given the prevalence of C programs on vulnerability disclosure lists.
Idris.
Basically, in Idris I can have a function concat which takes a Vector<N, T>, a Vector<M,T>, and produces a Vector<N+M, T>. You can have arbitrary expressions there, and even things like a function which returns a different type based on its boolean argument.
So most signatures will just carry dependencies through, but some things are dependent on runtime input and that's when the "N" part of the type will be tracked at runtime. At least, that's how I think it works.
Dependent types help eliminate large swaths of bounds checks. Not all of them, and I'm not sure how much better it does than a good optimizing compiler.
It can be done by static analysis by compilers in other languages also. Knowing size at compile time is easy most of the time and we are not talking at all about this case. Your answer that Idris solves that is wrong in the context that this discussion started (dynamically allocated memory/not knowing the size compile time). You still need to trade time/space comparing to solution without bound checking, doesn't matter which language you use.
I am getting info that I am submitting too fast? so answer to your other comment will need to wait.
The nice thing about dependent types is that in all the cases where you don't need a bounds check -- where a program invariant enforces that you are within bounds -- your code will not contain a bounds check regardless of optimization. Only in the cases where there are runtime dependencies that can't be resolved will there be issues, and in thee cases you need a bounds check anyway.
It's not perfect/complete, and like I said optimizers can get most of the way there anyway. I just wanted to note that there are better solutions for bounds checking.
I think the history of exploits proves that it's always acceptable for server code. The only possibly justifiable place to omit bounds checking is isolated, high performance numerical code, as used in scientific simulations. But this code isn't exposed to the public, so its vulnerabilities aren't important.
You need unsafe to do some things in Rust, sure. But usually this code is isolated and auditable, and nowhere near the linecount of the rest of the application. Even the operating systems written in Rust have pretty conservative use of unsafe (redox, Phil's OS, etc). Most of your code won't be working with raw memory operations, most of it will work with zero-cost abstractions on top of raw memory.
And that's what I wrote.
> Most of your code won't be working with raw memory operations, most of it will work with zero-cost abstractions on top of raw memory.
For some web applications sure but industry is not constrained only to writing web applications especially if you want to get into C market space. In HPC or time/space constraint environments this matter.
You might want to define further what you mean by "raw memory here". Rust lets you work with arrays and vectors and the heap safely just fine. These are designed as zero-cost abstractions over raw memory (slices, Vec<T>, Box<T>). The equivalent C (with relevant bounds checks) wouldn't be any faster. Rust does not let you do things like call out directly to malloc/free safely. But that's okay. The existence of these abstractions means that you rarely have to do this.
> And that's what I wrote.
.... huh? that's exactly why I put "sure" there, that means I agree with that statement, what followed was why I disagree with the conclusions.
Trading efficiency for safety for unaudited code should be difficult. C does not make this difficult. Rust makes this difficult.
Auditing should be supported for this kind of code. C does not support auditing in any meaningful way. Auditing much easier in Rust since it explicitly identifies code that may be unsafe.
In conclusion, I disagree emphatically that C and Rust are somehow equivalent even when you are dropping down to unsafe code.
Check out https://en.wikipedia.org/wiki/Return-oriented_programming for more information.
Why allow programmers to make mistakes? That was fine in the 70's when resources for compiler execution were limited. I don't see any reason for it today.
I mean, just look at the underhanded C contestants and especially winners for ways in which your program can completely blow up for extremely subtle reasons.
Isn't this also true of Rust, with its unsafe keyword? None of these languages are completely safe against programmer mistakes.
It's used sparingly. Not as sparingly as I'd like, but sparingly enough.
I think that makes a very significant difference.
Yes, carelessness in either could amount to trouble. No, they are not close to the same amount of risk because the relative amount of code in each is orders of magnitude different, and unsafe blocks can be heavily reviewed.
> None of these languages are completely safe
No, but the first time I encountered an unexplained segfault in Rust was the first time I've ever found & fixed the cause of a weird segfault message in just a few minutes, because the mistake was in the 12 lines I had written inside a clearly marked unsafe block rather than the other ~3000.
Without something like unsafe blocks, finding known bugs and auditing code for other "weird" memory issues means looking much wider (and often, screaming "what the hell are you talking about?!" at the screen on a routine basis).
For a philosophical counterpoint: Why allow anyone to do anything that might possibly be incorrect, harmful, or otherwise perceived by some to be negative?
I've looked at a lot of the talk surrounding "safe/secure languages", "safe/secure programming", etc., and yet every time I've heard people preach about the benefits, I feel like I just vehemently disagree. At a very deep and fundamental level, I feel like somehow we are sacrificing something more important in the pursuit of this "safety", this seemingly overpowering desire to make everything completely safe, mindlessly constricted, and stifling. It's not just software; the whole "war on terrorism" irks me in the same way. I imagine a "completely safe" world, the ones these "safe software" proponents appear to be striving for, would be rather dystopian.
A quote that immediately comes to mind is: "Freedom is not worth having if it does not include the freedom to make mistakes."
My serious point is that in practice, the example of C shows that if it is available and people understand that it is "performant" then you will see it all the time, including in libraries you are forced to use.
This isn't an accurate analogy at all because I can compile thousands of lines of Rust code without even thinking about needing to use `unsafe`. How many lines of C code can you compile without `cc`?
> My serious point is that in practice, the example of C shows that if it is available and people understand that it is "performant" then you will see it all the time, including in libraries you are forced to use.
Except this is already demonstrably false in the Rust ecosystem.
And there is no shortage of mistakes available for you to make, so don't worry about that.
Are engineering standards for public buildings evil because they stifle architects' freedom to design whatever crazy structures tickle their fancy? Should power tools not include safety features like blade guards because their users' freedom to accidentally kill or maim themselves must be held sacrosanct? Probably to both questions, the reasonable answer is "no".
If you're building something for yourself, and offer it to others only with clear warnings, then go nuts and make all the mistakes you want. Nobody's saying you can't do that; that sounds rather dystopian, a society where pointer arithmetic is illegal! But you can't treat a project people are meant to trust and rely on as your personal art project.
The physical analogy is good because even there one can see that there are different standards --- and, unlike what the "safe software" community seems to promote, engineers are not doing the equivalent of making every building strong enough to withstand a nuclear war and calling anything less "unsafe".
This also brings up another difference with software: the "absoluteness". In the physical world, no security is perfect. Locksmiths exist, and with enough determination, essentially anything can be broken into. But in the software world, with good encryption, that can never occur. Provably correct software can be employed for effectively unbreakable DRM and un-rootable/jailbreakable user-hostile devices, of which there is (fortunately) no similar real-world analogy I can think of.
Nobody's saying you can't do that; that sounds rather dystopian, a society where pointer arithmetic is illegal!
Given that there are attempts at even prohibiting Turing-completeness[1][2], I would not be surprised if that eventually happens. As it is, I'm sure there are already people who would consider you suspicious if you write software in an "unsafe" language, and from there it is not far to complete prohibition.
"The road to hell is paved with good intentions."
Yes. Except the equivalent in software is using something like Coq, where your intention is proven to be correct. (You even mention provably correct software in the next paragraph...) And yes, it would be awesome if we could prove every program, because all we would be doing is proving that the programmer's intent is correctly coded into the software, which doesn't limit freedom of expression at all. (I would argue that such is actually an even stronger form of expression, as it's afterwards impossible to misconstrue your intent.)
> Given that there are attempts at even prohibiting Turing-completeness[1][2], I would not be surprised if that eventually happen...
Rust is not one of these, and no one but you seems to have this in mind. So it's a bit of a strawman to be throwing it out here, no? But yes, obviously such a language would limit your freedom of expression.
You can extend that absolutist argument to say that inventing anything that could be used as a tool of oppression is unethical. Like, inventing plumbing may have done great things for human society, but it was ultimately unethical because when the police came to take away your general-purpose computer they used a pipe to hit your kneecaps until you told them where it was. Or, how about this: developing any society beyond the level of the most primitive hunter-gatherer tribe is unethical, because what are governments if not the agents of oppression themselves?
This is the sort of abstract position that can't really be argued with in a vacuum, so I'm not going to try. But it's also not a useful ethic for building a modern society free of large-scale oppression, because it completely ignores the practical realities of doing so.
In the mean time, buggy, exploitable software is out there in the real world hurting real people every day.
There is no philosophical beauty in typos or off by one errors. You are not being stifled when your compiler points out to you that, no, "coutn" is not a variable currently in scope. If we have a static proof that what you just did will never work, why would you still want to do it?
And if you had a reason, you still have a way to do it.
If you're trying to be clever, you need to prove that you know what you're doing.
We're talking about eliminating a class of programmer error through static analysis. (If you're writing a large C program, you arguably should be using static and runtime analysis tools on it anyway.) This is no different than using a strong type system. Yes, it limits your "freedom", but sending a "banana" to a `sine` function is nonsense anyway. Well, it also ends up that reading 10 bytes into a five byte array is also nonsense, and we now have tools to detect that nonsense and tell you that it is, in fact, nonsense.
However the point was of course to make a safe-land layer which contains the needed functionality. The point was never to have a function "safeioctl" in the first place.
And, honestly, one of the features is "vibrant community of developers." Even if Go and Rust were bad languages, which they're not, they'd still be better choices than C-with-custom-in-house-restrictions-and-libc-wrappers, simply because of the communities around them. If you're writing in Rust, someone else has already written the safe C wrappers, and if they haven't, there's a community of people who will code-review your wrappers for safety and merge them into a centrally-maintained project, which is extremely useful.
To answer your question though, the advantage of using C is certainly not the notorious bad-habits standard library. C is fun and productive exactly where there's just you and some bits and bytes to bang around. Coding in the small. Not platforms and architectures.
Just as you can statically link libc, without it showing up in the symbol table.
tangentially, ats has had some success in that area since it compiles via c. but I agree that c is less work to get set up with.
All the way through the 80's up to the early 90's, C toolchains only mattered to those that were working on UNIX workstations and servers. On home computers it was just, yet another language, with compilers generating slower code than junior Assembly programmers.
So of course Rust has to catch up a few decades of market use.
1. Buffer overflows aren't considered the most insidious issue in C nowadays. That award would probably go to use after free, which is not so easy to fix.
2. In C, it is easier and faster to do the wrong thing. Compare "char buf[256]; strcpy(buf, foo); ..." to "array_t buf = array_create(strlen(foo) + 1); strcpy(buf.ptr, foo); ... array_destroy(buf);"
3. Buffer overflows do not in fact come up routinely in Rust the way they do in C.
That it's not a by-default and forced language-feature and that most developer aren't going to spend those 10 minutes when they need an array.
They'll just use the language-provided array-implementation instead. Which in C is very, very unsafe.
Also it is actually a joke, since it still separates the pointers and length in two separate variables, instead of using some kind of struct.
The only thing it does is have functions with better semantics on the terminating nulls.
What you are talking about is value range analysis. Rust does a limited amount of it, and if it can prove safe indexing it will remove the runtime checks.
Moving to an abstraction would potentially solve some overflow issues and requires constant attention. Moving to a safe language likely solves all of them and you get checked by the compiler instead.
Rust and Go are not your only choices here. C++ with aggressive use of modern features and aggressive non-use of everything inherited from C is also a fine choice, but enforcing that discipline is hard, and the compiler won't help you a lot. Honestly, a dynamic language like Python or Ruby is also a fine choice in many cases, although NTP might be too latency-sensitive for that to work. (But it might not be! Premature optimization and all that.)
Scenarios where I frequently end up fixing other people's memory errors:
1. No Error Handling: not checking an error condition on a function that allocates, then using the uninitialized pointer anyway
2. Sloppy Error Handling: jumping to abort from an error without freeing what has already been allocated
3. Faith in \0: still using the old string functions
I'm on the fence about the whole thing, so others may be able to field something more compelling.
2) I can't remember if a panic in Rust calls destructors, which would clean up that memory. Can someone answer that please?
3) is only relevant in FFI scenarios in Rust, and everywhere else is irrelevant because Rust does not require the use of such footguns for basic string manipulation.
More exotic panics (like an infinite loop panic on a microcontroller) would not call destructors though.
Most panic impls will either unwind (which will call destructors) or abort/exit, so this is usually not a problem.
Y'all please, please note that dreta said "... an array implementation that prevents that from ever happening..."
When the figure of "ten minutes" was used - that's really what it should be as a mean or median figure, with some long-tail outliers for knotty cases.
While I am (somewhat) sympathetic, I don't miss what it is - it's make-work. There's not some body of material on language, O/S and software design available today that wasn't available 40 years ago. People new better back then, and didn't do better.
But the endless, circular, Utopian arguments just don't cut it. It ain't the crate, it's the pilot. I am sorry that oh so many are put upon to actually - gasp - test their code but that's what this is.
The economics of a Rust or Go or whatever make perfect sense - in the long run. Just recall what Keynes said about the long run.
Just to be clear - I don't suspect that the things that pretend to replace C are bad, or no good. I've just seen multiple "pretenders to the throne"[1] and what would appear to have happened is that we just moved the pathology around.
[1] please excuse the horrible metaphor.
After a few iterations of that, one begins to think that perhaps this is a distinctly human problem that is not necessarily addressable by improved language systems.
My greatest concern is that I keep seeing people learn the same things, over and over, on the job. There's no real repository of literature to actually address any of this - each engineer appears to have to learn it mostly from scratch.
There's a famous bon mot from pubic choice economics - "something must be done; this is something; this must be done." I just hope all the new crop of languages are not that.
Finally: If you can, and if you can learn the right patterns ( isn't that true of all languages, though?), the thing I have gone to, again and again that is better than test equipment in terms of reliability continues to be Tcl.
It is called university and engineering degree in software development.
So long as one can convert it in a straight-forward way that preserves the current meaning of the program's statements. If not, the port might introduce new problems.
Edit: Btw, send me an email (see profile). Got something you might be interested in. One or two things.
That (email thing) would be awesome, Nick. Thanks.