So in that sense, it's somewhat counterproductive to just say "somebody oughta rewrite this stuff", unless (like RMS) you're willing to dedicate a good chunk of your life to that mission - or think that your post will inspire somebody else to do the same.
http://www.drdobbs.com/architecture-and-design/cs-biggest-mi...
it is easy to add bounds checking arrays to the C language. The trouble is, nobody is interested in doing it.
It'd be a heckuva lot easier than changing languages.
It isn't a magic bullet, but as buffer overflows are (I presume) the most common cause of C security exploits, this would help a lot.
I agree completely about C. I've been saying this for years. There are three big problems that cause crashes in C programs: "How big is it?", "Who owns it?", and "Who locks it?". The result is over three decades of segfaults and buffer overflows.
There have been three or four variants on C which address some of those issues. I've proposed one myself. None got any traction. The only thing that might work is if someone developed a safe variant of C which could be machine-generated from existing C code, and didn't add significant overhead. GCC already has a fat-pointer subscript checking option, but nobody uses it. That approach is usually slow, with a subscript check on every reference. If you do it right, most subscript checks get hoisted out of loops. Go does that for many FOR statements.
Rust is one of the very few languages which addresses all three of those issues without resorting to garbage collection. I really hope the Rust crowd doesn't screw it up.
My experience with adding extensions to C++ is that nobody will use them, not even the people who proposed the extension, unless it is adopted by the Standard. The same goes for C.
The feature I proposed for C has been in D since the beginning, and has a very strong track record of success - both in user acceptance and in eliminating bugs. Whether the runtime bounds checking is actually done or not is controlled by a compiler switch - but most users choose to leave it on.
Things that are definitely used are __attribute__'s and labels-as-values.
Or, more generally, memory safety.
Of the three you described, Rust makes the strongest memory safety guarantees. Neither D nor Rust require a complicated runtime. I'd say this isn't really Go's intended area of usage.
Also, Go is missing things like ASLR and DEP (last time I checked), which means that if you link to any vulnerable non-Go code using cgo (which is almost inevitable when writing core utils), or if you find a good bug in the Go runtime, it's trivial to get code execution.
Unfortunately, it's a dead project, and as far as I recall, never compiled on 64-bit architectures.
1. A language that is low-level and safe but also gives you enough interesting & new to build some buzz/interest, rather than "just" safety. Rust is a candidate here, perhaps.
2. Static analyzers in C advance to the point where a subset of C large enough to be useful can be routinely checked for common types of errors. And it then becomes socially expected that at least core OS stuff will be written in that "checkable" subset of C, treating "unable to prove safety" warnings as errors, or at the very least as suspicious.
3. Mitigate it at the OS level with finer-grained access controls. Utilities like strings(1) or objdump(1) are the easy case here: they do not need to actually have permissions other than "read a file" and "print to output". Even in the worst case, arbitrary code execution in objdump(1) should not be able to delete your home directory, join a botnet, or email your ssh key somewhere, because objdump(1) does not need those permissions. FreeBSD's libcapsicum looks promising, in the sense that it is actually being implemented in the base system, rather than just being yet another ACL proposal going nowhere (Solaris/Illumos also has an actually-shipped privileges system, but I don't know how extensively the base install itself uses it).
2. I find compiler instrumentation (think AddressSanitizer and Mudflap) to be more promising than static analysis. Much of the latter is still stuck in the lint era and give out too much noise. That said, tools like Coverity have come a long way and I know a lot of FOSS projects use them frequently. I personally haven't.
3. Capsicum is quite promising, indeed. I like that it extends the existing file descriptor metaphor and offers sandboxing based on namespaces instead of system calls (unlike seccomp), as opposed to the crufty POSIX 1003.1e capabilities which are underdeveloped and still limited to executable processes, AFAIK. That said, we shouldn't just rely on sandboxing, jailing and capability-based security. We need to fix underlying application bugs, as well (the applications that implement the capabilities and sandboxing themselves, particularly so!)
Which of these are you describing as a buffer overflow?
http://www.cvedetails.com/vulnerability-list/vendor_id-9069/...
(This evening, I'm struggling with a broken "gedit" on Ubuntu. It turns out that editing a sufficiently large file will break "gedit". Not just crash it once, mess its configuration up so badly that future uses of "gedit" hang the entire GUI.
Known bug since 2012. https://bugs.launchpad.net/ubuntu/+source/gedit/+bug/1021720 Status: unassigned.
If the program terminated with "subscript out of range at line 1354 of 'editmain.c'", this would have been fixed by now.)
Ugh, wouldn't that fall in "just because you can do something doesn't mean you actually should" category? :) Suppose I wrote the library in Go and exported some type of C wrapper and then I loaded that in Rust using Rust->C mechanism - that will load the Golang runtime in Rust! And you still got non-trivial C code to deal with anyways!
Of course, it will also produce C headers to give the type information of the exported symbols, but I wouldn't call that code.
Lately I've even been playing with a Ruby gem that's a C extension that calls into a Rust .a that exposes itself via extern C.
But, we already have a huge number of system tools and utilities running in a huge number of languages. Python has become the lingua franca of Linux management tools; Perl was in that role in the past, and still exists in a lot of places; bash and shell are integral. All have separate runtimes, and sometimes interact with system-level C libraries, either directly or through command line interfaces.
Is it really that big of a deal to have a Rust or Go shared library in a system that already has a half dozen different languages and a variety of shared and unshared libraries existing at once? I suspect it would be unnoticeable on a modern system. I'd choose safety over shaving a few kilobytes of memory used.
There would be a cost in having that interface friction for developers...and we'd be paying that cost for years. But, it seems like both Rust and Go have planned for the languages to be integrated with C libraries from the beginning, so it seems less of a problem. If I were going to attempt such a thing, I'd probably start with the front end utilities that use the libraries and then work my way down to converting the libraries (even though the libraries, in this case, are where the problems lie). But, maybe it's possible to make Go or Rust code that provide the same C interface as the existing libraries...I don't know enough to even guess.
Sure, you don't get to call strlen (actually, you do if the compiler treats it as an intrinsic), but that's no big deal.
That is to say, the only overhead/problem with calling using C vs. using C and Rust is the increased binary size of having additional code included. There's no loss of flexibility (as long as you're not trying to run on a tiny microcontroller).
The point is there is no 'runtime' penalty with Rust†. It doesn't have a compulsory garbage collector, it doesn't have a compulsory complicated IO manager, it doesn't have a compulsory multiplexed threading system. In some circumstances lacking any of those will be a downside, of course, but not having it built-in and compulsory means Rust gains flexibility. In several ways, there's actually lower overhead with Rust, because it is a modern language designed around the advances and research that has happened, e.g. dynamic (de)allocation can be faster due to putting some sizedness requirements on the very lowest level APIs (which one rarely calls directly, so the programmer will never notice the restrictions, just the improved speed).
My comment was trying to address the Go <-> C <-> Rust bridge, pointing out that C and C <-> Rust are essentially the same, other than the extra code one must have from writing in two languages rather than one. If one was to forgo the C (which is possible, Rust can easily be used to write a library directly exposing a C ABI), there won't be much difference at all between C and Rust.
In summary, in future, any restrictions required would not be very significant, and, even now before Rust's runtime has been excised, the restrictions will be things like "no network IO", one still gets access to all sorts of fancy iterators and data structures without requiring any runtime.
†Strictly speaking, this isn't quite true right now, but there is a concrete plan currently being executed to make it true.
Thanks for taking the time to explain.
And, being a young language, it's not too hard to be doing something no-one else has ever done, especially with this low-level stuff. :) (Meaning compiler bugs and some "I don't know" answers.)
Once Rust settles down, the next step is academic work to automatically translate existing C into Rust.
Sadly, the list of better system programming languages that the industry decided to ignore is quite big.
Ada? You're trying too hard. It used to be a government standard and wasn't ignored in the least (see: http://www.seas.gwu.edu/~mfeldman/ada-project-summary.html)... much to a lot of people's chagrin. It's widely regarded as an example of a monster language, but it is still absolutely used where the safety is critical.
What overhead? The one spread by C crowed without experience in the said languages?
I used a few commercial compilers for those languages. They were quite comparable in terms of generated code quality to C compilers of the same generation, back in the day.
C was also a research language until AT&T made the code available.
> It's widely regarded as an example of a monster language, but it is still absolutely used where the safety is critical.
Actually, I would dare to say that Ada 2012 is smaller than C++14.
Its use has increased in Europe thanks to what is being discussed here, namely the amount of money lost in security issues thanks to the industry adoption of C due to its relation to UNIX.
I included Ada in the list, because most developers aren't aware that it still exists and is being used. Or that GNAT is just one of many compilers that are still available.
This lack of references means code is forced to do more copies or data-structure lookups.
(I dont actually know Ada or those other languages, but I did do some research a while ago and discussed this publically here and on /r/programming a few times, and have never been corrected, so I guess it is close to correct.)
Yes, this is a problem, certain data structures need to be intrusive in order to be fast. Rust for example is a highly memory-safe language but its "unsafe" mechanism needs to be used in order to implement performant data structures for this reason.
http://www.reddit.com/r/rust/comments/2jec05/problem_with_im...
let map: TreeMap<uint, SomeHugeThing> = make_map();
let value: Option<&SomeHugeThing> = map.find(&123);
`map` is a (recursive) tree that associates a `uint` key with a value of type `SomeHugeThing`, imagine that is, e.g. 1KB, or otherwise expensive to copy. The `value` is a direct pointer to the memory of the value associated with the key 123 (if it exists), that is, it is extremely cheap to manipulate `value` because it is basically a machine pointer directly into the memory of the TreeMap. Rust gives you the power for that to be perfectly safe without a GC: there's no risk that changes to the `map` structure will cause the `value` pointer to be invalidated.As far as I know, Ada etc. do not allow for this without a GC. That is, there's no way to have a safe pointer directly into the memory controlled by some dynamic data structure. Instead, one would have to search the `map` each time the value is wanted, or turn on GC.
> its "unsafe" mechanism needs to be used in order to implement performant data structures for this reason.
Implement some performant data structures. There's a lot of data structures that can be implemented performantly without begin intrusive (sure, one might want to use `unsafe` occasionally to optimise them fully, but the `unsafe` is almost always not being used to make it intrusive).
Also, I don't understand the relevance of that link, the top comment (which is mine, btw) clearly demonstrates that `unsafe` is not necessary.
Oberon and Modula-3 have GC and had real usable OS implemented in them, not just some kind of concept OS.
However those languages already offer the following in terms of security over C:
- String data type
- Open arrays (aka slices nowadays)
- Reference parameters (no need to pass pointers to functions/procedures)
- Bound checked arrays (you can bounds checking off, if you really need it)
- Pointer arithmetic is explicit operation
- Casting between types is explicit
- Enumerations are their own types, there is no implicit conversion to/from ints.
While they still don't cover all use cases in terms of memory safety, they already cover quite a few scenarios that in C just lead to unsafe code without the help of a static analyser.
https://metacpan.org/pod/PerlPowerTools
I'd rather see that than a language such as Rust, D, or Go.