Why Systems Programmers Still Use C (2006)
bitc-lang.org
bitc-lang.org
sort :: Ord a => [a] -> [a]
which just says that it takes a list of comparable things and returns another list of the same type. Here's the type signature for sort in C++:
template <class RandomAccessIterator, class Compare> void sort ( RandomAccessIterator first, RandomAccessIterator last, Compare comp );
The difference in terseness and clarity is a big reason why I use Haskell when I have a choice.
template<class I, class C> void sort(I a, I b, C cmp);
Now, the concepts aren't 1:1; Haskell for obvious reasons doesn't represent the idea of operating on storage directly, so you can't have an iterator and need to return a "new" list. C++ makes you write out the types you are parametrizing instead of getting it implicitly. And there are no doubt some really good arguments why Haskell is terser and clearer than C++.
But this isn't one of them. Come on.
Likewise, if you're truly confused about C++ STL iterators you're just waving your own ignorance around. They're a simple concept pervasively applied in the library. No experienced programmer is going to be confused by that function declaration.
Look, very good cases can be made for functional languages. But this is just surface-level stuff that frankly isn't going to help anyone. Expressing a sort simply isn't a complicated thing in C++ or Haskell and trying to claim otherwise is just dumb.
HBase, Hadoop, Cassandra, GWT tools, MQ, App Servers (Jetty, Tomcat, GlassFish, etc), EhCache.
Unless if people categorized the above software as non-system-programming.
Database engines highlight some of the reasons that Java is a poor systems language relative to C/C++. High-performance database engines these days are usually bottlenecked by memory I/O performance and efficiency, or storage I/O if you under-provision the machine resources for the workload. This is why you see I/O optimizations applied to in-memory databases, for example. Java may be fast at many things, but it is much slower than C/C++ for codes that are bound by memory performance and efficiency. You have to start doing very awkward things in Java to even come close to what comes naturally in languages that explicitly manage memory behavior and structure. As someone who has written a lot of database engine-y code in both C++ and Java, the difference in absolute performance for nominally equivalent code is not small, and is generally easier to achieve in C++. So for low-level performance-oriented systems, there is a pretty strong bias toward C++, particularly now that there is a lot of experience trying to implement the same systems in Java.
For performance sensitive codes, systems programming is really about carefully managing resources to optimize for characteristics of the system. C makes this very easy because it exposes all of it. To the extent that CPU architectures are increasingly bound by memory performance, I would not be surprised to see C++ supplanting Java for certain purposes.
This makes sense. Java was designed more as an application language that works well for codes that are unlikely to make heavy use of the memory system. Its popularity and generality has caused it to be widely used but that does not always make it a good choice outside of its original design case (see also: Perl).
So writing a JVM is systems programming. Writing a Java-based data store is not, even though other stuff then sits on top of the store.
Following your example there are also JVMs coded in Java.
What you do to effect a syscall is to call a JNI function to do the work. JNI is a C (!) API, defined in terms of the C (!) ABI for the platform.
And sure: you can generate native machine code in Java, just as you can in python or bash or even BASIC. But you can't call it.
One such example is the Jikes VM:
http://jikesrvm.sourceforge.net/apidocs/latest/org/vmmagic/p...
Or the Sun's research in writing drivers in Java http://labs.oracle.com/techrep/2006/abstract-156.html
Have a look at:
Modula-3 + Spin
Oberon + Oberon System
Spec# + Singularity
Haskell + Home
OCaml + Mirage
Probably the case, networked services represent an area where performance is at least somewhat a concern, but BitC is trying to address situations where one is writing a memory allocator, not building a server on top of a GC'd language known for having sloppy memory usage.
C may not be the easiest language to learn, but you wouldn't want newbies messing with systems programming anyway.
Higher level languages give you a more abstract view, but when you are doing systems programming that's not what you want, you need to be in full control. Only C gives you precise control of what the machine is doing all the time. You don't want a garbage collector to kick in unexpectedly, you don't want data structures to be allocated in mysterious ways.
The fact is, C is broken, and we all suffer every day because of it. From a security perspective, C is a nightmare. It is not just an unsafe language (in the PL sense of having undefined behavior), but a rampantly unsafe language. Null pointers (which C.A.R. Hoare famously called his "billion-dollar mistake"), buffer overruns, unrestricted pointer arithmetic, arbitrary casting, manual memory management; all of these are the cause of innumerable bugs in real code, often critical security vulnerabilities. C is also not the nicest of languages to code in. It is terribly verbose. There is no good way of writing generic code in it. The preprocessor is a gigantic ugly hack. Writing portable C is a pain in the ass, and building it portably is doubly so.
The worst part is, we know reasonable ways to solve most of these problems. In many cases we have known them for decades. Unfortunately, how to best integrate these solutions into a language and still keep it useful for systems programming was (and is) an open problem. I can't speak for Shapiro, but as I see it this is what BitC was aimed at doing.
All of that (and everything else C is attributed with) can be accomplished without using an arcane preprocessor/include system and you can have niceties such as a saner type system, generics/macros, namespaces, etc.
C is a language stuck with the design decisions that reflected the programming environments in the 60's and 70's but make absolutely no sense in modern context and now we are just stuck with it because of inertia.
It wasn't realistic to have the whole compile process done in memory in the early 70's, and OS virtual memory wasn't realistic either; but we've long surpassed those concerns, and that means that all the thinking about the program namespace, modules, linking, pre-processing, etc. is worth re-evaluating.
As well, we can do better static checking now, without including any notion of GC; in an imperative execution model, leakage of memory, handles, processes etc. remains orthogonal to type safety. C has heavily bottlenecked program control logic from the beginning - a callstack, looping constructs, etc. - and allowing some equivalents to exist in its data structures would make user code tremendously more reliable.
These changes, often paired with some more attention to concurrency, show up in all the newer system language designs - D, Go, Rust, Clay, BitC, etc.
I teach C, and I would love a "cleaner C" mode, which just got rid of lots of the bizareeness of C, several of which you mention. The fact that many compilers will warn with '-Wall' about code which is clearly incorrect, but due to the rules of C they cannot simply reject, is irritating.
Spin, Oberon, Singularity, Home are just a few of them.
The main problem is that for a systems programing language to be used a such, there much exist a successful operating system that uses it as its main language.
In Windows 8, the main systems language is C++/CX (C++ with reference counting extensions), so the time will come.
The rules of C are actually stricter than many reople realize (add -std=c99 -pedantic-errors to gcc and you can get an idea about that -- this, however, still won't catch semantic errors like aliasing violations).
Personally, I'm using clang with -Weverything and remove warnings as necessary. On gcc, there are a lot of useful warnings which are not included in -Wall. My current warning levels look like this:
-std=c99 -pedantic -Werror -Wall -Wextra \
-Wmissing-prototypes -Wmissing-declarations -Wshadow -Wpointer-arith \
-Wcast-align -Wwrite-strings -Wredundant-decls -Wcast-qual \
-Wnested-externs -Winline -Wno-long-long -Wconversion -Wstrict-prototypesMaybe in 1979, it did. Modern C compilers alter and modify code quite radically, to the point where it's quite difficult these days to go back from optimized assembly back to the original C code. These days, C gives you a good illusion of control, but all too often, C programmers mistake illusion for reality.
What you're describing is the difficulty in understanding what the guarantees of control are that you get form your compiler (and yes, that's a much more involved issue than simply reading a language spec). Yeah, C is hard. But the fact that you don't understand how the optimizer works (or how to read the generated assembly) isn't the work of an "illusion", it's just your own inexperience.
His answer was of course EROS.
Microsoft explored Singularity.
Some of his colleagues went and wrote Go as a systems language (if not a kernel-side language).
- low level memory management
- faster execution (in most cases)
- cache line optimization
- avoiding language implementation magic
The author definetely has no understanding of PL issues or has a huge bias toward C++ insane syntax.
https://github.com/bitc-repos/bitc
Projects like this make me wonder if they would have gotten further had they put it on a more social oss hub, be it Sourceforge in its day, or Github now.
The plain truth of the matter is that none of these things are due to a lack of special language support. It's that the whole system is too complex, and the complexity isn't even quarantined in such a way as to be harmless. We need to be able to start over (something the author acknowledges). We can't do this if we build yet another complex system in the belief that we'll get it all right this time round.
But if you're looking for a replacement for the entire software stack we use today, one that tries to shun complexity (20k LOC for everything from kernel to common GUI apps), here you go: http://www.vpri.org/pdf/tr2011004_steps11.pdf
More info (see previous STEPS reports): http://vpri.org/html/writings.php
>Can you imagine not having Unix-like systems, and not being able to use C and all the languages built around that ecosystem?
Yes! Charles Moore has been essentially living this since the 1970's. No many have his level of courage though.
Can you give some more detail? What does he use?
I'm an unbounded admirer of Charles Moore. But courageous? He's a software solipsist. His approach is certainly bold but I doubt he sees himself as courageous.
(Although having been a voyeur of his work for over a decade now, I'm sure he would be quick to point out that much of EROS was based on formalising ideas from GNOSIS and KeyKOS, and the work of Hardy, Franz, Landau et al. over 20-30 years earlier.)
BitC evolved out of a need to prove that the implementation of the confinement mechanism (prototyped in EROS) matched the model (proven in his PhD thesis), and so the CapROS system (built in BitC) was born (well that plus some architectural changes based on lessons learned from EROS). From that perspective it's much more than just another systems programming language - it has a definitive purpose to advance the state of the art in practical and theoretical computer science. (FWIW, they never did accomplish this goal [http://www.bitc-lang.org/docs/bitc/bitc-origins.html]).
Is there still any forward motion here? It seems like a lot of the sites haven't been updated for a good number of years. A notable exception was BitC, which was news from 2010 - still not that recent.
I suspect there are two forces at play:
1. Building a pure capability-based operating system has minimal payoff. While personally I believe the result will be an amazingly reliable, high performance, secure system, the fact is the operating systems we have today are apparently "good enough" that nobody is interested in funding further work in this area (AFAIK Shapiro and others did form a venture in this regard; what came of it I don't know). Keep in mind that much existing software will have to be re-engineered, and a good part of the OS utilities redesigned since if you're going the pure capability route the significance of files becomes pretty much purely a user thing.
2. At a higher level, in my opinion (as an amateur capability-based systems theorist) the benefits of distributed capability-based systems have already been realised as the shape of the evolving web, albeit on much cruder foundations than those designed as part of the literature. Cookies + URLs are pretty much capabilities, and web services (including web sites!) are effectively distributed objects. Javascript + HTML have fulfilled the dream of being able to ship and run data, code and user interfaces remotely, which is the foundation for an unplanned human + computer usable distributed ecosystem.
Javascript, while not the cleanest language, is a solid language for a capability-based system (in terms of capability rules: everything is an object, objects can only be accessed via references [capabilities], objects references can only be acquired by (a) creating the object or (b) by receiving a capability), i.e. capabilities cannot be forged. If you're interested in what an even more pure approach, designed by people who really know what they're talking about, would look like then take a look at http://www.erights.org and http://www.waterken.com .
In this sense, the first true modern capability-based operating system will probably be the first true web operating system.
I think it's pretty cool, and a confirmation of the ideas in Gnosis, KeyKOS, EROS, Coyotos, CapROS, E, etc. that the natural evolution of the largest distributed system on the planet (the Internet) effectively took the form of a distributed capability-based system.
Building a pure capability-based operating system has minimal payoff.
Given the massive security problems we see today, if capabilities are the right solution (are they?), it seems like the payoff could be massive. It seems like "mass adoption" would be hard to achieve (except in the really long term), but it seems like it would be possible to find early adopters who "really really" need good security (e.g. certain military applications, maybe?).
At a higher level, in my opinion (as an amateur capability-based systems theorist) the benefits of distributed capability-based systems have already been realised as the shape of the evolving web, albeit on much cruder foundations than those designed as part of the literature.
This is extremely interesting. Reminds me of the idea behind the "separation kernel" which is supposed to mimic in software the security that can be achieved by connecting systems only over extremely well-defined channels (e.g., the systems are physically disconnected except for an ethernet port that is very well controlled). I can find a refence on security kernels if you're interested and not aware of them already.
Anyway, it does seem to me that there are lots of applications (e.g., embedded applications) where distributing everything over the Internet isn't really going to work (not that you're suggesting that). I'd really like to see people tackle security for this kind of system in an entirely new way. Maybe capabilities is part of that.
Now it may be that you have some great insight but your post doesn't seem to give me any epiphanies.
You probably don't want that either because the cost of moving data between cores will be so high from both a performance and a power consumption POV. The days of free coherency between cores will eventually come to an end; you can already see evidence of that in multi-socket Intel and AMD machines (if you try to ignore the NUMA nature of memory, you swamp the link between sockets and perf tanks in a variety of apps).
All this throughput nonsense is benchmrk lies. Websites are not as a general rule highly responsive. They have huge variability in response times. Who cares if it costs a bit more to run the system if it means having a nice predictable system built for simplcity and safety from the ground up?
Any system you design will have to tackle a lot of the hard problems that kernels deal with. Now maybe you have a simpler system in mind, but it has to be drastically better for anyone to consider giving up binary compatibility.
You can't call something that's been around for less than 100 years "tried and true". I'm not saying we can't have "virtual CPUs" for tasks that can handle a performance hit. But it is becoming an unavoidable impediment. "Simple" doesn't mean "easy". It means predictable, small and undoable. The exact opposite to where we are heading.
The point here is that there are good reasons to want to use different systems together and then you end up putting in lot of effort working around the limitations of the specialized environments.
Also, you can ask a lot of people to give up binary compatibility and be okay, but asking people to give up source compatibility and change applications is a much, much harder pill to swallow. Specialized environments (like the syscall-shipping CNK on Blue Gene like jedbrown mentions) are usually seen as something to work around because they are impediments to productivity. A more robust system wouldn't be considered useful if everyone has to start from scratch.
If you're running a server application you split it into some reasonable number of tasks and allocate cores to them as needed. Nothing else runs on the system at all. If you need more cores either upgrade your hardware or you build another identical machine and talk to it over the network. If you look at the way Moore designs, he's crafting a complete artifact to solve a specific problem, not just throwing generic parts together.