write(fd, buf, size);
Can you list all the caveats, possible side effects, reasons of failure of this simple C function call? Hint: remember, fd can be a file, as well as a NFS-mounted file, pipe, network or local socket, FIFO, device, etc. Hint2: I doubt anyone can give a comprehensive description of the possible consequences of this call.Bottom line being, C can be incredibly hard to understand, or C++ can be clean, succinct and straightforward. Personally, I'm for the second, although I know probably 99.9% of all existing C++ code is just awful.
Stripping off ambiguity is almost the write syscall handler's only job.
It's worth mentioning here that not only is the convenience vs. explicitness tradeoff different in application code than in bare-metal systems programming code, but that that specific example of convenient app-level API is also an infamous Unix failure. In reality, app code cannot safely assume that an fd is just an abstract bucket you can read, write, close, and seek in.
1. General Commands
2. System Calls
3. Subroutines
4. Special Files
5. File Formats
6. Games
7. Macros and Conventions
8. Maintenence Commands
Really? write a block of data of size on bytes "bytes" on the file-pipe-network... pointed by the file descriptor "fd"??
It is extremely simple to understand, it makes a very simple thing, always right. I doubt you can make it simpler.
I have been using this all my life for NFS-mounted file, pipe, network or local socket... Never had any problems, and I have done very complex things.
Of course, in c++ you will use the very same function under a different calling, you can make it part of a class, whatever, but is going to use the same backend internally(because the OS only uses one), only that you can add a lot of abstraction(complexity) from c++ code.
Of course C++ could be clean, but the question is: what will happen if we make it into the kernel?. Linux thinks it won't work. I agree with him.
(Note that my point here is that for sane/good code, you normally can be sure that it does what it looks like -- for C++ just the same way as for C.)
Even mathematics has different processes for the same operators on different types. If you see 'a * b' you should be able to assume that it's multiplying two variables, but if those two are real numbers it's different than if they're two matrices of real numbers, for example.
Implementing a Matrix class that overrides * to imlement matrix multiplication seems perfectly fine for me - It's just the same as implementing a Multiply(a, b) method. Just as someone who writes a Multiply method could in theory make it do anything, it is assumed by reading the word 'Multiply' that that is what it does. Just as someone who makes a Multiply() function that does something other than multiply is an idiot, so is someone who overrides operator * to do something other than multiply.
* If you write to an NFS-mounted file and write() returns an error, does it mean the block wasn't written to the remote disk?
* If you write to a TCP socket and write() returns OK, does it mean data was received by your network peer?
* UNIX pipes: how many times is your block being copied before it reaches your peer's buffer supplied to read()?
And don't get me started on signals.
My favorite part of the whole man page is the Linus quote:
"The thing that has always disturbed me about O_DIRECT is that the whole interface is
just stupid, and was probably designed by a deranged monkey on some serious mind-con-
trolling substances." -- Linus
But thats not a language failing. A language that actually prevents you from writing complicated interfaces is probably entirely unsuitable for... almost anything.Overcoming the obvious difficulties, including "new" being a variable name in a library, there were no technical difficulties. (Had no one in history EVER written in C++ on linux before?)
The argument FOR C++ in the linux kernel is the same argument for using C++ anywhere else - object abstractions are useful. The project took far less time. The bugs were fewer. The code was of course more readable - because of the tremendous wealth of context that C++ provides thru strong typing.
C people argue its easier to use simple constructs, since they are instantly understandable. What is NOT understandable is WHY the code is putting an int into an array. Those types have no obvious meaning beyond the line of code they are in.
In an object-typed language, you CAN learn the scenery and rapidly become familiar with the object set. The argument above (write) is largely silly - right-click and GoToDefinition works in almost any IDE - yes there may be more than one possibility, the IDE will display them all. So navigation thru code is vastly simpler than grepping for names.
Anyway, in 18 months we had an Infiniband layer integrated into linux, not too hard. Had to invent fundamental kernel/user page primitives that were missing. Interesting to note: Windows had all the driver support we needed, didn't have to invent anything.
int i = 42;
foo->bar(i);
In C, we can tell that foo is either a "struct S* foo" or "union U* foo" which has a member "X (*bar)(Y)". Type X is unknown from this context, but Y is some type compatible with int. We will call the function which bar points to with a single value 42. The value of i will be unchanged.In C++, it could be the same. Or foo could be not a pointer at all, but some object with "operator->". bar() might take its first parameter by reference and end up changing i. bar() might have additional parameters with default values. bar might not even be a function, but some object with "operator()". etc. etc.
Most of the code I write is C++, so I'm not against it -- but I definitely understand the point that C requires less context.
Of course you can not look at the function and immediately know all consequences. The real point is that there is a procedure you can follow which will let you determine the answer.
1. Locate the "write" function. There will only be one, because C has no namespaces.
2. Read the "write" function.
3. Repeat recursively as needed (including for macros).
For C++, the equivalent would be something like fd->write(buf), and the procedure is:
1. Determine the type of 'fd'.
2. Determine what subtypes you could have there and which you might actually have in hand. Or is the class not virtually inherited in which case it doesn't matter?
3. Determine what the write method does. In order to do so, have intimate knowledge of all operator overloading any value used in the write method may have.
4. Figure out what "buf" is and whether it magically overloads other operators.
write(fd, buf, size) resolves to one function with three arguments that themselves can't be that magical, and usually one basic approach to the question of memory management. fd->write(buf) involves classes, inheritance, potentially overloads, interfaces that 'buf' may correspond to and the potential need to follow a chain of some number of functions just to see whether that was constructed automatically into another type, endless permutations of how memory may be handled, and so on and so forth. The assembler that the C code will generate will basically push three arguments and call a function; the assembler the C++ generates is effectively unbounded in complexity. This assembler complexity directly corresponds to complexity that must be understood in order to understand the line of code.
In general I'd prefer the C++, but when writing a kernel where every twitchy detail counts for everything and the slightest bit of "wrong" could be a rootable security bug, I see the counterarguments.
Also, generic functions in Common Lisp don't suffer from nearly the same amount of complexity as can be found in C++: there is always only one signature (lambda-list) for any function name, whether generic or not, there are no implicit conversions, no memory-allocation details, value/pointer/reference distinction, constness, virtualness, overloading of assignment operators etc.
I was under the impression that the function signature was not the issue raised... but the fact that the same function name could lead to completely different behaviors depending on a context. In the case of generic functions, the specific function implementation is usually (ignoring some of the other matching features) determined by the input types. The input types define the context.
All that said, I realize now that this might be less of a problem in CL if you just group your defmethods in the same area. If the problem is grepping for the definition, it is easy enough to rearrange things so that the related functions are in a close-enough place. In C++, the classes try to own their methods -- with the exception of warily-regarded "friend" methods.
http://www.cs.cmu.edu/Groups/AI/html/hyperspec/HyperSpec/Bod...
Though grep would probably do just fine too, since method definitions are defined with 'defmethod' instead of 'defun', so you can see at a glance whether something is generic or not.
Package is essentially a collection that maps symbols to values. In CL packages have (at least) two namespaces: one for functions, other for any kind of variables. The reason for this is that when you write the form
(fun #'a b)
you and the compiler can be sure that 'fun' and 'a' are both meant to be functions. There are both advantages and disadvantages to this.Edit: Well, seems you don't believe it.
Take a look here: http://www.google.com/codesearch?hl=en&lr=&q=lang:c%...
And here: http://www.google.com/codesearch?hl=en&lr=&q=lang:c+...
Just browse a bit through both, check a few different projects. Check also some of the big projects. (Apache, GCC, glibc, Linux, LLVM, clang, WebKit, Chromium, etc.)
C++ has too many rules and exceptions, while giving some sort of control to a genius - in hands of a good natured well meaning coder, it all goes to hell. That is so not in the beginning but in the long run.
All misgivings noted, compilers gone to future and back. They help alot.
C has few rules, and whats more if you use UNIX / ISO convention there are very few ways you can go wrong. C++ is an ugly duckling of era of 586 type machines, when you have tried to do something in binary , efficiently and with a style and yet still was found wanting.
Explicit is always best, because you always refactor. Say what you mean, write what it does. This way its easier to read, and easier to coax it to do something else. In the end good coders achieve things by correcting few things here and there. If you have some sort magic that developer can't trust to be exactly what they expect it to be - they can't code with honestly and without fear. And fear as we know is a mind killer. So it is a catch 22.
I am totally 100% with Linus on that one.
* If some moron created a "write()" macro, it means expand that macro (preferably it should also mean email the director of HR to suggest the author of the macro consider having the company pay for their MBA, to make sure they aren't allowed to touch code again)
* Otherwise, if a header has a definition for write use that. Otherwise, emit instructions to call put "fd, buf, size" on the stack and call write.
In C++ it could mean:
* There's a class "write" which has a constructor that takes three arguments
* There's a function in the global namespace called "write" that takes three arguments
* There's a macro called write
* There's method called write in the local class. That method could be invoked virtually or non-virtually. It could be inherited from a parent class.
* write could be an instance of an class that has "operator ()" defined
I probably left out some more. In fact, there's likely a firm somewhere that thinks "what could write(a, b, c) mean in C++" is a wonderful interview question (perhaps they could ask it to all those losers who don't know what "explicit" keyword does but still have the gall to think they're competent enough to work for them!)
In a UNIX system, with right headers included, what write does is specified by the POSIX API. POSIX is one of the best defined and cleanest APIs. It has existed before the world of IDEs and plug-and-play libraries. I can write non-blocking C network code for a UNIX OS with vi, from muscle memory after reading through man pages and working through Richard Stevens books.
Writing Java NIO code requires: an IDE to prevent RSI, traversing Javadocs to understand the non-intuitive APIs and searching mailing lists through Google to e.g., find out that Java NIO selector doesn't let me use edge triggered epoll because it would involve "tight coupling" (read: it might be difficult for somebody writing code on an AS/400 to use it).
JDK7 NIO2 potentially changes this (I can implement internals of a selector myself). I'm also sure I'll be able to target JDK7 with a Perl6 compiler... that I'll use to implement the firmware for my flying car, in which I'll travel pick up Hans Reiser when he's paroled from prison.
Note: in this case, it has nothing to do with the language. java.util.concurrent API is very well defined because it was written by great programmers (Doug Lea and Joshua Bloich) who are apt at API design. Unfortunately, Java doesn't "force you" to create a clean API like C does: there are no design patterns available, no IDEs, no built-in tools for literate programming in C; you either build a clean API, or no one will use it. While I'm in no way of Joshua Bloch's caliber, I can relate to him when he says that he stayed with imperative C until finding Java.
While Java doesn't force you to build a clean API, C++ almost makes it impossible to build a clean API: witness boost::spirit ("generic programming", C++'s idiom for extending the language), compare with lex/yacc (external DSLs) and parser combinators (internal DSLs) in Haskell or Scala.
I want to like C++. I have programmed it for a living before and will almost certainly do so again: there's a certain combination that requires low-level code with no memory management and Object Orientation; my chosen specialty (distributed systems) often requires that combination (fortunately, not always: Erlang and JVM languages have been used to build some incredibly impressive systems).
Perhaps Go and D could come along and step up to that challenge, but I am skeptical: Modula-3, despite influencing other languages hasn't been able to step up to that plate. I feel C++ 0x, Intel collections for C++ and some parts of boost, STL and tr1 e.g., tr1::unordered_map, boost::scoped_ptr are very cool and useful. Boost Graph library is simply awesome and has no equivalent.
It seems, though, as if C++ was built by warring hordes: one horde that wanted generic programming and thought OO was pointless, another horde that wanted OO but thought generic programming was pointless and yet another that hated both. Neither won nor lost, each side said "mission accomplished" and developers were treated as "collateral damage". I could care less about which one of those to use: I am productive doing OO programming in Perl, Python, Scala and Java. I am also productive doing generic programming with CLOS in Common Lisp or with type classes (or their equivalents) in statically typed languages. Likewise, I am perfectly productive writing imperative code in C. I just want clean, easy to program to APIs with understandable error messages (either at run time or compile time, I am not picky about dynamic vs. static typing -- they're tools, means to an end) and C++ doesn't allow for that.
http://blog.objectmentor.com/articles/2009/07/13/ending-the-...
"The fact that it took decades for the industry to arrive at something as useful as ActiveRecord in Rails is due primarily to the attitude that some language features [in this case, meta-programming] are just too powerful for everyone to use. "
There's also lots of subtle ways for the DB and a particular object or set of objects in memory to get out of sync and cause hard-to-understand bugs.
ActiveRecord is definitely one of those tools that "make easy things easy". But, at the same time, it's not hard to get overly clever and make a mess with it as well.
What's more, I finally gave up after a year of waiting for this pretty fundamental bug in the postgres driver to be fixed and finally just hacked around it: https://rails.lighthouseapp.com/projects/8994/tickets/2622-p...
This was how I felt when I first scratched the surface, but by using find(...) instead of raw sql you leave open other options for the future like using :include
Yeah, a lot of times you're using the same keywords you might use in SQL, but the new chainable extensions are really nice for building up queries and reusing common chunks (either in LINQ or in Rails 3)
I should note that this new relational syntax is provided by a library called Arel (http://github.com/nkallen/arel)
In a huge, sprawling project in which the relational database is a small portion, it may not be worth trading the power of ActiveRecord for the accompanying loss of explicitness and clarity.
-- > That doesn't mean Ruby is bad. Ruby is great. ActiveRecord is a great model for MVC web development. The standards of Ruby aren't good for more performance-oriented, "hard core" development - Ruby's a pretty mature language but Rubinius, the Ruby-in-Ruby project still isn't production-level. One might guess that compilers/interpreters require a different standard than web frameworks. Neither is bad though (writing a website in c/c++ would be onerous).
If the space shuttle code was written accord to the standards of Linux code, it would come apart completely too.
It's shame that Linus had to frame things as good language versus bad when it's a matter of the right tool for the right job.
Still, the proofs in the pudding and anyone who can write a kernel in C++ or Ruby would sure give those languages a boost (and these are the languages I like best).
∀job, ∃L | L(job) > C++(job)
This, plus stating that C++ is very complicated amounts to say that C++ is bad. Because even if C++ is "good enough" for a wide range of jobs, it's complexity makes it longer to learn than several, simpler, more specialized languages.My personal opinion is that C++ is best only when legacy code is involved, or when the team just won't learn other languages. I can understand them. Learning C++ is such an investment that they are more likely to "throw good learning after bad", or may think that learning another languages will be as difficult as learning C++.
I may change my mind when I bother to look at LLVM or V8. Perhaps.
It means that by their being an explicit exception a rule must exist for their to be an exception. For example:
"Special leave is given for men to be out of barracks tonight till 11.00 p.m."; "The exception proves the rule" means that this special leave implies a rule requiring men, except when an exception is made, to be in earlier. The value of this in interpreting statutes is plain.
As others say in this thread: ActiveRecord makes easy/simple things very easy. Unfortunately, it makes harder things damn near impossible. To aggravate matters, people that only use it for what it was designed for, tend to blame you, instead of acknowledging that it may very well be that their favorite tool doesn't support a particular kind of use.
I much prefer a few different tools that do a few different things really well. I don't use active record if I have to deal with legacy schemas. I'll use data mapper or roll my own. If I get to design my own schemas then I love the benefits AR gives me.
You should love (parts of) Haskell then. You even need to specify whether your code can have any side-effects there.
(Haskell's Type-classes are awfully implicit, on the other hand. Though not nearly as bad as overloading in C++. You do not need to abuse bit-shifting for sending stuff to streams in Haskell. If you want/need to, you can just make up your own new line-noise operator. That new operator won't be used anywhere else, and is thus perfectly grep-able.)
On the other hand, you can write obtuse and difficult to understand code in expressive and non-expressive languages. I suppose a point could be made that bad code in say Haskell or Lisp might be worse than bad code in Java.
Clojure in it's current incarnation is a great example of this problem, IMO. It provides a very expressive and highly abstract interface to the JVM and java libraries, but once you hit a stack trace the abstraction comes tumbling down and you have to start picking through the mixed java/clojure stack trace to figure out what went wrong. Throw macros into the mix and things get even hairier. I've found that in some cases it's better to just bang out vanilla java. Sure it takes more LOC to get the same things done, but the result is often dead-easy to understand and debug.
Like everything else in engineering though, it depends a lot on exactly what you're trying to build.
Good stack traces are about the maturity of the compiler and tool support- they have absolutely nothing to do with clarity of the language. Clojure is simple and Clojure opts for explicit context over implicit context almost everywhere (dynamic binding being a big exception)- this is exactly what Linus argues for. All you have in Clojure are functions and values - how much straightforward can you get? No objects, no hidden behaviors, no private variables, everything in namespaces, etc.
After two years of hacking on Clojure, I haven't really found macros to obfuscate anything. But that's because I learned not to use them unless I need them.
But, yeah, I'm looking forward to the community growing and providing better error messages and tools for deciphering raw Clojure stacktraces.
From another point of view, Clojure, like all higher level languages, is just an abstraction over an underlying machine. From this angle and in it's current state it's (IMO) a pretty leaky abstraction.
I think it's fair to say that this has more to do with the implementation than the design though.