Why is BIND 10 written in C++ and Python?
isc.org
isc.org
"As of right now, it ends up that about 75% of our code is C++ and 17% is Python (link) since it turns out that a lot of BIND 10 is performance-critical."
Which could easily be taken the wrong way. I believe the right way to think about it is "How much _more_ C++ code would there be, if there wasn't that 17% in Python?".
> [Python] has all of the features that we were looking for… except performance.
Now your interpretation is also valid since they could have written everything in C++ but they didn't.
Here is the benchmark result: http://speed.pypy.org/timeline/#/?exe=3,6,1,5&base=2+472...
0.002 for pypy VS 0.015 for CPython, which is 6.66 times faster.
So, for the last couple of years they have been wrong - since PyPy has shown pretty good performance in a DNS benchmark, so Python can have pretty good performance ;)
Minor pedantry: PyPy is not a language. It's an alternative compiler/runtime for Python.
If one language is more verbose, counting lines of code will throw off conclusions about how much functionality is implemented using each language.
It would be very interesting to see the code paths being run in python vs C. I suspect that the 17% python code is actually is around 80% of all possible code paths, but that the 75% C code is just like 5-15% of possible code paths. A possible way to check that could be by looking at the test suites and see which one is bigger, python tree or the c++ tree.
It seems that all apart from the performance critical parts are written in Python 3.1
I'm not saying the separate-process design of BIND 10 is bad (to the contrary it's good for security), but using multiple processes internally is only Unix-y in a superficial way.
It's odd that he's such an accomplished engineer but every presentation I've seen him give involves saying outrageous things while starry-eyed geeks stare at him adoringly. If it weren't for the fact that he's arguing from a position of (very great) authority, he would persuade very few people of anything.
He's loud and opinionated and one of the tech folk heroes, his opinion is going to be listened too and often parroted due to it, even if it's just his opinion rather than a technical fact.
Why is this? Because the man delivers.
If Linus is coding, it is probably more on his scuba diving tool than the Linux kernel.
First, by delivering working code that ended by running the majority of phones and smart devices out there. Also, by revolutionizing source code control. No, he didn't invent (almost) any of the concepts behind git, and by now the majority of the code was not written by Linus. However, he was able to strike a balance between features, usability, speed and working model that DID revolutionize version control. Monotone pioneered a lot, but was too slow and cumbersome; so was Bazaar without pioneering much. BitKeeper had a lot of things going for it, but freedom and price working against it. is more or less on par with git, but it's git that brought the revolution.
Second, by being able to successfully manage more than one huge project with hundreds of contributors, all of whom he can fire at, but which he didn't actually hire (nor can he, if he needs more work).
That is objectively incorrect. There is much less that can go wrong with C, as there is much less period. Everything that you can do incorrectly with C, you can do incorrectly with C++, plus 10 times more things that C++ introduced. The idea that C++ is safer because "we just won't do dangerous stuff" is silly, as any language is safe if you "don't do dangerous stuff", including C.
All languages are safe if you've written perfect code, but no one is perfect, C++ does try to catch some of the lower hanging fruit problems but if you're writing hoary code you're going to blow your foot off eventually.
- use safe arrays with bound checking if you feel like to
- use automatic memory management
- pass arguments by reference and being sure they point to valid data
- use proper strings without caring if the null character is missing
C++ is only dangerous thanks to the C compatibility legacy.
Don't use C'isms and the application will be a lot safer than doing pure C coding.
All of them would go away if C++ wasn't made to be C compatible.
You can bound-check your arrays at run-time, by wrapping them in a struct and only accessing them with functions.
There's even a proper (although conservative) garbage collector for C, while C++ usually boils down to reference counting.
Even in C you are not restricted to the standard library to handle strings. C doesn't have to mean null-terminated strings. (And in C++, too, string literals give you the old-time bad strings.)
This is no longer a data type seen as a language type, but an Abstract Data Type as known in Computer Science.
You are no longer using arrays, but a data structure made by yourself.
> There's even a proper (although conservative) garbage collector for C, while C++ usually boils down to reference counting.
If you are referering to Boehm-Demers-Weiser GC, it also works in C++.
C++11 also has a GC API.
> Even in C you are not restricted to the standard library to handle strings. C doesn't have to mean null-terminated strings.
The moment you do this, you are the strange kind in town as all libraries expect C style strings as input. So it is conversion party any time you need to call those functions.
> (And in C++, too, string literals give you the old-time bad strings.)
Yeah, this is a consequence of C's compatibility that infected C++.
Yes, it's no longer a built-in data type. But non-builtin-types are perfectly fine, too. By the way, how do you get bounds checking in C++? I guess you use the same technique, but C++'s overloading makes it syntactically easier to hide that you are not using a built-in?
Thankfully C++ is powerful enough that user types have the same rights as built-in ones.
This isn't true. C++ doesn't allow certain implicit casts that C allows (e.g., from void* to T*).
Uh. Everything that can go wrong in C can go wrong in C++ by definition.
C++ also adds a lot of extra things that can go wrong which don't exist in C. There are plenty of arguments to be made for C++ over C but "so much more that can go wrong in C" is not one of them.
The C++ FQA is strongly recommended reading: http://yosefk.com/c++fqa/
How much more work these 17% done in Python would take to be (re)written in C++? Would it make the code that much worse in terms of quality and managability? Having coherent, one-language codebase, is a good feature on its own too, often improving above mentioned factors.
I don't understand either, randomly opening parts of both the Python and C++ and it's certainly not complicated and they write the Python in a C style anyway.
I'm not a C++ coder and I only play with Python now and then, but can't say the code impressed me much. It's actually pretty hard to skim the code because it's massively over-commented and there are far too many tiny 2-line private functions that are called by exactly one other 2-line private function that is called by exactly one other 2-line function. Or Holographic code as John D. Cook called it[1].
And the copyright notice in every source file is extremely irritating.
Far too high noise to code ratio for my tastes, so I got bored before I could really 'see' how the Python had helped.
but I can't judge that well as I'm not sure if that's all just a bi-product of using C++ and being an open source project that comes out of committee.
[1] http://www.johndcook.com/blog/2012/01/09/holographic-source-...
That is because functions are a well understood abstraction. And having code separated into functions makes it easier for the reader to deduce the coupling points: two blocks of code after another can have all kinds of weird dependencies, e.g. the first blog might set some local variables that the second one relies on. But functions just have arguments and return-values.
I've found in my experience working with other people's code and maintaining large code bases you find that overly nesting functions causes a lot of problems when you're trying to read or debug code.
You often also see problems where the essentially dependant functions start to separate in the code as people accidentally add new functions between them.
The article I linked is good, John describes it well. It's a nightmare to work with when you get triple or quadruple nesting of tiny functions, like in the code of this program. It's totally unnecessary.
"The language had to address most of the problems with C. Ideally this meant something with good string handling, garbage collection, exceptions, and that was object oriented."
> String manipulation in C is a tedious chore.
Use a dynamic strings library, like Postfix for instance, and everything else.
> C lacks good memory management.
So strange that you went for C++ for most of your code that is not immune of problems from this point of view. I could understand that point if you were opting for a language with GC support. With C you can easily get better (that is, safer) than C++ native MM just building a reference counting system on top of your C "objects". This is trivial and it is what Redis, Tcl, C-Python, and many others are doing.
With Redis memory leaks or memory management never was a big issue.
> Error handling is optional and cumbersome.
Exceptions mostly suck, and in system software the only sane way to deal with errors is C-alike IMHO, that is, check the return value / error returned by every function and act accordingly.
> Encapsulation and other object-oriented features must be emulated.
This is not an objective point since many thing that this features are actually a problem.
Weak points IMHO, and C++ and Python with the minority of Python looks like a design error.
Why build your own when C++ has std::shared_ptr.
C doesn't have a destructor that gets called when something goes out of scope. That's taken advantage of in C++ (RAII) to implement various things (like scoped_ptr that helps avoid leaks).
Refcounting is also not perfect for every case.
But I like and prefer C. ;) Just pointing out one thing different in C++.
I've worked on C++ apps that had built-in garbage collection (basically asset (geometry) paging) for huge amounts of data that would allocate/free/page on demand based on what was going on.
Unique_ptrs are a total game changer, and the ability to use closures and lambdas when you want to set a callback function instead of a function pointer with a context pointer you have to cast and decode is absolutely huge for readability. maybe we aren't getting every last bit of performance out of it that we could with C, but it works at our pretty ridiculous scale, so I think C might have been a premature optimization for us.
And layers upon layers of object oriented crap is also a problem in C (I've seen it). At the end of the day I've just been burned more by the complexities of building large things in C (particularly when people do reference counting in C) than I have been by the complexity of c++ in general.
You meant that a linear type system is a total game changes. Unique_ptrs are an ugly hack to emulate a linear type system in an inadequate language.
As for MUMPS I know it is something used only in US it seems.
To nitpick a bit, compiled or not is more a property of the implementation than of the language itself. Of course, some languages are more commonly compiled than others. But, don't Facebook have a PHP compiler?
Yes, but I don't see PHP as a possible systems programming language, even with a compiled implementation.
Uhm, dreams of device drivers written in PHP...
C lets you manage memory without possibly inefficient or even broken magical black box automated processes or garbage collectors. Maybe what the author meant to say was "C lacks easy memory management."
There are two ways of allocating memory directly from the kernel (IIRC), brk and mmap.
There's some management malloc does, and if it's good or not depends on you application.
For example, size, number and behaviour of your allocations. Depending on your situation you may want to do your own memory management.
However, a call to malloc() is pretty straightforward, you expect it to return a pointer to the allotted memory, or not. Using a GC or the boots libraries is not quite as straightforward, and the black box is a lot bigger.
malloc and friends really are a quite poor way of managing a memory space if you want to do even slightly clever things I find, unfortunatly
Seems to me that you're making a weird strawman argument there.
You can't use realloc on C++ types (which typically need their constructors/destructors running without their memory address moving underneath them). I've written C types which had similar behaviour, and were not happy about being moved. Of course you can (and people do) write code which will after the move go through and do fix-ups, but it is often move pleasant to do the move yourself, if an in-place move isn't going to work.
Atleast i wouldn't sleep well if i know that the heart of the internet was running Java :P
Well, GCC (as in the compiler collection) has had a ahead-of-time compiler for ages:
There is also work in the LLVM camp on AOT Java compilation:
Atleast i wouldn't sleep well if i know that the heart of the internet was running Java :P
Well lucky you, only many other vital body parts are running on the JVM via Java, JRuby, and increasingly Scala ;).
- IBM J9
- Aonix PERC
- Oracle Squawk VM
- Oracle Embedded Java
- Excelsior JET
- Avian
- RoboVM
- ...
Apparently you are talking about the "Web", i was talking about the internet or networks in general.. I'm aware that tomcat may have a very large installation base. There is no use for your tomcat if the DNS is down, though :P
http://www.militaryaerospace.com/articles/2009/03/thales-cho...
The truth is, Java is viewed as "not Unix friendly" by many people, whether deserved or not. I suspect the ISC folks fall squarely into this camp. My personal and thus purely anecdotal experience is near-universal hostility from network admins.
Despite this, there is a fair amount of network management software written in Java. OpenNMS and Cisco Prime Networks (an absolute beast) both spring to mind.
I convinced unless a big OS vendor forces changes for their default system programming language, nothing will change.
This is why I like Microsoft's decision to drop C on Windows.
What I would like is to have a proper systems programming language in the lines of Modula-3/Active Oberon/C#/D/Rust or similar, being used.
Time will tell when this happens, but it will require a few generations of developers for the mentality change to take place.
C++ is a large, multi-paradigm language. It gives you choices and one is free to abuse those choices. But having more options gives you more power to express.
I think with modern C++ style, boost and C++11 its "difficult-to-work-with" reputation is massively overstated. It is possible to write succinct code with good design and get massive performance benefits.
"Whenever possible, we use Python"
"When necessary, we use C++"
"As of right now, it ends up that about 75% of our code is C++ and 17% is Python (link) since it turns out that a lot of BIND 10 is performance-critical."
Adding to this, the fact that BIND is something pretty performance-intensive, I don't think this makes it look bad at all.
This also clashes with the belief in the Python community that performance is not an issue because you just rewrite those few critical sections in C/C++. How's 75% as one possible definition of "few"?
As another comment puts it: How much more C++ would there be without python?
Of the total code of a very specific application, with very specific performance needs.
For others applications, including servers, the ratio could vary widely.
The percentages may or may not be simply a case of C++ verbosity, but without actually browsing the source, we shouldn't jump the gun on exactly what percentage of "heavy lifting" C++ is doing vs. Python.
Because it directly contradicts the "only 20% of your code is performance sensitive and the other 80% can be scripting language X" nonsense that scripting language apologists constantly parrot with no evidence. This is a good example of how scripting languages are in fact not well suited to application development, and should instead be used for scripting.
To put it more concisely: "only 20% of your logic is performance sensitive and the other 80% can be encoded in scripting language X leading to a significant reduction of your total code."
Expansion Expanded 1 to 1: 75/ 17 --> 75% 17% 1 to 3: 75/ 51 --> 56% 38% 1 to 5: 75/ 85 --> 45% 51% 1 to 10: 75/170 --> 30% 67%
The LOC doesn't measure how many features are implemented in each language. If a language is more verbose or need more detailed management the same features will appear to be longer.
Several boost libraries are inspired by python and C++11 features all help to write C++ code that is surprisingly similar to python, with a bit of extra type sepcification. If you think otherwise I'd suggest you're probably thinking of the C++ of the 90s - a more C with classes, and not modern C++.
I often find it useful to develop and prototype in python and convert to C++. Most often I'm doing this because my python simulations can take days and the C++ versions hours. Often I find it is not just the performance critical areas, but it is easy enough just to wholesale convert the lot.
I agree, especially with C++11 and (besides Boost) Qt. It's often the header/code separation that makes things a bit tedious, having to keep function and method signatures in-sync. Of course, if you are template-land that is not that much of a problem.
We're also replacing core op-graph and geometry processing from Python to C++, and again it's close to 1-1 - there's a bit of overhead in the loops - in our coding style we're caching begin iterators on the line before the loop, but other than that it's very close.
And the speed of the app is so much faster it's not even funny.
Can you give an example of what this looks like?
python:
for face in mesh.faces():
faceCentre = Point()
for v in face.vertices():
faceCentre.add(mesh.getPoint(v))
faceCentre.div(len(face.vertices()))
C++:
std::vector<Face>::const_iterator itFace = mesh.getFaces().begin();
for (; itFace != mesh.getFaces().end(); ++itFace)
{
const Face& face = *itFace;
Point faceCentre;
std::vector<unsigned int>::const_iterator itVertex = face.vertices.begin();
for (; itVertex != face.vertices().end(); ++itVertex)
{
const unsigned int& pointIndex = *itVertex;
faceCentre += mesh.getPoint(pointIndex);
}
faceCentre /= (float)face.vertices().size();
}
So the C++ is longer, but you've got braces, and the references to the Face and unsigned int pointIndex are placed as a local variables, which makes debugging much easier - they could be inlined.It's possible to get that down even more using more modern C++ - using auto variables and not declaring the start iterator on it's own line.
So yes, counting braces, it can be a lot more lines, but if you don't count braces, it's generally not that much more.
Also, as of C++11 you can do this:
for(const Face& face : mesh.getFaces()) {
Point faceCentre;
for(size_t pointIndex : face.vertices())
faceCentre += mesh.getPoint(pointIndex);
faceCentre /= static_cast<float>(face.vertices().size());
}We sometimes hoist the end iterator as well, as often g++ can't optimise out the call to .end() each iteration - it can if there's a ref to a const item and you call end() on that const ref, but otherwise, it generally doesn't as it can't guarantee the item hasn't been modified.
We're still stuck with CentOS 5.4, so g++ 4.1 for us as that's what we've got to deploy to (although we use ICC for production builds, building off the g++ 4.1 standard headers)...
Basically, we want top possible speed - if that means the code's a bit more verbose than it can be, so be it.
And just look at the difference in readability/simplicity. I can explain the Python code to my 12 year-old cousin, the C++ version though...
(https://plus.google.com/115212051037621986145/posts/HajXHPGN...)
(And as others have said, 17% of LOC in Python might still mean majority of the functionality in Python).
What I'm saying is that "amount of code" is not a valid magnitude when comparing different languages.
In my experience converting some unmaintained Perl tools to Python, the Python version was almost as long as the Perl one, but it included more functionality, comments and error handling (I'm not criticising Perl here but the unmaintained code that I had to convert).
Read: "This is one of the cornerstones of the internet that we didn't want to piss away on novelty".
That sounded rude, and I'm sorry for that, but I don't know how else to make that point cogent.