Cap'n Proto v0.2: Compiler rewritten from Haskell to C++11
kentonv.github.io
kentonv.github.io
Seems fair enough to me really?
Been down that road a few times myself.
C++ is a language only a mother could love, and only if that mother is a systems programmer who wants more abstraction than you can get with C. But for some reason that I don't think anyone has ever been able to explain, you can do some things in C++ that would be prohibitively difficult in any other language.
Interestingly, the Xen folks are now using OCaml for a bunch of system stuff. That could be worth checking out.
Yeah, most (perhaps all) design challenges where I want vtables can be solved elegantly in some other way in Haskell, but the solution depends on the use case and isn't a simple bijection. It's just tough to rewire the way the designs form in my head before I write the code. If I had infinite time...
I think this is debatable part.
First, calling Haskell from Python is quite easy. I've wanted that once and things went surprisingly well, see yourself: https://github.com/drdaeman/haskell-library-ffi-example
Second, not everyone has C++ compiler installed and readily available (although, it's true that it's far more likely that there's already a C++ compiler than Haskell one). But in the end, I don't think it really matters whenever one does `$package-install-command g++` or `$package-install-command ghc`, as, I believe, neither comes on out of box installs on most popular non source-based GNU distros, and certainly not on Windows. I don't know how things are in *BSD land.
Seriously, though, yeah, it's definitely debatable. If I had been loving writing my code in Haskell, the challenges of FFI probably wouldn't have been enough to sway me.
E.g. RMagick was famous for suffering from problems of that kind.
Sounds reasonable to me, but I think it's just personal preference. Without actually seeing the code (which I'm sure makes a huge difference), the imperative algorithm actually doesn't read as easily as the functional one does to me. We've become accustomed to imperative since it's what we all learned on, but there is a massive amount of book-keeping (although I'm sure the execution speed is much faster). At any rate, here is Scala with imaginary API's detailing how I would envision the functional algorithm...
def encodeFields(fields: List[Field]): String =
fields.map { f => f.position -> f.computeValue }.sortBy { _._1 }.map { _._2 }.mkStringThat said, given my knowledge of how the hardware actually works, the purely-functional approach makes my efficiency sense tingle. A lot.
Of course, whether or not it's a good idea to scratch it is something we've been debating for decades...
Very fair point about efficiency; for all my love of the FP style, I'll never claim it to be faster or leaner w.r.t memory usage. I'd be curious to see if the simple parallelization made possible by the functional approach could close that performance gap, though.
...and I'm almost positive the difference is subjective. =) Reading over your website, it's clear that you are a very talented engineer, especially in areas where C++ shines. Given this plus your experience at Google developing Protobuf, I'd hazard a guess that you have a significant amount of experience writing OO or imperative code, and probably somewhat less writing in the functional style.
Raw line count:
destroyer:alphaHeavy$ find . -name "*.hs" | xargs wc -l | tail -1
132272 total
Things that look like bang patterns or strict fields: destroyer:alphaHeavy$ grep "\!" -rI . --include "*.hs" | egrep -v "\\$\!|\.\!|\!\?|\!\!" | wc -l
1344
Strict returns: destroyer:alphaHeavy$ grep "\\$\!" -rI . --include "*.hs" | wc -l
762
Unboxed fields: destroyer:alphaHeavy$ grep UNPACK -rI . --include "*.hs" | wc -l
298However, strictness annotations in functions is a relatively rare thing.
import qualified Data.Vector.Unboxed.Mutable as V
encode len fields = do
buffer <- V.replicate len 0
forM_ fields $ \(fieldPos, fieldData) ->
V.copy (V.slice fieldPos (length fieldData) buffer) fieldData
return bufferObviously the functional parts of the language still applied with primarily functional algorithms, like parse-tree traversal and transformation, which is an excellent example of monads being useful and handy outside of pure and lazy language.
Note that the author also notes two significant problems that have nothing to do with functional programming or purity per-se:
1) Calling language-x from Python/Ruby/whatever is difficult unless language-x is C/C++. I fully agree here. At one employer we've built custom bindings from Perl to OCaml (for a DSL we've built) and it was troublesome (e.g., it was broken during 64-bit conversion and our patch fixing it was not initially accepted, although the Nria team later came up with the same approach -- which involved "burning" a register -- later anyway).
This is less of a problem for large scale distributed applications that talk via RPC, but this create a problem of building and re-building first-class RPC libraries in each language -- which (as in this very example!) is again best done by building the core RPC protocol implementation in C/C++ and linking it into the applications (with perhaps a native JVM implementation as JNI is a pile of manure? [Personal curiosity: with Google having a huge Java presence, how does Google handle maintaining up-to-date Java clients for BigTable/Spanner et al?]).
2) Lack of OO.
OCaml has OO but it's never used. OCaml has a marvelous module system that is far better than Haskell's. Unfortunately it lacks type classes (hence -- very annoyingly -- no generic "print"/"show" function), but first class modules and campl4 (equivalent of template Haskell) can help there. I am guessing that it doesn't solve the author's problem (which I honestly can't claim I have, but that is like my opinion, man) -- lack of dynamic dispatch.
Problem is inheritance based dynamic-dispatch polymorphism and generic polymorphism are hard to reconcile. I am very impressed by the work Odersky has done with Scala as well as C++11 (and boost/tr1 before it), but these are by no means simple languages (by simple I don't mean simple like Visual Basic, I mean simple as in "implementing a compiler for it is feasible in an advance undergraduate or beginner graduate class").
When I write golang, Haskell, Erlang, or OCaml I honestly don't miss OOP as modules, interfaces/signatures, type classes, and the like provide what I want out of OOP. However, I've never maintained (as opposed to wrote) commercial code in these languages.
tl;dr That's all folks, I'm now switching to an emacs window to write more C++ code -- while (as a big FP enthusiast) hoping practices for programming-in-the-large would emerge for non-OO statically typed functional languages (from places like Basho, Galois, and Jane Street who have built impressive multi-year projects in those languages), OCaml gets multi-core support, and Scala succeeds in its mission.
[edit]: English.
Instead successful RPC wrappers tend to follow two patterns:
1) Thin client, with heavy lifting done on a server side proxy written in the same language as the original client. This is the pattern we followed with Voldemort when I was at LinkedIn -- there was a heavy Java RPC client (which had a lot of client logic) and thin clients that (like parts of the Java client) used protobuf over the wire, but contained much less logic.
Problem with this is that it adds an extra hop (latency issue) and at times interface mis-match viz. local clients.
2) Only expose the protocol via the RPC, write a first-class implementation (given the full blown protocol) in major languages. Practically this means native "fat" client (whether in Java or C++) uses the full RPC wire protocol (but may use a higher performance server implementation) and works in the case where major languages are Java (or JVM-based) and C/C++: Python/Ruby/PHP/etc... use Swig or custom extensions to use the C++ client, Java has its own native client built in parallel. If you add another language without FFI to either Java or C/C++ to the mix (or there's significant velocity mismatch between team working on the Java-based vs. C++ based clients/server) maintenance becomes an issue.
This is an approach I was following on my last project at FB with HBase (I've since left FB and am elsewhere, but the work is continuing and will likely be open sourced) for another distributed system -- where the proxy Thrift service always lagged behind the native client and where most heavy C/C++ based clients built their own thrift proxies to talk to HBase. Upstream HBase (and HDFS) -- FB has its own (open sourced) branch -- has also followed a similar approach with converting the native protocol to use protocol buffers for SerDe.
Note that Cap'n Proto's RPC will support promise pipelining. This means that if one RPC returns a pointer to another interface, you can start making RPCs to the returned interface before the first RPC has actually returned. Basically you're saying to the server "When you finish this RPC, invoke this method on its result.". The hope is that this will make it possible to avoid a lot of round trips even when using a call-return-style interface, thus making it reasonably possible for clients to use the Cap'n Proto interface definition directly without the help of a fat client library. And that, in turn, gives you freedom to use any language that has Cap'n Proto bindings, rather than the specific languages the server chose to support.
Hah, dealing with native-encoded binary data in Haskell is rather cumbersome. I asked on r/haskell about decoding a native representation of float/double to corresponding Float/Double Haskell value, and this is the thread it generated:
http://www.reddit.com/r/haskell/comments/1k56i1/how_to_read_...
People have generally come with helpful suggestions, and rwbarton's excellent answer [check out the whole thread] has fully explained to me how to essentially reinterpret_cast a byte buffer.
I have been successfully playing with Haskell to implement some nontrivial _algorithms_ and liked. Like the OP, I'm now rethinking whether I want to invest more time into Haskell after seeing how cumbersome it can be to coerce the compiler into "DO THIS" where THIS is not a pure computation.
Now I'll take a pause from Haskell and start reading about Scala as it seems to provide a mix of FP and imperative (OO) that is more palatable to me.
I tried to figure out by myself how to reinterpret_cast a ByteString to something else (Float/Double in my case). I used ByteString because it was recommended by #haskell for binary IO.
I failed in my endeavor because the key function, which returns the pointer underlying ByteString, is in the hidden Internal subpackage. So I asked on reddit.
toForeignPtr should get you there: http://hackage.haskell.org/packages/archive/bytestring/0.9.1...
Had I known about the existence of toForeignPtr [the link to Internal module doesn't even appear in the locally installed library docs], I would not have had started the reddit thread.
Again, how was I supposed to know about the existence of Internal submodule and toForeignPtr function?
* You need to unmarshal hardware floats
* You visit Hackage, and use the IEEE754 package. http://hackage.haskell.org/package/data-binary-ieee754
If that is fast enough, you stop. You are done.
* IEEE754 is too slow for some reason, so you decide to do your own unsafeCoerce of bytestring
* You look at the source for IEE754 and decide to marshal memory yourself: http://hackage.haskell.org/packages/archive/data-binary-ieee...
* You read the bytestring package docs: http://hackage.haskell.org/packages/archive/bytestring/0.9.1... and pay attention to the "unsafe" functions that cast raw pointers.
* That lets you re-implement the IEEE754 package your own way
The mistake here was not using the highly-optimized existing library - http://hackage.haskell.org/packages/archive/data-binary-ieee... - which already does the correct thing. You jumped straight to "have to roll my own low level binary parsing" and then started a reddit thread instead. None of that was necessary.
But that was the whole point from the beginning! Find out how can I do it myself.
> You read the bytestring package docs:
I repeat: the Internal module isn't listed in the locally installed documentation, and the link on Hackage is broken. Do I really need to produce for you the two top-level links where I was looking for the documentation of Internal?
I am using JSON to save structured data for an application, and this is interesting, although my app doesn't spend alot of time encoding/deconding the JSON since the data structures are fairly simple.
If I wanted to use this for Java, wouldn't there still be decode step to convert the fields into java objects?
But... It's a cerealization protocol. :P
If I wanted to use this for Java, wouldn't there still be decode step to convert the fields into java objects?
Not necessarily. Your generated Java object could just contain a (ByteBuffer, int offset) pair. All the getter/setter methods then just write through to the underlying ByteBuffer.
In all seriousness, if you're familiar with protobuf the name seems reasonable enough. If you aren't, I guess you're not entirely the target audience.
Haskell having a high barrier to entry likely helps that along a great deal though.
Compare the quality in commentary on nontrivial subjects between /r/programming and /r/haskell for instance.
There will be an enormous amount of disinformation in the former, a state that is incredibly confusing for someone that lacks familiarity in the specific domain as discerning noise will not be easy.
On the other hand, /r/haskell largely in part does not comment as though they were an authority when they are not sufficiently competent to comment on a subject, so the signal remains high.
The subbreddit is fairly active given the niche current user base, and it is big enough that there are a decent amount of domain experts that comment.
The subreddit is also very polite.
I do not recall ever seeing a flame war break out there, which is notable compared to how often any discussion on the topic will break down into inanity, something which even occasionally happens here (such as https://news.ycombinator.com/item?id=6190005 from a couple of days ago).
> The subreddit is also very polite.
Yes, it is very refreshing: it is a complex language, but the community has not become elitist, and is instead polite, considerate and mature.
(Just kidding, I love GHC and its precise & helpful error messages. No sarcasm implied.)
"What an ignorant fool?! How could he abandon the holy grail of Haskell?!" -- like that? :)
For someone who wrote a lot of compiler and program analysis code in Haskell I can hardly imagine myself writing the same in C++. That said, I can totally understand the author's perspective. As the author writes, "I really wanted to like Haskell. [...] But when it comes down to it, I am an object-oriented programmer, and Haskell is not an object-oriented language." And I think that's the core of his troubles. I can imagine it's hard to write Haskell with an OOP mindset and expect it to be easy. I'm not saying pure functional programming is superior to imperative object-oriented programming -- they are just different and require different mindsets. Before I learned Haskell in the university the only language I had done any substantial amount of programming in was C++. And I had to make quite a switch in how I think of algorithms. I remember a lot of my classmates struggled with writing even the simplest programs in Haskell. But after some time I've really started understanding the Haskell way of thinking and, to me, it's easier to program in this language than in C++ which I regard as unnecessarily verbose and baroque.
Personally, I wouldn't even think about writing a compiler in C++ unless I was feeling masochistic. That said, I have never written a compiler in C++. So, should my experience in writing compilers in C++ have been greater and my experience with Haskell lower, it would have made sense for me to use the tool I knew how to use best (C++). Since I don't know how to write a correct compiler in C++ as fast as in Haskell, the only feeling I can express towards the author is pure and unadulterated awe.
PS: On that note, I've been shocked to discover that a horrified response from Haskell fans is expected. Haskell community is one of the friendliest and non-opinionated I've known, so I wonder why Haskellers are perceived as being akin to closet trolls.
The down side is that taking full advantage of metaprogramming in C++11 requires learning a very large and complicated language. But once you've done that, it's really not so bad.