Flatbuffers by Google – CapnProto alternative
github.com
github.com
http://kentonv.github.io/capnproto/news/2014-06-17-capnproto...
[1] http://kentonv.github.io/capnproto/news/2014-06-17-capnproto... (see feature matrix)
[2]http://google.github.io/flatbuffers/md__benchmarks.html (introduction/top of article)
MinGW and limited MSVC support will come soon.
0.5.0 will also feature cmake support (this is already in git).
e.g. if I make an app using 10 FOSS libraries, then I wouldn't want my app reporting to 10 different places everything which the user is doing.
Also, on the actual homepage for it (http://google.github.io/flatbuffers), the only mention of this call-home feature is buried at the bottom of the "building" page.
[EDIT] This is incorrect; see below comments about this tracking not being a call-home feature but instead just Google scanning apps submitted to the Play Store
Could it instead be just that Google scans Google Play apps for a string in the binary that matches the Flatbuffers version string format? That seems more likely given what the README does say about this. And it also seems more useful in general; Google would benefit more from knowing how many applications use the library than knowing how popular these applications are.
Of course that's what's happening: https://github.com/google/flatbuffers/blob/master/include/fl...
I actually find it kind of interesting that I mentally turned "tracked" into "calls home".
It seems like a very good way to jauge the interest for this lib on Android in order to decide how much resource they will allow to its dev.
It would be really interesting (and possibly more relevant for HN) to have benchmarks based on one dynamic language - say Python.
Oh and @kentonv - I'm not a native American English speaker (rest of the world really). I really, really have trouble pronouncing Capn'Proto. Even more difficult to pronounce it in a meeting and have people recall/Google it.
The real problem in dynamic languages is that they tend to be worse at inlining accessor functions. This is not really because inlining is impossible -- v8 can do it -- but because most dynamic languages don't prioritize performance in the first place and so haven't implemented such optimizations. This is actually a problem in Go as well, weirdly. Because of this, if you actually intend to consume most of the content of a message, it may make sense to parse it into a language-native data structure up front so that access doesn't need to go through accessor functions. Most Cap'n Proto implementations support this. Doing this will still be much faster than using Protobufs because the Cap'n Proto format is naturally faster to decode.
As David says, "Cap'n" should be pronounced like "happen", though pronouncing it as "captain" is OK as well (and will still get people to the right place if they Google it).
I dont know the answer, ergo the question.
Of course, in cases where Cap'n Proto has a more-than-constant speedup, such as reading a single field from a large message (O(1) in Cap'n Proto, O(n) in Protobufs), then the difference will still be huge regardless of language.
If you're looking for specific benchmark numbers, I don't have any handy, sorry. (But benchmarks can be manipulated to show any result, so you shouldn't trust any author-provided numbers anyway.)
I was thinking in that context.
[1] http://kentonv.github.io/capnproto/news/2013-09-04-capnproto...
I actually think it's likely that a pure-Python version of Cap'n Proto would be significantly faster than the pure-Python protobuf implementation. Parsing Protobufs in Python is really horrible performance-wise since you have to inspect and branch on almost every byte. The way to make Python fast is to delegate as much work as possible to the built-in libraries that are written in C. But, there's just nothing that can be delegated in the case of Protobufs. In contrast, a Cap'n Proto parser could pretty easily leverage the existing `struct` module.
That said, if you enable Cap'n Proto's "packed" mode, then this advantage is lost, since that's another byte-by-byte algorithm that will perform poorly in pure Python.
Strings are simply a vector of bytes, and are always null-terminated. Vectors are stored as contiguous aligned scalar elements prefixed by a 32bit element count (not including any null termination).
So... does the count include the null terminator byte or not?
The spec is a bit confusing though, because of the statement that "Strings are simply a vector of bytes". The way I understand it is that a string is a vector FOLLOWED BY a null terminator. The spec should probably say that rather than the current wording.
This would appear to be necessary so that a string can be treated as either a vector like any other (with the correct number of elements) and can also be accessed directly by a (char *) pointer without things going awry.
Disclaimer: I haven't read the whole spec yet; this is my off-the-top-of-my-head interpretation and I may have misunderstood it completely.
Technically speaking, you an implement a c-style string using a STD:Vector by ignoring the length preamble and ensuring room is made for the null character. I got away with this in my intro to c++ class after showing the teacher that I already knew how to implement strings in C from a previous class.
C-style strings: http://www.learncpp.com/cpp-tutorial/66-c-style-strings/
C++ Vector class: http://en.cppreference.com/w/cpp/container/vector
Java Vector Class: http://docs.oracle.com/javase/7/docs/api/java/util/Vector.ht...
I think the big draw for the flatbuffer system is that it can stream data in with a low memory foot-print.
Are you finding this particular tracking code acceptable? Apparently not.
So there was never any coherent whole that could have found something unacceptable to begin with and in the end there are still disparate parts that continue to find it unacceptable.
I guess this is an obvious and tiresome answer, but I'm not sure what else you would expect anyone to say.
Just because I review code doesn't meant I accept or use it so your assumption is wrong.
There should be nothing tiresome about calling out for discussion of evolving trends in software or anything else.
If you see nothing negative about this practice then speak out constructively.
What I find tiresome is insisting that the use of some license or the other is a statement of values (your phrasing also implies that history clearly agrees with you, which I tend to find tiresome).
If the readme and other materials made repeated attempts to invoke some set of values and the source was contrary to that, fine you have a valid gripe, but the readme doesn't mention the license and the homepage ( http://google.github.io/flatbuffers/ ) keeps it to "It is available as open source under the Apache license, v2 (see LICENSE.txt)."
https://github.com/google/flatbuffers/blob/master/include/fl...