Seems like it might take more than structure tweaking to overcome the quoted delta. I'd still like to see a general comparison too though, especially to Cap'n Proto on the one side and HDF5 on the other.
I haven't bench marked either of those 2 against each other but would not be surprised if they were very close to each other in the C++ implementations. That said, SBE does provide a non-JNI java implementation which as far as I know is not available for Capnproto yet. SBE also provides a C# implementation which Cap'n proto doesn't. That in and of itself may be enough for many people (Cap'N seems to have a lot of languages under development that SBE doesn't fwiw).
https://github.com/haberman/upb
http://blog.reverberate.org/2011/04/upb-status-and-prelimina...
The stock protobuf implementation is pretty close to optimal given what it does. Where upb is faster, it is faster by doing less.
For example, I can beat protobuf if you don't write the parsed data into a tree structure (or only write a few fields). I can beat it if your input has unknown fields and you don't care about preserving them. But if what you need is a tree structure that contains 100% of the input data, protobuf is hard to beat speed-wise (though there are still a few tricks, like arena allocation, that can beat it given some usage patterns).
This is just from one minute of looking. There are probably more problems. Unfortunately, I find this is typical of benchmarks that claim something is massively faster than Protobufs. Admittedly, this may be indicative of the Protobuf interface being unintuitive.
Not that I doubt that SBE is faster; the design should obviously make it so. But it's probably not 25x faster.
Update: Bug filed: https://github.com/real-logic/simple-binary-encoding/issues/...