A new ProtoBuf generator for Go
vitess.io
vitess.io
http://www.brendangregg.com/blog/2017-05-09/cpu-utilization-...
A much better way to test the influence of the new compiler would be to test the actual throughput at which saturation is achieved (which is what the benchmark in the C++ grpc library measure to assess their performance).
> The maintainers of Gogo, understandably, were not up to the gigantic task.
I'm 99% sure they are "up to" (as in "capable of") doing so, they are just not "up for" it (as in, "will not do it").
That said, I love the detailed post and the interesting solution, and the commitment to performance!
Making vtprotobuf an additional protoc plugin seems like the Right Thing™, although it's a shame how complicated protoc commands end up becoming for mature projects. I'm pretty tempted to port Authzed over to this and run some benchmarks -- our entire service requires e2e latency under 20ms, so every little bit counts. The biggest performance win is likely just having an unintrusive interface for pooling allocated protos.
(Holy shit, who is downvoting this? It's literally the whole article!)
Just in case you may be unaware, the latest GCs for Java (Shenandoah, ZGC) are miles ahead of anything available for Go due to sheer age and manpower. Parallel and Pauseless are easily achievable in most cases.
The point is, neither is "five orders of magnitude" below 20ms. And neither needs zero CPU even if it doesn't block other threads.
Beyond hyperbole, do you have any actual comparison of Go vs Java GC performance?
okay, you've done this, three years later and it's the same thing again since you need to accomodate the new features. your users haven't upgraded their computers. what do you do ?
Then just like C, writing a tiny set of functions in Assembly is always an option.
> If your path is sensitive to 200us of latency you should probably optimize your application and tune your GC.
> Don't guess, measure.
yes, I am saying that the original code has already gone through a complete optimization process, everything is written in written in assembly, and you are at 970us on your 1ms time budget. (I'm not pulling those out of thin air, I was literally at a client yesterday with some real-time code with a 1ms deadline on a desktop OS and we have to cram as much things as possible in that millisecond)
There is nothing one can do to a, say, a 1 kilo byte buffer that will cross 1 ms in any language. My own Go code doesn't cross more than few micros per message.
Perhaps they have significant external (network) latency leaving only a few ms budget for the application stack - so they could easily be up against a wall.
I agree it's unlikely the difference here will be solely responsible for tipping the GP's request above 20ms, but the memory problems could reasonably ruin tail latencies.
Why are you using Go then?
EDIT: Fixed unfortunate typo
Go doesn't give you control over inline vs indirect allocation, instead relying on escape analysis, which is notoriously finicky. Seemingly unrelated changes, along with compiler upgrades, can ruin your carefully optimized code.
This is especially heinous because it uses a GC; unnecessary allocations have a disproportionately large impact on your application performance. One or the other wouldn't be nearly as bad.
Time and time again we see reports from organizations/projects with perfectly fine average latency, but horrendous p95+ times, when written in Go - some going as far as to do straight-up insane optimizations (see Dragph) or rewrite in other languages.
https://medium.com/a-journey-with-go/go-introduction-to-the-...
No different than running other kinds of static analysis for well known languages, unsafe by default.
P99 < 1ms, that's when you're going to want to switch it up.
You could start with twiddling some of the GC knobs Go gives you, but you're still working against an SLO. If you need stronger guarantees you'll look at languages that completely eschew GC, because Go's GC still has STW bits. Climb the ladder further and you're reducing allocations, eventually avoiding any malloc() beyond what it takes to get an arena and doing your own bookkeeping. I've never been near the top of the ladder when you have hard real-time constraints, but I've heard it involves paying Wind River for VxWorks licenses ;)
- If you write the same message multiple times, protobuf implementations should merge fields with a last write wins policy (repeated fields are concatenated). This includes messages in oneofs.
- For a boolean array, you're better off using a packed, repeated int64 (if wire size matters a lot). Protobuf bools use varint encoding meaning you need at least 2 bytes for every boolean, 1+ for the tag and type and 1 byte for the 0 or 1 value. With a repeated int64, you'd encode the tag and length in 2 varints, and then you get 64 bools per 8 bytes.
- Fun trivia: Varints take up a max of 10 bytes but could be implemented in 9 bytes. You get 7 bits per varint byte, so 9 bytes gets you 63 bits. Then you could use the most significant bit of the last byte to indicate if the last bit is 0 or 1. Learned by reading the Go varint implementation [2].
- Messages can be recursive. This is easy if you represent messages as pointers since you can use nil. It's a fair bit harder if you want to always use a value object for each nested message since you need to break cycles by marking fields as `T | undefined` to avoid blowing the stack. Figuring out the minimal number of fields to break cycles is an NP hard problem called the minimum feedback arc set[3].
- If you're writing a protobuf implementation, the conformance tests are a really nice way to check that you've done a good job. Be wary of implementations that don't implement the conformance tests.
[1]: https://github.com/protocolbuffers/protobuf/tree/master/conf...
[2]: https://github.com/golang/go/blob/master/src/encoding/binary...
[3]: https://en.wikipedia.org/wiki/Feedback_arc_set#Minimum_feedb...
The solution for this is to subtract 1 from the integer every time you encode a byte (since the existence of the next byte you're adding already indicates that the intermediate value isn't 0)
If you are willing to use cgo, google already implemented one for gapid.
https://github.com/google/gapid/tree/master/core/memory/aren...
There is still so much education to do.
Most of the time, their non-presence is due to general pools being just as good most of the time, or people simply not needing them that much with modern GC
Probably lack of experience with machine friendly code.
From the same author, a zero-unsafe arena allocator: https://github.com/fitzgen/generational-arena
There are many, many arena implementations available with varying characteristics. It's disingenuous to act like Rust requires the author of an arena library to write "unsafe" everywhere.
In what concerns me, although I like Rust, I only see it for scenarios where any kind of memory allocation is very precious, Ada/SPARK and MISRA-C style.
I have been using GC languages with C++ like features, or polyglot codebases, for almost 20 years to think otherwise.
Most of the time developers learn about new and miss out on the low level language features.
It is a matter of balance, either trying to do everything in a single language, or eventually write a couple of functions in a lower level language that are then used as building blocks for the rest of the application.
No need to throw away the ecosystem and developer tooling just to rewrite a data structure.
For example, you can do codecs in C# on WinRT with .NET Native,
https://docs.microsoft.com/en-us/windows/uwp/audio-video-cam...
In the context of protobuf,
https://devblogs.microsoft.com/aspnet/grpc-performance-impro...
Back to Rust, yes it is a good option, I just wouldn't write the whole application on it, just specialized libraries.
Hence why I am looking forward to Rust/Windows efforts.
I don't know how you get that from the thread:
> Arenas are, however, unfeasible to implement in Go because it is a garbage collected language.
> If you are willing to use cgo, google already implemented one for gapid.
> there are other garbage collected languages like D, Nim and C# that offer the language features to do arenas without having to touch any C code.
It seems like the above statements implicitly or explicitly claim that this isn't feasible in Go without C.
Both of us are dismissing the assertion that "Arenas are, however, unfeasible to implement in Go because it is a garbage collected language."
You can do manually memory allocation via a syscall into the host OS, use unsafe to cast memory blocks to the types that you want and then clean it all up with defer, assuming the arena is only usable inside a lexical region, otherwise extra care is needed to avoid leaks.
I _think_ allocating a slice of contiguous bytes and using unsafe pointers should work fine as long as you are very cautious about structs/vars with pointers into the buffer getting freed by the GC.
Go's GC is conservative, so I don't think you need to take any special caution in that regard. I would expect that you just need to take care that your casts are correct (e.g., that you aren't casting overlapping regions of memory as distinct objects).
Disclaimer: Google had a lot of internal stuff they considered important to their core tech competencies. For example, no open source about Google paxos APIs and infrastructure, networking, etc.
For example, it looks like pooled decoders could be implemented by setting a custom unmarshaller through the ProtoMethods[2] API.
I wonder why not? Did the authors of the vtprotobuf extension not want to bite off that much work? Is the new API not sufficient to do what they want (thus failing some of the goals expressed in golang/protobuf#364?
[1]: https://github.com/golang/protobuf/issues/364
[2]: https://pkg.go.dev/google.golang.org/protobuf@v1.26.0/reflec...
I'm struggling to understand what the rationale _for_ doing it is though. Maybe it's to avoid an import cycle?