Is this typical?
Is this typical?
The Rust code and the blog post, on the other hand, seem to be written by someone less familiar with Rust and high-performance parsing. I think they would have avoided all their problems if they just used lifetimes to safely avoid copying from the start, instead of relying on increasingly elaborate workarounds like the `Bytes` crate. Apart from one "forces one to deal with the contagion of lifetimes" comment in the conclusion, they never mention why they didn't do this, even though it's clearly the idiomatic Rust solution. Maybe they had technical reasons for not doing lifetimes, but to me it just seems like unfamiliarity with Rust.
This one is actually out of the box in Golang, it's called sync.Pool and is accessible in the standard library. It's very easy to use and not error prone, I've used it many times without any issues.
But creator of VictoriaMetrics is indeed someone very knowledgeable and known in Golang community in context of optimization.
I guess that must be the logical consequence of the “async function are contagious” meme… I wonder if at some point we end up with people arguing that dynamic typing is obviously better because it avoids “contagion with types”.
For a language designed for fresh-out-of-college engineer to pick up in a few weeks and be effective, it is very easy to squeeze out a lot of performance.
* built-in profiler. * built-in escape analysis tool. * It's easy to pass pointers instead of copying data. * []byte is sub-slice-able, with a backing array. This does throw people off occasionally, but the trade off is performance. * Go lets you have real arrays of structs, optimizing cpu caches. * Built-in memory pools
And more.
And if you look at "non-idiomatic" performance code, they are surprisingly legible by the said fresh engineer. It's as if the designers didn't want to give up all the usual C performance tricks while making a Java/Python kind of friendly language, and this shows.
Of course Go can go only so far, due to the built-in runtime and GC. But it gets very far. Much farther than at first glance, or second glances that language snobs would give credit for.
Go has a good perf story, but typically rust or c++ would be faster after heavy optimisation; and should be more or less on par with typical applications. This isn’t a critique of go, and shouldn’t surprise anyone.
Typically go also has unexpected optimisation hoops to jump through and problems related to the heavy use of channels (see the well documented answer here: https://stackoverflow.com/questions/47312029/when-should-you...), so you would generally expect it to be slower…
…but, naive implementations are always slower, and really, it’s probably much of a muchness out the box for most day to day uses.
In almost all situations (even python or Java) you can get great performance if you invest time and effort in it.
But idiomatic code typically faster than rust? No, not really.
This is based on my first hand experience, but YMMV.
Unfortunely, it seems stuck in Go 1.18, the last pre-generics version, with no roadmap for moving forward.
Given Go's folks stance on generics, still a lot of Go code is compilable with gccgo.
it is concerning how much digging was required to optimize rust code in this case.
I don't see why did you label Aliaksandr Valialkin, the author, an "expert". I mean, he's no dummy but what exactly makes him an expert on optimizing Go code?
As someone who also writes Go, I don't see any "hyper optimizations" in the code. It just decodes the bytes of Protocol Buffer using straightforward code that I would expect a competent developer to write.
It really is just: read bytes from memory ([]byte) and interpret them according to PB spec.
There's only one trick there: unsafeBytesToString() that does no-allocation conversion of []byte to string. This is unsafe in general but safe in their specific case. And I've seen this trick before so it's not some secret, expert-only knowledge.
Most comments here are like bad LLMs: hallucinating opinions without bothering to spent even few minutes acquiring the data to base those opinions on.
That said, I agree with your assessment about this particular code. It’s fairly straightforward idiomatic go.
I was trying to convey the meaning of "far more experienced than the blog post authors", but without having to insult the authors. It's a good writeup after all, and I'm glad they took the time.
We must have some different interpretations of what "optimized" means. This is the very first piece of code in the file you linked:
func (fc *FieldContext) NextField(src []byte) ([]byte, error) {
if len(src) >= 2 {
n := uint16(src[0])<<8 | uint16(src[1])
if (n&0x8080 == 0) && (n&0x0700 == (uint16(wireTypeLen) << 8)) {
// Fast path - read message with the length smaller than 0x80 bytes.
msgLen := int(n & 0xff)
src = src[2:]
if len(src) < msgLen {
return src, fmt.Errorf("cannot read field for from %d bytes; need at least %d bytes", len(src), msgLen)
}
fc.FieldNum = uint32(n >> (8 + 3))
fc.wireType = wireTypeLen
fc.data = src[:msgLen]
src = src[msgLen:]
return src, nil
}
}
// ... function continues beyond this point
As far as I can tell, this entire codepath exists solely as an optimization. I spent many years working on a chess engine for fun, so I'm pretty well versed in bit twiddling, but I'm seriously struggling with this. Like, is it doing `(n&0x8080 == 0)` to check to whether length is less than 0x80? Is that even correct?I think "hyper optimized" is a completely fair characterization. But we clearly work in different industries.
If 0x8080 is not set, then the tag-value record is 2 bytes. Left byte has tag. Right is value. Then they're masking with 0x0700 to get the type of record, which should be LEN.
So if it's a single byte LEN record, they can take that single byte as the length (they mask with 0x00ff, but really it's 0x007f. They already know the 0x80 bit is zero, and the value is contained in the least significant 7 bits). Otherwise they have to do some fiddly logic to decode the variable length integer to figure out the length (length here being the L in TLV).
https://victoriametrics.com/team/ - Let's see. Author of multiple performance-optimized libraries with a masters degree in computer software engineering and a background in highly scalable systems (as needed for adtech). Sounds pretty much like an expert for optimizing code to me.
> And I've seen this trick before so it's not some secret, expert-only knowledge.
So, if you know it it's not export knowledge or what's your argument here?
Essentially it's a comparison between reference counting and tracing garbage collection based on reacheability.
Worded differently: does the strict concept of ownership/lifestimes in Rust bias a default (naive) implementation towards lower performance (eg due to required copying) when compared to a naive Golang (or even Java) implementation?
I have no doubts that after heavy optimization, Rust beats languages such as Go & Java.
When I learned Rust, I actually never went with the "use clone or Arc to make your life easier while learning" recommendation but always used references and learned how to use lifetime declarations and program design to go as far with them as reasonable. TBF I had experience with C and C++ already. But once reasonably experienced working in Rust (after a year?), your code should be faster most of the time the way you write it on the first try without needing optimization work.