https://github.com/protocolbuffers/protobuf/commit/50f9ac3ba...
On a sufficiently anemic CPU doing enough fast allocation you can get measurable improvements from that kind of thing, but for a superscalar CPU and a general purpose allocator I agree you'd be hard pressed.