Go-benchmark: Golang benchmarks used for optimizing code
github.com
github.com
A lot of Go programs spend a lot of time converting between byte[] and string. But the conversion is actually really slow and keeping your string data as byte arrays is much faster and since many go libraries can work with either, the conversation is not necessary.
Also, memory allocations and garbage collection are not free. If you have a hot path that gets hit every request or in a loop, it can save a lot of time reusing the same slices assuming you can do so without introducing bugs and/or multi-threading issues.
I've got a project where strings aren't exactly crucial, but we're storing a ton of them in memory. We store them as string and, frankly, rarely access most of them - though we do do some compares of strings on the hot path.
Using []byte sounds interesting, but troublesome at the same time. Namely the fact that byte is useless for reading characters from, as you'd eventually have to get runes from it anyway no?
Thoughts?
[1]: https://play.golang.org/p/5DEzw85J5Ob Note this is a conversion; the string is UTF-8 and the runes are 32-bit ints, so this creates a new string.
Tons of string manipulation just needs to split, etc on \n, ., ;, - which don't need runes.
Besides that, string comparisons don't need runes either (assuming what you're comparing is normalized to the same bytes, which if you do the same input processing to anything you store and to the strings you query on, it would be).
That hash map looks good and I'm thinking we'd probably benefit from using it on the hot path of some code we have that needs to be highly scalable.
(Both links were submitted to HN, but only this one seems to have landed on the front page.)
>>BenchmarkAtomicInt64-8 5000 354907 ns/op
that ns/op looks too big to be only for an int64 ..
BenchMarkSize = 1 << 10 // 1024
BenchMarkSizeLong = BenchMarkSize << 5I'm not sure what the idea is there, the benchmark system can accommodate small operations, I've used it before on things like single increments or single interface calls and gotten reasonable answers.
Also benchmarking atomic int increment without any contention is not necessarily "useless", but certainly not a full picture if one is investigating using it in a contended data structure. (My usage of it is mostly just counters, where I expect thousands upon thousands of instructions to be between the increment of any given one of them, meaning there's probably no contention to speak of even on production multi-core servers, so I have no idea how they perform under serious contention. But AFAIK it's just the hardware-supported atomic instructions, so what I don't know about is the hardware, not really Go per se.)
Yeah, I've been wondering if I'm missing something, because adding a second redundant loop to b.N is common in go repos, but seems pointless and needlessly obscuring.
Benchmarks can measure execution time of algorithms or if not careful, the ability of the ability of the compiler to optimize away unnecessary code.
(IMO the 'performance comparison' that was vastly in favor of Rust was a bit disingenuous.)
Use defer to make code readable. Go code should be readable, so you can figure out if it is correct, and only then should it be fast.
When benchmarking/profiling leads you to a hot spot in the code, rewrite it to remove defers.
This is typical: https://github.com/cornelk/go-benchmark/blob/master/defer_te...
Yeah, it's in the class of "things you should know about when optimizing tight loops", but not "things you should always be worrying about"; remember 5ns is on the order of a branch mispredict or an L2 cache miss [1]. If you haven't already squeezed out all the main memory accesses you can from your algorithm at 100ns a pop, or if you're in code dealing with networks (routinely micro seconds, even within datacenters; milli seconds if you have to leave), or files, or anything else like that, micro-optimizing defers is not going to have any visible results to your speed but can very badly damage your code's correctness and ease of writing and modifying.
The other thing to know about defers is that they are not declarations, they are instructions; every time the program counter moves past one, another defer call is added to the stack of calls to make while exiting the function. So you do have to be careful about deferring in a loop, not so much because it's "dangerous" but just because it's easy to be unaware that the defers are not declarations. If you intend to do it, it's fine, it can just bite you if you do it accidentally.
Don't use defer when there is only one branch in your code.
Remove defer when optimising hot code paths.
Code will get changed over time and one code path often becomes two or more. Chances are, the next developer who adds code will not realize that the resource needs to be freed, or closed and simply forgets to add the defer. The only thing where defer really gets into your way is for error handling (because it can get verbose if you handle errors in defer).
In the grand scheme of things, defer is not noticeablily slower.
What good is fast code if it is unmaintainable?
Depending on industry, time is relative. These optimizations will save you milliseconds on tight loops. In reality most apps aren't spending their time on tight loops. They are spending their time waiting for I/O or using inefficient algorithms. Neither of which these optimizations solve for.
If you are underperforming spec by 100ms it is likely you have one big change you need to make to save 100ms not 100 small micro-optimizations.
Only because most people don't have to do this there are still some people that have to do optimzation where possible.
I hope you understand that optimization without profiling is worthless. Once profiling is identifying something like your datastructure as some kind of bottleneck it might be worth a shot investigating that hint.
And these kind of articles are a nice summary of work being done, giving an overview over techniques and small tweaks that you might not have thought about previously.
It is good work by the author and he deserves being recognized for it. Not being dismissed by the bland statement don't waste your time with optimization.
The author's use case is very beneficial to optimize in this manner.
At work I just had to implement a heap just because the heap provided by the standard libs wasn't fitting our problem.
Please don't discourage people creating this kind of content. It matters to far more people than you might have in mind.