The State of Caching in Go
blog.dgraph.io
blog.dgraph.io
The GitHub stars obsession is too much, though. I read the request in the main content area, then a nag banner popped up at the bottom of the page requesting it again. How annoyingly desperate and uncool.
Enough already, I'm trying to read!
Or was, anyway..
Plus I'm not even logged in to GitHub at present and cannot without getting to another device for 2FA.
No offense, but now I'll likely never give that star.
Clicking on the [x] in the corner of the pop up (to get rid of it) then launched a new window (well, new tab in this case).
Didn't read the rest of the post.
(i'm the person who implemented that widget)
With guidance from Ben Manes, we're planning to work on a new concurrent Go cache based on Caffeine (Java). If you're interested in helping (or already have something similar), do reach out to us!
Meanwhile, check out Dgraph, our distributed graph database which is what keeps us busy: https://github.com/dgraph-io/dgraph
If it were an on disk cache []byte would likely make sense but given it’s in memory I’m not sure. If you use interface{} you’ll need to measure that cost as well.
If you are just going to send that stuff to nic/disk eventually anyway its perfectly fine (which is why something like Redis which is an out of memory cache can get away with it).
When de-serialized, in-memory structs take up more memory on heap and it gets hard to compute that -- we've tried that in vain in the past.
Having said that, there's still a general need to store pointers as values and that's something we'll look at once this byte-based implementation is ready.
It is a completely lock free key value store, though does not have eviction built in.
Its memory management is done using a free list of blocks, which is part of how it maintains lock free concurrency.
> FreeCache and GroupCache reads are not lock-free and don’t scale after a point (20 concurrent accesses). (lower value is better on y axis)
If you look at the graphs, they level out and seem to work best with more concurrent load, the exact opposite of what they say. Am I missing something?
Also, I think it would be valuable to see higher concurrent requests: with what I'm used to, 60 concurrent requests would be low - I'm interested in maybe 1-5k concurrent requests.
In other words, 9 concurrent pregnant mothers would show 1 month/baby on those charts.
If I have that right, their comment makes more sense. After 20 concurrent requests, the overall throughput of operations/second does not increase with additional concurrency.
There is a cultural skepticism about the overuse of packages. There are occasions where the STD lib copied a function instead of adding an import for example. Like many things in Go this pushback is needed but sometimes goes too far.
But maybe single function packages in npm we're a bad idea.
Anyway caching is a mixed bag in my opinion. Local per-thread caching is often better than a global cache as it avoids contention. It's also trivial to implement with a map and requires no special coordination.
It does require rethinking how you design a solution to a problem. FWIW that design process tends to lead you down a better direction anyway for producing distributed systems.
For example one giant, randomly distributed Kafka topic with a global redis db for a cache is probably a lot worse off than a system with more predictable data locality on the consumers.
Libraries are hardened by repeated usage by many different users who each bring their own special use cases and improve them to bring them to production quality. Go has no platform where certain well-written libraries can be recommended and get more exposure. Thus, they don't tend to mature more than the specific use case they get written for, provided they are still being maintained.
Arch Linux is a great example of giving well-doing packages more exposure by upgrading an AUR to community to core/extra. That way, more users gather around packages improving them even further.
To the second point about per-thread caching, Go does not expose threads to end-users. So, there's no concept of thread-local. What you're describing results in lock striping, which has contention issues as described in the post.
Godoc.org is that platform, it serves the purpose. Others in the module universe are under development.
> To the second point about per-thread caching, Go does not expose threads to end-users. So, there's no concept of thread-local. What you're describing results in lock striping, which has contention issues as described in the post.
In Go you'd have to orchestrate this locality yourself, i.e. a fixed worker (goroutine) pool each with its own cache. This is probably less work than it sounds.
Generally, Go does encourage you to author solutions to your specific problem, rather than adapting a general-purpose library. This is almost as much a part of the ethos of Go as implicitly-satisfied interfaces, or "share memory by communicating", or any of the other proverbs. Maybe Go takes it too far. But effectively no other language exists at this point on the spectrum, and I'm happy that we have at least some representation over here; I think lots of programmers live here, too, and appreciate the tradeoffs.
Godoc.org is great. But (beyond documentation) it is at best a search engine, providing equal platform to all libraries however production ready or broken they might be. Unless I'm missing something, it does not intend to promote certain libraries over others, the same way as AUR -> community -> core works in Arch (which is the model I think is missing in Go).
> In Go you'd have to orchestrate this locality yourself, i.e. a fixed worker (goroutine) pool each with its own cache.
Having many small caches within the same process would result in more misses per key, which if it results in disk accesses would not be ideal or might be worse than contention.
Moreover, being able to spin Goroutines as and when required to branch off a big job into smaller tasks is the beauty and benefit of Go compared to other languages like C++ or Java, where you must start a thread pool upfront and shoot tasks off to it.
Arch packager here.
It doesn't quite work like that. The distinction between community and extra is largely based on who packages it, it used to be defined as above 5% usage. The distinction between that and core is that core is considered essential to the distribution. No package is going to go from AUR to core without replacing some integral part of the system.
The underlying issue is that building a production ready library doesn't just happen -- it needs one or few initial authors and (along with a large group of users) a group of dedicated power users who can then continue to maintain the library and optimize it for their usage, in turn developing it enough to be considered production ready.
This is all a bit abstract. There are plenty of go libraries which are suitably mature for production use. For example Google's cloud libraries or Amazon's AWS libraries. There's a platform for communication of issues and releases: Github. I guess I haven't really had an issue with it... The only problem I've had is the messiness of versioning, but there's been solutions to that for a long time (godep, govendor, glock, etc...)
> To the second point about per-thread caching, Go does not expose threads to end-users. So, there's no concept of thread-local. What you're describing results in lock striping, which has contention issues as described in the post.
Sorry for the confusion. I did not mean thread-local variables. I meant breaking up your workload across mutiple goroutines so that you don't have to use shared state.
For example suppose you receive a message with an org ID. You start 8 workers, all data for (id % 8 == 0) goes to worker 0, (id % 8 == 1) goes to worker 1, and so on.
Each worker maintains a local map of data and no mutex is needed to coordinate access, because only one goroutine can use the data at a time. So there's no lock-striping and no contention issues.
This is the recommended approach for Go concurrency:
> Do not communicate by sharing memory; instead, share memory by communicating.
I understand that this sort of approach doesn't work for every problem, there's not always an easy way to split up the data like this (you might end up with hotspots for example). But I feel like that conclusion needs to come after a bit of reflection, instead of reflexively reaching for a global object + mutex. (Not saying that's what you guys are doing... just a common issue I've seen working on distributed systems with developers new to Go)
When Google needs something, it gets built, corrected and optimized quickly (e.g. Go context library). Open source / crowd sourcing doesn't work that way.
The author did indicate that they had problems with oom and memory eviction so it probably is a memory limited use case.
Juggling caches on databases is a challenging thing, and there has been some back-and-forth on best practices. MySQL for the longest time shipped with a query cache. As of MySQL 5.7.20 it was deprecated, and has now been removed in MySQL 8, largely because it was as likely to hurt you badly as help you, particularly with correctly sized InnoDB buffering.
A database will often scan many records which would flush an LRU. Postgres uses small LRU buffer caches in your per-thread model, backed by a larger LRU-like cache. The buffer caches are easily flushed by scans, but protect the shared cache from this noise. That shared cache could probably benefit from a smarter policy and this is an on going topic.
Go programs are written in a synchronous, thread-per-request model. You'd end up with _tons_ of small, very cold caches and a miserable hit rate.
You could approximate the former in Go by just having an array of N caches and picking from them randomly. This is similar to their "lock striping" with less contention (no stripe is a hot spot) but a lower hit rate.
What is this? Google thinks it's some sort of metal stamp.