Introducing Varnish Massive Storage Engine
varnish-software.com
varnish-software.com
If/when the kernel decides it needs to use RAM for something else, the page will get written to the backing file and the RAM page reused elsewhere.
When Varnish next time refers to the virtual memory, the operating system will find a RAM page, possibly freeing one, and read the contents in from the backing file.
And that's it. "
https://www.varnish-cache.org/trac/wiki/ArchitectNotes
I'll try and hold back the snark but I find it interesting that after attacking '1975 programming' and Squid's deficiencies, here we are 8 years later and maybe the kernel doesn't know best.
So it looks like they gave up on the "kernel knows best" approach quite a while ago. But then they show a graph, where the older mmap() approach is more than twice as fast as the malloc() approach across the entire range of the graph. They explain this in a sentence below the graph saying "Malloc suffers quite a bit here as the swap performance on Linux is rather abysmally bad. " Well, yes. But presumably that is the case we're interested in, as a cache that has room for everything is a much easier problem.
So what are we to make of the earlier contention that malloc() trumps mmap()?
However, the quality of the sysadmins that prefer FreeBSD is pretty high, though. And FreeBSD is pretty neat.
Physical memory is RAM memory. Their initial observation was:
1) If you need less memory than you have RAM installed - use malloc. Faster than file-backed mmap().
2) If you need more memory than you have RAM installed - use file-backed mmap(). Faster than malloc() + swap.
Edit: you can even see this on the graph in the article. Around time 0 it is malloc that is faster than file, then malloc sinks and file gets faster.
Varnish (w/MSE) still relies on the kernel for this. However, as you cannot atomically write a whole page through a mmap (thereby avoiding the page fault) we have to use write(). It still is read through mmap and the kernel still drop pages whenever it feels like it.
No, it isn't. That's very much the point. They have spent all this time realizing that doing it wrong is bad and rewriting it to actually work.
>do you have any data to prove this?
No, that was 2 jobs ago. I'm sure you can get a hold of the source code of old varnish and test it out yourself.
Troll much? Nope. That is not accurate at all. Varnish, with it's current storage engines manages to power sites such as Wikipedia and NYT. This announcement was not about that, it was about enabling caching of really large datasets. Datasets that run in the hundreds of terabytes.
And we did tens of migrations from Squid -> Varnish. Usually it involved scaling down the number of servers to 1/6 whilst seeing a massive decrease in response time. So it kind of surprising that your findings are completely different and somewhat disappointing that you cannot back them up.
1) Compiling config file into an .so object.
2) Using innovative priority queue (heap) implementation: http://queue.acm.org/detail.cfm?id=1814327
Varnish userbase is the best proof for its quality though.
That is neither new, nor interesting "engineering". Lots of software that needs significant customization uses the host programming language for configuration.
>Using innovative priority queue (heap) implementation
How is "we used an existing data structure and deliberately misrepresent this as some amazing discovery" an amazing feat of engineering? This supports the notion that varnish's biggest achievement is an amazing feat of marketing.
About the alternative heap implementation: can you point out who else has described this aggregated-heap data structure? I don't even know if it has any name, but sure as hell it is a good idea.
> Taking advantage of the fact that Varnish is a cache and letting it actually eliminate objects that are blocking new allocations simplifies the allocation process.
The page fault is a synchronous read. Horrible for performance.
I agree on the fact that using mmap() reads vs read() reads leaves the kernel with no possibility to reorder requests, as any thread trapped in a pagefault is obviously unable to generate more I/O requests. Reordering can lead to better performance. This is not the case with asynchronous read().
I've seen this mentioned in the context of RocksDB; but contradicted by e.g. SQLite. The case for mmap has always been that one avoids the overhead of a system call & some double-copying, and in either case it just dirties the page cache and is only "really" written in periodic flushes (assuming it's not writing via direct IO). Can someone explain what the bottleneck is on the mmap side and why write() might be faster?
mmaps are still great as the page cache still manages the split cache between disk and memory.
[1]: http://linux.die.net/man/2/posix_fadvise [2]: http://linux.die.net/man/2/madvise
It's also pretty common for people to enable file system caching plugins in Wordpress, Drupal, etc. and forget to leave the system enough RAM to actually do its business. D'oh.
Really good performance needs doing things differently, not the same thing faster. Yet most organisations don't want to try something different or give the space to try it.
Varnish was built for caching web apps. Squid is a forward proxy that can be configured to work as a web app caching program. So, when Varnish was designed we where able to disregard a lot of stuff that isn't needed when caching in reverse mode. On the other hand, squid has been around for ages and is a very, very mature product with a very well known set of strengths and weaknesses. Varnish is only 5 years old.
Personally, I have more than enough RAM now and I no longer need virtual memory. I do not need a "disk" or other secondary storage in order to retrieve and consume data.
I consider virtual memory a relic from an earlier era of limited computing resources, like "user accounts" designed for an era of time-limited use of prohibitively expensive, shared computers.
We now all have our own _personal_ computers and GB's of RAM, but we still have use kernels with builtin solutions designed to address issues of scarce, expensive, shared computers and scarce, expensive RAM.
If it is difficult, then why is this the case? Are languages lacking in this respect?
C++ is especially appealing, because you can embed file-backed mmap allocations in chosen classes by overriding their new operator. So you could for example create a float array class that automagically allocates itself in a file-backed mmap region by a simple new Array() call.
Edit: Python's numpy has a file-backed array: http://docs.scipy.org/doc/numpy/reference/generated/numpy.me...
thus, every pointer dereference needs a base address. how is the support for this?
.so files are loaded through mmap, and the pointers problem is solved there through the mechanism of relocations. But please don't write raw memory objects to disk, use Google Protobuf or ASN.1
For example, try storing a C++ std::map inside an mmapped file. Practically not possible AFAIK.
I really hate it when companies feel they cannot be transparent about pricing, it is such an obvious strategy to work out how much they can squeeze out of each potential customer. How are you meant to trust a company like that?
I don't remember the prices offhand but it was $57K US for a 3 node license at platinum support level and around $14k per node after that. He did provide that over the phone (but as I check my email, not in written form) on our first call.
I can't find it now, so my comment doesn't particularly add much value, but hopefully someone else can provide the link.