Another reason why Docker containers may be slow
hackernoon.com
hackernoon.com
I asked the engineer in question to investigate, but after looking he said, "It's nothing I could be doing."
So I sat with him and used git-bisect to prove to him it was his commit: he had added trace logging within a couple of tight loops in the hottest parts of the code base. I smiled.
"But it's trace. That's disabled in production. It can't be that," he said. But we had already proven it was that commit, and the only thing that changed was additional logging.
Long story short, the logging library was filtering calls by level just before actually writing, rather than as close as possible to the call site—a design bug, for sure.
I had him swap out the library everywhere it was being used.
Moral: logging is not free.
It's not easy problem to solve. The best I've come up with so far, is using a decorator in conjunction with a DI framework that supports it. It's still only a wrap around method calls, but it enables completely turning off logging in production, and enable it when it is needed without making a special build, or wasting resources during normal execution.
The compile-time cost of stripping out preprocessor macros disabled via #ifdef is close enough to free you'd have a really hard time measuring it - O(n) on number of lines of logging code, with n smaller than the number of lines of actual code for any real scenario, and a very small multiplier.
The "cost" there is in the visual appearance of the source code.
I can see how this is even true for a sniffing hardware bridge but isn't it free when your source is fiber?
(Solaris lives on in Illumos et al)
cgroups, Jails, and Zones all suffer from having to cover an immense surface area. Contrast with VMs, which only require managing a few, significantly simpler interfaces. There are definitely differences in quality, but they all use a similar approach.
Oracle used excuse, that it's current OpenJDK license (GPLv2) is incompatible with license, used by Google's runtime (Apache 2). If Google re-licensed it's Java implementation under GPL, some of arguments, used by Oracle lawyers (code reuse and patent (?) violations), would have been void, and arguing about reuse of APIs would have been a lot harder.
Of course, this does not really matter, because the whole lawsuit is just excuse for power games between corporations. Oracle's goal wasn't about Java licensing, it was about gaining some degree of control over emerging Android ecosystem.
If you are itching for some slides, http://bhyvecon.org/bhyvecon2018-Gwydir.pdf
We have some work to do before we are ready for widespread use. As those pieces fall in place, we'll announce on smartos-discuss@lists.smartos.org.
Their whole premise is that their containers are equivalent to/better than VMs[0].
0: https://www.joyent.com/blog/understanding-triton-containers
I am guessing it is to avoid to learn a whole new ecosystem, tools, environments, rules, package system etc. It's just simpler to stick to Linux. But it all depends, let's say if my application on Illumos shows a 60% performance improvement, well I can see spending time learning it and using it as a base. But it would really have to be large benefit to justify switching OSes. Of course how would I even bother benchmarking to start with? I'd probably have to hear other stories or mentions on HN and such...
For instance, the directory name lookup cache can greatly impact the performance of file operations. This is what makes it so that when you open /a/b/c/d/e/blah, the OS has a good chance of knowing how to open "blah" directly without first searching /a, /a/b, /a/b/c, /a/b/c/d, and /a/b/c/d/e. I don't have performance numbers handy (left at $job - 1 and $job - 2), but the default size is not great for a system that handles hundreds of thousands of files. If that sounds like a lot of files, count how many files are read as part of a reasonably large build. Then imagine there are dozens of them running concurrently. Or imagine that you are hosting git repos or a bunch of static web content.
The obvious answer is to just increase the size of the DNLC. The problem is that there are other parts of the system that behave poorly with a large DNLC. For instance, whenever anyone tries to unmount a file system, dnlc_purge_vfsp() is called. This walks the cache, looking for entries that are associated with the file system. Who cares, right? We hardly ever unmount file systems. Well, if you use the automounter (by default Solaris uses it for at least /home/*), every few minutes it is trying to unmount all of the automounted file systems. The purge happens before EBUSY can be detected so hot cached entries may be purged while adding contention on dnlc-related locks. What's worse, the automounter doesn't reset its inactivity timer when it hits EBUSY, causing more frequent attempts to unmount than you would otherwise expect. There are other DNLC bottlenecks on a Solaris NFS server when a client removes a file.
And then as you get to larger systems with more NUMA effects, these type of operations become even more expensive.
Operating systems are hard. Making them scale to an infinite number of processes and processors is impossible. At a certain point, it becomes beneficial to use smaller hardware and/or add VMs into the mix.
I'm not entirely clear as to why a logging library needs to call fadvise; a log file, is, I presume opened in append-only mode. Isn't "append" sufficient advice to the kernel? Also, fadvise needs byte ranges, and I have no idea what you'd pass for a log file…
More importantly, if the disk can not catch up, the log data is going to end up waiting in page cache anyway (typical case of bufferbloat). Linux kernel does not have telepathic abilities to balance needs of crazy logger and other applications in system, so without resolving underlying issue (bufferbloat), those writes would take up too much cache, potentially bringing down disk performance of other applications.
fadvise() may schedule quicker eviction, effectively acting as syscall version of vm.dirty_ratio. Of cause, that does not resolve the problem, — just moves it to different layer. The real solution is either
1) blocking the apps until their logs are fully written (for example, by using O_DIRECT)
2) showing those apps middle finger and throwing away some of their logs (AFAIK, this is occasionally done by syslog).
This specific bug was caused by putting high load on "kernel dentry cache", e.g. a contention for memory structure, present in kernel memory. Guests normally don't share memory, so contending for it was avoided.
Incidentally, there are situations, when different guests can compete for same memory — when VM uses so-called "memory deduplication" techniques. Which is why enabling that stuff on production systems may be a bad idea.
> I agree that piping all logs to stdout would be the best solution in case of Dockerized microservices. It’s just that in our case we were porting an existing system, which a) already heavily relied on logging to files b) consisted of many microservices itself, which we couldn’t yet split into separate Docker containers but also couldn’t pipe all their logs to the same stdout.
but using Dinghy greatly helped sped everything up due to it using nfs. just in case anyone wanted to know.
So itnisnsafe to define these in any comoose file no matter where it runs, just need to make sure that the app in the container can actually deal with the consistency level applied to the particular mount.
+1 perf. -1 stale library use. -1 misdirected learning