Apache Traffic Server
trafficserver.apache.org
trafficserver.apache.org
In his benchmarks he was running with a single volume configured for cache, which would have a global lock on cache. If he partitioned the cache into multiple volumes (something we do by default now) he would have had much lower cache hit response times.
The majority of our cache hit response times in production are less than 1 ms.
In the benchmarks I have run ATS has always been faster then Varnish and NGiNX. If they weren't I would have made changes to ATS to make it faster.
I'm assuming as a top 5 site you probably deal with a large library, and appear to be ok with ATS despite that. Comments?
The benefit is that writing is fast and it's constant time since you don't have to do an LRU lookup to pick a place to store the object. The downside is that you are creating cache misses unnecessarily.
It's never really been a problem for me in practice. If you have a lot of heartache over it, I would suggest putting a second cache tier in place. Very unlikely to strike out on both tiers.
Statistically, it balances out just fine. It turns out that just by controlling how objects get in to the cache, you can effect cache policy enough that eviction policies don't much matter, or at least, a "random out" isn't much different from a "LRU".
https://github.com/apache/trafficserver/tree/master/plugins/...
Cache insertion research is generally focused on use cases like small associative hardware caches. There's very little applicable public research for larger software caching systems, that Ive found. Probably the best would be Gil Einziger. He appears to have found it as an application of his work on extremely space/time efficient counting of sets, http://www.graduate.technion.ac.il/Theses/Abstracts.asp?Id=2... and http://www.cs.technion.ac.il/~gilga/. Of notable mention is TinyLFU http://www.cs.technion.ac.il/~gilga/TinyLFU_PDP2014.pdf. Gil submitted it to Caffeine (Java caching library) last summer, https://github.com/ben-manes/caffeine/pull/24. It got some traction over the winter and is now showing up in other places (https://issues.apache.org/jira/browse/CASSANDRA-10855). In fact Ben Manes just had a guest post on High Scalability the other day http://highscalability.com/blog/2016/1/25/design-of-a-modern....
PS: If anyone is interested in these problems, We're Hiring.
edit: https://aws.amazon.com/careers/ or preferably drop me a line to my profile email or my username "at amazon.com" for a totally informal chat (Im an IC, not manager nor recruiter nor sales)
mutli1 trace - W-TinyLFU: 55.6% - TinyLFU+Random: 53.8% - TinyLFU+FIFO: 48.6% - TinyLFU+LRU: 54.8% - FIFO: 41.0% - LRU: 46.7%
FIFO's poor choice of a victim has a big impact, so it doesn't seem promising. A random policy looks like an attractive fit, though.
We have looked at not evicting objects on disk if they are in the RAM cache and that has a LRU like eviction algorithm (really it is a CFLUS). Doing this would help in not evicting really popular objects.
FIFO has advantages over LRU for disks. It is very efficient with writes since they are all sequential. We use rotational disks when building out very large second tier caches.
There are other things to consider when looking at cache in a proxy server. How many bytes does the in memory index take per object in cache (for ATS 10 bytes and that is extremely efficient). Also, does the cache use the filesystem and/or use sendfile for HTTP (like NGiNX), but can't use sendfile when using HTTPS or HTTP/2. Netflix is experience this pain when moving to HTTPS with NGiNX.
Every proxy server some advantage, easy of use, well supported APIs, flexible configuration, dynamic loadable modules, HTTP specification compliance, HTTP/2 support, TLS support, performance, etc. It really depends on what you are looking for when choosing a proxy server.
Fwiw, it does support cache pinning, but that's rarely used nor necessary.
... the results indicated that Apache Traffic Server reached better cache hit rates and slightly better bandwidth throughput with the cost of higher system and network resource usage. Varnish on the other hand managed to response higher request rates with better response time, especially for the cache hits. The findings in this thesis indicates that Varnish seems to be more promising reverse proxy.
I've used both Squid, varnish and nginx plenty, but traffic server beat them in our benchmarks and the built in ssl termination + plugin api makes it extremely powerful...
[2]: https://www.iispeed.com/pagespeed/products/ats-pagespeed
You mean as a reverse proxy / cache server?. I don't have statistics but I would think that most people use apache as a regular http server (serving files or as part of a *AMP stack)
Edit: I was serious about that; I was under the impression that most stacks use much lighter weight http daemons than Apache these days. I understand that legacy apps are still out there and not everyone is going to refactor, but anyone developing web applications under Apache in 2016 is just a glutton for punishment...
I think people are being to hard on Apache, it's still a great webserver for a ton of applications. That being said I prefer to configure Nginx over Apache.
Change this to "People still use X?" where the value of X is pretty much any technology you've ever heard of, dating back to the 1970's (if not before). And the answer will, to a first approximation, always be "yes".
Now the number of people using X might be small, but you can all but bet your life that somebody, somewhere is, indeed, still using it. And depending on what it is, you might be surprised as how large the number actually is. Keep in mind, HN and Reddit, etc., comprise something of an echo chamber, where people of a certain mindset and orientation flock. The world is MUCH larger.
No, RPG on iSeries /AS400 machines isn't "cool" and you won't see it mentioned on HN much (if at all) but this stuff is still used all over the place. LAMP stacks? Yeah, still widely used. OS/2? Not exactly "widely" used, but still used. COBOL? Yep. Fortran? Yep. MVS? Yep. And so on. Now, granted, this stuff isn't used by "hip startups" or "unicorns", but the world is a lot bigger than the SV startup scene.
Technologies die VERY slowly for whatever reason, at least in regards to the "long tail" (so to speak) of the usage curve.
Most of it is used by business that literally manage your life. Banks, electrical and water, manufacturing plants, telephone exchanges, everything.
ATS scales orders of magnitude better than Apache, due to its process model. Whereas at Yahoo we would budget between 30-200 simultaneous connections per Apache server (prefork), the proxy service which I ran using ATS was budgeted for over 100,000 concurrent connections per machine.
It's significantly less featureful than Apache, but it does caching substantially better than any other cache server commonly available (nginx, apache, squid, varnish).
Again, the pricing wasn't the issue; it was the fact we would have had to go through procurement - which involves a few weeks/months of process at any decently large company. So we ended up using Apache because the team was familiar with it and knew it could support our use case. While Nginx probably would have performed better, in the end it was just easier to use Apache because it was a mature Open Source project and throw a few extra VMs at the caching tier to make up for the performance gap.
Seems to me that the above is all based on reflections from decades ago w/ Apache httpd 1.3. Right now, all web servers can handle similar levels of concurrency with the bottleneck being the network pipe itself.
ATS is a great platform; using it in combo w/ Apache httpd (2.4) allows a pure open source implementation with all the power, speed, reliability one could want, and protection against Open Core business models.
Nevertheless, I don't think that nginx / Apache / Varnish / haproxy / etc. are able to handle similar concurrent connection levels as ATS without significantly impacting 95th percentile latency due to their core architectures.
http://www.slideshare.net/bryan_call/choosing-a-proxy-server-apachecon-2014(I worked on a team closely related to Bryan's team at Yahoo)
Some features that I preach are:
- good turnkey default values
- lua support
- config options galore (Bryan labels it a con, but if you want control it's perfect)
- good logging
- historically proven scalability on large smp, xxlarge memory, multi nic systems
Edit: forgot to add one more thing though maybe not worthy a bullet point. If possible, a preference for physical rather than virtual is where I've seen performance with ATS shine. That is one reason why you would want as much config control possible.https://github.com/apache/trafficserver/tree/master/lib/atsc...
You can do stuff like:
> registerHook(HOOK_READ_REQUEST_HEADERS_PRE_REMAP);
> registerHook(HOOK_READ_REQUEST_HEADERS_POST_REMAP);
> registerHook(HOOK_SEND_REQUEST_HEADERS);
> if (transaction.getClientRequest().getUrl().getQuery().find("redirect=1") != string::npos) {
> cout << "Sending this guy to google." << endl;
> transaction.getClientResponse().getHeaders().append("Location", "http://www.google.com");
> transaction.getClientResponse().setStatusCode(HTTP_STATUS_MOVED_TEMPORARILY);
> transaction.getClientResponse().setReasonPhrase("Come Back Later");
> }
> transaction.resume();For one thing, not having to copy response data between processes improves throughput. Since Varnish is so resistant to supporting SSL natively, you'll always have to place something in front of it to use it with the modern web. Whether it's haproxy, Apache or nginx, that's just one more thing to deal with.
I have some other beefs with Varnish, but the most annoying one is the absence of a persistent disk cache. If the Varnish process dies, there goes your disk cache. Even though cache data is written out to disk, Varnish punted on saving an index and re-using an old process's cache, so it writes the cache to an unlinked file.
Imagine a bad code push or new traffic pattern that causes core dumps across your entire service footprint -- and now it isn't just a problem of getting the process back up and stable, you have also lost hundreds of terabytes of cache data. Or something as simple as rolling out a new version. You can architect around the problem, but why should you even have to?
ATS also (recently) supports Lua for plugins, which is way more powerful than VCL. It is a finicky piece of software though, and there are a lot more sharp edges that you're likely to cut yourself on during the initial honeymoon period versus Varnish.
I wish more companies would model themselves after Percona: charge for support, custom engineering and on-call -- don't fork or paywall any code.
ATS suffers by comparison, since there is no "ATS Inc." to provide support and engineering work. There's OmniTI, but I don't have first hand experience with their service to say if it's worthwhile or not. They did get paid to write the current ATS docs, so presumably they know what they're doing.
I wish ATS got more attention, but it is after all a bit of a niche product hidden away in the Apache Foundation with a bunch of unrelated Java projects. It's too fiddly for small scale use, and once you hit large scale you're pretty much hiring someone from Yahoo or elsewhere that has experience running and developing it (for example: I'd like to hire ATS people). Doesn't give it a lot of opportunity to trickle into smaller shops and grow with their service.
I am so tired of developer entitlement and complaining, and not wanting to pay for software. Most frustrating thing being a founder and engineer.
Now think about services that use nginx on every machine as a general purpose URL API interface. It's not uncommon, why bother re-inventing the HTTP server wheel. At a previous company, a service I ran would have cost $3.6 million a year in nginx plus licensing fees. Almost none of the added 'plus' features would have been at all useful.
If you see nginx plus as a way to pay for the core software then sure, maybe that cost is appropriate. I will not support them with per-server licensing of gated off features. Oh and by the way, nginx plus is closed source and is only available on a small handful of platforms, and doesn't always maintain the same release schedule as the open source version, all as a way of supporting their licensing scheme.
I would support nginx via professional services and support fees, but only in conjunction with the open source release. So it's up to them if they want that money or not.
For instance, the last nginx plus release (R8) is based on nginx 1.9.9 (plus was released 40 days later). The previous is based on nginx 1.9.4 (plus was released 25 days later). It isn't a matter of life and death, but it is an annoyance and unnecessary.
At least with Varnish Plus, it's still the same open source server with proprietary modules added in. I understand that nginx wants to go that route eventually, but they're hamstrung by the lack of dynamic module loading for the moment.
It took us some digging and work to get configured exactly right, especially since we were using it fairly nonstandard - as a caching forward proxy to external data sources.
Good stuff.
I never use it, but from what I heard, it performs & scale really well too.
The reason to do so is because you want the rest of what nginx provides, not to get its caching module. It is an extremely barebones solution that only solves the most basic requirements. I can only presume that CloudFlare and others have written their own caching modules for nginx.
A short list of annoyances:
1. No support for multiple disk devices. Files are written to a fixed temp path and then renamed to their real destination. So you need to use RAID to present the disks as one logical device, which is a wholly unnecessary expense in a caching environment.
2. No support for purging in open source. This is an nginx plus feature, which starts at $1900/server/year.
3. Because of the temp file / rename thing, support for streaming subsequent requests off of the first request that is filling cache is janky. Subsequent requests have to acquire a lock.
4. No support for any fancier cache setups, utilizing ICP / HTCP.
VCL in particular is kind of a trap. Early on it can do what you need -- remove a header, set a header, basic branching. Then you want to do basic arithmetic, or validity checking, or anything that isn't suitable for string assignment or regex and you straight up can't do it. VCL makes me long for the power of bash scripts.
Ultimately though, it's really not a great solution to separate these concerns between multiple applications. You're going to get bitten somewhere, even if it's just the old ephemeral port exhaustion problem.
[0]: https://github.com/pintsized/ledge
[1]: https://github.com/wandenberg/nginx-selective-cache-purge-mo...
Storing externally in redis (for ledge) seems like the wrong approach to me. Better to store metadata externally and generate the purge URLs based on that. It's not ideal, but it's the best option I've come up with.
And right or wrong, there are not that many implementations to choose from in the Nginx land, unfortunately.
Initially created by Inktomi, bought by Yahoo! and got open sourced and brought to Apache Foundation in 2009 because of the good experience of Yahoo! with Hadoop in 2008/09.