Why Loggly Chose AWS Route 53 Over Elastic Load Balancing
loggly.com
loggly.com
Formally they had all of their EC2 instances configured to run without swap and didn't use EBS such that instances would crash 1-3 times a day and lose all data which would require 1-2 day customer restores of data.
Additionally, this Java shop oversubscribed threads on every Solr box which made them restart each Solr instance every hour. To think any revolutionary engineering ideas come from an former Apple marketing wannabee who puts outsourced Indian engineering in place as yes men is a huge stretch.
Let's be honest, Loggly is in huge trouble and can't hire quality engineering talent and as a result is trying to remarket themselves as an engineering driven company as they outsource to India.
Key question isn't..do you use DNS or Elastic Load Balance...it is...what is your VOLUNTARY RATE OF ATTRITION? Hint, really bad!
Then reading this fluff piece made me glad I never even thought about working there after that phone interview. Whoever claims a DNS round robin is a good way to handle fail over doesn't really know what they are talking about. I have dug into the how something like rsyslog handles a dns request. My guess is it just passes it off to the OS.
But what I got from this is loggly is ok with losing customer data
This is unrelated to the loggly bit, it would just be interesting to know for a non-ops-guy.
This is a reflection of the social and communication issues in typical outsourcing setups, not a reflection of the talent of the outsourced team.
A logging platform that lists 1 of their 2 major requirements as "To not drop any data, ever" is using round robin DNS for fault tolerance? I can't see too many people on HN upvoting this for being insightful or impressive.
Edit: I just can't help myself. How are you going to send syslog when any server fails and not "drop any data, ever"? Even over TCP the in transit messages are lost when the connection is broken. So like, their business is basically syslog and they don't know that?
EC2 instances running haproxy would mitigate a number of the problems they discussed with using ELBs but the inability to use VIPs (with vrrp or ucarp) in AWS means that a failure will always boil down to the same pattern: a key front end instance dies, client traffic keeps being directed to it for 5 minutes (at best), and that's life.
TL;DR Loggly can't promise no data lose in its current incarnation.
Except when for example rsyslog caches DNS resolution forever. Or the log forwarded doesn't have a buffer and logs get lost.
EDIT: To reply to myself http://docs.aws.amazon.com/Route53/latest/DeveloperGuide/hea... http://docs.aws.amazon.com/Route53/latest/APIReference/API_C...
Complementary to that it's also possible to assign reserved IPs to machines so that at least the set of IP is always the same even if the hosts get rotated. Assigning IPs to hosts is not instantaneous either and depends on an API call. Also IPs are tied to a specific region.
In a real datacenter other options are possible like sharing IP addresses between multiple devices and having lower-latency failover. These aren't perfect either and have different failure scenarios.
That chances of failure go up dramatically if 2+ hosts behind round robin fail, etc.
Not to mention once hosts resolve this to an IP they will re-use the route. This approach is not balanced.
I don't want to be /that/ guy but if they can't scale with ELB they should invest in a dedicated load balancer infrastructure that can offload requests to their cloud instances.
This is a really bizarre post.
The 2+ hosts failing should not be much of a problem if you have a separate health checker host which does nothing except gathering heart-beats from all the hosts in your fleet and updating the DNS periodically.
If you have 3 hosts and 2 of them go down, in this setup there is more than 50% chance that cached hosts will be trying to connect to a non-existing server.
Also expecting client to perform a DNS lookup every time there is an outgoing log packet is pretty shitty for performance. You can't guarantee near instant DNS server availability for every client.
In their docs, Loggly only gives out one API endpoint: logs-01.loggly.com.
It is referenced as the endpoint for HTTP, HTTPS, syslog and syslog TLS. These seem to be the only methods available to send log data to them.
There is the obvious problem that a DNS record with a 60s TTL cannot possibly receive every single packet sent to it in the event of a server failure. Even if the returned IP address is an elastic IP, it takes a substantial amount of time to move to another instance in AWS.
I don't know why you would use the same service hostname for all of these endpoints. Separate names for each endpoint, even if they all pointed to the same pool of hosts, would at least give some flexibility in the future when they have enough traffic to get desperate about capacity. I would also think they might want to segregate native syslog from HTTP traffic, since I presume it uses different processes on the backend.
It's also curious that they chose to return only one A record. DNS RR is a poor substitute for real load balancing, but it's better than nothing. With multiple A records, there is at least a chance that some of their traffic will go to other servers -- rather than all of it potentially going to one as it is now.
While they made no claims about using Route 53 for its geo DNS capabilities, I still found it amusing that I was sent to a US East IP from California. Not that it's super critical that my log lines get delivered quickly, but it is ideal to shorten the path of an insecure and unreliable transport in order to improve durability. Although I would never ship syslog out to some host on the Internet, a host 16 hops away is even more ludicrous.
I think their article says a lot more about how poorly ELBs function when you exceed the low traffic threshold it is seemingly designed for than about how well Route 53 works (and it is a decent static DNS service). The inability to robustly direct incoming traffic is the achilles heel of AWS.
I'm actually in the middle of deciding between Loggly, Papertrail, and Logentries for centralized log management. I guess that cuts it down to two.
The transmission type is still up to the client/receiver.
Background and setup: http://help.papertrailapp.com/kb/configuration/java-logback-...
GitHub repo: https://github.com/papertrail/logback-syslog4j
Papertrail also works with the standard Logback SyslogAppender.
I recall having some trouble with that when using syslog with Loggly, before switching over to the json appender.
Long answer: logback and both appenders can accept pattern formats to adjust how they're formatted. How useful the end result is depends a lot on the receiver, though, and more than that, there's no one implementation that's great for everyone -- that is, there's no right way to "handle large, multi-line log messages," only attempts at making them more useful.
An easy example is searching. Some people want to see the entire message, others want only the matching portion of a stack trace, others want some combination, and others - probably most people - just want something that's useful, however the actual UX works.
In Papertrail's case, our sender-specific context links (think grep -A/-B/-C) were designed for navigating multiline output from a single sender: https://papertrailapp.com/tour/viewer/context. It's basically pivoting from a single entry in a stack trace to the entire stack trace.
The setup is straightforward [0], and you'll find that the plans are very competitive.
[0] https://www.scalyr.com/help/install-agent
(Full disclosure: I work with them.)
* HTTPS termination.
* Autoscaling group management. By connecting an ELB to an autoscaling group, the logic of registration and deregistration is fully managed behind the scenes. With route53, you have to implement it yourself.
* Minimum autoscaling group size. If you enable ELB health checks, you can rely on the ELB to maintain a group of instances of constant size.
Not to mention it still needlessly includes a ton of dangerously insecure ciphers just begging to be misclicked.
[1] http://docs.aws.amazon.com/ElasticLoadBalancing/latest/Devel...
[2] https://wiki.mozilla.org/Security/Server_Side_TLS#Amazon_Web...
[3] http://mir.aculo.us/2014/04/04/how-to-get-an-a-on-the-qualsy...
Even with a short TTL, are there still servers out there that don't respect all TTLs, or has that been eliminated by now?
>If you’ve ever used the Internet, you’ve used the Domain Name System, or DNS, weather you realize it or not.
Interesting article, wrong weather used in this sentence.
Route 53 Health Check: either 10s or 30s