When your IP traffic in AWS disappears into a black hole
engineering.clever.com
engineering.clever.com
The solution for us was to set this in sysctl.conf:
net.ipv4.neigh.default.gc_thresh1=0
https://bugs.launchpad.net/ubuntu/+source/linux/+bug/1331150... https://bugs.launchpad.net/ubuntu/+source/linux/+bug/1331150...
Love the technical write-up -- thanks!
More modern and highly specific technologies like HAproxy, Salt, Node, Redis are not going to be known by every experienced, competent applicant though, and while you can often pick them up on the job if you have general experience in the area, there's always going to be a learning curve.
Instead we try to maintain a SOA which is agnostic to what underlying tools back our services, something that we've found Docker is useful for.
That being said, we generally just use Ubuntu, and have some other AWS-specific setups around networking. If you like hacking on Cisco routers, I think it's perhaps not the right position :)
Seems they are more dev then ops for devops role.
He ended up disabling arp caching though, rather than clearing it.
Yeah, that seems like a really bad idea. That will drastically increase the latency and cut throughput.
Nice write-up. Too bad I'm not in SF. :P
1. something silly like: "arp -a | awk '{print $2}' | xargs -n 1 arp -d". It's been a while and I don't think I was the original culprit.
# arp -s 192.168.0.1 00:18:4d:f8:a4:6e
Wait your working over SSH? Oh ummmm, talk to the data center guy and have him break out his USB to serial dongle and a RS232 cable.
The intermittent connectivity to server A is what makes this problem so fun to diagnose. :-)
As part of the cloud-init scripts on new machines we would install arping and have the newly provisioned VM automatically send gratuitous arps so that arp caches would get updated.
I haven't seen the issue on my RHEL 7 VM's though, nor in my current environment, but we don't spin up/down as many VM's as we used to.
Just today: Servers with a static IP address getting it changed to an APIPA just because some other 3rd party device sent a unicast ARP request to it that should've been a broadcast as per RFC4436. I personally find ARP to be one of the funniest protocols out there and I love the faces people make when they understand how ARP glues L2 and L3 and suddenly everything makes sense.
It is funny to see articles like this, because the network seems to be the "last frontier" for IT companies/workers. Only a handful seem to be brave enough to work on it (and actually like it). It is supposed to work like the lights when we switch them on and any interaction with a network engineer is just to tell them about an issue :)
In my opinion, the network guys are the geeks among the geeks. Don't get me wrong, it is easy to understand networking, it just takes dedication.
We only saw the issue when using the "iSCSI Auto-Config" mode. Manually configuring the switches with the same config but entered by hand resolved the issue.
Basically, we had multiple teams all launching/terminating web servers. Unfortunately, they were all in the same EC2 deployment, and more often than not our load balancers from one team would send traffic to the web servers of another team. Furthermore, our setups were similar enough that this would sometimes cause bad results for users. We fixed it by making sure that our web servers on every team spoke on different ports. Not elegant, but effective (until two teams accidentally picked the same ports).
These days we have good enough infrastructure tools that this problem should never happen. But in 2009, at a company that was overwhelmed with growth, those sort of things happen.
As for other cloud providers, I'd be really interested to hear from Digital Ocean, Rackspace, and Google!
I'm more surprised that AWS isn't sending a gratuitous ARP when IPs are re-cycled to mitigate issues like this.
1) Amazon sends GARP on behalf of the guest OS, from outside the kernel 2) Amazon has hooks into the guest OS to instruct the GARP to occur
Both are very bad things with regard to customer / Amazon separation of privileges and control.
I think every network engineer was probably yelling ARP at the screen early on in this write up and is likely one of the first things those seasoned in L2/L3 would look at first.
Also, as others have stated IPv6 removes this problem altogether. With the complete gutting of the concept of a broadcast and replaced by announcements within multicast groups which are much more efficient when you're talking about "small" prefixes that house 18,446,744,073,709,552,000 hosts (/64). ARP tables wouldn't be able to scale or work efficiently.
Overall this makes me think that it would be interesting to build a best practices guide around IPv4 and IPv6 networking within cloud provider environments. I think the gratuitous ARP on boot is a relatively safe practice, especially with regard to environments that are in continual flux.
$0.02.
eth0 Link encap:Ethernet HWaddr 42:01:0a:f0:40:dd inet addr:10.240.64.221 Bcast:10.240.64.221 Mask:255.255.255.255
This isn't an intentional feature, it's just a property of how our Andromeda SDN[0] is wired up. In particular, we lift the business end of figuring out where the other hosts on your LAN are up into the SDN rather than relying on the guest to cache ARP entries (hence the /32 netmask).
[0]: http://googlecloudplatform.blogspot.com/2014/04/enter-androm...
(I'm the original TL for GCE)
Perhaps AWS is blocking some class of packet?
> [...] These workers all connect to a MongoDB replicaset to update and read data [...] > 27017 = port mongod is listening on