HTTP/HTTPS not working inside your VM? Wait for it
rachelbythebay.com
rachelbythebay.com
It's a special kind of horror to find, after hours of high-end-googling, the one thread where someone reports the same problem you are experiencing, and it's just the question, and then one other person asking if the problem has been solved because she/he is having the same problem.
The one thing that is worse is if the OP then makes another post that simply says "Solved it! =D", without giving any explanation on how they solved it.
Now that I think of it, I have seen that happen, but the reply was inevitably "I have already googled the question to death, without finding any useful results."
The person who had posted the answer was, of course, also me.
No, I am not bitter. Learning where their FTP server is and how to navigate, well, that was gold.
https://web.archive.org/web/*/
to the now broken URL.https://addons.mozilla.org/en-US/firefox/addon/resurrect-pag... https://github.com/arantius/resurrect-pages
Here you go:
ip6tables -A OUTPUT -t mangle -m multiport -o eth0 --protocol tcp \
--tcp-flags ALL SYN --dports 80,8080,443 -j MARK --set-mark 6
tc qdisc del dev eth0 root
tc qdisc add dev eth0 root handle 1: htb default 1
tc class add dev eth0 parent 1: classid 1:6 htb rate 10000Mbps
tc qdisc add dev eth0 parent 1:6 handle 6: netem delay 100ms
tc filter add dev eth0 protocol ipv6 prio 1 handle 6 fw flowid 1:6
Although, I would personally throw in an extra class to add 200ms delay to legacy IP protocol: iptables -A OUTPUT -t mangle -o eth0 -j MARK --set-mark 4
ip6tables -A OUTPUT -t mangle -o eth0 -m multiport --protocol tcp \
--tcp-flags ALL SYN --dports 80,8080,443 -j MARK --set-mark 6
tc qdisc del dev eth0 root
tc qdisc add dev eth0 root handle 1: htb default 1
tc class add dev eth0 parent 1: classid 1:4 htb rate 10000Mbps
tc class add dev eth0 parent 1: classid 1:6 htb rate 10000Mbps
tc qdisc add dev eth0 parent 1:4 handle 4: netem delay 200ms
tc qdisc add dev eth0 parent 1:6 handle 6: netem delay 100ms
tc filter add dev eth0 protocol ip prio 4 handle 4 fw flowid 1:4
tc filter add dev eth0 protocol ipv6 prio 6 handle 6 fw flowid 1:6
...just to be compliant with draft-howard-sunset4-v4historic-00 once it becomes RFC.1. http://www.loopinsight.com/2016/01/28/vmware-abruptly-fires-...
Rachel, could you write a small sneaky program (using eg libpcap) to see if the TCP handshake has completed by the time connect(2) returns control to your program, before your first write(2)?
From what I know about tc, to only delay ip6 traffic you've got to create a root qdisc that has multiple subclasses (like tc-prio [0]), and attach tc-netem to one while passing the other straight through. Then classify packets between the two, although I'd do that with iptables rather than figuring out any more of the tc workings that necessary.
[0] The default pfifo_fast has multiple subclasses, but from what I remember it had some problem with child qdiscs?
Also does anyone know what is the reason behind this peculiar behavior? A bug or something more fundamental ?
And I've never seen a VPS provider that doesn't offer v4. Even a $5/mo DigitalOcean box has it.
https://lowendbox.com/ has many options cheaper than that which do have real IPv4.
It's not that I need to restrict access to them, just that they don't need IPv6, since anyone accessing them usually already has IPv6, and dual-stack is extra work.
BTW, if you use Firefox there is an addon called FlushDNS that shows the IP address of the web server. It's main purpose is to remove an address from the DNS cache inside the browser but actually it's more useful as an inspection tool.
tc qdisc add dev eth0 root netem delay 100ms
and:
printf 'HEAD / HTTP/1.0\r\n\r\n' | nc -6 rachelbythebay.com 80
returns nothing whereas
printf 'HEAD / HTTP/1.0\r\n\r\n' | nc -4 rachelbythebay.com 80
returns as expected.
In my case, firefox has no problem if invoked from command line, but "sometimes" it just hang up when invoked from script.
[1] VMware's network driver does not handle TCP, or IP. It's just layer 2; it
implements one of a couple kinds of network hardware, that's it.
[2] VMware Guest Tools does install a para-virtualized network card driver
- vmxnet2/vmxnet3. It communicates with the physical network device by
communicating with the host OS, rather than emulating a network driver. That
potentially may do something wonky with something above layer 3, even though
it really should not be.
[3] VMware does have a virtual network switch, which forwards frames between
the physical NIC and virtual NIC based on MAC address.
[4] VMware may handle moving frames from a virtual NIC to a physical differently
than moving it to another virtual NIC.
[5] VMware provides VMDirectPath I/O, which allows the guest to directly address
the network hardware.
[6] TSO/LSO/LRO can have a negative impact on performance in Linux (though
supposedly, LRO only works on vmnet3 drivers, and from VM-to-VM,
for Linux).
[7] Emulated network devices may not be able to process traffic fast enough,
resulting in rx errors on the virtual switch.
[8] Promiscuous mode will make the guest OS receive network traffic from
everything going across the virtual switch or on the same network segment
(when using VLANs).
[1] You can try changing the VMware guest's emulated network card (vlance, e1000) and trying your thing again, but I doubt it will change much.[2] Try installing or uninstalling VMware Guest Tools and corresponding drivers.
[3] Nothing to do here, really. If you have multiple guests sharing one physical NIC, try changing it to just one?
[4] Try your test again between two VMs on the same host.
[5] Try this, or not?
[6] Try enabling or disabling LRO. Or play with all three settings and see what happens. https://kb.vmware.com/selfservice/microsites/search.do?langu...
[7] Try increasing buffer sizes. https://kb.vmware.com/selfservice/microsites/search.do?langu...
[8] Disable promiscuous mode on your NIC.
Other non-VMware things to investigate:
[1] Your guest OS may have bugs. In its emulated network drivers, in its
tcp/ip stack, in its applications, etc.
[2] An intermediary piece of software may be fucking with your network
connection. IPtables firewall, router/firewall on your host OS, after
the host OS/before your internet connection, at your destination host, etc.
[3] Sometimes, intermittent network traffic makes it look like there is a
specific cause, when really the problem is hiding in the time it takes
you to test.
[4] The Linux tcp/ip stack (and network drivers) collect statistics about
erroneous network traffic.
[5] Network traffic will show missing packets, duplicate packets, unexpected
terminations, etc.
[6] Your host OS or network hardware may be buggin'.
[1] Try a different guest OS.[2] Make sure you have no firewall rules on the guest, host, internet gateway, etc. Try a different destination host.
[3] Run tests in bulk, collect lots of samples and look for patterns.
[4] Check for dropped packets, errors on the network interface, in tcp/ip stats.
[5] Tcpdump the connection to see what happens when it succeeds or fails.
[6] Try a different host for your VM.
edit one more idea: Look at the response headers for the request to the site. The content length is 1413 bytes. Add on the TCPv6 and IPv6 header overhead (and http headers, etc) and this is probably over 1500 bytes, the typical MTU maximum. Try requesting a "hello world" text file and try your test again.
Maybe I was just missing a setting somewhere, but couldn't find it then.
Once upon a time it was common to think that we can design software without bugs, or at least almost. That didn't work at all! What did work is redundant systems, invariant testing and fail-fast with restarts. This is how reliable systems are written these days.
Bugs are common; we have to learn to work around them.
void segfault_sigaction(int signal, siginfo_t *si, void *arg)
{
//Pretend it never happened
return;
}Idiomatic Erlang doesn't differentiate between "system" / "environment" errors and local bugs. If it has failed — restart it!
The only way I can imagine this working is if Erlang is so buggy and nondeterministic that it inserts crashes sometimes but not all of the time. But that's obviously absurd.
"3.4 The Restart Frequency Limit Mechanism"
It seems actively worse to allow users to retry requests that are doomed to failure than to put up a fail-whale or similar while the ops team is being paged.
Also, the bug discussed in this article wasn't causing crashes. What would you propose be crashed and restarted in this case?
If it quickly repeats, you've isolated the failure to happening within a narrow scope.
This part isn't really Erlang magic, apache in pre-fork mode has a lot of the same properties. There may be some magic in supervision strategies, but I think the real magic is the amount of code you get to leave out by accepting the possibility of crashes and having concise ways to bail out on error cases.
For example, to do an mnesia write and continue if successful and crash if not, you can write
ok = mnesia:write(Record)
Similarly, when you're writing a case statement (like a switch/case in C), if you expect only certain cases, you can leave out a default case, and just crash if you get weird input.I also find the catch Expression way of dealing with possible exceptions is often nicer than try/catch. It returns the exception so you can do something like
case catch Expression of
something_good -> ok;
{'EXIT', badarg} -> not_so_great
end
and handle the errors you care about in the same place as where you handle the successes.Edited to add, re: failwhale, your HTTP entrypoints can usually be something like
try
real_work_and_output()
catch
E:R ->
log_and_or_page(E,R)
output_failwhale()
end.
As long as the failure in real_work_and_output is quick enough, you'll get your failwhale. Of course, if the problem is processing is too slow, you might want to set a global failwhale flag somewhere, but your ops team can hotload a patch if they need to fix the performance of the failwhale ;) case catch Expression of"
Something to be aware of is the cost of a bare catch when an exception of type 'error' is thrown:"[W]hen the exception type is 'error', the catch will build a result containing the symbolic stack trace, and this will then in the first case [1] be immediately discarded, or in the second case matched on and then possibly discarded later. Whereas if you use try/catch, you can ensure that no stack trace is constructed at all to begin with." [0]
Stack trace construction isn't free, so it makes sense to avoid it if you're not going to use it. I know that in either Erlang 17 or Erlang 18, parts of Mnesia were slightly refactored to move from bare catch to try/catch for this very reason.
[0] http://erlang.org/pipermail/erlang-questions/2013-November/0...
[1] He's referring back to an example in the email
Or we could, you know, fix them.
I wasn't asking for a justification. I was just asking why this is occurring. If you don't know, that's cool. I mean, one of the reasons I ask is because I'd like to know if VMWare are going to fix this bug.
So thank you for explaining that software has bugs. I'm sure I'll remember that the next time I fix a regression in LibreOffice, as I did with the issue with EMF dashed lines not displaying correctly or when I fixed the issue where JPEG exports didn't export the DPI value correctly...
I have not seen the issue described here in any configuration - neither on clients nor on servers. I wonder whether this is an issue with VMWare running on a specific host?
edit: From the forum post in the linked article, I'm gathering they are using IPv6 NAT. So this might be a problem with the VMWare NAT interface - my configurations are all bridged.