But is it a compromise of the app? I'm failing to see the nexus between Houseparty, the app, and these subdomains. How does someone using the app end up on the subdomain?
But is it a compromise of the app? I'm failing to see the nexus between Houseparty, the app, and these subdomains. How does someone using the app end up on the subdomain?
- Houseparty spins virtual servers up and down quickly to respond to changing demand.
- They have a DNS TTL that's longer than the draining time for those VMs. Like, maybe they it takes them five minutes to kill a VM but DNS records are cached for an hour.
- The attacker takes advantage of that window to spin up a bunch of tiny VMs and see if any of them are allocated one of the IPs that Houseparty recently abandoned. If so, then they fire up Nginx and serve poisoned PDFs to anyone who connects to it.
- An end-user's app connects to "ephemeral-vm-837.houseparty.whatever" and gets the cached value that their ISP is still serving, because Epic configured their DNS to tell the ISPs to cache it for an hour, except now that IP is served by the attacker instead of by Epic.
Mitigating this could be as blunt as setting DNS TTLs to 5 minutes, then putting "sleep 300" at the start of each server's shutdown script. When you decide to kill a server, first kill its DNS record so that nothing refers to it anymore, then start the waiting period before actually removing it from rotation.
Even if the attacker redirects the traffic, it still needs to serve the HTTPS/SSL certificates. How will the attacker do that?
There are a lot of ways this could go wrong.
EDIT:
Specifically: https://letsencrypt.org/docs/challenge-types/
It won't give you a wildcard certificate, but you don't need one for the type of attack we're talking about.
You get a bunch of elastic IPs, you point your domains to those, you point those to your instances.
That's secure, no one can use those IPs until you release them. Is this rocket science? On AWS I don't think you really get charged for this even (or very little) from what I can tell if you have them all pointing to your instances.
[1] https://aws.amazon.com/premiumsupport/knowledge-center/elast...
[2] https://aws.amazon.com/ec2/pricing/on-demand/#Elastic_IP_Add...
One could argue that this is pocket change for a large customer to do the right them, but from a PM's perspective that $40K per year they could be spending towards another engineer's salary.
[0] https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/elastic-... [1] https://aws.amazon.com/ec2/pricing/on-demand/#Elastic_IP_Add...
If I had a dollar for everytime traffic stopped to an IP after the DNS ttl expired, I wouldn't have any dollars. After taking IPs out of DNS for www.popular site, I would continue to see traffic for weeks, and the TTL was set for somewhere like 5 minutes, an hour tops. Edit: also, when we moved our authoritative DNS, the old provider kept seeing queries for at least 6 weeks.
If the only thing preventing you from sending private data or receiving trusted data is that it came from an IP you got from DNS, it's not secure.
If they do, then the procedure above (and doing everything over https) should be enough, I think.
What I proposed was a blunt force tool to try to help mitigate the problem, but you're 100% correct that it's not sufficient. Key pinning would go a lot further towards fixing it. What I suggest could be rolled out pretty quickly and without risk, though, so IMHO it's worth adding to the to-do list.
They are NOT then checking that the subdomain was released and got a new IP recently (or didn't get an IP).
These PR people love to split these hairs. Our app is secure - but we leaked 100% of your data out of our back-end servers when they were hacked for 3 months - is another one that's common.
"Certificate Pinning" and/or HSTS both help, and there are other measures.