Multi-Region Latency Based Routing now Available for AWS
aws.typepad.com
aws.typepad.com
Come join us and be part of the fun that is AWS: http://aws.amazon.com/careers/
If you can name it, we are probably hiring for it.
As one of my colleagues mentioned, we are hiring (currently 516 open jobs on the AWS team). Here's a list that's easier to read:
http://awsmedia.s3.amazonaws.com/jobs/all_aws_jobs_list.html
(Sorry Jeff, I couldn't resist)
http://www.amazon.com/gp/jobs/ref=j_sq_btn?keywords=&cat...
(sorry Jeremy, coudln't resist) :-)
I'm looking for engineers and UX people to come help me make the Heroku Add-ons platform even more amazing. If putting these cloud services into the hands of developers and changing the way people think about provisioning these services sounds like something you'd like to be part of: glenn at heroku dot com
We build software for AWS and are looking for all sorts of engineers: from kernel development to building great web front-ends for our customers. Check out http://www.amazon.co.za/ for more info.
I was under the impression that end users typically talked to DNS cache servers rather than directly to the authoritative servers in the domain's registration. If that's true, how can AWS provide different records based on the requesting user?
If end users are talking directly to one of the ~4 authoritative name servers listed in the registration, how does that scale to billions of queries?
Then, that DNS server can decide where the other end is and provide them with the correct IP.
You're right that most users use a caching DNS server, so it is actually the location of the DNS caching server. Of course, since most users choose a DNS server close to them (usually, their ISP), this should still result in a correct approximation. If you're using, say, 8.8.8.8 (google public dns), that's (probably) more than 1 server -- you're using the closest one to you. So then AWS will provide you with the closest AWS region to the closest Google DNS server, which is hopefully the closest one to you.
You are right that a DNS server sees the query coming from the resolver instead of the user. So how does it pick the region closest to the user? As the blog post describes, we measure latencies from client networks to AWS regions and we also have a mapping of which resolvers are used by which client networks. If you put both of them together, you can compute which region is closest to the users of the resolver.
How did you build this map?
I can think of a few complicated ways to go about it, but I'm wondering if there is something easy I'm missing.
Full-disclosure: I work on Amazon Route 53, and although we don't quite use that same method - it will give you an idea of what's possible. PS; we're hiring.
1) I visit the site, it gets my IP address
2) Magic happens
3) It displays my nameserver address.
What is going on in step 2?
Edit: worked it out.
For those interested, it uses a Javascript include from a unique subdomain name. Because the subdomain is unique the app can work out the relationship between client IP and resolver.
Fetching http://whatsmyresolver.stdlib.net/resolver/ triggers a 302 redirect to a url of the form;
http://$guid.nonce.stdlib.net/resolver/
The DNS server authoritative for nonce.stdlib.net has a simple wildcard configured, so *.nonce.stdlib.net all resolve to the same web-server. Obviously the DNS request for the globally unique id domain name has to come before any HTTP request to the guid url, so when the DNS request comes in the authoritative server can record it in a simple lookup store (guid -> resolver source ip).Then, when the HTTP request makes it to the web-server, it can inspect the Host: header to determine what the guid was. It then uses this guid to correlate the HTTP request it is handling and the resolver source ip, and generates some javascript with the data we need;
var resolver="192.0.2.53";var edns=true;
It's just a hack I wrote up for my own reasons years ago. But if you'd like to avail of it for any reason (ie helping end-users debug things), feel free to embed; <script language="javascript" src="http://whatsmyresolver.stdlib.net/resolver/"/>
and use the variables it populates. No warrantees or guarantee implied :-)http://www.facebook.com/note.php?note_id=10150212498738920
If I had to guess, I'd say that since route 53 is a dns host for many domains, they might be able to work out the user ip / resolvers map passively. Pretty awesome stuff, Amazon! This is a big deal for your customers.
Nearest in terms of AS path. This can be wildly different than latency as BGP isn't really implemented with cost/performance in mind. Anyone relying on anycast to answer their latency story is doing a lazy/poor/ignorant job of it.
> You're right that most users use a caching DNS server, so it is actually the location of the DNS caching server.
And another case where end users would be better served by providers implementing edns client-subnet extensions. This would allow savvy content providers to route to the actual end user, and not their resolver.
It sounds like here Amazon is doing unicast DNS, but dynamically changing the response depending on the IP address the query comes from.
Since they have a big database of IP addresses and latency, they can choose the best response.
The missing thing here is that they don't get the IP address of the client computer, only the DNS server. In many cases that (along with geo-location of the DNS Server) might be enough, but it is possible to do better. I'd speculate what they do here is correlate the IP addresses with the DNS server either by subnet or by using active measurement techniques.
There's still no good high-availability story for region-wide outages, and there won't be until they do this.
They need to just stop having region wide outages...
I guess an alternative might be to do some kind of NAT for unbound anycast addresses which forwards packets to an available region, but that is hugely complicated.
The only thing I'm curious as to is what kind of measurements is Amazon gathering, and how is it gathering them? Is it using ELB, and looking at TCP latency (delta time between SYN/ACK? Curious minds want to know...
If this is so implausible then why can't they provide a cap? It feels to me that they are purposefully trying to capitalize on people's mistakes. Call me paranoid but it is honestly the primary reason why I've been wary to play with AWS.
My point is that it makes business sense for Amazon to lure this old-fashioned crowd, and it could do so by implementing a simple feauture such as spending control.
One thing that'd help dramatically would be exposing the various account activity and cost data via API. There's currently no API for accessing one's monthly spend to-date.
From a comment elsewhere, looks like they're sorta working on it, finally (CloudWatch metric).
Doesn't it seem bizarre to give your credit card number over and agree to pay some undisclosed sum with no cap? If it's my money and I'm paying for some nonessential service, I expect to never be surprised at how much I'm charged, even if its $10 when I expected $2.
> I'd challenge your assumption that Amazon is intentionally trying capitalize on mistakes
I actually worked at Amazon (not on AWS) and I've worked at a few other BigCo's and you might be surprised by how many things exactly like this they do have the time and desire to worry about. You have it backwards; for a company as big as Amazon it's worth paying an entire salary just to worry about things like this, since a fraction of a percent of increased AWS income will more than pay for itself. It's extremely likely they are at least trying to capitalize on people who sign up for the free tier and accidentally go over their limit, or else they would have a "trial" mode that only confirms your credit card and won't ever charge it without your approval.
http://blog.bitnami.org/2011/12/monitor-your-estimated-aws-c...
[1] https://forums.aws.amazon.com/thread.jspa?messageID=249000