How CDNs Generate Certificates
fly.io
fly.io
In principle they can all generate those statistics because they (are supposed to) log enough information to identify what went wrong when, inevitably, something is misissued. Logically that also includes at least which method was used to verify domain authorization or control.
One of the things wrong at Symantec is that it turns out some of the records were notionally kept at CrossCert, a separate Korean company. CrossCert simply did not keep any records (or if it did they were in such disarray that it seemed less likely to attract retribution by refusing to disclose them) and Symantec had seemingly never checked.
Knowing which methods are popular with Subscribers, and whether that varies considerably between CAs would be valuable in trying to figure out how more of the worst Blessed Methods can be deprecated or improved, and who we need to be talking to about that.
For example maybe Let's Encrypt is doing almost all the 3.2.2.4.19 ("Agreed Upon Change to Website - ACME") then there's no point ragging on other CAs for the shortcomings of relying on plaintext HTTP in this method. Or maybe DigiCert are doing a lot of 3.2.2.4.15 ("Phone Contact with Domain Contact") so they are the people to talk through any proposed improvements around stuff like leaving a Voice mail.
This delighted me.
How is the Fly proxy implemented? Are you using rustls and/or any of the available ACME crates?
I've been wanting to implement tls-alpn-01 support for rustls (although it might be possible to do this just by mutating the ServerConfig over time).
Also interested to hear your general impressions of Rust so far (I think I read some Twitter grumbling...).
I think I came across as grumbling about Rust when my real perspective was much more subtle. My take on Rust so far is that it has been, for me, a vindication of a lot of decisions the Go team made, because I've been directly exposed to some of the downsides of the opposite decisions. But, while that sounds like a critique of Rust, it's not! Rust is the way it is for real reasons: zero-cost abstractions and no runtime GC, which are, right now, requirements for some application domains.
For me, right now, writing in Rust feels almost identical to how writing in C++ felt 15 years ago. But I'll keep writing in it, and it'll get faster for me. We're a Rust-on-the-data-plane shop!
Fly's proxy uses a mix of tokio, hyper and rustls. We don't need to use a crate that handles ACME because we're processing all the validation and certificate authorizations from a centralized, boring, Rails application.
We've had to submit a PR to the rustls project a few months ago to handle different ALPNs. Instead of resolving a certificate only from a SNI, the crate now provides the full ClientHello which contains negotiable ALPNs. With that information you can respond to the tls-alpn-01 challenge.
"I absolutely understand what y’all like so much about Rust, but I have to say that as an auditor, my blood pressure drops and my shoulders relax the moment I switch from reading a Rust project to reading a Go project."
(I will however miss match expressions when I return to my home planet.)
It feels like the Internet is so fragile.
Stories like this are stories like those mentioned in the post: everything on the Internet is awful, and you work with what you have and work around things by adding JPEGs of cats and silly session IDs (As TLS 1.3 does) and whatever it takes to “make it work”.
DoH/DoT, if it’s able to be adopted before ISPs and middleboxes ossify it with subpar implementations as part of their effort to avoid net neutrality ( https://www.theverge.com/2019/10/4/20898779/fcc-net-neutrali... ), is the first time in a long time that innovation and experimentation in DNS might happen again. This is why we see things like ESNI and SVCB ( https://tools.ietf.org/html/draft-ietf-dnsop-svcb-httpssvc ) having a chance at being deployed. At least, if clients can get there before the middleboxes ruin the party, like they’ve tried to do for QUIC.
Fly.io apps can define different service "handlers" (like TLS and HTTP). If you want to, you can accept TCP connections and bypass our logic. Which is great and flexible.
The problem is, when someone is deploying a new version of their app where they _change_ one of those things, we have to e really careful about how we (a) load balance and (b) decide to do things like TLS. If we're not careful we can end up sending the wrong type of connection to a new VM that's expecting something else.
SRV records sound like they'd have it worse. If you do a DNS lookup to detect something like http2, the IP you connect to _can't_ do anything else. It's much simpler / safer to negotiate stuff like that at connection time.
I think they'd end up with 4 lookups; A, AAAA, SRV (_http2._tls), and SRV (_http._tls).
Though perhaps you are suggesting DoH could mean the resolver also returns SRV records if you request A or AAAA? i.e. proactively point out there's an HTTP server?
I regularly work with load-balancers such as Citrix ADC (NetScaler) or F5 BIG IP. These do DNS-based load-balancing, dynamically returning "A" records to that the browsers so that they can get the "single working IP address" they're expecting. The browsers don't try very hard to fail over to secondary IPs because this is the established standard architecture, but they don't need to because of this common setup.
Sounds like an optimal solution, right? It does at first glance anyway, as long as you ignore the eye-watering price tag on those load balancer boxes.
The subtle but critical issue is that by returning "A" records, the load balancers have to use a short time to live (TTL)! This is because there's a trade-off: You can have fast failover, OR long-lived DNS caching. With A records you can't have both!
Typical response TTL times are 5-30 seconds, 5 minutes tops if you hate your users. This means that many browsers will be forced to repeatedly re-query the DNS servers on every page load for typical end-user workflows. It also means that for all but the biggest, most popular sites, the ISP DNS cache does practically nothing for these records.
Meanwhile with SRV records the TTL times can be much higher, hours even. This is how Active Directory works, for example, all of the Domain Controllers add themselves to various SRV records so that if you query "_ldap._tcp.dc._msdcs.test.com" you get back all the DCs. These records include priorities and weightings, so you can pull tricks like incrementally demote a DC or prioritise the shiny new one.
If you watch the AD connection traffic in WireShark, it's incredible. It very quickly steps through alternate services and then reorders the successful hits in front of the failures so that subsequent queries are lightning fast. It is astonishingly tolerant of partial networking failures, yet still fast to connect despite that!
The key mistake made by the original DNS design working groups was that SRV records should have returned a list of IP addresses instead of a list of host names.
I am more interested in the mesh. Do you have more details on that? Specifically why this architecture was chosen, what kind of latency does WireGuard add, etc.
We picked it because it's really simple to manage, and we wanted to ensure traffic between datacenters was always encrypted. We have a little tool called "flywire" that keeps wireguard peer configs updated from Consul. Once we accept a connection from a user, we pick a target VM, and then connect them over the wireguard mesh.
For our purposes, it basically doesn't add any noticeable latency. I think when we tested we say something on the order of 0.1ms of added latency over wireguard, but I don't quite remember. It's never been the source of latency problems when we do have them, at least!
P.S: I’ve built a service using Fly and can’t recommend it enough!
https://fly.io/docs/app-guides/run-a-private-dns-over-https-...
(Beyond that, for whatever it's worth: you can skip our HTTP/H2 termination entirely and speak TCP directly to your VMs).
... and use Caddy to do the heavy lifting. (I'm biased, yes. But the linked doc is multi-authored and applies to every sysadmin or developer who needs to manage certs, regardless of your software choice.)
I also love Caddy. In fact, you can run it on Fly.io (and even opt out of our TLS/cert stack). I would love it if it could just put certs in Vault, though.
micro-2x shared 512MB $0.000003044 $8 VS t3a.nano 2 Variable 0.5 GiB EBS Only $0.0031 per Hour
I'm missing something? cause seeing the pricing I still feel AWS is cheaper.
(For the unfamiliar reader: Firecracker is a micro-vm system that sits sort of in between a fully virtualized host, like an EC2 instance, and a container like Docker; you get the security isolation of a hypervisor but the speed/simplicity of Docker. It's the engine that powers AWS Lambda and Fargate. The Usenix paper is a pretty great read, and the code [it's all in Rust] is simple and easy to follow.)
Disclaimer: interned with the team