Stories like this are stories like those mentioned in the post: everything on the Internet is awful, and you work with what you have and work around things by adding JPEGs of cats and silly session IDs (As TLS 1.3 does) and whatever it takes to “make it work”.
DoH/DoT, if it’s able to be adopted before ISPs and middleboxes ossify it with subpar implementations as part of their effort to avoid net neutrality ( https://www.theverge.com/2019/10/4/20898779/fcc-net-neutrali... ), is the first time in a long time that innovation and experimentation in DNS might happen again. This is why we see things like ESNI and SVCB ( https://tools.ietf.org/html/draft-ietf-dnsop-svcb-httpssvc ) having a chance at being deployed. At least, if clients can get there before the middleboxes ruin the party, like they’ve tried to do for QUIC.
Fly.io apps can define different service "handlers" (like TLS and HTTP). If you want to, you can accept TCP connections and bypass our logic. Which is great and flexible.
The problem is, when someone is deploying a new version of their app where they _change_ one of those things, we have to e really careful about how we (a) load balance and (b) decide to do things like TLS. If we're not careful we can end up sending the wrong type of connection to a new VM that's expecting something else.
SRV records sound like they'd have it worse. If you do a DNS lookup to detect something like http2, the IP you connect to _can't_ do anything else. It's much simpler / safer to negotiate stuff like that at connection time.
I think they'd end up with 4 lookups; A, AAAA, SRV (_http2._tls), and SRV (_http._tls).
Though perhaps you are suggesting DoH could mean the resolver also returns SRV records if you request A or AAAA? i.e. proactively point out there's an HTTP server?
I regularly work with load-balancers such as Citrix ADC (NetScaler) or F5 BIG IP. These do DNS-based load-balancing, dynamically returning "A" records to that the browsers so that they can get the "single working IP address" they're expecting. The browsers don't try very hard to fail over to secondary IPs because this is the established standard architecture, but they don't need to because of this common setup.
Sounds like an optimal solution, right? It does at first glance anyway, as long as you ignore the eye-watering price tag on those load balancer boxes.
The subtle but critical issue is that by returning "A" records, the load balancers have to use a short time to live (TTL)! This is because there's a trade-off: You can have fast failover, OR long-lived DNS caching. With A records you can't have both!
Typical response TTL times are 5-30 seconds, 5 minutes tops if you hate your users. This means that many browsers will be forced to repeatedly re-query the DNS servers on every page load for typical end-user workflows. It also means that for all but the biggest, most popular sites, the ISP DNS cache does practically nothing for these records.
Meanwhile with SRV records the TTL times can be much higher, hours even. This is how Active Directory works, for example, all of the Domain Controllers add themselves to various SRV records so that if you query "_ldap._tcp.dc._msdcs.test.com" you get back all the DCs. These records include priorities and weightings, so you can pull tricks like incrementally demote a DC or prioritise the shiny new one.
If you watch the AD connection traffic in WireShark, it's incredible. It very quickly steps through alternate services and then reorders the successful hits in front of the failures so that subsequent queries are lightning fast. It is astonishingly tolerant of partial networking failures, yet still fast to connect despite that!
The key mistake made by the original DNS design working groups was that SRV records should have returned a list of IP addresses instead of a list of host names.