We need better support for SSH host certificates
mjg59.dreamwidth.org
mjg59.dreamwidth.org
Avoiding Trust on First Use is potentially a big benefit, but the workflow improvements for developers, and especially non-technical people, is a huge win too.
At work, we switched to Step CA [1] about 2 years ago. The workflow for our developers looks like:
1. `ssh client-hosts-01`
2. Browser window opens prompting for AzureAD login
3. SSH connection is accepted
It really is that simple, and is extremely secure. During those 3 steps, we've verified the host key (and not just TOFU'd it!), verified the user identity, and verified that the user should have access to this server.In the background, we're using `@cert-authority` for host cert verification. A list of "allowed principals" is embedded in the users' cert, which are checked against the hosts' authorized_principals [2] file, so we have total control over who can access which hosts (we're doing this through Azure security groups, so it's all managed at our Azure portal). The generated user cert lasts for 24 hours, so we have some protection against stolen laptops. And finally, the keys are stored in `ssh-agent`, so they work seamlessly with any app that supports `ssh-agent` (either the new Windows named pipe style, or "pageant" style via winssh-pageant [3]) - for us, that means VSCode, DBeaver, and GitLab all work nicely.
My personal wishlist addition for GitHub: Support for `@cert-authority` as an alternative to SSH/GPG keys. That would effectively allow us to delegate access control to our own CA, independent of GitHub.
[1] https://smallstep.com/docs/step-ca
[2] https://man.openbsd.org/sshd_config#AuthorizedPrincipalsFile
How.
Specifically, what I cannot determine from their docs is how the VM obtains a host key/cert signed by the CA. How does the CA know the VM is who the VM says it is? (I.e., the bootstrap problem.)
(I assume that you also need your clients to trust the CA … and that has its own issues, but those are mostly human-space ones, to me. In theory, you can hand a dev a laptop pre-initialized with it.)
Because each of our servers are bespoke, we can use the admin provisioner when the server is first being set up (and actually, Ansible handles this part).
I don't have experience with it, but StepCA also has Kubernetes support, and I imagine the control plane could authenticate the pod when a cert needs to be issued or renewed.
an alternative example: our org solves the issue with TOTP, required every 8 hours for any operation; from ssh/git CLI based actions (prompted at the terminal) to SSO integrations. decoupling security from unrelated programs. simple and elegant.
Besides, neither me nor GP is saying this needs to be a universal pattern. We are saying that it’s a viable pattern for a lot of orgs.
Plus, as others have said, you can jump through SSH sessions with on the client ssh command (ie without having to manually invoke ssh on the jump box).
I do, however when I do this I make sure the certificate is signed with permit-agent-forwarding and demand people just forward their ssh agent on their laptops.
This also discourages people from leaving their SSH private key on a server just for ssh-ing into other servers in CRON instead of using a proper machine-key.
It's better to configure jump hosts in your local ssh config.
For example, if I've SSHed from my laptop to Host A to Host B to Host C then need to authenticate a CLI program I'm running on Host C, the program can show a link in the terminal which I can open on my laptop.
If key forwarding works, that might be workable.
I'm extremely wary of non-standard ssh login processes as they tend to break basic scripting and tooling.
To be clear, I’m not suggesting the GPs approach is “optimal”. But if you’ve gone to the trouble of setting that up then you should have already solved the problems of data sharing (mitigating the need for rsync), network segregation and secure access (negating the need for jump boxes), etc.
SSH is a fantastic tool but mature enterprise systems should have more robust solutions in place (and with more detailed audit logs than an rsync connection would produce) by the time you’re looking at using AD as your server auth.
[1] https://developer.hashicorp.com/vault/docs/secrets/ssh/signe...
1. `ssh client-hosts-01`
2. Browser window opens prompting for AzureAD login
3. SSH connection is accepted
How is that simple, compared to `ssh -i .ssh/my-cert.rsa someone@destination --> connection is accepted, here's your prompt` ?If you deactivate someone in AD, poof, all their access is magically gone, instead of having to remove their public key from every server.
Your command is disingenuous in that it only works if the certificate has already been issued to you. If you were to include issuance, your command would very much turn non-simple.
Which would mean you have one single point of attack/DOS/failure that needs to be kept utterly secure at all costs?
Ever since I saw it in action with Tailscale I've always wondered how it actually works, and I guess if anyone would know they'd be on HN
https://en.wikipedia.org/wiki/SSHFP_record
Of course this has its own challenges, but if you automate your DNS, this can be neat!
Cheers!
For ssh that doesn't really apply as clearly though.
https://datatracker.ietf.org/doc/html/draft-zhang-trans-ct-d...
https://www.huque.com/2014/07/30/dnssec-key-trans.html
https://datatracker.ietf.org/doc/html/draft-ietf-dnsop-deleg...
Do all CAs implement multi-prespective validation these days? Let's Encrypt implemented that only in 2020 and they believed they were the first ones:
https://letsencrypt.org/2020/02/19/multi-perspective-validat...
Do you think that DNS-based proof of ownership is not something that CAs would use for SSH certification?
In contrast, there is no unique "root CA" that can fail.
DNS root key compromise breaks the entire system until it is replaced.
Not seeing a huge difference.
It is not possible to revoke the DNS root, and there are no widely deployed alternatives. The incentive to do the right thing isn't as hard: it's just "good" guys doing the right stuff. If something wrong happens, where will you go ? Nowhere else.
A malicious CA in one country can issue a fraudulent certificate for a site in another country, whereas the people operating .ru can't affect the records for example.us so the blast radius is limited by design.
Moreover, no one is required to use a ccTLD, and there are hundreds of gTLDs to choose from, or you could even run one yourself if necessary.
Sure and they'll be quickly mistrusted. You can't really revoke DNNSEC trust of an ccTLD operator.
> Moreover, no one is required to use a ccTLD, and there are hundreds of gTLDs to choose from, or you could even run one yourself if necessary.
This is bypassing a dangerous design, at best.
But you don't have to, because the blast radius is so much smaller, and the incentives are aligned better. The reason why CAs require such extreme punishment for misbehaviour is that one bad CA can break the trust for every site on the web.
If a country decided to invalidate the security of (predominantly) its own citizens' websites then that wouldn't harm anyone who used any of the other ccTLDs in the world (not to mention the hundreds of gTLDs).
Also, I think you are over-estimating the ease with which a CA can be "quickly mistrusted". What is the record for how quickly a CA has been taken out of browsers' certificate stores, measured from the time of their first misissuance?
And I would argue that revoking CA trust to Let's Encrypt / IdenTrust would be much more disruptive than revoking a single ccTLD operator, since that would mean breaking most sites on the web. So DNSSEC is actually better in terms of the "too big to fail" problem.
> This is bypassing a dangerous design, at best.
But that's my point; DNSSEC lets you bypass the danger of a rogue issuer, by swapping to an alternate domain in the worst case, whereas with CAs you have to hope that the rogue issuer doesn't decide to target you, and wait for the bureaucratic and software update processes to remove that CA from all your users' browsers.
There are definitely limitations to the DNSSEC system as currently deployed, just as there were with the web PKI system before browsers started to patch all the holes in that, but I don't know why my position on this technical question is so controversial. Nevertheless, I really appreciate you taking the time to offer intelligent counter-arguments in your comment, thank you.
Entire countries best case is a small blast radius? A small CA going rogue would have a much smaller one, when we're talking about best case. Worst case is massive either way (say LetsEncrypt and .com). People also buy a lot of domains ignoring the fact that they're ccTLDs. The mere implication that people should choose their domains considering this fact is terrible.
> The reason why CAs require such extreme punishment for misbehaviour is that one bad CA can break the trust for every site on the web.
They can, but it'll be discovered really quick, especially with CAA violations. This can't be said about DNSSEC, any key compromise and abuse is difficult if not impossible to detect. Imagine that but with DANE, indefinite MITM, scary.
> DNSSEC lets you bypass the danger of a rogue issuer, by swapping to an alternate domain in the worst case, whereas with CAs you have to hope that the rogue issuer doesn't decide to target you
That's an insane bypass though. "Just cut your arm off, then it won't hurt." Change your email, figure out how to patch millions of devices out in the wild, so many problems.
A rogue issuer is much less hassle short- and long-term to deal with. Most browsers ship CRLite or similar and can revoke the root quickly. You can resume operation with a new CA rather fast.
DNSSEC is a nice complement to WebPKI and vice versa, but for our all sake, it can't be the only source of trust.
CloudFlare added support in 2018: https://blog.cloudflare.com/additional-record-types-availabl...
AWS still doesn’t support it: http://web.archive.org/web/20210429183447/https://forums.aws...
Namecheap doesn’t: https://www.namecheap.com/support/knowledgebase/article.aspx...
GCP doesn’t: https://cloud.google.com/dns/docs/records
Azure doesn’t: https://learn.microsoft.com/en-us/azure/dns/dns-zones-record...
Definitely annoying, but does ensure that the returned SSH key is always correct and not from a forged record.
I don't believe that that is true. DNSSEC RRSIG records are created over the entire result set. So even if there are numerous records returned, you should still be able to verify the signature. Also, there is nothing stopping you from also returning multiple SSHFP records in a single query.
However, SPF does have a design flaw (amongst many other) that the record is placed under the domain root, which is often already polluted with other records. This is why other standards that use TXT (DMARC, DKIM, BIMI, MTA-STS, TLSRPT, etc.) use a specific label prefix, or a selector. But this is not because of DNSSEC.
> SSHFP:
> https://www.rfc-editor.org/rfc/rfc4255
>> Re SSHFP:
>> Regarding DNS as a database for keys... Please stop this madness.
>> DNS isn't a database.
>> It's not a configuration store.
>> It was meant and should be used only for name resolution.
[0] https://mjg59.dreamwidth.org/65874.html?thread=2106450#cmt21...
>> It's not a configuration store.
Its quite literally both those things.
There may be practical reasons to question storing high value keys in dns. But not being a database of configuration info isn't one of them.
People have been using DNS for all kinds of configuration for decades now, with great success. So what, other than a pedantic and grumpy programmer’s perspective, should keep us from using DNS for configuration records?
> I must say that I am not a fan of these floppy-based routers. Essentially, you are taking one of the most unreliable pieces of storage known to man, and trying to build security infrastructure on it. That's madness.
dns is not the foundation upon which you want to build your secure infrastructure.
Besides, we’re not talking about secure infrastructure per se, but distributing public metadata for IP addresses. That sounds like the prime thing DNS has been invented for to me.
2. Cache. People are caching DNS aggressively. Often more aggressive than what the TTL allows.
You don't want to save the fingerprint in a stale, insecure database like that.
If you start using other ways of using DNS, then you might as well not use DNS and instead develop something more suited to the purpose.
I'd be interested to know where you get your data from. By my count, there are 142 ccTLDs that support it and 106 that don't, out of 248 ccTLDs.[0]
That's already more than half, but if you include the gTLDs then the number of TLDs signed in the DNS root goes up to 92% according to the best data I can find.[1]
I'm trying to figure out how having a stale fingerprint would be an automatic bad thing.
Let's assume you have a server with a fingerprint stored in DNS. Something happens and the server's certificate/key needs to change. So now you push out a new key fingerprint to DNS.
The failure mode for an out-of-date fingerprint would be to not trust the new server's key. In this case, the default failure mode is to fail safely. The client could then have a few new options, like querying the authoritative DNS server or prompt the user.
You can argue that you wouldn't want to have stale DNS caching in an automated system, but in a user-interactive mode, it's not the worst thing.
And for an automated system, the system should be robust enough to manage a bad fingerprint (fail safely again), or be in control of the entire infrastructure, including you DNS cache.
Or am I missing something?
But its kind of a strawman because nobody is suggesting putting unauthenticated keys in dns with no dnssec. The suggestion is either using dnssec, or have some sort of CA system.
We ended up assessing the current host certificate situation as bringing some benefit, but it was difficult to assert that the benefits were commensurate with the costs, especially considering the many other quirks, and we judged it as distinctly possible that if we did try to roll them out we'd find some other "quirk" was actually a significant stopper.
It's not the same scale, but it's a very similar problem to what would happen if TLS could only sign at one degree of remove like that. If you occasionally take the time to poke through the padlock icon in your browser, you'll find the bare minimum trust chain you'll ever find is three; a root cert in your browser trust store, some signing cert, then the cert for the domain. I've seen some with more layers than that, but 3 is fairly common. Those 3 would be 2 in the SSH context since the last one would be the key rather than a cert. Can you imagine how much rougher it would be on the web if you could only have two levels, if a root cert could sign your site but that was it?
Accept new host key without prompting the usual `Are you sure you want to continue connecting (yes/no/[fingerprint])?` (great when scripting on 100+ servers). Still reject known hosts mismatch (so shity Wifi can't inject their ads). Everything can be setting up again by just removing `~/.ssh/known_hosts` (instead of that GitHub messy blog post about curling and seding output to `~/.ssh/known_hosts`).
The point that mjg59 points out, just not super explicitly and using many more words, is that the confirmation step wouldn't be necessary, if only all of our tooling were just a little bit better. Instead of ssh $ip_address and getting that prompt, instead, you'd do ssh $hostname, where hostname can be generated and contain numbers, and then the same mechanism that checks SSL certificates then also checks if the host's key actually matches what the rest of the system (the configured CAs) is saying it should be. If it does, great! No need to ask the user "Are you sure you want to continue connecting (yes/no/[fingerprint])", with fingerprint being the one thing to check out of band, but few people are really that fastidious.
Competent Security/IT departments are able to get this to work going from corp managed endpoints (laptops) going to corp managed servers. It's just that the rest of the world hasn't caught up yet.
That's not what mjg59 is suggesting, to establish a trust chain to system shipped CAs. In fact that wouldn't be so easy anways because SSH certificates aren't X.509 compatible at all but use their completely homegrown specification. They explicitly say "please do not do this" about this idea.
mjg59 suggests to make the TOFU approach on the client trust certificates if a certificate is provided. That's a way narrower suggestion:
> OpenSSH has no way to do TOFU for CAs, just the keys themselves. This means there's no way to do a git clone ssh://git@github.com/whatever and get a prompt asking you to trust Github's CA. Instead, you need to add a @cert-authority github.com (key) line to your known_hosts file by hand, and since approximately nobody's going to do that there's only marginal benefit in going to the effort to implement this infrastructure. The most important thing we can do to improve the security of the SSH ecosystem is to make it easier to use certificates, and that means improving the behaviour of the clients.
> "Chained" certificates, where the signature key type is a certificate type itself are NOT supported.
https://cvsweb.openbsd.org/src/usr.bin/ssh/PROTOCOL.certkeys...
I pine for a world in which we had made it more approachable for the people who just wanted to quickly build an application without learning about the underlying infrastructure.
I've been dealing with authn/authz for a long time and kerberos is still one of the best protocols in existence.
Now we're swinging back to recognizing the benefits to some level of centralization in authentication.
From a historical point of view, this all seems very familiar.
But the hard part is getting everybody to support this. We need to start somewhere. Maybe GitHub can use this bad publicity and turn this into one good
If you're doing TLS then you might as well do HTTPS. GitHub already supports HTTPS and way more features for it than SSH, and HTTPS works over more networks than SSH does. Continuing to use SSH is literally just being obsessed with a backwards old protocol for nostalgia reasons.
Honest question:
Client authorization (for private repos) is an important feature of SSH.
How does this work with HTTPS?
EDIT: I just remembered the existence of TLS client certificate authentication. I wonder if it's possible to use an SSH client key for this (making the server accept a self-signed client certificate from the list of keys in authorized_keys).
This is a good point actually. It's kinda funny how even Microsoft's own GUI IDE uses an underlying ssh protocol with ssh keys that the end user doesn't even need to see or know about.
Now that we're mostly using ssh to push source code, or connect to remote systems, why not just use tls instead? Aside from its general availability, feels like inevitable.
As it is, I‘d argue that straightforward protocol layer client public key authentication is one of the biggest benefits of using SSH, although I do wish X509/PKI server authentication could optionally be used with it.
I disagree, it seems not straightforward at all.
First as a user you have to generate a key. You need to install software and look up a manual. Then you have to give a server admin your public key, which for many users is often confused for the private key and thus defeats the purpose. Then when you connect to the server you have to manually validate the host fingerprint, which again requires more commands and documentation.
The alternative (in HTTPS) is to use Let's Encrypt on the server, and enable HTTP Basic authentication from either a web server or a reverse proxy. The admin gives the user their login and the client just connects.
With those two features you get 1) strong encryption that your user does not need to manually validate, and 2) a non-bruteforceable authentication token. Not only that, but it's compatible with a thousand other applications and modifiable in a million ways.
I think an ideal solution would offer some hybrid of TOFU and PKI for server authentication, and self-signed keys for client authentication (like FIDO and WebAuthN, for example), but at the protocol level.
HTTPS really falls short here and means that everybody has to implement something custom on top of it, and too often that is a login form driving OAuth.
[1] https://docs.aws.amazon.com/Route53/latest/DeveloperGuide/Re...
https://roumenpetrov.info/secsh/index.html
Roumen has been regularly updating this for a pretty long time now.
Many would say x509 is the real abomination.
However, for all its warts, x509 due to hardware implementations, seems a great deal more secure than sitting on the FS SSH host certificates.
Source: I work for Teleport so a little biased, but love the fact that more and more orgs are going away from static creds/keys and using short-lived certs and/or passwordless solutions. Spent many years in a production role and wish I had these tools (or was aware of these tools) back then.
At the moment, there are too many strong opinions about 100 different "best" ways to do things leaving most of the rest of us just using whatever is easiest to find on Google rather than what the industry has discussed and approved.
> UpdateHostKeys is enabled by default if the user has not overridden the default UserKnownHostsFile setting and has not enabled VerifyHostKeyDNS, otherwise UpdateHostKeys will be set to no.
It's not that useful for actual key rotations, though; the use case seems to be more geared towards switching to newer key types (e.g. adding an Ed25529 key seamlessly) than replacing a potentially compromised key. At least I couldn't find a straightforward way for the OpenSSH sshd to actually provide two keys of the same type to my client.
All in all, I‘m not a fan of the feature – it seems to complicate key security for pretty marginal benefits.
That said, it did auto-upgrade the Github key for me as far as I can tell!
Can someone elaborate why? We are already depending on lists of known good CAs for everything (even banking). Why not leverage the same for SSH as well?
Wire-level network protocols like Wireguard are somewhat useful, but largely a large step away from the modern best practices. We need more Zero Trust, Federated Identity, Fine-grained Access Control, and Least Privilege. Those solutions exist, but they are almost always for-pay, because SSH is always used as the default option, and so no more effort is put into better security practices.
And even without putting any real effort into a better protocol, you can just implement on top of HTTPS. Look at HTTPS+Git, compared to SSH+Git. First and foremost, this RSA key leak bullshit just wouldn't happen. Even if the TLS key on the server got leaked (and why the hell would it?! it gets generated automatically by Let's Encrypt), revoking and issuing a new one happens automatically for the client with no fraught extra options, and nobody's client configuration by default disables validating the certs! Then there's the fact that you can use a variety of AuthN+Z options, it goes over standard ports, and most providers give more fine-grained access control for it.
Nerds love SSH. But it is literally worse than the alternatives, and is in practice often not used in a secure manner. Kill your darlings and use something demonstrably better.
Yes! Please invent something else, and leave SSH alone.
Plenty of "neckbeards" understand containers and probably understand the underlying technologies (cgroups, etc) as well as comparable (or historical) approaches (jails, zones, etc) better than most.
But, because something doesn't cater to your whims, must complicate things that actually work.
What's a standard port? AFAIK 22 is a standard port. https://en.wikipedia.org/wiki/List_of_TCP_and_UDP_port_numbe...
Are you saying that everything should only go over 443 (=HTTPS)?
SSHv1 is indeed an outdated protocol but nobody uses it anymore.
What alternative is there to SSH? You want to go back to telnet?