Microsoft Teams outage due to expired certificate
techcrunch.com
techcrunch.com
https://gfycat.com/fortunateyawningcopperhead
My router is blocking this domain on DNS level in OpenWRT, techcrunch.com is redirecting me there, so I can't visit this page.
Edit: I literally can not visit any techcrunch.com article, all of them redirect me to this doggy ads+tracking domain. It doesn't matter if I came from google, DDG, reddit or HN.
No it isn't.
GET https://techcrunch.com/2020/02/03/microsoft-teams-has-been-down-this-morning/
HTTP/2 307 Temporary Redirect 157msThe redirect you’re getting is to a “consent” page that’s not actually compliant with the GDPR as it takes a good 5 minutes to opt out of all the bullshit tracking and even then I don’t trust it to actually opt me out and not track.
Yahoo/Verizon is cancer and should die in a fire.
It's a feature, not a bug.
Now I avoid TC at all costs, it's contaminated with vile ad trackers.
More recently, it happened with Ericsson: https://www.theverge.com/2018/12/7/18130323/ericsson-softwar...
This article has some information about how Let's Encrypt enabled an "automated process that handles renewals": https://duo.com/decipher/proposal-to-make-https-certificate-...
I wonder if such a process should be made an industry standard? Does anyone know if there are any proposals for it?
https://en.wikipedia.org/wiki/Automated_Certificate_Manageme...
ACME would be ideal, but the official response of Let's Encrypt is that ACME is overkill for corporate environments and you should roll your own certificate automation.
At one place we aggressively policed external facing certificates. Don’t follow the process, your service gets whacked.
It’s a process you should look into, because the compliance regimes will start paying attention to it someday soon.
In private PKIs either proof of control over the name is considered an out-of-band problem or it's elided altogether. I wrote about this when Peter Gutmann was being angry about ACME years ago, Gutmann saw ACME as redundant because of existing protocols like SCEP. If you want certificate issuance automation but you don't need proof of control over the names then you don't need ACME - SCEP (or half a dozen others) are fine for this purpose and you should use those.
Now, of course if you don't want proof of control but you do want automation it's worth taking a moment to reconsider exactly what your goals really are: What is a certificate doing for me here? What exactly am I certifying, if anything? But if the technical requirement is there, regardless of whether it makes philosophical sense, SCEP delivers and ACME is overkill.
As I wrote in my thread with Peter, it's usual with things like SCEP to provide a default implementation in which the part where you only give certificates to the Right People is left marked /* TODO */. Thus often the result is pure security theatre, certificates are issued equally to bad guys or good guys without distinction and nothing at all is really being certified.
Manual processes may or may not be better in terms of actually certifying anything.
Even with letsencrypt you still MUST have a monitoring system, notifications going to the right person, and in general an organization that can act on this. The problem is often not technical: your organization must be structured in a way that those notifications are acted upon. If anything, letsencrypt lessens the frequency of those notifications so I posit that it’s even worse from the point of view of validating your organization: with regular certs, you get notified once every year or two.
Anyway, letsencrypt is a good thing: it’s just not a solution to this problem.
But in a non-technical organisation, who should those messages go to?
Often the initial LetsEncrypt setup will be handled, correctly, by some IT staff. Then it might break several months or years later for some odd reason.
The organisational challenge is to get the message through to someone who understands it and will act on it.
Yes, and to keep a very infrequently used channel working and up-to-date.
Hence efforts like https://tools.ietf.org/html/rfc8701 - designed to make sure extensibility options don't get messed up.
This time the problem was LE all of a sudden decided to start storing my certificate in a directory called mydomain.com-0001 instead of mydomain.com, breaking the rest of the setup that relies on things being in the right directory. Automation is only useful when the software behaves predictably and consistently.
So if renewal fails you should have ~30 days to fix it.
But this does work best if your tools try the 2-month renewal. (I'm looking at you, wpengine!)
But as a person who rolls LE certs across a very non-happy-path environment (many SAN domains, edge nodes which are geo-balanced).
I have a lot of issues automating this process, right now I have a HTTPd which reverse proxies the .well-known back to a central place where I run certbot and then I push out the cert to the nodes, however, sometimes one of our SAN domains will need to be removed and the whole universe comes crashing down.
DNS-01 Challenge is "nice" (although, doesn't feel super well supported); but requires domains registered with some kind of DNS server that accepts API's to change records, so Amazon route53- but it's exceptionally hard to roll your own DNS in this case. :\
1. HTTP-01 challenges have a "correct" answer for a given Let's Encrypt account which depends only on the challenge (ie the part in the HTTP GET request) and on knowing which account you want to use.
Silently Certbot creates you an account, with a private authentication key and so on, for the Let's Encrypt service. When it gives you a file to prove control by placing it in /.well-known/acme-challenge/ the content of the file is always the same as the filename, plus a suffix that depends on your key.
So long as you use the same account you can thus bake this suffix into the web server, essentially causing it to answer any request from anybody: "Hey, who is allowed to issue for somename.example ?" "dijit is allowed to do that". Bad guys can't use this because they don't know your private account key, but for you now magically everything is authorised, since when it is asked your server will answer "dijit is allowed to do that" to any question it's asked.
2. DNS-01 can be redirected using CNAME. Add a CNAME, once, manually if necessary, to redirect the DNS-01 checks to a DNS server you've set up for this specific purpose.
I'm looking for a program that is going to connect to all my SSL sites, and report back a problem if the cert is within 30 days of expiration. It is easy enough to write a client that will fail after the cert has expired, but I want one that will warn me ahead of time. And I don't want to mess with the system clock or something like that.
It shouldn't be a big program, I may just write it myself.
Also my employer (well, until the end of this week) Kynd does this as part of its broader "Check cyber-security stuff" offering for non-technical people. https://www.kynd.io/
define command {
command_name check_certificate
command_line $USER1$/check_http -H $ARG1$ -p $ARG2$ -4 -S -C 21 -t 20 --sni
}
A bigger challenge is getting a complete list of your https websites, and an even bigger one is finding and monitoring all those non-https certificates, eg payment gateway certificates.is totally free, superfast and checks your TLS settings too (it's my baby)
Challenge is discovering all the certs that exist across the org (10k+ often) and then having a fully automated renewal process requires a pretty advanced / complex platform - don’t know anyone who has built this tech successfully in house - curious if anyone has though in the HN community?
Disclosure I’m an investor in AppViewX
Has anyone even felt that Teams is a heavy app that consumes a lot of time to come alive?
Even during calls, the quality is horrible that I don’t even want to describe the pain I go through. There’s strong distortion and voices will never be heard clearly.
I don't think it's possible for Slack to be better enough to justify that kind of a price difference.
There’s a famous saying “no one gets promoted for fixing bugs”. I imagine it’s why Google has 10 messaging products that are half assed too.
It’s really hard to say No and focus on making the best essentials. It’s something that even Apple is having trouble with.
I think this is where Startups sometimes shine and why Slack is here to stay if they don’t also fall in the same trap.
It seems like it's a clone made just to get into this space and everything looks like a generic attempt to throw money into an app to take on Slack.
Whenever Teams comes up on HN, there are people complaining about how much infinitely better some cooler product is. I don't doubt that you see a huge difference, but I feel like I'm listening to a wine expert explain how some $1000 bottle of wine is better than a $20 bottle, when to me they taste the same.
My main worry was about voice and video quality as I’m full time remote, but it’s more or less perfect I think. Haven’t had a disconnect or poor audio/video.
The screen sharing does disconnect some times which is annoying to say the least.
I just don't like the fact I can't control my notifications and chat history well, probably because it's enterprise. There's also a creepy 'get notification when colleague is online' feature that I can't turn off in my privacy settings. Otherwise quite nice.
I interface with too many people, it has brought the dumpster fire of sharepoint to a chat client. I’m now enrolled in 96 teams, 94 of which have no activity.
I think there is a universal law of chat client entropy. All chat clients only get worse over time. The category hit peak UX in 1999 with AIM/Yahoo, and gets worse every day. Eventually, Teams will merge into exchange, and your chats will be specially formatted emails, stored in Sharepoint.
I would think at this point they would be their own major certificate authority and maybe domain registrar.
From experience this probably wouldn't fix things.
What often happens is that somebody creates a system that uses a certificate, doesn’t automate renewal, and then the person responsible for renewing it changes teams or leaves the company. Email reminders only go so far—they not only need to go into the right inbox, but the person watching that inbox has to care.
If it's in production, just buy a 10 year cert. This virtually guarantees an outage after 10 years but virtually guarantees it won't be your fault when it happens...
I really like that Let’s Encrypt certificates last only 90 days.
That said, the CYA aspect of “10 years, somebody else’s problem” is really appealing. If only I believed that it wouldn’t be my ass on the line 10 years from now!
New certificates in the Web PKI ("SSL certificates") have a maximum lifespan of 825 days. This is enforced (if a CA were to issue a certificate with a longer lifespan Chrome for example would just treat this certificate as invalid). The commercial CAs mostly offer one year or two years, with renewals using the 825 day limit to offer renewals in the overlap, so e.g. you buy two years in June 2018, in April 2020 you can pay for two year renewal and the new certificate expires in June 2022 not April 2022.
If you're using certificates in your own PKI (as it's likely Microsoft actually was in this particular incident) then there's no need to buy them and it's up to you what your appetite for risk is on when they expire.
I honestly thought this was why Teams was down for me.
For future reference, it was probably this one: https://news.ycombinator.com/item?id=22222121
Is there a reminder service out there that specializes in your long-term expiring things? I'm not sure what would be different about it than a regular calendar, but it seems like many of us need something that makes this easier.
It would be nice if they prompted you to 'add to calendar' when you are creating certs.
That would actually be a really _awesome_ integration of ACME client software with CalDAV [0] and other ticketing software APIs like Jira: set up calendar events and/or tickets with increasing priority to verify that the certificate is updated.
At larger companies a lot of the issue isn't literally generating reminders, it is making sure they're sent to the correct people/departments and are actioned by anyone.
For example you sometimes have reminders sent to ex-employees, or sent to a mailing list and everyone assuming everyone else is going to action it. Or the reminder gets ping-ponged between multiple managers via email with nobody either able or willing to deal with it.
None of these are tech' issues, and they don't have technology solutions as a consequence. So whenever I see an embarrassing expired cert, I don't assume technical malfunction, I assume political malfunction.
import ssl
import OpenSSL
from datetime import datetime
cert = ssl.get_server_certificate((www.google.com, 443))
x509 = OpenSSL.crypto.load_certificate(OpenSSL.crypto.FILETYPE_PEM, cert)
expiry_dt = datetime.strptime(x509.get_notAfter()[:8].decode('utf-8'), '%Y%m%d')Also, no need for python:
echo | openssl s_client -connect google.com:443 -servername google.com -showcerts 2>&1 | openssl x509 -text -noout | grep "Not After" | cut -d : -f 2- | xargs -I {} date -d "{}" +%Y%m%d
tcp_client = TCPSocket.new(domain, 443)
ssl_client = OpenSSL::SSL::SSLSocket.new(tcp_client)
ssl_client.hostname = domain
ssl_client.connect
cert = OpenSSL::X509::Certificate.new(ssl_client.peer_cert)
ssl_client.sysclose
tcp_client.close
certprops = OpenSSL::X509::Name.new(cert.issuer).to_a
issuer = certprops.select { |name, data, type| name == "O" }.first[1]
results = {
valid_on: cert.not_before.utc,
valid_until: cert.not_after.utc,
issuer: issuer,
days_left: (cert.not_after.utc - Time.now.utc).to_i / (24 * 60 * 60),
}I appreciate the sentiment but I think it's fine for them to just say "we examined our processes, found out what led to the issue, and have modified procedures". I worry a detailed post mortem would just throw specific folks under the bus (most likely some low level employee who isn't actually the root cause)
if free, simple and it checks your TLS settings too
The turnover insures that nobody in the department was there when the process was started/last interacted with, and so it is off the collective organizational radar so to speak.
Historically it was common to issue 3 year certs, and five year certs weren't rare (until 2015). But whilst it's reasonable to expect microsoft.com or bbc.co.uk belonging to the same outfit in five years, it's hard to be as sure about say jsnes.org (currently a Javascript NES emulator) or catandgirl.com (a web comic by Dorothy Gambrell) which might well entertain offers from somebody else who wanted those names.
The underlying domain name is typically on an annual renewal cycle with perhaps just 14 days grace if you stop paying, and individual FQDNs might have even shorter turnaround. With a five year certificate this means you could buy a certificate the day before your renewal payment is due, and then still have an apparently good, working certificate for that name five years later when it's owned by somebody else entirely who has no idea you once owned that name. Not great. Let's Encrypt's renewal cycle closes this gap considerably. The BRs were also amended, the limit is now 825 days instead of 39 months or (originally) 60 months.
You work in an industry whose entire purpose is to automate such things.
alert: TlsCertExpiringSoon
expr: (probe_ssl_earliest_cert_expiry
- time()) < (86400 * 14)
for: 10m
labels:
product: Name_of_Product
severity: page
annotations:
description: the tls cert for the URL {{ $labels.instance }} expires in less than 14 days!
summary: TLS cert for {{$labels.instance}} expiringNot like it matters, but it kinda does, because those tend to be private and internally generated, and not necessarily signed by an external certificate authority.
You have to rely on either the code itself checking each time it uses the certificate, and alerting.
Or (taken from elsewhere in this thread) you test it during your build, and hope that someone is still building the code by the time the cert comes up for renewal.
I'll probably be doing both.
~# dig -t A +short www.certera.io
certera-io.github.io.
185.199.108.153
185.199.110.153
185.199.111.153
185.199.109.153
~# dig -t A +short certera.io
185.199.108.153
185.199.109.153
185.199.110.153
185.199.111.153
Looks like https://www.certera.io is going to github of which is only returning a cert for itself, and not his domain name.https://www.certera.io fails the certificate check.
It's a good example of the difficulty of getting TLS perfectly right.
In theory this set up is fine; the default behavior of all the browsers when typing "www.certera.io" is to interpret it as a request for http://www.certera.io.
But if the client has anything in place that automatically upgrades http to https before submitting the request, you're going to need a valid cert for the www subdomain in place or you'll throw a cert error before reaching the redirect.
Even if your site omits the www subdomain in production (as certera does), a lot of users will just type it in anyway. So, you better be ready to handle that request via https.
$ echo | openssl s_client -connect "www.certera.io":443 -servername "www.certera.io" -verify_hostname "www.certera.io" 2>/dev/null | openssl x509 -noout -issuer
issuer=C = US, O = DigiCert Inc, OU = www.digicert.com, CN = DigiCert SHA2 High Assurance Server CA
$ echo | openssl s_client -connect "certera.io":443 -servername "certera.io" -verify_hostname "certera.io" 2>/dev/null | openssl x509 -noout -issuer
issuer=C = US, O = Let's Encrypt, CN = Let's Encrypt Authority X3https://github.community/t5/GitHub-Pages/Does-GitHub-Pages-S...
Hopefully when I make some money I can move to a hosted setup where I can control it all.
Any thoughts on the license? How is it working? Why did you pick that? I like that type of license, but it's not very common. Drone does it too, but I haven't seen many others. You don't have to answer if you don't want to, but it's nice to see people deviating from standard licenses like GPL and MIT since I feel like those make it too easy for large businesses to take advantage of small projects.
Your licensing and attribution pages look like a lot of thought went into it, so you probably have some decent insight.
I haven't been marketing at all, and I just recently finished the first stable release, so the jury is still out on whether this is all a good idea or not!
The docs are based on ReadTheDocs, but settled on a single file layout instead of having multiple pages.
Jokes aside, I don't understand how this problem hasn't been solved in the general case.
isn't that what ACME is suppose to do?
https://en.wikipedia.org/wiki/Automated_Certificate_Manageme...
Happens all the time, even if ACME is employed, and it's unlikely to ever stop happening.
Initially, I made light of a requirement that we send out automated nag emails starting a couple months before the keys expired. With a bit more time and observation it became pretty clear this was a valid concern. I eventually came around, and while I didn't implement the feature, I did create the integration tests.
Making certs is stupid simple. Maintaining them takes some support and we don't always have it at the ready. Remembering to do something once a year or two years isn't something we're particularly built for. We have a habit of forgetting these duties when we hand projects over to others.
I wonder, has Lets Encrypt solved this?
Would be interesting what CA they used for it and if it's a SAN certificate.
Edit: here's the certificate log of the teams subdomain but I couldn't find the one that expired today in it https://crt.sh/?q=teams.microsoft.com
All certificates in the Web PKI are obliged to use SAN (Subject Alternative Name). Although it is often mistaken for some sort of aliasing feature, SAN is the Internet's agreed alternative way to name things. X.509 is intended as part of the X.500 directory system, a global directory system which obviously was never actually built, and so its built-in name scheme doesn't resemble anything that actually exists.
When Netscape invented SSL back in the 1990s they hijacked X.509's Common Name field to write DNS names in, but this field is just defined as arbitrary human readable text. "news.ycombinator.com" is text, but so is "News·YCombinator,coM " and only one of those is a DNS name. So, when the IETF standardised PKIX it designed a dedicated schema for the various types of Internet names to put into certificates, Subject Alternative Names SANs.
Unlike the Common Name a SAN dnsName literally can't be anything but a DNS name using A-labels ("punycode"). So the opportunity for confusion is removed. Since it doesn't need to address universal human text it doesn't need to support weird encodings or anything else.
When PKIX was standardised the existing abuse of Common Name was grandfathered in. PKIX says all certs should list one or more SANs but they may continue to have a DNS name in the Common Name too. The Baseline Requirements, years later, explicitly tell CAs to only use a Common Name which matches one of the SANs they've also baked into the certificate, if there are no SANs they're doing it wrong. Alas, as so often, commercial priorities beat security and so even last decade it was common to find people trying to get away without SANs. However the roll out of Certificate Transparency allowed us to have a clear view, without waiting for incident reports from affected users, of non-compliant certificates still being issued, so we could address it as it happened. A few years ago Firefox and Chrome (and Safari I think?) were able to remove their code for trying to process the human readable Common Name as if it might be important, and so if anybody were to issue such a certificate with no SANs today it wouldn't "work" in popular browsers.
Heck - googling - https://tecadmin.net/auto-renew-lets-encrypt-certificates/
Now, with Let's Encrypt, there's no excuse to not SSL.
Even if your site is some static video game guide for Club Penguin, an attacker can inject some 0day, some privacy-invading analytics code, or even a dumb alert("Your windows is out of date").
All of them. Script jacking, ad insertion, redirection, tracking insertion, etc. are all done at scale by everything from national ISPs to coffe shop routers.
HTTPS provides authenticity of all transmitted data; this is more important than confidentiality because without authenticity you can’t tell that you are talking in secret with an atttacker.
apt install ssl-cert-checkIt's barely usable.
That's how brands are supposed to work, "I've had other things from this company, they were good, I'll get more".
Companies are usually designed to prevent single shitty employees from ruining things (and usually to prevent goods being "too good" as well!).
I've heard that MS is structured as highly separated departments, like separate single-product companies. Which goes someway to explain things - if they're bad at sharing best practice and senior management can't/won't control quality.
Personally, this seems highly interesting.