This incident was complicated by the fact that it was first noticed while investigating a particular CA (DarkMatter) that was already suspected of various shady activities. Most of the other affected organizations took the incident seriously, fixed the problem and revoked the affected certificates, as required by the letter of the rules, even though the actual risk was small.
DarkMatter chose to spend a lot of time complaining and making weaselly arguments on the mailing list (like "a 64-bit number where the leading bit is always zero is still technically a 64-bit random number") which did not do much to convince anyone of their trustworthiness.
> So, when I would walk backstage, if I saw a brown M&M in that bowl … well, line-check the entire production. Guaranteed you’re going to arrive at a technical error. They didn’t read the contract. Guaranteed you’d run into a problem. Sometimes it would threaten to just destroy the whole show. Something like, literally, life-threatening.
"I found some brown M&M’s, I went into full Shakespearean “What is this before me?” … you know, with the skull in one hand … and promptly trashed the dressing room. Dumped the buffet, kicked a hole in the door, twelve thousand dollars’ worth of fun.
The staging sank through their floor. They didn’t bother to look at the weight requirements or anything, and this sank through their new flooring and did eighty thousand dollars’ worth of damage to the arena floor. The whole thing had to be replaced.
Otherwise, "smash therapy" places couldn't exist.
If I were writing contracts for rock stars, I'd probably look to account for tantrums and raging parties in the wording.
It's easy to say that the brown M&Ms acted as a signal, but I can't think of any way to depend on that signal that isn't reckless. Just check the actual production, since that seems to be an option.
But, rock stars are not going to arrive early every night just to stay alive. The buck has to stop somewhere, and EVH decided it was in the bowl of M&M's.
For one, the same people in charge of safety like riggings and electrical are not going to be the same people that are supplying the refreshments.
I can also imagine a high number of people like myself reading that part and think it's a joke or purposely ignoring it because there are much more important issues to deal with and because I'm not much into indulging the ridiculous whims of self important people.
It doesn't say "no brown M&Ms;" - it says "no brown M&Ms, otherwise we cancel the show and you still pay us the entire amount".
If I saw something like that in a contract, maybe I wouldn't go looking for brown M&Ms, I'd at least ask my boss (who hopefully has a direct line to our lawyers) "is this for real?", and if my boss and our lawyers are any good they'd tell me that's as real as it gets.
Then you probably wouldn't be working at a concert venue / on a showbusiness production team anyway.
Edit: have you seen a Van Halen (or, really, any stadium-headlining) show? The entire performance is "the ridiculous whims of self important people". That's kind of why we like it.
So, the logic of this signal is like this:
- If a brown M&M HAS been found, you know for a fact they didn't follow the specs well enough;
- IF a brown M&M HAS NOT been found (denying the antecedent), you can't state anything about they following the specs or not.
In my experience it’s a pretty good heuristic. It fails badly if the organization knows what pieces will be inspected (queue Goodhart). It can also be fantastically inefficient if there’s a ton of checks or the checks are ambiguous.
Being a stickler at inspections sets a useful precedent that organizations will do a better job for you in the future.
Napoleon took advantage of this during his first inspection of a combined arms unit (Napoleon was an artillery officer). He inspected the shit out the artillery unit first. Then he could be reasonably confident the light infantry and Calvary units would take the inspection seriously.
The Paris MOU randomly inspects vessels entering European ports with the rate of inspections based on their flag. So if you're a UK registered ship, you might expected to go several years without inspection, while if you're Albanian registered, Albania is blacklisted, you're getting inspected twice a year or more, and if you fail several inspections you're getting banned from all European ports.
How did Albania get blacklisted? The Paris MOU tracks the results of inspections. Each time a ship flying some flag is inspected, the result of the inspection changes the calculation of how much apparent risk there is of ships with that flag failing inspection. Risk too high? Greylist. Risk higher still? Blacklist.
Why flags? Because before Port State Control existed, the Flag States (Country where the ship is registered) were responsible for periodically inspecting ships. And they still are, it's international law, the Europeans just got sick of all these ships with a "flag of convenience" from some distant island that are clearly non-compliant and have never been inspected in their nice well-regulated ports. So they invented the Paris MOU and Port State Control.
And it worked, I mentioned the UK flag is whitelisted, but the Bahamas are a famous Flag of Convenience, and they're whitelisted too, because they took the job seriously. A Bahamas registration is a little cheaper†, but your ship will be properly inspected on their behalf, and so sure enough when it gets randomly inspected again in say, Antwerp, it'll likely pass. Meanwhile St Vincent, another Flag of Convenience, is on the Greylist and a few tiny islands that dipped their toes in the "Flag of Convenience" business are blacklisted. The promise of inspectors poking around inside your ships every five minutes is a bad deal even if Tuvalu shipping registration is half price.
† The other thing we don't like about Flags of Convenience is they make it easier to hide who owns things, which can serve to also avoid taxation. The Paris MOU doesn't fix that, that superyacht registered in the Bahamas might be safe but the owner likely didn't pay tax...
A big-name traveling band is kind of like a carnival without the ferris wheels. They have a lot less support staff than you probably think, and most of them are busy most of the time you're not driving.
If you have more than one person playing management/problem solver/coordinator/process lubricant, you've got it easy. Much more typical is everyone heads-down on their part of the event and one person trying to hold it all together.
That person has no free bandwidth, and depends heavily on proxy measures of the state of things.
I've played roadie before, not for a super big act and not for long. And I've ended up running point for a lot of conferences, more or less because I know how to run that sort of thing. Conferences are so, so much easier - they have resources, can assume workers aren't wasted, can assume the location at least vaguely cares about things like fire suppression, can assume the venue doing a revenue split actually has a license to sell alcohol...
I think that's what gets me. I'm pretty "Old Man Yells At Cloud" right now, criticizing an 80s metal band for not meeting the safety standard that I've imagined. It's just that people pull out this idea as being something neat and I think it's a completely inappropriate technique when other options are available and safety is on the line. If either of these things aren't true, then sure, try it.
You can't see how people follow orders in war without sending them to war. So you tell them "no food in the barracks" and when you catch them with a bagel you make them do a hundred pushups. And it's easy to toss out resumes with typos since they've indicated they don't actually have attention to detail. Those make sense because they're the best you can do.
But if someone dies at your concert, you'll probably think it would have been worth it to have more than one person playing safety checker. The contracts and M&Ms aren't actually going to make you feel better. (Or necessarily keep the lawyers away).
At its deepest, I think I'm reacting to a change that happened to me over time. I used to work at a ropes course. Think team building, high wires and zip lines. We'd give a list of instructions to our group contact to pass along to the group, including that people need to wear long pants and closed-toe shoes.
Well, about 20% of groups would hop out of their cars with a bunch of people in shorts and sandals. "Oh, no one told us." And we'd announce to ourselves that the group contact had failed to communicate and we should "watch out" for the rest of the day.
Looking back, I think we were the dumb ones. If 20% of groups are showing up unprepared to my ropes course, then that's my problem. I need to make different choices that lead to better outcomes. If I have evidence I shouldn't trust group contracts to make things happen (I definitely do) then I need to stop counting on them.
But we do this with all things all the time. You didn't take a ruler/micrometer and measure tolerances between parts of the last car you bought, nor the last cab/rideshare/friend's car you got into. Instead, you make some assumptions and use some cues to inform those assumptions.
Perhaps you're about to get into some ride share vehicle and you notice the exhaust looks a little smokey for the age of the car, and think maybe they aren't taking care of it. Then maybe to take a closer look at the tires and notice they're bald, etc. This is definitely something that would be considered "literally life-threatening", but there's a statistical likelihood that makes it something that not everyone checks every time.
> But if someone dies at your concert, you'll probably think it would have been worth it to have more than one person playing safety checker.
But it's not "your" concert, as much as it's billed that way. It's a joint operation between the artist and the venue, and each have their own responsibilities. It's the venue's responsibility to do certain things. This is just the artist trying to use one technique of many that are likely employed (like outright asking) to gauge how well the other party has fulfilled their responsibilities. You can't check every single thing, otherwise there's no point in there being another party (and they may not give you that access to check), but there are things you can do let you know a closer look is warranted.
> Looking back, I think we were the dumb ones. If 20% of groups are showing up unprepared to my ropes course, then that's my problem.
To some degree, maybe. I think the best outcome would be both cases, "look out" when stuff looks awry, but also examine why it doesn't seem to change. But what if you can't eliminate the problem? What if, no matter what you iterate on and try to get the contact to correctly relay your information, you can't always get he info to all the relevant people either yourself or through the contact, and 5% still show up like that? Do you stop with the "watch out" notice? No, most likely you still do that because it's useful and better than nothing, and still provides a little benefit on top of everything else that was done.
That's what we should assume the rock band was doing. They have contracts that stipulate how stuff is supposed to be, and different staff to coordinate with the venue reps for aspects of the production, and they should be doing what they can to make sure stuff is set as expected. But if brown M&M's show up, maybe it's worth paying those people a little overtime to grill the venue on what they did and didn't do that they said they did and check their work, because who knows, maybe this is the time it saves a life.
On your rideshare example, I don't do a preflight check when I ride an airliner. I did when I was a pilot, even if it had just come out of the shop.
> But it's not "your" concert, as much as it's billed that way.
This is it. This is the difference in mindset. I now think of it as "mine". Not that it's mine and no one else's. But I choose to take responsibility for anything that went wrong that I could have been prevented. I don't care how the responsibilities are divvied up or what the contracts say. So anywhere I'm in some position where people are counting on me, I'm either going to check things myself or know that someone I trust did the checking.
It’s a signal to a potentially larger issue.
> This is it. This is the difference in mindset. I now think of it as "mine". Not that it's mine and no one else's. But I choose to take responsibility for anything that went wrong that I could have been prevented.
That's fine, and you can take responsibility, but you can't do all the work, not in any way that scales. Even in the preflight checks for the plane, you're not disassembling wings and checking for cracks on internal struts I imagine. Someone else does that occasionally and you have to trust (or not) their opinions and that they've actually done the work. And you can't always spend a week doing a thorough background check of those people either (or for some reason you're forced to use someone you would rather not), so sometimes little tricks that might indicate that you should be wary are useful.
> So just as a little test, in the technical aspect of the rider, it would say “Article 148: There will be fifteen amperage voltage sockets at twenty-foot spaces, evenly, providing nineteen amperes …” This kind of thing. And article number 126, in the middle of nowhere, was: “There will be no brown M&M’s in the backstage area, upon pain of forfeiture of the show, with full compensation.”
So they had other simple-to-check signs; the brown M&M one is just the only one commonly cited as it's the funniest/most interesting.
So what I think what the contract was specifying was several 15-amp receptacles, spaced 20 feet apart, from which a combined total of not less than 19 amps may be drawn. A 20 amp breaker would do that nicely.
Furthermore, the brown M&Ms clause was buried deep in a list of technical riders. The person in charge of the green room isn't reading the technical part of the contract at all. Why would they? It's not their job. They're reading the other part of the contract that plainly says "please put a bowl of M&Ms in the green room". So the only two cases in which a venue is going to pick out the M&Ms is...
1. The technical director saw the requirement and told the green room guy to start picking out M&Ms or the show is cancelled
2. The green room guy somehow knows about the brown M&Ms thing and decides to falsely signal compliance
Presumably, the whole brown M&Ms thing was obscure enough that no venue actually decided to just check for odd requirements. Furthermore, if they did decide to just do that and only that, specifically knowing that they were interfering with such a heuristic, they'd almost certainly be liable for something. So I doubt anyone would deliberately interfere with this, and it's hard to accidentally comply with just the no brown M&Ms thing.
"Just line-check the production all the time" might not always be an option.
> The Zlint project only attempts to check whether a certificate exceeds the maximum validity allowed by the baseline requirements [398 days], and is not configurable.
To extend the 'no brown M&Ms' analogy, this would be like Van Halen learning the caterer was only checking for FDA safety requirements and not their specific needs. It's good that it's food safe, and a brown M&M (or one with a fleck of brown exposed) isn't going to hurt anyone, but it means that you've got a breakdown in your process.
Specifically, Let's Encrypt needs a review of their linting process. The requirement was that the lifetime be set to 7775999 seconds, or 90 days, this particular tool may not have thrown an error if the time had been 180 days or 360 days, and instead only thrown an error at 399 days. One second is not a problem, but 300 days would be a problem for sure! How were the lints which they're running selected, in particular how was the rule 'e_tls_server_cert_valid_time_longer_than_398_days' accepted, knowing that their validity was shorter than that? Are there other rules where the linter is checking against the most liberal specs but Let's Encrypt is offering something more precise? Fuzzing against a validity clock of 7775998 seconds, 7775999 seconds, and 7776000 seconds would find this issue, are there any other parameters that can be fuzzed? Questions like this need to be asked.
I see it more as a "no brwn M&Ms" violation.
We would all assume the intent was to say no brown M&Ms, rather than imagine that no action was required.
They really do _not_ want to be seen to be giving Lets Encrypt special treatment because of the connection between the CAB members and Lets Encrypt and the threat that Lets Encrypt poses to the business models of the traditional CAs.
So a violation of the rules by Lets Encrypt is being dealt with by the letter, even when the issue clearly does not warrant that level of reaction.
(See also the time Google Search banned Google Chrome from results for a month when one of their advertising campaigns breached the rules for paid SEO)
We've lived with certs that were valid for a second longer than they should have been since the inception of Let's Encrypt, and three months won't kill anyone.
edit follows:
When they revoked 3% of their certificates, not all of them were able to renew in time due to physical server limitations. The renewals required server administrators to forcibly renew their certificates, and the email address associated with the certificate was contacted to let them know. It would be an unmitigated disaster if 185 million certificates were suddenly revoked.
> In order to get all those certificates replaced, we need an efficient and automated way to notify ACME clients that they should perform early renewal. Normally ACME clients renew their certificates when one third of their lifetime is remaining, and don’t contact our servers otherwise. We published a draft extension to ACME last year that describes a way for clients to regularly poll ACME servers to find out about early-renewal events. We plan to polish up that draft, implement, and collaborate with clients and large integrators to get it implemented on the client side.
Running certbot daily will currently do nothing. It won't think that the certificate needs to be replaced since it has not yet reached the limit required to renew.
Yep, this. Clients need to be built to support it.
I know this is from the linked article and not from you, but:
> In order to get all those certificates replaced, we need an efficient and automated way to notify ACME clients that they should perform early renewal.
This is already possible. It's called OCSP stapling, and it's what Caddy (and CertMagic) does by default, automatically. When Caddy sees from OCSP that the certificate is revoked, it will automatically replace it.
I think the desire here would be for a mechanism to alert clients to obtain a new certificate before their current certificate is revoked and becomes invalid.
edit: formatting
A certificate that is going to be revoked is as good as revoked. There is no "almost untrusted, but not quite yet" gray area (unless you're talking about expiration dates, which some browsers allow leniency on; but we're talking about revocation, where we know there was a problem or misissuance, whereas expiration is mainly a passive safeguard against indefinite trust).
So, once a client sees a "Revoked" OCSP status, it can replace the certificate immediately, before the previous, valid OCSP response expires.
Google's current policy still seems to require them to use 1 Google log, and 1 other non-Google log given the lifetime. They seem to use 1 Google log, and 1 other random log, including their own.
I'm doubt the logs will be able to keep up.
The other CA (KIR S.A.) actually issued certificates with a validity of one year, so, for them, dodging this is not that easy.
KIR S.A. is another CA that issued certificates with the same one second issue (1 year + 1 second instead of 1 year) and reported that one month ago.
In practice, it's a small problem and revocation is an excessive response, but those are conclusions you reach after duly investigating the issue, not before.