779 karma · joined May 11, 2011
From the inside, it’s felt like steadily rolling up a larger and larger katamari ball, if you remember that game.
We have a very performant prototype!
One usde case I am excited about: taking the OpenTelemetry C++ SDK, and binding it to the OpenTelemetry interfaces in other languages, such as Ruby and Python.
Some runtimes may see a performance boost by running the observability code independently from the GIL, GC, etc.
You could shift the model around and maybe have a “station attendant” rather than a “driver,” but it’s an issue that I don’t see addressed in many autonomous conversations. The bus driver doesn’t just drive the bus, they also regulate and deal with all the bullshit happening on and around the bus.
Automation-in-public has so many issues beyond it’s closed-course equivalent, where you simply have to drive the vehicle. It will be a much bigger adjustment than just high quality autopilot.
For me, the question is not about whether high quality projects are going to be okay. The question is: on what basis do I believe there has been sufficient oversight to guarantee the soundness of new construction?
I can remember talking to SPUR members in 2001 about how problematic it was to extend downtown development into SOMA, etc, due to liquefaction and uncertainty. When the building craze hit, it felt like it brushed these concerns aside, rather than answer them. Now we have leaning buildings that don’t meet basic construction requirements. It doesn’t inspire a lot of faith.
For example, handguns and other concealable weapons were almost non-existent in Australia when I was living there. As a side effect, there were a lot of fistfights; far more than I've ever seen at equivalent bars/pubs in the US. That sucked, but I still found it preferable to being shot/mugged at gunpoint.
Likewise, acid attacks are terrible, and there are possibly more of them since guns and knives are less accessible. Doesn't mean that overall things are worse because of the ban.
And as for 2., I recall a hilarious incident with an email alert turning a deployment into a spambot, so... :)
Meanwhile, the crazy code that panicked is still running it's other goroutines - remember, you called recover! - so maybe now the webserver still has an open port and is allowing your users to access whatever strange state is left inside... gonzo things really can happen if you let a corrupted program stay on rather than shutting it down immediately. Even if nothing bad is happening, you still are out of commission for that entire period.
The point is, murphy's law always comes into play. So if we're talking about production best practices, consider that "most likely fine" means "definitely not fine at scale over time". Just make sure that whatever you're doing during shutdown can't block.
1. Don’t wrap errors, log.
Errors have two purposes - control flow for your application, and information for the developer and operator. If you are already logging the flow of your program, there is no need to wrap an error and create a call stack - you already have the call stack. If you cannot follow the control flow from your logs, you need to improve your logging, because you will also need to debug production problems that do not produce a literal error object.
And if you want more causality than you can get out of your logging, consider upgrading from logging to tracing. I contribute to http://opentracing.io so naturally that’s the tracing API I would suggest you look at.
2. Do not do anything with panics except recover from them or exit.
If you have a panic that you cannot recover from (and in general, that should be all of them), then it means your application is in a unknowable state. While Go is not as unsafe as C, any and all state in your program could be bad, and it is completely unclear what is still working and what is not. Nothing you do at this point is safe. For the love of god, do not hang your application and prevent it from exiting by trying to send a Slack message. Your goal should be to restart your program from a fresh clean state as quickly as possible, not to tie it up on the way out. Let the program exit and have an external monitoring program, whose state is NOT corrupted because its in a separate process, do all of the triage and reporting.