1,791 karma · joined June 13, 2014
> 1. Engineers shall be guided in all their relations by the highest standards of honesty and integrity.
> ...
> b. Engineers shall advise their clients or employers when they believe a project will not be successful.
There's nuance around how hard you should push back on bad requests and where ownership/accountability and decision-making responsibility ultimately lie but providing professional judgement/advise/opinions is definitely in bounds for engineers.
[1] https://www.exmark.com/electric [2] https://www.toro.com/en/professional-contractor/commercial-m...
It's the sort of thing where I'd advise exploring other options first and only using it if the whys[2] really resonate with you because it definitely comes with some overhead.
[1] https://smithy.io/2.0/index.html [2] https://amazon-ion.github.io/ion-docs/guides/why.html
It's slowly changing but I wouldn't consider Amazon's tooling to be industry standard by any definition.
A simple example is setting an alarm on your max load test traffic. When you get that alarm you know your system is now operating in unknown territory and it's probably time to do another load test or at least take a close look at scaling.
Lately for a 1-3 day horizon I've been adding everything I'm working on as blocked. That definitely helps me visualize my priorities for the next couple days and whether I have room to fit more meetings or tasks. The main downside is I'm used to a more free-flowing schedule where I pick a task based on how I'm feeling.
> Each Availability Zone is separated by a meaningful physical distance from other zones to avoid correlated failure scenarios due to environmental hazards like fires, floods, and tornadoes.
https://docs.aws.amazon.com/wellarchitected/latest/reliabili...
As a company grows sooner or later most of these features become pretty desirable from an operations perspective. Feature developers likely don't and shouldn't need to care. It probably starts with things like Auth and basic load balancing. As the company grows to dozens of teams and services then you'll start feeling pain around service discovery and wish you didn't need to implement yet another custom auth scheme to integrate with another department's service.
After a few retry storm outages people will start paying more attention to load shedding, autoscaling, circuit breakers, rate limiting.
More mature companies or ones with compliance obligations start thinking about zero-trust, TLS everywhere, auditing, and centralized telemetry.
Is there complexity? Absolutely. Is it worth it? That depends where your company is in its lifecycle. Sometimes yes, other times you're probably better off just building things and living with the fact that your load shedding strategy is "just tip over".
[1] https://www.youtube.com/watch?v=4o_jPzBZSIo (~15:30 in)
Our problems were caused by a mix of serialization format (JSON numbers) and not always converting into the language's date/time types at the boundary (sometimes raw epoch seconds/millis were passed around layers of code and only parsed into a date for display. That created opportunities for misinterpretation at every function call.
My general rules for non-performance critical code are
1. Always Parse into a first-class date/time/duration type at the serialization boundary.
2. Always use an unambiguous format (e.g. ISO-8601) for serialization
It's not the most efficient but lets you rely on the type system for everything in your code and only deal with conversion at one place.
I ran into this a couple years ago with partially duplicating a service that existed under another SVP org. It didn't feel good at the time but in hindsight it was probably less effort and higher chance of success than trying to align priorities across SVP orgs and making their thing work for our slightly different use cases.
Most of the harder bugs I run into fall into this category. Great list of techniques to use. For when I'm really stumped I go into a more formal scientific method to the splitting the problem space approach. Formulate a hypothesis that if answered will reduce the search space, either by ruling in or ruling out things. Write it down - this is the important part, test it, repeat.
It's slow but usually helps me keep forward momentum. Worst case it at least creates a list of things that need better logging to figure out what's going on
That's the point, right? Visually representing the complexity of the system. I've used IntelliJ to do this before to show why modifying certain behavior was so slow and error-prone. In that case there were 3-4 classes with heavily overlapping functionality because, surprise, in the past there were multiple teams contributing to the same codebase that all did their own thing.
[1] https://www.smh.com.au/lifestyle/health-and-wellness/toxic-s...
Fair, I was thinking of this in the non-cloud perspective where efficiency improvements can push you to upgrade even if the hardware still works. In cloud provider mode it makes sense to keep it around as long as it still works and it's not too annoying to run. It doesn't really matter how (in)efficient the hardware is because you set the pricing to keep it profitable as long as someone's willing to buy it.