Anyway my point with that is that there are a lot more high level things that needed to be fixed - and this is the case in most, if not all organizations I've worked at - before you get to the lower levels of performance. And this probably applies everywhere.
In this example, they have their software so optimized and tweaked already that it ended up being a low level JVM issue.
Mind you / on the other hand, they mentioned something about subclasses and the like, it might be that a change in their Java code would have solved or amortized this issue as well.
Couldn't agree more, except on the timeline. In my experience, the 90s and most of the 00s proper engineering like this was the norm in silicon valley and I loved it. It's the last 10-15 years that PMs have taken over and now it's all so superficial and uninteresting. I've moved away from hands-on development work because nobody is doing anything intellectually rewarding anymore. But I wish I was doing hard-core engineering like it used to be.
Because growth is most of the time translated as acquiring more users/customers. Eventually making more money out of them.
Somewhere in the distant future, making money becomes cost saving, this is when companies start thinking about how can we cut the cost and such projects/explorations become relevant.
Even then it is not easy to sell as a project. Because there are so many unknowns, discussion goes like:
me: hey, I see spikes in CPU/mem usage, want to investigate why.
Product: is it important now?
me: I don't exactly know, but it might end up saving our instance costs.
Product: how many instances we are talking about?
me: I don't know yet how big is the problem, depends on what would be the result
Product: how long does it take you to do a research?
me: I don't know what is the problem yet, so I can't evaluate how long it takes to find it.
Product: maybe we shouldn't do it then?
me: (again) but it might end up saving our instance costs.
Product: ok, do you think you can finish investigation in 2 days and share with me timeline and metrics this project would improve?
me: ok, let's forget about this shit we made. what is a next "metric improving, highly visible" project we have?
Product: cool, I want you to work on this super exciting "Move button from Page A to Page B" projectIf you always do the obvious things, you’ll end up as a mediocre company that might not survive. If you say yes to any project that comes to mind, you’ll end up burning all your money and end up with endless projects. It’s important to strike a balance and try to find the right investments others don’t see to create a competitive advantage.
It helps if the culture allows for that. The questions still are relevant, you just don’t always need precise answers, Probe 100 instances and see if it surfaces more than once. Set up a list of things you think you’ll need to be doing and an initial proposition what you think might change. It’s the PMs job to then convince others that it’s worth taking the risk.
My observations contributing to the frustration:
* Developers believe performance is speed. It isn’t. Performance is a difference in measure between two (or more) competing qualities, the performance gap measured in either duration or frequency.
* Developers will double down on the prior mistake by making anecdotal observations without any forms of measurement and will vehemently argue their position in the form of a logic based value proposition. Example: X must be fast because Y and Z appear fast.
* Developers will further hold their unmeasured position by isolating their observations from how things actually work. Software is a complex system of many instructions. Except for the most trivial of toy scripts nothing executes in isolation. For example developers follow the value proposition that a grain of sand is tiny and weighs little, so therefore does not contribute to the weight or size of the beach. An example is that developers will argue to the death using querySelectors to avoid walking the DOM but in Firefox DOM walking can be as fast as 250000x faster than querySelectorAll, which is indeed significant.
I suspect there exist a variety motives for why developers reason things in this way, but from the outside it looks like a house of cards based upon argument from ignorance fallacies outputting really slow software by really defensive people.
I do not think that’s restricted to front-end developers. Back-end developers rarely worry about inserting yet another http request in code called from user action in a web browser, for example. A tenth of a second here, a tenth there, it all adds up.
Product managers and users also don’t seem to care much or do not know how fast modern hardware is. I’ve frequently seen web page refreshes take 5 seconds or more without getting any user complaint, even when I explicitly ask them, and tell them how that could easily be halved.
> * Developers believe performance is speed. It isn’t. Performance is a difference in measure between two (or more) competing qualities, the performance gap measured in either duration or frequency.
Maybe I'm a dummy, but if someone says "make a fast webapp", I'll be doing things like reducing requests, optimizing queries, making things smaller, using the right data structures, manipulating data in a fast and/or memory efficient way.
This should result in lower loads, more stability, better availability, etc, so also "more performance". I do think of it as lots of focus on speed since, for the most part, I'm measuring how long it takes for a function/query/etc to run.
Happy to check out some articles/keyword suggestions.
I'd love to get back into it.
I have seen far more companies wasting resources on the projects of no value because the stakeholders believe themselves to be of same level as Netflix, Google.
This is a case of eeking more out, optimizing, but the general diagnostician view is deeply needed widely. This industry has an incredibly difficult time accepting the view of those close to the machine, those who would push for better. Those who know are aliens to the exterior boring normal business, and it's too hard for the pressing mound of these alien concerns to bubble up & register & turn into corporate-political will to improve, even though this is often a critical limitation & constraint on company's acceptance/trust.
There's a lot of ultra passing this off comments. No one seems to get it. Scaling up is linear. Complexity & slowness is geometric or worse. Slowness as it takes root is ever more unsolvable.
This rings immensely hollow to me, & borders on victim blaming. Oh sure; telling the lower rungs of the totem pole it's their fault for not convincing the business to care- for not being able to adequately tune in the business to the sea of deeper technical concerns- has some tiny kernel of truth to it. Maybe the every-person coder could do better, maybe, granted. But I see the structural issues & organization issues as vastly vastly more immense impediments to understanding-our-selves.
There is such a strong anti-elitism bias in society. We don't like know it alls, whether their disposition is classicaly braggadocious, or humble as a dove. We are intimidated by those who have real, deep, sincere & obvious masteries. We cannot outrun the unpleasantness of the alien, the foreign concerns, steeped in assurity & depth, that we can scarcely begin to grok. Techies face this regularly, are ostracized & distanced from with habit. Few, very few, are willing to sit at points of discomfort to trade with the alien, to work through what they are saying, to register their concerns.
> I’m not sure it’s useful to assign unilateral blame to management based on a failure to listen to the engineers
Again, granted! There absolutely are plenty of poor decisions all around. Engineers rarely have good sounding boards, good feedback, for a variety of reasons but the above forced alienation is definitely a sizable factor where engineers go wrong; being too alone in deciding, not knowing or not having people to turn to to figure shit out, to get not just superficial but critical review that strikes at the heart.
This again does not dissuade me from my core feelings on my core point. I think specifically most companies are hugely unable to assess their own products & systems health, unable to gauge the decay & rot within. Whether it's slog or real peril, there are few rangers tasked with scouting the terrain & identifying the features & failures. And the efforts are renewal/healing are all too often patchwork, haphazard, & random, done as emergency patches. These organisms of business are discorporated, lack cohesion & understanding of themselves & what they are. Having real on the ground truthfinders, truthtellers, assessers, monitors- having people close to the machine who can speak for the machine, for the systems, that is rarely a role we embrace, and so often we simply rely on the same chain of management which is also responsible for self/group-promotion & tasking & reporting which has far too many conflicted interests for it to be expectable for them to deliver these more unvarnished, technically minded views.
[1] https://www.usenix.org/system/files/1311_05-08_mickens.pdf
Cheap money and small, but easyly observable gains made all the ones not knowing better appreciate linear scaling improvements.
We did that (ops, not developing the actual app) few times where app scaled badly
We're working on applications that's either waiting for the disk, waiting for the db, waiting for some http server or waiting for the user.
None of our customers will notice the difference if my button click event handler takes 50ms instead of 10ms, or if the integration service that processes a few dozen files per day spends 5 seconds instead of 1 second.
I'll easily trade 5x performance in 99% of my code if it makes me produce 5x more features, because most of the time my code runs in just a few milliseconds at a time anyway.
Of course, I'm weary of big-Oh, that one will bit ya. But a mere 3.5x is almost never worth chasing for us.
> At Netflix, we periodically reevaluate our workloads to optimize utilization of available capacity.
The idea being that if you pay for fewer servers you spend less money
If you try to be good at waiting on many things, you can use one machine instead of a hundred
> None of our customers will notice the difference if my button click event handler takes 50ms instead of 10ms
They absolutely will, what
We can't run on one, because customers run our application on-prem.
> They absolutely will, what
How can you be so sure?
Sure, 10ms vs 40ms is measurable, and for the keen-eyed noticeable. But if you're only pressing the button once every 5 minutes, it doesn't matter. Similarly, if the button triggers an asynchronous call to a third-party webservice that takes seconds to respond, it doesn't matter. And so on.
Of course, for the things where users are affected by low latency, we try to take care. But overall that's a very, very small portion out of our full functionality.
It's a great tool with a low learning curve that just requires SSH.
In this case, the engineers had a clear expectation and good reasons that something was wrong and such throughput should be possible. The arguments for that are simple enough that every manager should also understand this.
The other aspect is also latency but also here I would assume that managers should at least care a bit. They also have the old baseline to compare to, and see that it is suddenly much worse, so they should care about why that is and how to fix it.
There’s a reason YCombinator (as an example) insists on at least one strongly technical founder.
Bean counters rot companies by not seeing the wood for the trees.
The same applies to MBAs (no offence to anyone that has one). Engineering companies need to either be led by people the engineers respect, which means an engineer - or at the very least someone who understands the engineering thoroughly enough to get out of the way.
The second you are subjected to one as an engineer you are being told you’re in the wrong place.
I don’t care what the stock options are, I don’t care what the free morning deviled eggs are like, that place is cancer and you need to escape it.
Mind sharing what field you work in that gathers so much talent like that?
I’ve worked in all the industries, including a long time in yours. We probably even overlapped or maybe worked together!
"My team makes your team apps and make them run cheaper"