How and why to make your software faster
incoherency.co.uk
incoherency.co.uk
Or you can make the logging part asynchronous (provided the server itself is not maxed out) and cut that 30% of request time.
Interesting approach. This is also nicely recursive, with each run being more expensive. I would prefer that statement to have a cost cap. Say "... in three days of work".
Thanks for reading.
At work we were debugging why a process that is usually quite fast (< 5min) was taking close to 2 hours once we got to the load we predict for this coming fall.
We noticed the box it was running on was running out of memory, so we assumed resizing it to have ~4x the RAM would help, but surprisingly it didn't. The actual answer turned out to be algorithmic, and we were able to rewrite it to use an O(n) algorithm rather than an O(n^2) algorithm.
For what it's worth, a quite staggering number of sites I've fixed up have had databases with varchar ID columns, and no indexes anywhere.
Thankfully the DB sizes have been small enough that fixing this has taken a few seconds to a few minutes max of DB downtime, with instant, seriously-impressive-looking results afterwards.
also, i simply don't buy the assertion that any ops or dev can be automated in any meaningful way. the only thing it's doing is making it easier for morons to shoot themselves in the foot by ignoring the experts and then run to the same experts like a little child when everything catches on fire. see: recent story of a guy who deleted 1500 systems at once. oops.
are we all doing more, or less work than 5 years ago? has all this cloud and deployment tech made any less work for anyone, or destroyed any jobs? lol no. i repeat emphatically LOL NO. such a claim is just absurd in my view. anyone who claims such a nonsensical statement has never done actual ops work where you get called if shit goes south.
the smartest thing amazon ever did was simply give people a button to push to give them more money when they're faced with their own overwhelming stupidity. ain't nobody to blame but yo'self then.
I'm genuinely curious as to why you feel this way?
On a more technical side as an example, the number of open files can only be so high. You really start to hit the limit when you do things like that, which is when the most interesting kinds application crashes happen. "Just increase the file count." results in "Our report shows that an OldServer did not include a necessary open file limit config change. Engineer #582 has been fired."
I suppose we are lucky to have a very consistent workload per logical database.
This allowed us to calculate our costs ahead of time. We found that we could keep DB costs at ~$1/customer/month. For a B2B SaaS product we were satisfied with that cost.
Most decent RDBMS have a good way to ensure isolated locks across tenant ids e.g. MySQL partitions.
Currently, I manage several hundred logical databases for a B2B SaaS product. This has the potential to grow to several thousand over the course of years.
My customers enjoy that (arbitrary) piece of mind that their data is logically separated from all else. I enjoy not having to remember to include a tenant_id clause.
Our traffic pattern is quite steady per hour of day and per customer. That, plus using RDS for horizontal scaling (fixed # of logical DBs per host) allowed us to take this route.
I suppose the decision is a matter of circumstance and preference, but I don't think this is necessarily "wrong" to do.
It's not just talent/skills - there's more communication overhead when dealing with a team, and when the team has a range of skill levels on it, the comm overhead can make things worse (if not managed well)
I have a longstanding suspicion that developer cost climbs slower than developer skill and productivity, meaning companies should always hire the best they can get, because they'll maximize productivity that way.
To the downvoter: Maybe in your fantasy world, throwing more computing power at problems solves Eeeverything, including P ?= NP. I do wonder - why people even bother offering such services as the OP, if it's so pointless?
I once spent weeks optimizing code on a "best in the world" (for my branch of work) open source project because it made the assumption that it'd run on SSDs, hence it made zillions of small reads from the disk. There were two approaches to solve that: buy an SSD, or refactor the code. But it was through profiling that I found the answer (and in my economic situation at the time, buying an SSD was definitely not viable, much less a step 0!).
Firstly, this doesn't help if you're CPU-bound or write-bandwidth-bound, both of which are quite common.
Secondly, you may not be able to simply get a bigger box. Most motherboards have effective RAM limits. Bigger iron gets exponentially more expensive. Jeff Atwood on this: http://blog.codinghorror.com/scaling-up-vs-scaling-out-hidde...
(I remember about a decade ago when our chip design startup bought a 128GB RAM box because that was the largest available at sane pricing ...)
* Database with homegrown software load-balancer in front of it, and the database replicated to a half-dozen subscribers. Load balancing required database calls, and the subscribers had their LUNs on the same array as the main database. So 7 different servers, but all beating the same storage. In the end, we ended up with no load balancing (just the single instance).
* Application server with lots of agents (one on every workstation). Agents were checking in every 200ms, even if they had nothing to report. Result was lots of tiny database transactions (for logging no activity). Throttled that down to 10 seconds (product managers agreeing that was the more realistic threshold) and batching the updates, made the logging asynchronous, and saw the database transaction log no longer waiting on log writes constantly.
* Bad clustering keys for tables. Some tables that were being searched constantly as heaps. Sharded tables with divergent schemas. This was fun, as had to conform everything while it was live (got agreement on what was reasonable delay for synch jobs). Chose correct clustering keys, compressed the tables, and then got rid of other indexes that were no longer necessary. Cut IO by 99%, and writes were now just appends.
* Badly written queries using inadequate indexes on badly structured tables. TL;DR: An analysis process took 24 hours, and was tuned to take 4.5 hours. I spent a day on it to cut it down to an hour while I worked on other things. We then spent 2 weeks refactoring it to cut it down to about 8 minutes, despite the underlying data model being bleh. We also added ad-hoc functionality to do stuff out-of-band of the regular scheduled job, with an average response time of 6ms.
* ORM lazy loading by default, including for things that constantly retrieved collections. Rewrote them to be programmatically eager loaded and reduced database calls from millions a day to thousands a day.
* And a little bit of hardware. SAN HBA links were moved form dual 1Gbps to dual 10Gbps. Some SSDs added. Ended up reducing IO latency by 96-98%, depending on the server. Needed to happen, as IO latency was in hundreds of ms, and we were able to get it down to where it should be (< 10ms for reads, 1-3ms for writes).
Non-scaling SAAS stack went from falling on its face with an entire rack of hardware, to being cut down to 1/3rd of a rack of hardware and the hardware basically twiddling its thumbs. The organization was in the process of building of entirely new SAAS platform in a completely different tech stack -- they immediately started questioning that decision when the old stack handled everything trivially, and they knew there was a lot more tuning that still could be done.
> "In the 1880s, James MacNeill Whistler, as plaintiff in a libel action, was challenged, "For two days' labour, you ask two hundred guineas?" "No, I ask it for the experience of a lifetime." That seems an apt summary of the message of this legend."
Maybe: s/practically nothing/less
[0]https://lists.freebsd.org/pipermail/freebsd-current/2010-Aug...
"no code is faster than no code"
[0]: https://twitter.com/davecheney/status/715421274722271232
Unlikely that the worst performing page is the most valuable page; which is to say that offers should focus on providing value.
I once worked at a company where I noticed an employee would come in over an hour earlier than normal on Monday mornings. It turns out she got in early to run a report that the CEO liked to see every Monday morning. An hour or two later I added an index, and refactored some SQL so that we could do more in a single query rather than N queries, and the report time went down from 45 minutes to 3 minutes.
A page that takes 10s to load that you never look at is a much worse candidate than one that takes 8s to load and you look at 20 times a day.
There is absolutely no call for a web app that badly written to exist in 2016. Period.
So how about this: You go over there (points to corner) and make angry and dogmatic proclamations of what should and shouldn't be, while the rest of us go out into the world, where these apps do exist in large numbers and see heavy usage supporting very valuable workflows, and collect obscene salaries making them better.
There are many scenarios where a page used 20 times a day takes 8 seconds to load. For example, an admin dashboard page that shows some statistics on some aspect of customer usage written early on in a startup's life might only need to be looked at once a month by one of the founders (since it app has little traction, they have bigger things to worry about than seeing statistics from an n of 3).
So, it's hammered out with inefficient queries, improper indexes, etc. (after all, we don't even have a clear enough view of actual usage yet to know what indexes or queries we should really be doing).
Time goes on, the startup picks up traction, and before you know it, it's April 15, 2016, and the application handles millions interactions a week across its users. Someone on the support team at some point in the near past realized that they should be looking at these usage statistics to help drive improvements to the interface of the application, but the page now takes 8 seconds to load, because of all the data stored in that table, and because of the inefficient queries, improper indexes, etc.
Due to the nature of the business and the aforementioned constraints of budgets, deadlines, and priorities, this person has spent the last month or so just dealing with the 8-second page load. But now, it's a priority, and someone can go in, look at the data that's actually there, and figure out the best way to refactor, optimize, and improve it.
The point is, there are many scenarios where a web app can have a page checked 20 times a day that takes 8 seconds to load without assuming it's so badly written, and it certainly has a reason to exist (after all, it is partially the success of the app and value it provides to people that has caused this issue in the first place), and it doesn't need to die, it needs to adapt.
I got it down to a few seconds, but that's not the important part: that was an application that was being used, daily. That page would just take 5 hours to load.
The backend team was more worried about keeping daily processing and offline reports below 24 hours. The offline reports were essentially SQL queries.
If you're just rendering a couple database fields to a page? Sure, that's easy
For example, I'm currently writing a book on web app performance (focusing on ASP.NET Core) and apart from the stuff built into Visual Studio there are also tools such as Glimpse[0] and Prefix[1].
Profile the application. Find out what it's spending too much time on, and figure out how to make it do less of that.
I wish you every success in this business - if it takes off, I might try it myself. I found `oprofile` to be my preferred Linux profiler, and haven't worked out what the corresponding Windows solution is yet.
YMMV of course; that might just be my biased view after weeks of trying to improve performance.
What makes MySQL better than MongoDB is that is older and closer to what a database should do. They still have their warts and some of the warts might not be removed, since it would break compatibility.
At this point MongoDB does not have anything going for them. Its key benefit was their speed, but as it turned out that was because data was stored mostly in RAM if Mongo crashed or there was a power loss then you most likely would lose significant (possibly all) of your data. Since then they fixed that and make it more reliable, it is worse in performance than a relational database [1] and it also doesn't scale well [2]
Essentially NoSQL databases were designed to be simple and without relational features in exchange for scalability and performance. Mongo you get neither. Mongo doesn't even try to benefit from the CAP theorem. It's neither always consistent nor always available [3].
You generally should always use a relational database, because in majority of cases you do have some schema and you expect data to be consistent. NoSQL databases (especially ones that are AP in CAP) are generally good for specialized use cases, things that have no relations and are acceptable to be wrong or missing occasionally. For example storing logs, or user sessions etc.
Lastly, regarding question about performance. You need to understand your data and what you are doing. At my previous job there was a database called region. It was intended for a task such as looking up latitude/longitude -> ZIP code, and also IP -> ZIP code.
That database was (and probably still is) running on 3 beefy machines running Mongo and contained data was about 13GB. One time I wanted to see how it would work in a relational database. So I loaded the dataset to a Postgres database and installed PostGIS and ip4r extensions. And you know what? The same data took only 600MB there and all queries took sub millisecond on smallest VM.
How come? Mongo does not understand IP addresses, so what they did is they converted an IP into a number and stored it as a 64bit integer in Mongo. Postgres on the other hand was simply storing IP ranges, and with an GiST index.
Why I'm telling you this? I think it is important to know that RDBMS databases have been used for a long time, and many problems were already solved there if you're having some performance problem chances are someone else did have it as well before you (in this case someone wrote an extension providing a new data type)
[1] http://www.enterprisedb.com/postgres-plus-edb-blog/marc-lins... [2] http://www.datastax.com/wp-content/themes/datastax-2014-08/f... [3] https://aphyr.com/posts/322-call-me-maybe-mongodb-stale-read...
The graph at the top shows you which methods took how much time. The wider a bar, the longer it took. So it's immediately obvious that there are two substitutions called from Kernel::Output::HTML::Layout::Output that take up nearly all of the time. That's where you can focus your optimizations on. The other methods only take so much time because they include calls to the offending methods.
You can click on the bars (or the the routines listed below the graph) to get a few of the code, annotated with call counts and time spent on each line.
This is from OTRS, when it still used its home-grown, hacky template system, which is the major offender here. After switching to a proper template system, this is what the same request looks like: http://moritz.faui2k3.org/tmp/otrs-nytprof/agentticketzoom-u...
You can see from the absolute numbers at the top that it's much faster (18.3s -> 2.12s), and there's no single spot anymore where all the time is spent.
It is rather typical that unoptimized code has one or more big time sinks, and profiling makes that obvious.
E.g. - "60% of the time is in the sequence main() -> process_unit() -> validate_input() and the things called from there" or "45% of the time is in all of the call sequences which then lead back into write_line()" or things like that. Usually, you see patterns of a small number of slow functions that everything depends upon and/or, arranging the call stacks into a sort of "tree" of calls, a branch of the tree that you spend an inordinate amount of time in.
There are newer tools that provide graphical representation of this data as well (e.g. - show the tree as sort of a topographical map, with hotspots as peaks, and the call tree as geological strata)
Of course, if your program has a giganto routine that it never leaves, you might not learn much -- "do_all_the_work_inline() is executing 95% of the time!", unless you start looking at things at the line number level. Blech.
It describes the order that tables are accessed, which indices (if any) are used for each access, "and so on".
At the highest level, you might just know "we spent 25% of time reading from the disk, 10% writing to network", etc.
At the lowest level, you can know exactly how many times each instruction was executed, and how long it took, and the call stack of each execution.
Normally somewhere in between. The appropriate level depends on what you're trying to optimise.
In C++, `unique_ptr` is an "abstraction layer" with no runtime overhead, that will probably make your code faster due to the better safety.
I was trying to say: "consider the abstractions, and make them more leaky when necessary" (e.g. allow a flag to say "don't do the expensive calculation"), not "remove the abstractions".
The traditional BLL/BOL/DAL architecture is probably the worst. It aims for some kind of fantasy "separation of concerns" between the business layer and the data access layer, but many concerns simply can't be classified as exclusively business concerns or exclusively data access concerns. When people attempt to introduce and enforce such a separation, they end up with all sorts of performance problems.
It can be looked at on a few levels.
One is the "simple things" that waste a few seconds here and a few seconds there and pretty soon we are talking about hours.
Another is the algorithm that does more work than it has to do, scales poorly, etc.
In my mind "encapsulation" and "abstraction" do tend to be leaky when it comes to things like: performance, security, thread safety, scalability, etc.