Dynimize: Speed Up MySQL with CPU Performance Virtualization
dynimize.com
dynimize.com
https://dynimize.com/blog/discussions/dynimizer-mysql-cross-...
It seems to me the only system with SSDs (Ivy Bridge) sees higher transactions per second improvement than the other systems.
It seems to be JITting the generic binary code.
Would these same gains be seen from compiling MySQL for the actual target architecture gentoo-style e.g. -farch=native ?
Has this been run in production anywhere? Were there any instances of corrupted writes? Inaccurate reads?
What’s your plan for monetization down the road?
In terms of innacurate reads or corrupted writes, that would be a bug if it ever happens. That would not be part of normal operation and would not be expected. That said, all software including MySQL, gcc, and Linux are full of bugs and Dynimizer is not immune to that of course. However it has been stress tested thoroughly with MySQL, MariaDB, and Percona Server up to MySQL 5.7, MariaDB 10.2.
Best of luck with this. Exciting stuff!
> 7. Workload Requirements
> To obtain benefit from the current version of Dynimizer, all of the following workload conditions must be met:
> A small number of CPU intensive processes - On a given OS host where the workload is running, the workload must be comprised of one or a few CPU intensive processes. Optimizing a large number of processes at once is not recommended.
> Long running programs - The processes being optimized have long lifetimes, and their workloads are long running in order to amortize the warmup time associated with optimization.
> x86-64 - Optimized processes must be 64-bit, derived from x86-64 executables and shared libraries, which must comply with the x86-64 ABI and ELF-64 formats. Most statically compiled applications on Linux meet this requirement.
> Dynamically Linked - Target processes must be dynamically linked to its shared libraries. Statically linked processes are not yet supported. Most Linux programs are dynamically linked.
> No self modifying code - The target application must not be running its own Just-In-Time compiler such as those found in Java virtual machines. This therefore excludes Java Applications.
> Front-end CPU stalls - The workload wastes a lot of time in CPU instruction cache misses, instruction TLB misses, and to a lesser extent branch mispredictions.
> User mode execution - Much of that wasted time is spent in user mode execution (as opposed to kernel mode), as Dynimizer only optimizes user mode machine code.
> Because of these requirements, Dynimizer takes a whitelist approach when determining if programs are allowed to be optimized, with MySQL and its variants being the currently supported optimization targets on that list for this early beta release. Other programs are not currently supported, and while they can be used with Dynimizer, they should be very thoroughly tested by the user or system administrator before being deployed in a production environment.
> Future versions of Dynimizer may eliminate many of these workload requirements, broadening the variety of applicable scenarios as well as further increasing the performance delivered in previously beneficial cases.
My educated guess is that it relocates the hot path of the text segment to better pack into the instruction cache. Cool.
"[T]he industry's first CPU performance virtualization software"
"A New Frontier For JIT Compilers… JIT compilers use as input a virtual machine code format… Dynimizer [uses] real machine code as input instead"
Just 20 years after Mojo, Dynamo, DynamoRIO, etc:
http://program-transformation.org/Transform/BinaryOptimisers http://www.dynamorio.org https://www.complang.tuwien.ac.at/andi/bala.pdf http://cseweb.ucsd.edu/~lerner/mojo.ps …
Maybe they should present themselves as "fire and forget TPS optimizer"
Dynamizer is translating x86-64 to faster x86-64, but the concept is the same.
DynamoRIO was actually talked about for application acceleration. There was at least a PoC that did dynamic function inlining.
I would be very interested to hear a sampling of the war stories that came out of building this. I had a friend working on the IBM zPDT JIT at one point, and while I unfortunately can't remember many of the details at the moment, I remember boggling (in that sort of emergently satisfying way) at some of the 'oh shit' moments that came up.
https://www.percona.com/live/18/sites/default/files/slides/A...
sudo bash c 'bash <(wget O https://dynimize.com/install) default'
Come on, please don't teach people horrendous security practices... :(This will sound clichéd but I blame the rise of “poor mans devops” whereby management fires all the ops, and lets developers manage infrastructure.
1: https://www.idontplaydarts.com/2016/04/detecting-curl-pipe-b...
"Oh, this is just for a test mock up. Nobody is supposed to actually use this to install it for real."
Well, to experienced people it makes you look moderately stupid, and to inexperienced people it looks like an elegant solution. It's actively hostile to secure system planning.
It reminds me of the NPM left-pad debacle[0] and some of the criticism[1] that came up from that.
0: https://www.theregister.co.uk/2016/03/23/npm_left_pad_chaos/
1: http://www.haneycodes.net/npm-left-pad-have-we-forgotten-how...
I can’t wait to see tc39’s response to the `is-even` shit show after they decided to just add leftpad to the stdlib.
Reminds me of: https://github.com/jezen/is-thirteen
This clown was using a bit wise operation in `is-even` “because everyone already knows about % 2 === 0`, and thus was hurting performance (on top of whatever extra memory is used for the module, function call overhead etc)
Beyond actual errors being produced, I'm wondering what'd happen in weird scenarios such as one where by primary gets heavily optimized for its write load and creeps up to, say, 80% CPU or so at peak.. What then happens then if my replica which has been heavily optimized for its read load gets promoted in a failure scenario and gets pegged at high CPU?
Final thought here is if this tech really is solid, when is AWS going to start shipping it with my VM?
I think you'd just have to measure scenarios before using it in production.
Dynimizer can detect a drastic workload change as you described and reoptimize in response. That’s the default operation which can be turned off. What has been observed is that if we optimized for say a write-heavy workload and then change it to read-only without reoptimizing (or vice versa), it will still show an improvement, just not as much. Hope that makes sense.
From the /product page:
> ...It profiles applications using the Linux perf_events subsystem and interfaces with a target application's machine code through the Linux ptrace system call. When optimizing a program, it loads a code cache into the target program's address space...
Couple of questions. Since this seems to be a very general technology, why the emphasis on MySQL (and DBs in general)? Marketing? Also, I found Dynimize vs Dynimizer confusing - is that company name vs product name?
Dynimize is the company, Dynimizer the product. We may ditch the name Dynimizer and just go with Dynimize to avoid confusion. Thoughts?
It is a general purpose approach to optimization and MySQL is just a starting point. It was chosen first because it has a broad user base and is relatively easier to support compared to many other Linux programs: single process architecture, long process lifetimes, OLTP workloads are known to spend much of their time in front-end CPU stalls on the CPU side which are effectively targeted with profile guided compiler optimizations, and it’s statically compiled. We’ve tried it on MongoDB and seen similar benefit but not supported yet. Coming soon. Windows will probably require some driver development for effective sample based profiling and will happen later on. We will improve the effectiveness of our other optimizations that don’t target front-end CPU stalls and better support multiprocess workloads with short process lifetimes, which will allow us to target many other types of programs in the future.
Hope that clarifies things a bit.
Wonder if it could be integrated into the whole OS/kernel, and if it can help with typical ML workloads like running a randomforest in scikit learn.