16x16, 8x8, or 4x4 would be the way to go. You'd want each RGB block to map to a single H.264 macroblock.
Using non order of 2 numbers means that individual blocks don't line up with macroblocks. Having a single macroblock represent 1, 4, or 16 RGB pixels would be ideal.
In fact, I bet modifying the original code to use a scaling factor of 16 instead of 20 would produce some significant improvements.
It would be better to use YUV/YCbCr directly instead of RGB.
There's a bunch of other things too, like YUV420p and TV colour range: 16-235, so you only get 7.7bits / pixel.
If anything you would want to encode your data in some way that abuses the P and B frames, and the macro block size of 16x16.
Coding theory for the data output at your end is only one side of the coin, the VP9 codec stupidly good compression is a completely different game to wrangle.
And I kinda doubt you'll get much better than your estimate of 1% from the original scheme.