I recently had a class project where we had similar issues. Being lazy students, we just let it perform slower for matrices of the wrong size; what is a good way to handle this sort of thing?
Unless you are a nutter like me, in which case you hand-roll everything in assembly.
I remember the first time I saw an image data struct that specified bits per pixel, image width and then also had bytes per row. I was a bit perplexed about why you'd store bytes per row if you could just compute it from bpp and width. Turns out you can fix/avoid a lot of performance issues if you can pad the data, so it's good practice to always save the number of bytes per row and use that in your mallocs/copies/iterators/etc.