While it's cool that this works at all, I wish we would stop using MNIST as a benchmark given how trivial it is.
Sure it's not great at differentiating between SotA techniques, but it's very useful for sanity checks like this one.
Even for SotA models, it's still useful to verify that you can get greater than 98% accuracy on MNIST, before exploring larger, more complex bench marks.
It certainly shouldn't be the only benchmark but it's a great place to start iterating on ideas.
It's a benchmark.
Not a real world problem.
That's why the traditional path has been MNIST -> CIFAR10 (optionally -> CIFAR100) -> ImageNet -> ????!?.
Because it gets gradually more complicated.
Your iteration time is the constraint to development progress.
Keep that down, and the bugs from your initial implementation will be significantly less impactful.