If you think about how the human vision system is constructed (similar to convnets in deep learning), the brain uses a hierarchy of constructs like circles, lines, corners, etc. Recognizing letters requires the triggering of multiple layers of these constructs.
A dot, and another dot 1 pixel higher or lower is easy to get conflated. There's going to be a high degree of uncertainty. We have a hard enough time distinguishing '1' and 'l', and 'O' and '0'. This makes it so every letter takes on that trait. We usually must rely on context of the surrounding letters and words to fill in the missing information.
I think the concept of this idea is valid, but we need more than moving pixels up and down to distinguish letters. If instead a form existed that mixed various small primitive shapes it might work.
I'm not an expert on the vision system but I would think, sticking to at most 2 different locations instead of 5, and mixing in vertical lines, horizontal lines, diagonals, corners, and circles might be work exploring.
There's also color but lots of people are color blind.