This is a common fallacy in computer vision. Often people get to a within a pixel and stop because they assume they can't get better. Most times you can get substantially better. I worked on a qr-code like system where the scanner could reconstruct the code from an image with (slightly) less pixels than 'pixels' in the code.
Not obvious a priori, but makes sense.
The light is blurred, with a fairly predictable pattern, so if you fit a function to the shape, you can find the peak of the function and that is the most likely center position for the point source.
There is the star tracker camera (widefield, fast), and the imaging camera (narrowfield, slow).
Star trackers use a fast widefield camera because it's easy to get a good signal:noise ratio (there is a lot of contrast between the stars and the background).
The imaging camera, on the other hand, generally takes much longer exposures (HST subjects are generally very dim, relative to stars). All those beautiful, nebulous HST photographs you see? Those have exposure times on the order of hours or days. In practice, being off by a pixel momentarily is not a big concern -- the amount of "bad" photons you collect during that time is very small.
The gyros are only used for coarse pointing, guide stars are used for much more accurate and precise position sensing during exposures (which is why they are called Fine Guidance Sensors).