To answer your question more helpfully, having closed loop machine vision stuff is quite computationally intensive, though there is already some on the rover (to help negotiating obstacles or stop if it sees something like a trench (MERs would famously stop at the same time every day, thinking that their now-infront shadow was a pit)) but it's power intensive and expensive (both mass and cost of a dedicated odometry camera, for example) and the lighting conditions are very variable. The MERs developed a technique whereby they would 'shimmy' their steering every so many meters, and then with the rear facing camera snap a photo of their tracks (with shimmies) and infer distance travelled and slippage and so on. This stuff is often easier with humans in the loop as the terrain still throws up unknown unknowns when driving. Here's a pic of the tracks:
http://photojournal.jpl.nasa.gov/jpeg/PIA14129.jpg
I guess another point is that while it might be a perfectly sensible suggestion technically, there's only so much 'new' they want to risk per mission (the Entry Descent and Landing phase contained a whole lot of 'new' as we know) and they maybe just wanted to carry on doing it in a way they know works ok from before.
With no patterns one could think of a situation where the surface sand/terrain would be displaced but your actual position was exactly the same.
The method hinges on establishing correspondence between visual features in before/after images. Then, since you know (through stereo ranging) how far away things are in both images, you can match points and see how far you've gone. This works better the more visual features (tiny edges and corners) you have to cue off of, and presumably these treads are generating more distinguishable features. They ran VO on Spirit and Opportunity, and they will on Curiousity as well.
A summary is in the first page of this little white paper:
http://www.lpi.usra.edu/meetings/marsconcepts2012/pdf/4053.p...
and one full description is here:
http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.104....
The good thing is that VO requires no extra hardware, because you already have stereo cameras.
By the way, they do make commercial velocity sensors of the kind described in the GP comment (special camera looking down):
http://www.corrsys-datron.com/optical_sensors.htm
These may not use VO, they might just use optical correlation. Sometimes they're used in robotics.
They could have one on each wheel and then another set of sensors pointing down, so wheel slippage vs ground movement at that axle.
Besides, a camera would work on hard ground, while this method wouldn't. I don't know if that's a problem in practice on mars of course.