Imagine a user chooses "Sort by rating", and they subsequently observe an item with an average 4.5 ranking above a score of 5.0 because it has a higher Wilson score. Some portion of users will think "Ah, yes, this makes sense because the 4.5 rating is based on many more reviews, therefore its Wilson score is higher." and the vast, vast majority of users will think "What the heck? This site is rigging the system! How come this one is ranked higher than that one?" and erode confidence in the rankings.
In fact, these kinds of black-box rankings* frequently land sites like Yelp into trouble, because it is natural to assume that the company has a finger on the scale so to speak when it is in their financial interests to do so. In particular, entries with a higher Wilson score are likely to be more expensive because their ostensibly-superior quality commands (or depends upon) their higher cost, exacerbating this effect due to perceived-higher margins.
So the next logical step is to present the Wilson score directly, but this merely shifts the confusion elsewhere -- the user may find an item they're interested in buying, find it has one 5-star review, and yet its Wilson score is << 5, producing at least the same perception and possibly a worse one.
Instead, providing the statistically-sound score but de-emphasizing or hiding it, such as by making it accessible in the DOM but not visible, allows for the creation of alternative sorting mechanisms via e.g. browser extensions for the statistically-minded, without sacrificing the intuition of the top-line score.
* I assume that most companies would choose not to explain the statistical foundations of their ranking algorithm.