Due to the regularisation yeah they all start at the same rating. But you don't need many votes to start getting good ratings.
I introduced this method to Dyson for objectively calculating very subjective measurements (e.g. "how frizzy does this hair look?"). We basically crowd sourced it to other engineers.
I did a load of studies on different methods by ranking something that's sort of hard to rank but you know the answer to - I used 10 grey squares that only differed by 2/255 and you had to pick the brighter one.
Some other things:
1. I don't remember the exact details but there's a slight extension of the method where you give each user a "how good are you" coefficient that you simultaneously solve for. This helps eliminate people that vote randomly, and also inverts the votes of people that deliberately pick the wrong answer (as long as they're consistently wrong).
2. You can put confidence limits on the values very easily too since it's a MAP estimate. Actually I showed curves for each item - basically how does the model probability vary as you sweep one rating up and down a bit. People didn't understand it at all though.
3. You can calculate the rankings incrementally very quickly (details in the answer) which means you can show users comparisons that give the most information. This usually means you end up showing users endless difficult choices which can frustrate them, especially if it's a forced choice.
4. I never found a principled way to incorporate a "they look the same" option. I tried some ad-hoc methods and IIRC a "much better, slightly better, can't tell, slightly worse, much worse" scale gave the fastest convergence but it was pretty unsatisfying that I just used some as hoc method to add the results.
It was all closed source and I haven't worked there for years so the code is lost to the wind unfortunately.