loopXX instructions do not use CPU LSD (Loop Stream Detector) while cmp/jnz construct takes advantage of it. This speeds up some small loops. Also, there are some rules in intel manuals for instructions within cmp/jnz loop like no mismatched push/pop, etc.
now, does anyone know why?