The number one thing in my opinion is that Stan's algorithm(s) for drawing from a posterior distribution produce samples that have much less dependence among adjacent draws than other simpler MCMC algorithms like Metropolis-Hastings or Gibbs samplers. Since it is often difficult to answer the question "How much dependence is too much in practice?", it is prudent to use the algorithm that yields the least dependence because the effective sample size (from the posterior distribution) per unit of wall time will usually be greater. PyMC3 has started to incorporate some of Stan's algorithms, although their implementations are not as far along.