1) One dimensional phase detect sensors, rather than the (more expensive) cross-hatch ones used in DSLRs, which don't perform as well with movement or low light conditions.
2) Reduced processing power, lack of dedicated autofocus processing silicon, and less sophisicated data pipeline in general.
3) Greatly reduced user control. DSLRs, pro ones especially, have enormously sophisiticated control systems, and the ability to completely switch setup very quickly. There are a huge number of ways I can influence what descions the autofocus system makes on my cameras, and I use them to make sure the camera corretly interprets what part of what object it should be focussing on. I probably want the camera to be focussing on the head of the bird (rather than the wingtip, which is typically what a m4/3 system will go for), but that's unlikely to be in the exact centre of the frame, and definitely isn't the closest or most distinct object in the scene, so I need to tell the camera (very quickly and accurately) what it should do. Small mirrorless cameras with limited physical buttons and less tightly controllable focus systems make this very hard.
Highend DSLRs can also beat that continuous shooting rate - a Nikon D5, for example, will do 12fps (or 14fps with the mirror locked up). However frames per second isn't always relevent - you're going to be at 1/2000s minimum for this sort of shot, which means that even at 12fps you're capturing only ~ 0.6% of the time. If there's one specific moment you're targetting, even at very high fps, holding the shutter down is a bad strategy. It's useful in some situations (fps is important for pro sports photography for example), but it's often less important than a lot of people seem to believe.