This one doesn't need a huge separate dataset for each animal (~200 frames), just user has to define parts of animal they wish to track. Most 3D pose estimators require a huge 3D pose dataset for training.
This is because of transfer learning: for example, by default, DeepLabCut starts with a pre-trained ResNet-50. Thus these trackers start out with an excellent "understanding" of the statistics of natural scenes out of the box.