Comments and documentation become stale because they require deliberate human intervention to maintain their truthfulness. They also rely on the author of the documentation or comments correctly mentally executing the code being marked up.
Tests are compared with the API every time they're run, as seen by the computer's model of execution, giving immediate feedback on detected discrepancies.
(I believe that Pythonistas have clever docs that can encode snippets of testing into documentation, which splits the difference a bit, but is no substitute for a robust test suite).