If you filter your dataset to failed ceasefires, why would anyone expect it to show that they work?
What I think the data the author used shows is "Of the ceasefires that failed, it would have generally been better not to have had a ceasefire at all".
Which is an entirely different thing from "You shouldn't try to have ceasefires at all."