They use a method called propensity score matching to try their best to match patients on both sides using a simple linear model with various features that try to ensure that only pairs of closely matched patient histories are compared.
Unfortunately this is rarely clean. Its also easy to make mistakes. Sometimes two arms are fundamentally incomparable. The quality and rigor of the comparison is often determined by a lot of extra checks and validations, and different journals demand different levels of rigor. I need to read it carefully to judge if this is good or not.