> well, we tried this procedure on 100 patients
"this procedure" is the most important part, and they describe it in detail in the paper, hopefully well enough that someone could attempt to replicate it, and some do attempt to replicate it. (Not that procedure descriptions in such papers are always sufficient for this.)
The difference between that and
> we tried this procedure on 100 datapoints
Is that it's nigh impossible to describe a ML procedure in enough detail to reproduce with just the description in the paper. Tiny changes in the parameters and construction can completely change the result; the only way to be able to reproduce it is if you had the source code. And also the source data which is just if not more important as the source code (see sibling thread).
The opportunity that academic CS has over every other science is that they could empower every reader with the capability to verify the results of every paper they read, and this is actually attainable. Reviewers of other sciences don't reproduce findings themselves for purely practical reasons that don't need to exist in CS.