> To evaluate Form Recognizer, I split the data randomly into 26 training documents and 25 test documents.
Training on just 26 documents seems woefully inadequate. I'm not a data scientist and have only cursory exposure to ML, but I'm not surprised to see terrible results with such a small training set.