A bunch of the projects described, and the (technical) difficulties encountered, make me wonder if GPT-2-style systems would have better odds. Is anyone looking into applying that to medical/scientific text NLP problems?
I mean, GPT-2 still often produces nonsense more similar to dream imagery than useful reasoning, but I gather most medical residency students are half-asleep most of the time anyway, so... :)