ML researchers are playing scientists: tweak a few parameters in an LLM, re-train it on a largish dataset (need access to $$$ GPUs), find metrics on which the tweaked LLM makes a barely noticeable improvement and make the other metrics where it actually gets worse look insignificant, write a paper, upload to arxiv, and update your resume.