Author here. As I noted in the post, this is an elementary post to help people understand the very basics. I didn't want to bring in anything more than a "101"-level view.
I do mention output sampling briefly (Cmd-F "self-consistency"). And yes, there are a lot of good techniques on the validation set too. At the most basic, you can sample, of course, but you can also perform uncertainty analysis on each individual test case so that future tests sample either the most uncertain, or a diverse set of uncertain and not test cases. I also didn't go into few-shot very much, since choosing the exemplars for few-shot are a whole thing unto itself. And this benefits from "sampling" (of sorts) as well. But again, a whole topic on its own. And so on.
As for top_p, for classification this is a very good tool, and I do talk about top_p as well (Cmd-F "confusion matrix")! I again, felt it was too specific or too advanced to dive in more deeply in this blog post, but I linked to various research if people are interested.
To the grandparent re: temperature: when I first tweeted about this, I noted in a tweet that I ran all these tests with some fixed parameters (i.e. temp) but in a realistic environment and depending on the problem statement, you'd want to take those into account as well.
There's a lot that could be covered! But the post was getting long so I wanted to keep this really as... baby's first guide to prompt eng.