But literally the first sentence of the readme is "Prompt engineering is kind of like alchemy."
That's empiricism. Scientific method implies formulating falsifiable theoretical claims, producing meaningful analysis of experimental conditions, publishing peer-reviewed and reproducible results.
You're trying to correct people on a subject you know nothing about.
https://www.nationalgeographic.com/magazine/article/leeches-...
Science would be replicable, ie demonstrate that this particular approach to prompt writing yields better prompts than some baseline prompt writing approach across an array of different problems.
But these models are changing literally every day, so there's no fixed thing to reveal.
So no one will be able to reproduce anything at all.
This makes all this "engineering" pretty ridiculous in my eyes, it's literally for one models bizarre emergent properties.
Regarding randomness, the initialization of the weights is random, and if they use dropouts that is random too, plus the order in which to process the text might be random.
In the 90s and even 00s, the theory first crowd was mainstream and the empirical first crowd was considered fringe. Very fringe.
Personally I appreciated it when LeCun was like: “you can’t find the solution if you only search where the lamplight is shining.” Or other early deep learning practitioners note that ML theory is usually so far disconnected from practice in terms of tightness and bounds that you might as well ignore pure theory completely.
Anyway, it wasn’t until deep learning methods really smashed benchmarks across the board did people give in to the black magic / alchemy driven approaches of empiricism based upon intuition and bias developed through long-held experience.