Auto-Differentiating Any LLM Workflow: A Farewell to Manual Prompting
arxiv.org
arxiv.org
This must be what AI hype actually is. Complete incoherent language to explain a very straight forward concept.
This is just: LLMs judging intermediate node outputs, and reverse traversing the graph while doing so until it modifies the original prompt.
Background difference I suppose.
> This must be what AI hype actually is. Complete incoherent language to explain a very straightforward concept.
True, a lot of papers overdo the jargon just for hype purposes. My favorite funniest example is this one from Google Research (and universities) (have linked the paper review video below)
See the YouTube chapter about "Multidiffusion" (around 38minutes)
They spent multiple paragraphs formulating an "optimisation problem" which when peeled down amounts to taking the mean, just to be able to superficially cite their own previous paper.
Quite the sorry state of things.
Damn, I knew we were lazy but describing prompting as labor-intensive is impressively lazy even to me.
Obviously reading the rest of the abstract was too labor intensive for me but I'm hoping I can just hook a probe up to my drool and it can infer my desires from that.
Setting parameters for any ml model is easy, but we'd call it labour intensive if we expected people to do it manually despite having evals. Instead we have ways of searching for and optimising settings. The methods for that are obvious for small cardinality discrete values or continuous variables. Less so for arbitrary text.
This work introduces a way to treat these prompts like trainable parameters, updating them through automatic differentiation of some kind of supervised training loss.
For me it kind of feels like deep dream or style transfer, which use autograd to optimize the model inputs (instead of the parameters) to achieve some goal (like mixing the style and content of two input images)
To the extent that you need to eke out reliability on the margins, one is vastly better served by actual fine-tuning, which is available both for open-source models and most major proprietary models.
I don't think anyone has pretrained a remotely-close-to-SOTA sized backwards model.
We are continuously adding more benchmarks to the paper with UTAustin.