Show HN: AdalFlow: The library to build and auto-optimize any LLM task pipeline
github.com
github.com
I will check more into the soft-prompt tuning.
For the current scope, we are focused on in-context learning, ways to improve model reasoning at the inference time.
We use auto-differentiative framework (backpropagation) to do zero-shot instruction optimization and few-shot demonstration. currently even just zero-shot can often surpass Dspy's few-shots (as many as 40 shots). And I have come up a training paradigm that will (1) start zero-shot (2) review performance from advanced teacher model to see if we can have a gap to gain from the teacher. (3) if there is a gap to teacher, we start to do low-shot demonstrations, and gradually increase the number of shots.
0.2.0 release highlight a unified auto-differentiative framework where you can perform both instruction and few-shot optimization. Along with our own research, “Learn-to-Reason Few-shot In-context Learning” and “Text-Grad 2.0”, AdalFlow optimizer converge faster, more token efficient, and with better accuracy than optimization-focused frameworks like Dspy and text-grad.
[0] https://github.com/stanfordnlp/dspy
Our benchmark has compared with Dspy and Text-grad(https://github.com/zou-group/textgrad)
We have better accuracy, more token-efficient, and faster convergence speed. We are publishing three research papers to explain this better to researchers.
https://adalflow.sylph.ai/use_cases/question_answering.html
We will compare with these optimization libraries but wont compare with libraries like LangChain or LlamaIndex. As they simply dont have optimization and it is pain to build on them.
Hope this make sense
Trainer.diagnose helps you get a final eval score across different splits of datasets: train, val, test, and it logs all errors, including format errors so that you can manually diagnose and to decide if the evaluation is too low that you need further text-grad optimization.
if there is still a big gap between your optimized prompt vs performance on a more advanced model with the same prompt (say gpt4o), then you can use our "Learn-to-reason few-shot" to create demonstration from the advanced model to further close the performance gap. We have use cases optimized the performance all the way from 60% to 94% on gpt3.5 and the gpt4o has 98%.
We will give users some guideline in general.
We are the only library provides "diagnose" and "debug" feature and a clear optimization goal.