DSPy – Programming–not prompting–LMs
dspy.ai
dspy.ai
classify = dspy.Predict(f"text -> label:Literal{CLASSES}")
optimized = dspy.BootstrapFewShot(metric=(lambda x, y: x.label == y.label))
.compile(classify, trainset=load_dataset('Banking77'))
label = optimized(text="What does a pending cash withdrawal mean?").label
What this does is optimize a prompt given a dataset (here Banking77).The optimizer, BootstrapFewShot, simply selects a bunch of random subsets from the training set, and measures which gives the best performance on the rest of the dataset when used as few-shot examples.
There are also more fancy optimizers, including ones that first optimize the prompt, and then use the improved model as a teacher to optimize the weights. This has the advantage that you don't need to pay for a super long prompt on every inference call.
dspy has more cool features, such as the ability to train a large composite LLM program "end to end", similar to backprop.
The main advantage, imo, is just not having "stale" prompts everywhere in your code base. You might have written some neat few-shot examples for the middle layers of your pipeline, but then you change something at the start, and you have to manually rewrite all the examples for every other module. With dspy you just keep your training datasets around, and the rest is automated.
(Note, the example above is taken from the new website: https://dspy.ai/#__tabbed_3_3 and simplified a bit)
Evaluations are first class and have a natural place in optimization. I still usually spend some time adjusting initial prompts, but more time doing traditional ML things… like working with SMEs, building training sets, evaluating models and developing the pipeline. If you’re an ML engineer that’s frustrated by the “loose” nature of developing applications with LLMs, I recommend trying it out.
With assertions and suggestions, there’s also additional pathways you can use to enforce constraints on the output and build in requirements from your customer.
Most important to me is that I can write evaluations based on feedback from the team and build them into the pipeline using suggestions and track them with LLM as a judge (and other) metrics. With some of the optimizers, you can use stronger models to help propose and test new instructions for your student model to follow, as well as optimize the N shot examples to use in the prompt (MIPROv2 optimizer).
It’s not that a lot of that can’t be done other ways, but as a framework it provides a non-trivial amount of value to me when I’m trying to keep track of requirements that grow over time instead of playing the whack a mole game in the prompt.
I have had precisely zero success with the LLM-prompt-writer elements. I would love to be wrong, but DSpy makes huge promises and falls painfully short on basically all of them.
I do not see any reason to use it.
Behind the scenes it's using LLM's to find the proper prompting. I find that it uses a terminology and abstraction that is way too complicated for what it is.
I've been using guidance, outlines, GBF grammars, etc. What advantage does dspy have over those alternatives?
I've learnt that the best package to use LLMs is just Python. These "LLM packages" just make it harder to do customizations as they all make opinionated assumptions and decisions.
Are there any beginner-friendly Python packages that you would recommend to facilitate fast experimentation with such ideas?
For your example; Write a python script with requests that hits the OpenAI API. You can even hardcode the API key because its just a script on your computer! Now you have the GPT-4proLight-mini-deluxe response in JSON. You can pipe that into a bazzillion and one different places including another API request to Anthropic. Once that returns, you can now have TWO llm responses to analyze.
I tried haystack, langchain, txtai, langroid, CrewAI, Autogen, and more that I am forgetting. One day while I was reading r/Localllama someone wrote; "All these packages are TRASH, just write python!"... Lightbulb moment for me. Duh! Now I don't need to learn a massive framework to only use 1/363802983th of it while cursing that I can't figure out how to make it do what I want it to do.
Just write python. I tell you that has been massive for my usage of these LLM's outside of the chat interfaces like LibreChat and OpenWebUI. You can even have claude or deepseek write the script for you. That often gets me within striking distance of what I really want to achieve at that moment.
I've settled on the Mirascope library (https://mirascope.com/), which suits my use cases and lets me implement structured inputs/outputs via pydantic models, which is nice. I really like using it, and the team behind it is really responsive and helpful.
That being said, Pydantic just released an AI library of their own (https://ai.pydantic.dev/) that I haven't checked out, but I'd love to hear from someone who has! Given their track record, it's certainly worth keeping an eye on.
I'm working on a naive approach to identify errors in LLM responses which I talk about at https://news.ycombinator.com/item?id=42313401#42313990, which can be used to scrutinize responses. It's written in Javascript though, but you will be able to create a new chat by calling a http endpoint.
I'm hoping to have the system in place in a couple of weeks.
https://github.com/jackmpcollins/magentic
It's based on pydantic and aims to make writing LLM queries as easy/compact as possible by using type annotations, including for structured outputs and streaming. If you use it please reach out!
In langroid you set up a ChatAgent class which encapsulates an LLM-interface plus any state you'd like. There's a Task class that wraps an Agent and allows inter-agent communication and tool-handling. We have devs who've found our framework easy to understand and extend for their purposes, and some companies are using it in production (some have endorsed us publicly). A quick tour gives a flavor of Langroid: https://langroid.github.io/langroid/tutorials/langroid-tour/
Feel free to drop into our discord for help.
I tried many python frameworks but the lack of customizations and observability limited the utility. Now, I only use Instructor (Jason Liu's library). That, and concurrent futures for parallel processing.
e.g Tests I want applied to anything retrieved from the database. What I'd like is to optimise the prompt around those (or maybe even the tests themselves) but I can't seem to express that in DSPy signatures.
I'm new to the library but from what I can see the Chat adapter will do this automatically if I use the forward call
Quoting the docs: "Though rarely needed, you can write custom LMs by inheriting from dspy.BaseLM. Another advanced layer in the DSPy ecosystem is that of adapters, which sit between DSPy signatures and LMs. A future version of this guide will discuss these advanced features, though you likely don't need them."
I could be wrong, I could be looking for complexity where there is none.
Have a fundamentally misunderstood how this all works, it sometimes feels like I have?
We took this kind of concept all the way to making a DSL called BAML, where prompts look like literal functions, with input and output types.
Playground link here https://www.promptfiddle.com/
https://github.com/BoundaryML/baml
(tried pasting code but the formatting is completely off here, sorry).
We think we could run some optimizers on this as well in the future! We'll definitely use DSPy as inspiration!
Even your toy examples look bad - wouldn't want to see what an actual program would look like.
Hopefully this, dspy and the like that have poor design, inelegant won't become common standards
I’m genuinely curious since if we can convince someone like you that BAML is amazing we’re on a good track.
We’ve helped people remove really ugly concatenated strings or raw yaml files with json schemas just by using our prompt format (which uses jinja2!)
- The GitHub page is very busy
- A clear example should come up early on the page. It's only when I got to the fiddle I could see a motivating example i.e the extractions, functions and tests.
- Then a section for running tests/evaluations
-Then deployment or run with/without the baml cli
- I do wonder if all the functions have to be so tightly coupled with the model. In dspy my modules are model agnostic and I can evaluate behaviour across different models.
Will incorporate this feedback.
- The GitHub page is very busy
- A clear example should come up early on the page. It's only when I got to the fiddle I could see a motivating example i.e the extractions, functions and tests.
- Then a section for running tests/evaluations
-Then deployment or run with/without the baml cli
- I do wonder if all the functions have to be so tightly coupled with the model. In dspy my modules are model agnostic and I can evaluate behaviour across different models. It's not so clear how to do this
What experiences/evidence do you have that informed your opinion? It sounds like you've had pretty negative experiences.
https://www.darinkishore.com/posts/mcp#building-tools-that-l...
I'd like the modules (dspy speak for inference ) to be aware of the chat history. Without this I'm not able to test the accuracy of the responses in the various chat contexts they appear.
It's been asked for in the TextGrad & Dspy GitHub issues.
https://colab.research.google.com/drive/1obuS9cEWN9MT-MIv5aL...