The Problem with LangChain
minimaxir.com
minimaxir.com
Frankly I could see langchain is garbage software just by looking at the code.
It still helps me get shit done fast to figure out how things are supposed to work. Sort of a cookbook of AI recepies. Once I have an approach narrowed down I'll rewrite everything on top of stuff langchain is supposedly wrapping. For now it's faster than tracking down individual libraries and learning the apis. It will stay in notebooks basically.
The goal is to capture all that knowledge that langchain has, into consistent legos that you can combine and parameterize with the prompts, without all the complexity and boilerplate of langchain, nor having to learn all the Python libraries and their APIs. Perfect for prototypes and experiments (like a notebook, as you suggest), and then if you find something that really works, you can hand-off a single text file to an engineer and they can make it work in a production environment.
After scoffing at your shameless plug I read some of the code and was surprised to find it was not written in perl. Maybe some other options for the acronym could be:
* LASI (Language Artificial Specific Intelligence) * LLMML (Large Language Model Markup Language) * LLMMA (Large Language Model Metalinguistic Abstraction [this one is a stretch])
At this point Langchain is almost required to accept any PR. They raised money, and now their growth metric is Github stars.
A simple wrapper around the APIs (I used llamaflow, which is now llm-api https://github.com/dzhng/llm-api) and a templating engine is most of what you need.
AI is not at a point where generalist prompts to do agent/memory/search things is a good idea for a real product. You need to integrate procedural guidance unless you want your UX to be awful.
Working on https://github.com/eth-sri/lmql (shameless plug, sorry), we have always found that compositional abstractions on top of LMQL are mostly there already, once you internalize prompts being functions.
import llm
model = llm.get_model("gpt-3.5-turbo")
model.key = 'YOUR_API_KEY_HERE'
response = model.prompt(
"Five surprising names for a pet pelican"
)
print(response.text())
Or you can stream the responses like this: response = model.prompt(
"Five diabolical names for a pet goat"
)
for chunk in response:
print(chunk, end="")
It works with other models too, installed via plugins - including models that can run directly on your machine: pip install llm-gpt4all
Then: model = llm.get_model("ggml-vicuna-7b-1")
print(model.prompt(
"What is the capital of France?"
).text())
It also handles conversations, where each prompt needs to include the previous context of the conversation: c = model.conversation()
print(c.prompt("Capital of France?").text())
print(c.prompt("what language do they speak?").text())
I wrote more about the new plugin system for adding extra models here: https://simonwillison.net/2023/Jul/12/llm/I went back and forth on a bunch of different designs, but eventually decided to try to make it so that each plugin that implemented a new model would only have to subclass Model and add new methods.
You can see all of the design arguments I had with myself about this here, across the course of 129 pull requests comments: https://github.com/simonw/llm/pull/65
Is support for embedding and querying a corpus of custom text planned at all? 99% of what I wanted to use langchain for was building a chatbot that can answer questions about my own documents.
Alex Garcia built a better one here as a SQLite Rust extension: https://github.com/asg017/sqlite-vss
Hearing this perspective helps put my frustration in context. I need to lower my expectations and just get used to its quirks. Despite its issues I've had a ton of fun building with langchain and will keep using it.
Read this page and mentally swap "chain" for "function": https://python.langchain.com/docs/modules/chains/foundationa...
They build all these adapters and integrations and make it seem like they're helping you piece together a solution, but in how many cases were they necessary as a middleman? Like I really don't need a wrapper around the openai client, and for the more complex stuff like Agents, isn't this the most critical part of your app, if it's production? And for notebooks, is it any better than coding directly against your llm api? You probably won't swap llm backends, and if you're using the notebook for education/documentation, I think I'd want to show the actual openai api calls.
OpenAI's official documentation gives you the code for a tool-running agent anyway. Taking that and editing it can probably be done faster than pip-installing langchain and navigating its docs.
Some prefer React, somet NextJS. LangChain has its place.
By the time you have built custom chains, custom prompts, and custom agents to support all that, you basically are using their interface and not their code. At that point, eh, it's pointless. It's great for demos, I give you that, but every time I tried to coherce it into a product, it fell short.
I wish the code quality was better, but poking around their docs does give pretty interesting ideas you can build yourself, even if you don't use LangChain. I think the release of OpenAI function calling has also just kinda sideswiped the need for large parts of these kind of frameworks — you don't need much help in coercing to JSON or parsing anymore if you use the function calling API.
1: https://python.langchain.com/docs/modules/model_io/prompts/e...
Also: before using any new technology, go to hn.algolia.com and search HN for comments about it first.
(def chatgpt
(-> (endpoint :chat "gpt-4" auth)
(chat/set-opts [:endpoint] {:stream true})
chat/retry
chat/catch-unkown-commands
;; chat/history
chat/string-input
(chat/file-io "file.txt" :stream true)
#_(chat/preserve-ctx chatgpt-ctx)))
These higher-level functions take as input the next function in the chain, and return a function that is responsible for calling it and takes a context hashmap as argument. Additionally, you can pass messages to these functions like :history/reset, impacting the atom defined in the let part of lolol. I use the same pattern for higher level constructs:[1]https://letoverlambda.com/index.cl/guest/chap2.html#sec_6
I was thinking about doing something similar with dynamic scoping in Emacs Lisp.
Here's how level-1 functions are defined. 'ctx-fn' is just an indirection macro for 'fn' I haven't really used yet.
(defn preserve-ctx [f & [at]]
(let [mem (or at (atom {:ctx true}))]
(ctx-fn
:preserve-ctx [ctx]
(if (->> ctx :it (cmd? :preserve-ctx))
(do (case (:it ctx)
:preserve-ctx/get @mem
:preserve-ctx/reset (doto {:ctx true}
(->> (reset! mem))))
(println (-> ctx :it) "done."))
(-> ctx
(when-not-> (-> :states :preserved)
(as-> $ (ctx-it (merge @mem $) (-> $ :it))))
f
(assoc-in [:states :preserved] true)
(doto (->> (reset! mem))))))))
Edit: reading what I wrote, 'endpoint' is in fact a function returning function, except that since it's the "end" of the chain, it doesn't take a function to call next as argument.The most advanced lib for dealing with LLMs is
https://github.com/zmedelis/bosquet
There is also https://github.com/cjbarre/multi-gpt/tree/main but it hasn't been update in 3 months and seems rather basic.
Alternatively, you can shoot me an email at
(->> '(102 117 110 116 97 105 110 64 109 101 46 99 111 109) (map char) (apply str))
and I'll prepare a repo for what I've been working on. It's usable but I wanted to clear some things up before a public release.
I mean, it starts by giving the full name of the original author… calls using his free software a waste of time… destroys the design as if tearing down someone’s work who made bad decisions at every turn when making a brand-new thing that is FOSS.
Did anyone else find that a bit strange? Did I miss the light-hearted tone buried somewhere?
That said, strangely… useful examples and up-to-date detailed technical analysis that I like. I just don’t get why this sector gets so nasty between people… fight the robots, guys! ;)
Anyway, for those of us who've used LangChain a bit, I don't think anything in the article is revelatory, the codebase is structured as a first mover project, with all of the copy+paste and poor abstractions that come with it. It's absolutely great for prototyping and simple scripts, but I don't think I would use it in a production codebase, I found it very difficult to work around the bugs and wonky abstractions.
> No one wants to be that asshole who criticizes free and open source software operating in good faith like LangChain, but I’ll take the burden. To be clear, I have nothing against Harrison Chase or the other maintainers of LangChain (who encourage feedback!). However, LangChain’s popularity has warped the AI startup ecosystem around LangChain itself and the hope of OMG AGI I MADE SKYNET, which is why I am compelled to be honest with my misgivings about it.
My main headfake in this actually revolutionary-incremental advance over the past 8 or so months… has been more similar to yours than I read at first. Nice post!
"simpleaichat is a Python package for easily interfacing with chat apps like ChatGPT and GPT-4 with robust features and minimal code complexity. This tool has many features optimized for working with ChatGPT as fast and as cheap as possible, but still much more capable of modern AI tricks than most implementations"
https://github.com/minimaxir/simpleaichat
Separately, in the article, typo expect->except here:
"LangChain uses about the same amount of code as just using the official openai library, expect LangChain incorporates more object classes for not much obvious code benefit."
https://blog.langchain.dev/announcing-our-10m-seed-round-led...
Anyone who's been burned by LangChain, especially now that it has VC funding, has to be worried that LangChain will become the cross-LLM standard library™, and they'll be dealing with it and endless patches to it for the rest of their lives. (Think systemd or NPM or Python packaging.) If it's as bad as described, the time to stop LangChain is to strangle it in the cradle, before it can get too far or risks finding a killer-app/niche which will immortalize it no matter how bad it is.
> The SK extensible programming model combines natural language semantic functions, traditional code native functions, and embeddings-based memory unlocking new potential and adding value to applications with AI. > SK supports prompt templating, function chaining, vectorized memory, and intelligent planning capabilities out of the box.
I can't speak from experience whether it's better than LangChain, however.
I’d probably spend 15% too many manhours making it bespoke…
But The Creator of whatever just solved a problem before I did. Of course I’m going to use that until a better solution is practical.
People are taking it personally because there's shared a realization that he represents the Crypto-ization of raising for AI based startups: going full hype leader over substance.
I do recommend prototyping; even if you think your idea is feasible, you want to see it in action before committing significant resources. Personally, my prototype showed me that a single run of my idea takes more time and money than I initially expected.
To illustrate the complexity of this, here's a list of things that you might have to do if implementing a document bot:
1. Handle uploading or storing documents somewhere and keeping track of the location.
2. Handling different document types. Sticking to PDFs for this list.
3. Manually or using the PDF to augment the documents with tags, keyterms, titles, etc.
4. At this point you need somewhere to store the metadata. Maybe a DB or using the vector store.
5. Just dealing with PDFs requires some type of PDF library. Other documents may or may not require an additional library.
6. Extracting text from the PDF with something like pdf2image. Not all PDFs have extracted (selectable) text in them. Also, PDFs have images, which sometimes have text.
7. Doing some sort of OCR is very likely. Think about OCR'ing the whole thing to deal with no extracted text/images with text.
8. Assuming that, converting pages to images. Also, consider images have data in them, so extracting an image from the image and running some type of detection on it...
9. Using some OCR model to extract text, or figure out how to extract them from the PDF data.
10. Cleaning up that text, then parsing it cleanly. nltk comes into play here.
11. Fragmentation/windowing of text. How long to create the fragments? Or should it be variable?
12. Using the text to get more text via a prompt to a model. Here we can get additional keyterms, or perhaps a summary or question about the text fragment. (we'll use something we write for prompting the LLM here in a second)
13. Storing the fragment. Most people use a vector database for this now, so we can use Weaviate or Pinecone, or ???. Also, consider a moderate amount of fragments and their vectors can be stored in a pickle format with manual dot products for ranking.
14. Figuring out where you are going to get a user prompt. Assuming the easiest thing, collect input from the user in a command prompt.
15. What to do with the user's prompt once you get it. Do you ask an LLM for more info on the prompt? Or do you just jump to...
16. Embed the user's prompt to get a vector back. (Weaviate does this transparently, but you can easily do it yourself using the ada-002 endpoint from OpenAI)
17. Taking that vector (from the embedding/inference to the embed model) do a comparison to other vectors/text you've stored.
18. Think about what text is important for a new prompt to the LLM. Should it contain directives? How much reference text from the documents does it need? Is a cosign distance or some approximate nearest neighbor match going to be enough?
19. Think about augmenting the vector search with keyterms that were extracted earlier (by both the PDF itself + any LLM inference step you impelment)
20. Take the text you pull back from wherever you stored the vector/text and then build a long string to stuff into a prompt.
21. Consider some type of template structure for the prompts, so you can tweak them without losing your mind. String templates for files in Python are great for $this.
22. Calling the various LLM endpoints. There are multiple models, in a variety of API endpoints, with tokens usually for auth.
23. Consider you may just want text back from the LLM, or maybe you want it to complete or write a dict or array (in which case you may want to make this configurable). You may want to eval things that the LLM writes too.
24. Consider the LLM (ChatGPT for example) may do a completion that contains a block delimited by ```python or similar.
25. Consider those two things may require different completion endpoints, and some endpoints may be deprecated by the provider later.
26. Think if you need function completion calling. GPT-X supports this, so you need a function and a way to pass that function's parameters to the LLM.
27. Build the prompt and submit it. Don't forget to protect your tokens, using env or config.py files.
28. Take the response and do something with it that makes sense. Maybe give it to the user, or use it to build another prompt.
29. Loop back to interact with the user. If the interaction is complicated, like with Discord integration, you may have to do this asynchronously.
27. Think about storing the interaction for use in building future prompts. Hack this into #15.
28. Always consider optimizing your prompt length.
29. Consider how many tokens you are chewing through doing all this.
30. Consider questions by the user about "what is on page 2?" need context. Another good one is "how many pages is this document", or "what is the title of the document?". A hard one would be "how many images are in this PDF?", meaning how many illustrations...
31. If the document discusses code, and the model outputs code, or SQL, do you run it and if you do, how?
Example of most of this in action: https://github.com/FeatureBaseDB/DoctorGPTI tried experimenting on building a library that makes it easy and transparent to use LLM https://github.com/adityapurwa/jehuty and tried the middleware approach that might be more familiar in general. Its an experiment so the API might changes a lot until we find a sweet spot. If you have an advice or suggestions it would be helpful and appreciated.
I'm especially looking forward to playing with the structured data models bit: https://github.com/minimaxir/simpleaichat/blob/main/examples...
Well done, Max!
Disclaimer: I am the author of txtai
As long as you're cycling data between LLM unstructured text is going to be just fine.
Have your tools accept natural language as well, and you're golden.
In between you may want some form of data storage, and natural language can be that as well, as it makes retrieval trivial for LLM.
It also feels like you can keep it more on track by telling it to put the data into specific fields with very specific descriptions.
My personal preference is Llama Index by Jerry Liu. It excels at clearer docs + better abstraction.
It’s still very early days for software composing AI models and we almost certainly don’t have all the right metaphors yet. And I think there is a lot to be said for strong typing and simple, robust code!
I also played with huggingface's transformer agent (https://huggingface.co/docs/transformers/transformers_agents ) and thought it was a lot easier to useas far as the tools go, though is perhaps less capable for other things. I may go back to playing with that actually.
This is an example of using pgvector that might be relevant for your use case:
SELECT d.id, d.doc, 1 - (embedding <=> (SELECT embedding FROM documentation_embedding WHERE id = 1)) AS similarity FROM documentation_embedding de join documentation d on de.documentation_id = d.id ORDER BY similarity desc;
be cautious with pgvector - the recall can be extremely bad (% retrieved nearest neighbors vs ground truth)
However, I think that langchain's minor version should be increased to 1 or more. The current latest version of langchain is 0.0.234. I think it's a problem that the use in production is very peaky because all the changes that should be used properly by minor, patch version, etc. are all lumped together.
But as always I trust the OSS community to make it better or replace it with something else
Decided to reimplements using embeddings and my own glue code. Took like a week (much less than the langchain work), and it's cheaper and better
With OpenAI functions though, I find it easier to just make a local sequence of function executions than work w langchain abstractions.
If one can somehow reuse those integrations without even care about/use langchain, it's a massive time saver.
I haven't looked at the code so not sure how reusable they really are.
If I want to do something I have to rely on other sources to use it. That's just an overkill.
HelloWorldPrint(Baseprint): @validators def input_variables...
bros really just need llm.call(), I ended up rewriting my own tools which goes 10x faster for me personally.