Building Boba AI: Lessons learnt in building an LLM-powered application
martinfowler.com
martinfowler.com
Most of the time on that project I've just played around with different prompts to make it do what I'm hoping for, with almost no mental process for understanding problems and solutions, mostly just randomly coming up with experiments and reading the results very carefully, checking how consistent they are.
I moved towards finding and isolating the parts LLMs are good (and reliable!) at and using deterministic approaches for everything I possibly can. That part is not too tedious, but all this black box trial and error (with all the waiting and errors)...
Luckily the client doesn't expect LLMs to be what the hype says and mainly just wants some reasonably useful features so they can say they use AI - wouldn't enjoy dealing with someone thinking it's super easy and I just need to write a few scripts because they tried this in ChatGPT once.
I see a lot of companies that are experimenting with it, and there are a few press releases here and there, but there is a next generation that will be more ambitious. You just can't build more ambitious in the timeframe yet.
There's a insidious reason for this. It's because LLMs are one of the few technologies we don't fully understand and we can't fully control.
It's a stark contrast with traditional engineering.
I would urge companies and people to build their own stuff because there is so much value in learning the tech. For off the shelf LLM tools, it is hard to beat OpenAI’s web app, Microsoft Bing+ChatGPT and Office 365, and Google’s beta Bard integrations with Google Docs, etc.
Too short to call it a book.
The average non tech person doesn't know what an http API is.
Maybe there's a killer unicorn app hiding in one of those wafer thin wrappers, but OP's point is precisely that it's disappointing how wafer-ish all these "exciting applications" actually are, relative to the massive sense of expectation that existed only a few months ago.
All these projects are desperately slapping buzzwords such as ‘LLMs’, ‘AI’, ‘ChatGPT’, etc to pretend that their product is somehow revolutionary despite sitting on someone else’s AI model in the cloud via an API that they do not own.
They are slowing realizing that there is little to no moat with these LLMs let alone any serious use cases other than summarizing / rewording text.
Everything else requires the triple checking of another human to review the bullshit that the so-called black-box AI outputs.
This was an extreme bizarre example, but most mobile apps from the first wave were also pretty dull
Give it some time. I think something similar will happen: most LLM apps from this first wave will be forgotten soon, and when people shift the focus from the technology back to the users and their real pain points, the good stuff will surface
Does it cross the threshold?
I'm going to check out you're browser thing later tonight, it looks good!
https://news.ycombinator.com/item?id=36529885
I realized how primitive even pure prompting is. Stable Diffusion is kind of primitive too, but the input/prompting methods it has are lightyears ahead.
NLP folks don't know shit about prompt engineering, which is ironic.
https://gist.github.com/Hellisotherpeople/45c619ee22aac6865c...
I still can't emphasize certain tokens in ChatGPT, or mathamatically average them. Not sure why NLP folks don't bother implementing these things, even in the oogabooga frontend (which is supposed to be the automatic1111 of LLMs)
But recently I threw a just a bit of similar-ish stuff as you describe there into a TTS model, barely knowing anything, and yeah it's totally works and is fun and cool. The stuff that doesn't work fails in interesting and strange ways, so it almost STILL works. (Well, it gives people really bizarre speech impediments, at least...)
I was just working on prompt editing actually. Which is weird to imagine in a TTS model. It makes sense for the future tokens of course, for words the model has not said yet. I think it even makes sense for the past right? You can rewrite the past context, and it still changes future output audio model. In bark it's two different things: one is the text prompt, and one is the generated audio tokens/context, which is not the same. (The text and the past audio is concatted in the Bark prompt, so this idea makes sense in Bark but not in other models. You could change either text OR 'what was generated with the text' independently.)
As long as you don't rewrite the time touching the last token, at 0 seconds - if it's like a segment 2 to 4 seconds in the past, it should influence future output but not cause a discontinuity in the audio. I think?
BTW an easy and fun thing - just let generation parameters be dependent variables. Of anything.
A trivial example: why is temperature just a number, why not a function? Like the temp varies according to how far long in the prompt you are. For music, just that is already a fun tool. Now as a music segment starts or ends the style transitions. Or: spike the temperature at regular intervals - like use a sine wave for temp, input is current token position. You can probably imagine that works great in music model.
Even in a TTS model this you can get weird and diverse speech patterns.
The thing is: I really very a low level of competence. Total monkey hitting keys and googling, and even I can make it work, easily. Sampling is just a loop, okay, what if I copy logits from sample A and subtract them from sample B. What if take the last generation, save the tokens, ban then in the next. Really just do anything and you end up in interesting places in the model you didn't know existed and are often cool. (Recently, TTS output overlapping speech, for example.)
Like I recently generated french accents from any voice in the Bark TTS model, with no fine-tuning, no training, actually not even really any AI. Just by counting token frequencies in the french voices, and having the sampler loop go, "Okay let's bump these those logits up a bit, and the others down" and it just somehow works. No Loras, no fine-tuning, not stats, it's like middle school level math, but sounded great.
(I'm in a bit of a stream of consciousness ramble mode from lack of sleep, but I'll keep going on this message anyway so I don't forget to come back to your post when I'm back at normal capacity. And just hope I don't cringe too hard reading this when better rested.)
Oh I'd love to hear your thoughts on negative prompts in LLMs.
1) What does 'working correctly' look like?
For an audio LLM, I'm thinking something like: a negative prompt "I'm screaming and I hate you!!!" makes the model more inclined to generate quieter, friendly speech, in your positive prompt. Something like that?
2) How to make it work.
This is probably very model dependent and fiddly. My first thought was generate two samples in sequence. The first sample is the negative prompt. Save all the logits and tokens. Use them as a negative influence in the second prompt. At least in Bark you can't just like flat subtract them or what you actually get is more like 'the opposite of speech' than 'the opposite of your prompt' but when I did french accents I basically just fiddled with a bunch of constant values and weights and eventually it worked. So I'm hoping the same applies. I can imagine a more complicated versions where you do some more math to figure out what's unique about a text prompt, versus 'a generic sentence from that language' and only push on those logits. I suppose that might be necessary.
I started off with something to create AI art, but that didn't really take off. I'm also disappointed with Dall-E 2 lagging behind the others in terms of image quality. So now I'm focusing on code generation.
* Use a text template to enrich a prompt with context and structure
* Tell the LLM to respond in a structured data format
* Stream the response to the UI so users can monitor progress
* Capture and add relevant context information to subsequent action
* Allow direct conversation with the LLM within a context.
* Tell LLM to generate intermediate results while answering
* Provide affordances for the user to have a back-and-forth interaction with the co-pilot
* Combine LLM with other information sources to access data beyond the LLM's training setThis is a game changer to the UX
I’d like to incorporate this in a production workflow for generating schema-compliant test data for use in few-shot promoting - would you mind saying a few words about your medium term plans for the library? The LangChain API is changing all the time at the moment so we’re trying to figure out where it’s safe to stand. No expectations, of course, just curious.
I noticed some occasional funkiness from GPT-4 around sending back properly formatted dates yesterday but haven’t yet dug into it properly. Might be a good candidate for a transformation.
why though?
If AI can build me a useful marketing plan in 15 mins with multiple agents doing work this seems fine to me. It’s going to take much longer to get a human involved.
It still feels like a bit of a Wild West for patterns in this area as yet, with a lot of people trying lots of things and it might be too soon for defining terms. A useful resource is still things like the OpenAI Cookbook, that is a decent collection of a lot of the things in this article but with a more implementation bent.[1]
The area that seems to get a lot of idea duplication currently is in providing either a 'session' or a longer term context for GPT, be it with embeddings or rolling prompts for these apps. The use of vector search and embedded chunks is something that seems to be missing so far from vendors like OpenAI, and you can't help but wonder that they'll move it behind their API eventually with a 'session id' in the end. I think that was mentioned as on their roadmap for this year too. The lack of GPT-4 fine tuning options just seems to push people more to look at the Pinecone, Weaviates etc stores and chaining up their own sequences to achieve some sort of memory.
I've implemented features with GPT-4 and functions and so far it's feeling useful for 'data model' like use (where you're bringing json into the prompt about a domain noun, e.g. 'Tasks') but is pretty hairy when it comes to pure functions - the tuning they've done to get it to pick which function and which parameters to use is still hard going to get right, which means there doesn't feel like a lot of trust that it is going to be usable. It's like there needs to be a set of patterns or categories for 'business apps' that are heavily siloed into just a subset of available functions it can work with, making it more task-specific rather than as a general chat agent we see a lot of. The difference in approach between LangChain's Chain of Thought pattern and just using OpenAI functions is sort of up in the air as well. Like I said, it still all feels like we're in wild west times, at least as an app developer.
By far, the best resource I've found is the Prompt Engineering Guide: https://www.promptingguide.ai/
> you can't help but wonder that they'll move it behind their API eventually with a 'session id' in the end
For in-context learning, I think it is fair to expect 100k to 500k context windows sooner. OpenAI is already at 32k.
Agreed, that is a good resource for sure. For tooling I like https://promptmetheus.com/ but any pun name gets bonus points from me.
> For in-context learning, I think it is fair to expect 100k to 500k context windows sooner. OpenAI is already at 32k.
It has been interesting to see that window increase so quickly. For LLM context the biggest thing is the pay-per-token constraint if you don't run your own, so have to wonder if that is what will be around in the future given how this is trending? Just in terms of idempotent calls, throwing everything in context up every time seems like it makes it likely that OpenAI will encroach on the stores side as well and do sessions?
Learning implies that the underlying weights of the LLM changed. They didn't.
If the probability of the model spitting out something bad is 0.01% will my testing find it? Probably not.. but my users certainly will.
Closed feedback loops to improve the quality of reasoning for LLM's requires human-in-the-loop tools.
Prompting methods are often a hit and miss as the same prompts can lead to variable quality outputs.
I think that is wrong - it is a config management problem. Think Prompts X Chains X LLMs. Your prompts wont work across everything and everything will break on model change. Coding this into ur classes is what everyone does.
Instead we pull out the prompts X chains as jsonnet code. Call it trauma & learnings from the K8s/Borg world. We have formats that have evolved as a result of millions of lines of code wrangling clusters/terraform/etc - so we decided to build a SDK over it.
that is what we did here - https://github.com/arakoodev/EdgeChains/releases/tag/0.2.0
EdgeChains is basically Generative AI prompt engineering modeled as config management. Funnily people organically build this out over 6 months of engineering AI applications. That's what Martin did with templating this as text !!
I think a lot of people are learning these lessons in isolation, I do wish there was a centralized place where people working on UX-focused LLM based apps were exchanging lessons
HN has been a pretty good source of exchanging knowledge so far, every couple days or so there's a write up like this that has some new tidbits or confirmations of ideas. If everyone keeps doing that we're doing great in my opinion. Looking forward to seeing your write up on here!
Multi-modal models are going to change things even further.