Hard stuff when building products with LLMs
honeycomb.io
honeycomb.io
The engineers I know whose job it is to implement LLM features are much more skeptical about the near future than the engineers who are interested in the topic but lack hands-on experience.
All to figure out the singular most important thing: chat interfaces are the worst.
Having an LLM spell-checker that would autocorrect my spelling as I typed, based on the context of what I was typing? That would be magnificent.
It does have much more context to work with (open tabs, other files) vs. a single/short chat session.
Since it's just OpenAI's text completion model with a code finetune and without the chat/assistant RLHF, it works much better as an "advanced autocomplete" than ChatGPT or even OpenAI's Turbo model via their API. I can be much more surgical with how I use it (often accepting just a few words at a time), and it's good at following my usual tone.
At my org we use a chatbot for pull requests — you get pinged by the bot when the PR is ready to merge, with a button in the chat interface that merges the PR — no need to open GitHub and locate the big green button yourself.
That won’t 10x your productivity or whatever, but it does make it slightly more pleasant.
Can’t agree more. As an user, a chatbot makes me think the company has put some kind of dumb parrot in front of me in order to avoid giving actual support.
There were search engines. Plenty. And indexes. And ... That wasn't it.
Every site had search. That also wasn't it.
The problem was that these things were incredibly low quality versions. Exactly what you're complaining about now with chatbot interfaces. In fact that these were such low quality is what made Google such an incredible opportunity: centralization came almost built in. They never needed to fight books.com: their search sucked. Even now finding a book using Google's interface works better than all search on the internet, except perhaps Amazon. And Amazon's "shittifying" it's search engine too now ...
Google was also never "the best". It was simply consistent quality that worked everywhere. And because all other search engines were pretty far along in their enshittification cycle, with MBA's unwilling to go back, nobody even made a serious attempt at fighting them. Except, perhaps, and very late to the party, Microsoft.
Google was better, but not incredibly better (and it's been going downhill for like 5 years now). It makes me think that if you could make something that would advise 10% of the world population on how to boil eggs, that would be incredible.
I also hope there will be some whistleblower within OpenAI (and others like it) that exposes its internal practices and all of the hypocricy surrounding it. And, usually the fish rots from the head, as they say.
To be brutally honest, I expected better from Sam, but he has lost all credibility in my eyes based on how they chose to roll out ChatGPT. I now see that he’s even more hawky and daring and manipulative than Zuckelberg ever was.
This might result in a sort of transformation that engineers and power users aren't geared to appreciate. You might look at a natural language log query and say, "that would actually slow me down". But if it makes Honeycomb suddenly useful to stakeholders who couldn't before, it could lead to use cases not on the radar right now.
- How easily can you query stuff when you're interested
- How easily can you get other people on your team to use the product too
- How quickly can you narrow down a problem (e.g., during an outage) to something you can fix
- How relevant is your alerting (i.e., SLOs) to the success or failure of something business critical
Our bet here is that the first two could potentially be improved by using LLMs, since we hypothesized (and confirmed in some new user interviews) that there's an "expressivity gap" in our product. A lot of people who aren't already observability experts, but do have some vested interest in observability, often know what they want to look for but get confused by a UX that's tailored for people who are more familiar with these kinds of tools.
It's only been 3 weeks so it's too early to tell, but we're seeing some signs that the needle is being moved a bit on some key metrics. We're not betting the farm on this stuff just yet, and it's really cool that there's technology that lets us experiment in this way without having to hire a whole ML engineering team.
And agreed re: development effort - compared to other hype cycles of AI, it's important for folks to understand that the results they see are coming at a fraction of the experimental budget.
Those who figure out how to do that well will have quite a legacy under their belts, and money if they're a profit-making company and handle it well.
It's not about whether or not it can be done, it's not hard to lay out the math and prove that pretty trivially if you know where to look. Just actually doing it is the hard part in a way where the inductive bias translates appropriately to the solution at hand.
There will be a lot of wasted effort in pursuit of this unreachable goal (with LLM technology), an effort better spent elsewhere, like solving cancer or climate change, and stealing young and naive people’s minds away from these problems.
If you had an intern at your beck and call, who was obedient but knew nothing about anything except how to use a search engine, would that make you a "10x" at whatever you do?
That said, my beef is not with its use for this purpose - whatever makes people happy and feel more productive, it’s all good and fun. The problem is when the CEO of a company like OpenAI starts talking with a straight face in his Congress testimony how this not-really-new tech has the potential to run afoul and enslave us (I’m heavily paraphrasing, obviously). That is the problem - this particular CEO suddenly feels entitled enough to make a giant PR stunt by using his official testimony to endow his product with the qualities it cannot possibly have - that is the problem. The blatant lies and deception is frankly disgusting.
In response I said, “While we’re at it let’s encourage people to stop making music and peanut butter as well.”
Then dinvald said, “Your point?”
Can you explain to dinvald what my point was?
ChatGPT:
Absolutely. Your response is a form of reductio ad absurdum, a strategy where one extends an argument to an absurd extreme to point out its flaws. In this case, you are challenging dinvald's claim that efforts towards developing Large Language Models (LLMs) is a waste of resources that could be better spent on solving other problems.
By suggesting we also discourage people from making music and peanut butter, you are demonstrating that dinvald's argument could be applied to any human endeavor that isn't directly tied to solving major global issues. Your point is that not every pursuit needs to be measured against solving the world's largest problems, like cancer or climate change. You imply that just like music or peanut butter, AI development also has its place and value.
If you want to eat a peanut butter sandwich that's marketed as a panacea that can cure all diseases, while not even knowing if the butter has long expired, then all props to you, but I don't.
We've all lamented the disinformation campaigns made by humans recently, well this is like a disinformation engine on atomic steroids! And no, they're not going to "solve it" by feeding it more data and parameters - this problem is inherent in LLMs, and in fact gets harder the more data there is.
Of course, as always, anything that comes out of Silicon Valley is a fad that is shoved down our throats as the best thing since sliced bread. But in reality, it's completely unnecessary, as it only marginally improves the quality of life of the 0.1% of the population, while simultaneously decreasing it for the rest. That is the sad reality we live in, my friend.
Hence, as a distilled set of operators in task space, that disambiguation then becomes "reasoning" for a number of people. We're not saying they're human. Under reasonable mathematical constraints you can show pretty clearly reasoning. Heck, even without it you can have good inductive tests.
I'm not sure the core pith of what you're getting at but I have a suspicion that there's something else under there other than the raw mechanics of how LLMs do or don't work. Would this be a correct assertion on my end? :) <3 :D :)
Instead, to get more educated on it, I'll read Gary Marcus's "Rebooting AI: Building Artificial Intelligence We Can Trust" that discusses research on AI reasoning in more detail ;)
I will probably also look into "Ghost Work: How to Stop Silicon Valley from Building a New Global Underclass" by Mary L. Gray and Siddharth Suri, which discusses already-existing huge ethical violations around building these AI systems by large corporations like Google and Microsoft-backed OpenAI :D <3
If you have any other good reads on this (particularly the 2nd topic) based on your experience, I'd greatly appreciate it!
If you're looking for the more technical side of things (reasoning, etc) then I'd recommend taking a look at the original Shannon and Weaver paper, plus information topology from there on out. That field alone is interesting enough to dive down to a Ph.D. in.
Re reasoning, I'm more curious what you think of the non-statistical school of thought, e.g. the arguments against LLMs as "reasoning" engines, as popularized by Noam Chomsky and Gary Marcus.
But I think we've only scratched the surface as to what LLMs fine-tuned on specific tasks, especially for abstract reasoning over narrow domains, could do.
These applications possibly won't look anything like the chat interfaces that people are getting excited about now, and fine-tuning is not as accessible as prompt engineering. But there's a whole lot more to explore.
“I tried a known-bad early iteration of a technology that has already been superseded — hence it is bad and will always be bad.” is not a convincing argument.
The first transistor was an ugly and impractical thing too.
It doesn't matter though. I would prefer you don't use it.
That to me feels like there's some prompting improvements that you could do. It could be that your problem is just harder for LLMs than ours, but 20-50% of the time isn't what we observed after several prompting changes. The other thing is that we do regularly get outputs that are "mostly correct" and we can actually correct them manually, so while the model may have a higher fault rate, the actual end-user experience has a lower one.
Isn't that always the case?
Photoshop: https://the-decoder.com/adobe-photoshop-can-now-modify-image...
Robotics: https://youtu.be/UuKAp9a6wMs
Home automation: ( couldn't find a good video, but it's with home assistant)
The lack of formal tooling for prompt engineering drives me bonkers, and it compounds the problems outlined in the article around correctness and chaining.
Then there are the hot takes on Twitter from people claiming prompt engineering will soon be obsolete, or people selling blind prompts without any quality metrics. It's surprisingly hard to get LLMs to do _exactly_ what you want.
I'm building an open-source framework for systematically measuring prompt quality [0], inspired by best practices for traditional engineering systems.
Any thoughts on managing costs? I've been developing against gpt-4, and it runs up charges quickly. I've been thinking I will need to be careful about adding live api calls in any sort of testing situations.
Wondering if your tool has any features to help avoid/minimize wasted api usage?
In my experience with cloud-based and local models, LLM chaining compounds errors; I would urge you to look at few-shot, single-query interaction models for business applications using LLMs as a new unit of compute.
We're going to likely settle on just storing whatever version of the prompt is considered "stable" as a source file, but for now this isn't actively hurting us, as far as we can tell, and there's a lot of prompt engineering left to do.
No idea if I’m doing it right.
The gotcha is that it’s a search problem. The article mentions embeddings and dot product, but that’s the most basic and naive search you can do. Search is information retrieval, and it’s a huge problem space.
You need a proper retriever that you tune for relevance. You should use a search engine for this, and have multiple features and do reranking.
That’s the only way to crack the context window problem. But the good news, is that once you do, things get much much better! You can then apply your search/retriever skills on all kinds of other problems - because search is really the backbone of all AI.
[0]: https://twitter.com/_cartermp/status/1657037648400117760
A lot of the problems are much more easily solved when you're working on a new product from scratch.
Also, I think LLMs work much better with structured data when you use them as a selector instead of a generator.
Asking an LLM to generate a structured schema is a bad idea. Ask it to pick from a set of pre-defined schemas instead, for example.
You're not using LLMs for their schema generating ability, you're using them for their intelligence and creativity. Don't make them do things they're not good at.
I don't go directly to an LLM asking for structured data, or even a final answer, so you can type literally anything into the entry field and get a useful result
People are trying to treat them as conversational, but I'd say for most products it'll be rare to ever want more than one response for a given system prompt, and instead you'll want to build a final answer procedurally.
But it was really tedious to prompt ChatGPT into being properly critical about an idea that doesn't exist: A basic "make me a persona" prompt will give you an answer, but if you can really break down the generation of the persona (ie. instead of asking for the whole thing, ask who are the people likely to use X, what's the range of incomes they have, etc) you get a much better answer
The site just automates that process and presents chats that are seeded with the result of that process so the LLM is more willing to imagine things. For example, if a persona complains about a feature, when can hit 'Chat with X' and interrogate them about it instead of running into 'As a LLM' you should get an actual answer
I fall into the category of developers using LLMs every single day, for both answering questions while working, and also for more exploratory “bounce ideas off the wall” exercises.
Everytime I find a new way to explain to the LLM how I want us to work together, I feel like Ive unlocked new abilities and usecases I didnt expect the model to have.
Some examples for those curious:
* I am interested in learning more about X because I want to achieve Y. Please give me an overview of concepts that would be useful to learn about X to achieve Y. [then after going back and forth fleshing out what Im interested in learning] Please create a syllabus for me to begin learning about X based on the information you've given me. Provide examples of additional materials I can study which already exist, and some exercises to test and operationalize my knowledge.
* [I find that the above can often make the model attempt to squeeze all the info into a single response, which compresses the fidelity of the knowledge and tends towards big shallow lists, so I will employ this trick] I want you to go deeper into each topic you have listed, one at a time. When I say “next” move onto the next topic
* You are my personal coach for X, here is context about the problem I want to work on and my goals. This is our first coaching session, ask me any questions you need to gather more information, but never more than 3 at once. Where should we start?
FWIW we have about a ~7% failure rate (meaning it fails to produce a valid, runnable query) after some work done to correct what we consider correctable outputs. Not terrible, but we think the above idea could help with that.
Maybe somewhat counter-intuitively to how most people view LLMs, I strongly believe they're better when you constrain them a bit with some guardrails (E.g. pieces of a query, a bunch of existing queries, etc).
Happily surprised you guys managed to get it down to only a 7% failure rate though! For how temperamental LLMs are and the seeming complexity of the task that's impressive.
Thanks! It, uhh, was quite a bit higher before we did some of that work though, heh. Since we can take a query and attempt to run it, we get good errors for anything that's ill-specified, and we can track it. Ideally we'd address everything with better prompt engineering, but it's certainly quicker to just fix stuff up after the fact when we know how to.
https://github.com/hellisotherpeople/constrained-text-genera...
They cite "two to 15+ seconds" in this blog post for responses. Via the OpenAI API I've been seeing more like 45-60 seconds for responses (using GPT-3.5-turbo or GPT-4 in chat mode). Note, this is using ~3500 tokens total.
I've had to extensively adapt to that latency in the UI of our product. Maybe I should start showing funny messages while the user is waiting (like I've seen porkbun do when you pay for domain names).
...also why we can't wait for other vendors to get SOC I/II clearance, and I guess eventually fine-tuning our own model, so we're not stuck with situations like this.
The description of how they're handling the threat of prompt injection was particularly smart.
This was so true that there was an obvious chunk of teams in the hack-a-thon who didn’t even bother doing anything more than a fancy version of asking ChatGPT “where should I go for dinner in Brooklyn?” or straight up failed to even deliver a concept of a product.
Asking a clear question and harvesting accurate results from AI prompts is far more difficult than you might think it would be.
Some things never change...
isn't this exactly the (theoretical) strength of a chatbot - asking follow-up questions to remove uncertainty?
What we're looking at doing is:
- Creating, storing, and updating an embedding of a schema that people query against
- Creating an embedding of the user's input
- Running a cosine similarity against the user input embedding and each column in a schema, then sorting by relevancy (it's a score from 1.0 to -1.0)
- Using the top n most "relevant" columns instead of passing the full schema
So far, there's some pros and cons. On the pros side, it's really fast and lets us generally be more accurate for schemas that are very large, since those can get truncated today. We've seen in some cases it can also help reduce LLM hallucinations. On the cons side, it's another layer of probabilistic behavior and still has the chance of "missing" a relevant column. We can't really say for sure if it's better overall in our test environment, so we're going to just test in production and flag it out if it's yielding worse results.
LLMs are purely suited for search precisely because they can't guarantee to "find" information that actually exists in the real world, but they do add a lot of unnecessary noise, making it harder to weed out the truth, not easier.
Working towards having a chat box on every website is not a useful outcome for users or businesses, just OpenAI.