Coping with dumb LLMs using classic ML
softwaredoug.com
softwaredoug.com
I had a similar experience few years back when participating in a ML competitions [1,2] for detecting and typing phrases in a text. I submitted an approach based on Named Enttiy Recognition using Conditional Random Field (CRF) which has been quite robust and well known in the community and my solution beat most of tuned Deep learning solutions by quite a large margin [1].
I think a lot of folks underestimate the complexity of using some of these models (DL, LLM) and just throw them at the problem or don't compare it well against well established baselines.
[1] https://scholar.google.com/citations?view_op=view_citation&h... [2] https://scholar.google.com/citations?view_op=view_citation&h...
I have a BERT + SVM + Logistic Regression (for calibration) model that can train 20 models for automatic model selection and calibration in about 3 minutes. I feel like I understand the behavior of it really well.
I've tried fine tuning a BERT for the same task and the shortest model builds take 30 minutes, the training curves make no sense (back in the day I used to be able to train networks with early stopping and get a good one every time) and if I look at arXiv papers it is rare for anyone to have a model selection process with any discipline at all, mainly people use a recipe that sorta-kinda seemed to work in some other paper. People scoff at you if you ask the engineering-oriented question "What training procedure can I use to get a good model consistently?"
Because of that I like classical ML.
If I am building a set of models for a domain, I might fine-tune the representation layer. On a per-model basis I typically just train the SVM and calibrate it. For the amount of time this whole pipeline takes (not counting the occasions when I fine-tune), it works amazingly well.
When I presented the proposal, nobody read it and the meeting immediately turned to the vp of engineering and the ceo discussing neural networks and some other ML system that they had read about on HN the day before. When I tried to bring collaborative filtering up again, the VP said "I don't know what that is", so obviously he hadn't read the doc that I was assigned to write over the last week
All the models I've tried (Sonnet 3.5, GPT 4o, Llama 3.2, Qwen2 VL) have been pretty good at extracting text, but they failed miserably at finding bounding boxes, usually just making up random coordinates. I thought this might have been due to internal resizing of images so tried to get them to use relative % based coordinates, but no luck there either.
Eventually gave up and went back to good old PP-OCR models (are these still state of the art? would love to try out some better ones). The actual extraction feels a bit less accurate than the best LLMs, but bounding box detection is pretty much spot on all the time, and it's literally several orders of magnitude more efficient in terms of memory and overall energy use.
My conclusion was that current gen models still just aren't capable enough yet, but I can't help but feel like I might be missing something. How the heck did Anthropic and OpenAI manage to build computer use if their models can't give them accurate coordinates of objects in screenshots?
There are some VLLMs that seem to be specifically trained to do bounding box detection (Moondream comes to mind as one that advertises this?), but in general I wouldn't be surprised if none of them work as well as traditional methods.
You may also be able to get the computer use API to draw bounding boxes if the costs make sense.
That said, I think the correct solution is likely to use a non-VLM to draw bounding boxes. Depends on the dataset and problem.
1. https://www.anthropic.com/news/developing-computer-use 2. https://huggingface.co/blog/paligemma
PaliGemma seems to fit into a completely different niche right now (VQA and Segmentation) that I don't really see having practical applications for computer use.
[1] https://huggingface.co/microsoft/OmniParser?language=python [2] https://github.com/browser-use/browser-use
This makes sense to me. These LLMs likely have no statistics about the spatial relationships of tokens in a 2D raster space.
[1] https://huggingface.co/osunlp/UGround-V1-7B?language=python
That is one of the implications of transformers being DLOGTIME-uniform TC0, they don't have access to counter analogs.
You would need to move to log depth circuits, add mod-p_n gates etc... unless someone finds some new mathematics.
Proposition 6.14 in Immerman is where this is lost if you want a cite.
It will be counterintuitive that division is in TC0, but (general) counting is not.
2. BigTech and SmallTech train their fancy bounding box / detection models on large datasets that have been built using classical detectors and a ton of manual curation
https://simonwillison.net/2024/Aug/26/gemini-bounding-box-vi...
Especially in programming it is fun. People spent hours over hours to come up with a prompt that can (kinda-of) reliably produce code. So they try to hack/program some weird black box so that they can do their actual programming tasks. On some areas there might be a speed up, but I still don't know if it's worth it. It feels like we are creating more problems than solutions
I recently was chatting with my friend that wanted to automate one of his tasks by writing a python script with AI -> because all the influencers said it was "so easy" and "no programming knowledge" required.
That might have been the single funniest piece of code I have seen in a long time. Didn't install the dependencies, didn't fill in the Twitter API key, instead of searching for a keyword on Twitter it just looked up 3 random accounts, 25 functions on like 120 lines of code?
Also, the line numbers in the errors weren't helpful because the whole thing lived in Windows notepad. That was a flagship AI and a (in my opinion) capable human not being able to assemble a simple script.
If you had no idea of what code looks like and poor critical thinking abilities God help you.
The finding surprises me. I would expect modern LLMs to be powerful enough to do well at the task. Given how much the data is processed before the decision trees, I wouldn't expect decision trees to add much. I can see value in this approach if you're unable to optimize the LLM. But, if you can, I think end-to-end training with a pre-trained LLM is likely to work better.
(However 'better' might be defined, I care more about the precision / recall tradeoff)
Performing feature engineering with LLMs and then storing the embeddings in a vector database also allows you to reuse the embeddings for multiple tasks (eg clustering, nearest neighbor).
Generally no one uses plain decision trees since random forest or gradient boosted trees perform better and are more robust.
As the foundational models can parse super complex stuff like dense human language, music, etc. with context - like a really good pre-built auto-encoder, which would be a nightmare with classic machine learning feature selection (remember bag of words? and word2vec?).
I wonder how such an approach would compare to just fine-tuning one model though? And how the cost of fine-tuning vs. greater inference cost for an ensemble compares?
The meta-strategy of combining LLM and non-LLM techniques is going to be key for getting good results for some time.
that's a classic strategy to solve problems
Query: entrance table Product LHS name: aleah coffee table Product LHS description: You'll love this table from lazy boy. It goes in your living room. And you'll find ... ... Or Product LHS name: marta coffee table Product RHS description: This coffee table is great for your entrance, use it to put in your doorway... ... Or Neither / Need more product attributes
Only respond 'LHS' or 'RHS' if you are confident in your decision
RESPONSE: RHS --- LHS is include. Hopefully this is a bug in the blog and not the code
LLMs are narrative machines. They make up stories which often make sense.
> Which of these furniture products is more relevant to the furniture e-commerce search query:
Fixed in the post. Thanks
I think agents will run into the same problem - if they will try to find a classical ML solution to verify what comes out of the LLM.
LLMs are amazing we are creating better and better hyperdimentional maps of language but until we have systems that are not just crystallized maps of the language they were trained on we will never have something that can really think, let alone AGI or whatever new term we come up with.
The thing about the NFL theorem is that it assumes an equal weight or probability over each problem/task. It's impossible to find a search/learning algorithm that performs superiorly over another, 'averaged' over all tasks. But—and this is purely my intuition—the problems that humans want to solve, are a very small subset of all possible search/learning problems. And this imbalance allows us to find algorithms that work particularly well on the subset of problems we want to solve.
Coming back to representation and maps. Human understanding/worldview is a good example. Human understanding and worldview is itself a map of reality. This map models certain facts of the world well and other facts poorly. It is optimized for human cognition. But it's still broad enough to be useful for a variety of problems. If this map wasn't useful, we probably wouldn't have evolved it.
The point is, I do think there's a philosopher's pebble, and I do think there's a few free bites of lunch. These can be found in the discrepancy between all theoretically possible tasks and the tasks that we actually want to do.
Language itself is a kind of map, and it has pretty universal reach.
"No Free Lunch (NFL) theorem" isn't quite mathematics, it is more in the domain of philosophy.
I thought that the reference was to the general 'no free lunch' assumption: https://en.wikipedia.org/wiki/No_free_lunch_theorem
Finetune the models to be better
Optimise the prompts to be better
Train better models
Codeium for example will absolutely bend over backwards to provide you with solutions to requests that can't be satisfied, producing more and more garbage for every attempt. I don't think I've ever seen it just say no.
ChatGPT is marginally better and will sometimes tell you straight up that an algorithm can't be rewritten as you suggest, because of ... But sometimes it too will produce garbage in its attempts at doing something impossible that you ask it to do.
Unfortunately this very often it gets wrong, especially if it involves some multistep process.
This is one of the clearest ways to demonstrate that an LLM doesn't "know" anything, and isn't "intelligence." Until an LLM can determine whether its own output is based on something or completely made up, it's not intelligent. I find them downright infuriating to use because of this property.
I'm glad to see other people are waking up
Hackernews has Intelligent people...
Q. E. D.
% LLMs can RAG incorrect PDF citations too
I don’t see any reason that an IDE especially with a statically typed language can’t have an AI integrated that at least will never hallucinate classes/functions that don’t exist.
Modern IDEs can already give you real time errors across large solutions for code that won’t compile.
Tools need to mature.
Secondly, if an llm is giving you the runaround it does not have a solution for the prompt you asked and you need either another prompt or another model or another approach to using the model (for vendor lock in like openai)
At the end of the day, this is because it isn't "writing code" in the sense that you or I do. It is a fancy regurgitation engine, that will output bits of stuff it's seen before that seem related to your question. LLMs are incredibly good at this, but that it also why you can never trust their output.
Three problems with this:
* salespeople constantly try to sell the automation as more complete than it is
* product owners try to push us developers into making it more fully automated
* users get lulled into thinking it's more complete than it is (and accepting suggestions instead of deeply thinking through the issues like they would if they had to think things from scratch)
Maybe fixing management is the more pressing issue then working on the task of selfreplacement in the name of profit for others. Thinking about it, the implications are interesting. What is the energyconsumption of a human thinking in comparison with the energy requirement of a possible machinic replacement?
I.e. for classification you can judge "certainty" by the soft-max outputs of the classifier, then in the less certain cases can refuse to classify and send it to humans.
And also do random sampling of outputs by humans to verify accuracy over time.
It's just that humans are really expensive and slow though, so it can be hard to maintain.
But if humans have to review everything anyway (like with the EU's AI act for many applications) then you don't really gain much - even though the humans would likely just do a cursory rubber-stamp review anyway, as anyone who has seen Pull Request reviews can attest to.
LLMs are able to counterfeit a truly impressive number of indirect signals which humans currently use to make snap-judgements and mental-shortcuts, and somehow reviewers need to be shielded from that.