One prompt? Fair. 10? Still ok. 100? You're pushing it. 10M - get help.
One prompt? Fair. 10? Still ok. 100? You're pushing it. 10M - get help.
So for personal, medium or even large workloads, I think it has killed it. It needs to be extremely large. If you are classifying or segmenting comments on a social media platform were you need to deal with billions a day, then LLM would be a very inefficient approach, but for 90+% of use cases. I think it wins.
I'm assuming you are going to run it locally because everyone is paranoid about their data. It's even cheaper if you use a cloud API.
Though LLMs sure have made creating training data to train old school models for those cases a lot easier.
20 bottles of ferric chloride
salesforce
...
I know it's hard to understand, but you can achieve a throughput that is a few orders of magnitude higher.
[1]: https://openrouter.ai/meta-llama/llama-3.3-70b-instruct
Why 2 days? Machine Learning took over the NLP space 10-15 years ago, so the comparison is between small, performant task-specific models versus LLMs. There is no reason to believe the "traditional" NLP pipelines are inherently slower than Large Language Models, and they aren't.
System message: answer with just "service" or "product"
User message (variable): 20 bottles of ferric chloride
Response: product
Model: OpenAI GPT-4o-mini
$0.075/1Mt batch input * 27 input tokens * 10M jobs = $20.25
$0.300/1Mt batch output * 1 output token * 10M jobs = $3.00
It's a sub-$25 job.
You'd need to be doing 20 times that volume every single day to even start to justify hiring an NLP engineer instead.
How much for the “prompt engineer”? Who is going to be doing the work and validating the output?
You glossed over the meat of the question.
At that volume you're going to use automated tests with known correct answers + random sampling for human validation.
Correct, everything is easy and simple if you make it simple and easy…
Simple things were not always easy. Many of them are, now.
Most classification prompts can be extremely easy and intuitive. The idea you have to hire a completely different prompt engineer is kind of funny. In fact you might be able to get the llm itself to help revise the prompt.
https://news.ycombinator.com/item?id=42748189
I don’t know the domain beforehand they are working in, I do validation testing with them.
I've heard from people running 100+ prompts in parallel against it.
https://ai.google.dev/pricing#1_5flash
4o-mini's rate limits scale based on your account history, from 500RPM/200,000TPM to 30,000RPM/150,000,000TPM.
Really it boils down to balance of time and cost, and the skill set of the person getting the job done.
But you seem really anti establishment (hung up over $25 cloud spend), so you do you.
Just don't expect everyone else to agree with you.
17 input tokens and 2 output tokens * 10 million jobs = 170,000,000 input tokens, 20,000,000 output tokens... which costs a total of $6.38 https://tools.simonwillison.net/llm-prices
As for rate limits, https://ai.google.dev/pricing#1_5flash-8B says 4,000 requests per minute and 4 million tokens per minute - so you could run those 10 million jobs in about 2500 minutes or 42 hours. I imagine you could pull a trick like sending 10 items in a single prompt to help speed that up, but you'd have to test carefully to check the accuracy effects of doing that.
So you'd have to account for the work of catching the residue of 2-8%+ error from LLMs. I believe the premise is for NLP, that's just incremental work, but for LLM's that could be impossible to correct (i.e., cost per next-percentage-correction explodes), for lack of easily controllable (or even understandable) models.
But it's most rational in business to focus on the easy majority with lower costs, and ignore hard parts that don't lead to dramatically larger TAM.
Like, lemmation is pretty damn dumb in NLP, while a better LLM model will be orders of magnitude more correct.
No matter how much energy you save personally, running your jobs on Sam A’s earth killer ten thousand cluster of GPUs is literally against your own self interest of delaying climate disasters.
LLM have huge negative externalities, there is a moral argument to only use them when other tools won’t work.
- It uses f16 for the data format whereas quantization can reduce the memory burden without a meaningful drop in accuracy, especially as compared with traditional NLP techniques.
- The quality of LLMs typically outperform OpenCV + NER.
- You can choose to replace just part of the pipeline instead of using the LLM for everything (e.g. using text-only 3B or 1B models to replace the NER model while keeping OpenCV)
- The (LLM compute / quality) / watt is constantly decreasing. Meaning even if it’s too expensive today, the system you’ve spent time building, tuning and maintaining today is quickly becoming obsolete.
- Talking with new grads in NLP programs, all the focus is basically on LLMs.
- The capability + quality out of models / size of model keeps increasing. That means your existing RAM & performance budget keeps absorbing problems that seemed previously out of reach
Now of course traditional techniques are valuable because they can be an important tool in bringing down costs (fixed function accelerator vs general purpose compute), but it’s going to become more niche and specialized with most tasks transitioning to LLMs I think.
The “bitter lesson” paper is really relevant to these kinds of discussions.
That’s obviously not sustainable indefinitely, but these kinds of exponentials are precisely why people often make incorrect conclusions on how long change will take to happen. Just a reminder: CPUs were 2x more performance every 18 months and continued to continually upend software companies for 20 years who weren’t in tune with this cycle (i.e. focusing on performance instead of features). For example, even if you’re spending $10k/month for LLM vs $100/month to process the 10M item, it can still be more beneficial to go the LLM route as you can buy cheaper expertise to put together your LLM pipeline than the NLP route to make up the ~100k/year difference (assuming the performance otherwise works and the improved quality and robustness of the LLM solution isn’t providing extra revenue to offset).
I think for the most part, casual nlp is dead because of LLMs. And LLM costs are going to plummet soon, so large scale nlp that you’re talking about is probably dead within 5 years or less. The fact that you can replace programmers with prompts is huge in my opinion so no one needs to learn an nlm API anymore, just stuff it into a prompt. Once costs to power LLMs decrease to meet the cost of programmers it’s game over.
Inference costs, not training costs.
> The fact that you can replace programmers
You can’t… not for any real project. For quick mockups they’re serviceable
> That’s sort of like asking a horse and buggy driver whether automobiles
Kind of an insult to OP, no? Horse and buggy drivers were not highly educated experts in their field.
Maybe take the word of domain experts rather than AI company marketing teams.
Appeal to authority is a well known logical fallacy.
I know how dead NLP is personally because I’ve never been able to get NLP working but once ChatGPT came around, I was able to classify texts extremely easily. It’s transformational.
I was able to get ChatGPT to classify posts based on how political it was from a scale of 1 to 10 and which political leaning they were and then classify the persons likely political affiliations.
All of this without needing to learn any APIs or anything about NLPs. Sorry but given my experience, NLPs are dead in the water right now, except in terms of cost. And cost will go down exponentially as they always do. Right now I’m waiting for the RTC 5090 so I can just do it myself with open source LLM.
While LLM’s can have their uses, let’s not get carried away.
Context for the problem space:
I did not make an appeal to authority. I made an appeal to expertise.
It’s why you’d trust a doctor’s medical opinion over a child’s.
I’m not saying “listen to this guy because their captain of NLP” I’m saying listen because experts have spent years of hands on experience with things like getting NLP working at all.
> I know how dead NLP is personally because I’ve never been able to get NLP working
So you’re not an expert in the field. Barely know anything about it, but you’re okay hand waving away expertise bc you got a toy NLP Demo working…
That’s great, dude.
> I was able to get ChatGPT to classify posts based on how political it was from a scale of 1 to 10
And I know you didn’t compare the results against classic NLP to see if there was any improvements because you don’t know how…
Lol
> I’m saying listen because experts have spent years of hands on experience with things like getting NLP working at all.
“It is difficult to get a man to understand something, when his salary depends on his not understanding it.”
Upton Sinclair
> Barely know anything about it, but you’re okay hand waving away expertise bc you got a toy NLP Demo working…
Yes that’s my point. I don’t know anything about implementing an NLP but got something that works pretty well using an LLM extremely quickly and easily.
> And I know you didn’t compare the results against classic NLP to see if there was any improvements because you don’t know NLP…
Do you cross reference all your Google searches to make sure they are giving you the best results vs Bing and DDG?
Do you cross reference the results from your NLP with LLMs to see if there were any improvements?
Great argument
> “It is difficult to get a man to understand something, when his salary depends on his not understanding it.”
NLP professionals are also LLM professionals. LLMs are tools in an NLP toolkit. LLMs don’t make the NLP professional obsolete the way it makes handwritten spam obsolete.
I was going to explain this further but you literally wouldn’t understand.
> Do you cross reference all your Google searches to make sure they are giving you the best results vs Bing and DDG?
…Yes I do…
That’s why I cancelled my kagi subscription. It was just as good as DDG.
> Do you cross reference the results from your NLP with LLMs to see if there were any improvements?
Yes I do… because I want to use the best tool for the job. Not just the first one I was able to get working…
It does seem likely we’ll soon have cheap enough LLM inference to displace traditional NLP entirely, although not quite yet.
False.
With all due respect, the fact that you're referring to natural language parsing as "NLPs" makes me question whether you have any experience or modest knowledge around this topic, so it's rather bold of you to make such sweeping generalizations.
It works for your use case because you're just one person running it on your home computer with consumer hardware. Some of us have to run NLP related processing (POS taggers, keyword extraction, etc) in a professional environment at tremendous scale, and reaching for an LLM would absolutely kill our performance.
Why does training cost matter if you have a general intelligence that can do the task for you, that’s getting cheaper to run the task on?
> for quick mockups they’re serviceable
I know multiple startups that use LLMs as their core bread-and-butter intelligence platform instead of tuned but traditional NLP models
> take the word of domain experts
I guess? I wouldn’t call myself an expert by any means but I’ve been working on NLP problems for about 5 years. Most people I know in NLP-adjacent fields have converged around LLMs being good for most (but obviously not all) problems.
> kind of an insult
Depends on whether you think OP intended to offend, ig
Assuming we didn’t need to train it ever again, it wouldn’t. But we don’t have that, so…
> I know multiple startups that use LLMs as their core bread-and-butter intelligence platform instead of tuned but traditional NLP models
Okay? Did that system write itself entirely? Did it replace the programmers that actually made it?
If so, they should pivot into a Devin competitor.
> Most people I know in NLP-adjacent fields have converged around LLMs being good for most (but obviously not all) problems.
Yeah LLMs are quite good at comming NLP tasks, but AFAIK are not SOTA at any specific task.
Either way, LLMs obviously don’t kill the need for the NLP field.
It seems like LLMs would be perfect for start-ups that are iterating quickly. As the business, problem, and data mature though I would expect those LLMs to be consolidated into simpler models. This makes sense from a cost and reliability perspective. I wonder also about the impact of making your core IP a set of prompts beholden to the behavior of someone else’s model.
No, you can't. The only thing LLM's replace is internet commentators.
If you're going for rough approximation, LLMs are great, and good enough. More care and conventional ML methods are appropriate as the stakes increase though.
Nah, searching Stackoverflow and Github doesn't take "weeks".
That said, due to how utterly broken internet search is nowadays, using an LLM as a search engine proxy is viable.
this is how you end up with 1000s of lines of slop that you have no idea how it functions.
This particular problem, at least to me, seems trivial, and to use an LLM for anything like this for more than a hundred cases seems incredibly wasteful.
I imagine this wouldn't be the case if we were to do more classification projects, but we aren't. We did try to find replacements first, but it was impossible for us to attract any talent, which isn't too much of a surprise considering it's mainly maintenance. Using external consultants for that maintenance proved to be almost more expensive than having two full time employees.
Things are such an opportunity cost now days. It’s like trying to capture value out of a transient amorphous cloud, you can’t hold any of it in your hand but the phenomenon is clearly occurring.
> One prompt? Fair. 10? Still ok. 100? You're pushing it. 10M - get help.
Assuming you could do 10M+ LLM calls for this task at trivial cost and time, would you do it? i.e. is the only thing keeping you away from LLM the fact they're currently too cumbersome to use?
I would believe that many NLP problems can be easily solved even by smaller LLM models.
For example, an approach that does me well is clustering then using LLMs on representative docs. Tools like bertopic are great for this.
I also don't see a clear cut difference between the two in certain areas. Embeddings are critical in LLM pipelines, but for me anyway, also "old school" tools.
I think NLP as described in the article is certainly under threat, but the tools and approaches compliment LLM use well, are far more efficient, and distinguish the pros from the neophytes.
If you're using LLMs for NLP-type tasks, but don't know the NLP tools, you're missing out.
This is a real problem I am dealing with at a library project.
Each document is between 100 to 10k tokens.
Most top (read most expensive) LLMs available in OpenRouter work great, it is the cost (and speed) that is the issue.
If I could come up with something locally runnable that would be fantastic.
Presumably BERT based classifiers would work if I had one properly trained for the language.
Claude Haiku is $0.25 per 1M tokens in, $1.05 per 1M out, so cost would be ~$35.
GPT-4o mini is even cheaper at $0.15 per 1M in.
Of course if your volume justifies the hardware cost you could always run Llama locally, for the cost of the electricity used.
Say Gemini Flash 8B, allowing ~28 tokens for prompt input at $0.075/1M tokens, plus 2 output tokens at $0.30/1M. Works out to $0.0027 per classification. Or in other words, for 1 penny you could do this classification 3.7 times.
There may be some use case but I'm not convinced with the one you gave.
What? That's simply not true.
Current embedding models are incredibly fast and cheap and will, in the vast majority of NLP tasks, get you far better results than any local set of features you can develop yourself.
I've also done this at work numerous times, and have been working on various NLP tasks for over a decade now. For all future traditional NLP tasks the first pass is going to be to get fetch LLM embeddings and stick on a fairly simple classification model.
> One prompt? Fair. 10? Still ok. 100? You're pushing it. 10M - get help.
"Prompting" is not how you use LLMs for classification tasks. Sure you can build 0-shot classifiers for some tricky tasks, but if you're doing classification for documents today and you're not starting with an embedding model you're missing some easy gains.