Speech and Language Processing (3rd ed. draft)
web.stanford.edu
web.stanford.edu
Maybe machine learning engineers want to explain what I am missing?
As an NLP researcher who is in contact with companies that demand NLP applications, I think we are not there yet. For example, a company that wants to extract information from medical records cannot use solutions like GPT-4 or Claude that involve sending protected data to third parties in foreign jurisdictions. Modest local models (like 7B models) don't work so well for things like named entity extraction yet. And furthermore, companies typically want some explanation and accountability for the results (especially in sensitive domains...) and sure, LLMs can explain, but you have no guarantee that the explanation isn't hallucinated. When I mention the possibility of hallucinations, companies typically balk and say that they prefer the classic way.
My (potentially biased) opinion is that classic NLP still has life left in it. For how long, I don't know. If small, locally-runnable LLMs get much better and more reliable, addressing the hallucination problems, what you mention might become largely true for most engineering applications.
Also note that beyond engineering, things like syntactic parsing are also useful for scientific pursuits too, and at the moment they seem to be out of reach of LLMs.
It’s magic stuff
10 years ago, I started an ML & NLP consulting firm. Back then nobody was doing NLP in production (SpaCy hadn't come out yet, efficient vector embeddings were not around, the only comprehensive library to do NLP was NLTK, which was academic and hard to run in prod).
I recently revisited some of our projects from back then (like this super fun one where we put 1 Million words into the dictionary [1]) and realized how much faster we could have done many of those tasks with LLMs.
Except we couldn't — the whole "in production" part would have made LLMs for the most minute tasks prohibitively expensive, and that is not going to change for a while, sadly. So, if you want to work something in prod that is not specifically an LLM application, this book is still super valuable.
[1] https://www.nytimes.com/2015/10/04/technology/scouring-the-w...
The problem is exacerbated by a very big token dictionary. Every new token you generate is independently sampled (from tokens outside of the context size). If you have any task that requires joint distribution modeling LLM will fail and HMMs/CRFs will succeed. Of course, it depends when the problem will manifest and I do not know that, but skipping the whole book is not recommended.
For example, there were approaches for many language tasks in that book that used CNNs (instead of something more principled) and were extremely successful (despite the lack of joint modeling). Who knows when this accumulation of errors becomes measurable.
As long as your generation of tokens fits inside the context size, you're modeling everything jointly. But I guess you need to be aware that when you drop the furthest tokens to generate next ones, the performance can drastically drop (multiplicative accumulation of errors).
Additionally, entity linking is an extremely common task, for which LLMs fail at pretty miserably (certainly if the dictionary is custom/private. Additional work must be done to somehow (!? Many options here ?!) perform EL.
So, in the end, making a silver corpus from an LLM may be an option for NER to train a much much smaller algorithm. But EL is _still_ not a 'plug and play' problem, and can actually be pretty difficult to do "well" (using the modern techniques of MHS, etc).
LLMs are powerful, but can be difficult to work with w.r.t. doing post processing or adding markup/annotations in the text. For example, asking it to label terms in a sentence it can produce varied output, e.g. sometimes listing "word or clause: description of the meaning" which is hard to parse in downstream tasks, or sometimes out of order.
I've also seen LLMs label split infinitives with the correct meaning at the preposition instead of the whole subclause. With NLP you would label either the start and span of the label (common) or the root of the subclause, depending on the application. The NLP approach would work for more complex nested and interconnected expressions where the structure is not flat text.
If you ask it to generate XML or HTML, it will usually invent its own markup, or sometimes generate invalid output. E.g. I've seen it output HTML using tags like `<Question Word>` in some cases.
Asking it to generate CoNLL-U (which it knew about from its response) -- it only outputted 3 columns as "in _ PREP" etc. which is not correct.
Asking `Can you lemmatize "She is going to the bank to get some money."` I get `"She go(ing) to bank(er) for money(y)"` which leaves out words, doesn't lemmatize some words correctly, and has a format that is difficult to parse/interpret reliably.
LLMs also have a limited context window, so if you are trying to process a large document, or have a large complex set of instructions (e.g. on how to label, process, or format the data) then it can lose those instructions and deviate from what you are doing.
LLMs are also susceptible to being guided by the input text as the prompt and input are taken together. Thus, if the text is talking about a different format or something else it can easily switch to that.
--
While NLP pipelines require more work to get right, they can often be more efficient computationally as they are often a lot smaller than the 7b parameters, or use other techniques that don't use ML/NNs.
You can build more custom pipelines by querying the different features (part of speech, lemma, lexical features, etc.) without having to reparse the data, and you can keep the annotations consistent across those different pipelines, e.g. when labelling, extracting and storing the data in a database for searching, etc. so a user can see all the places where a term is referenced in a given text.
I don’t think that book was ever aimed at application engineers in the first place.
But I also think when they created that book originally it was about giving ML practitioners tools they would use.
(Cf. BloombergGPT paper all financial benchmark tasks).
And that's not even taking into account inference cost, but that is a business case issue.
I'm not sure what BloombergGPT has to do with LLMs vs non-LLMs; BloombergGPT is an LLM [2], and it defeating other LLMs on financial benchmarks doesn't prove much about large language models other than "LLMs can be trained to be better at specific tasks."
1: https://arxiv.org/abs/2005.14165
2: https://www.bloomberg.com/company/press/bloomberggpt-50-bill...
When I used "specific task", I meant specialized, domain specific tasks like financial sentiment and event extraction in which I hold a PhD. As a matter of fact, for Fiqa SA finetuned Roberta scores 88% F1 while BloombergGPT scores 75% F1. [1] Still very impressive for zero/few shot learner, but depending on data availability, performance targets and inference cost tradeoffs, it might not need your meds.
My point was "small" masked encoder transformer LMs like BERT can still hold their own on narrow domain tasks. And what OP claims that all NLP is solved by prompting a general purpose LLM service is simply inaccurate.
I hadn't read the financial paper you linked, it's very interesting! One odd bit I did notice was they set the gpt-4 temperature to 1.0, which is... not a great setting for analysis, and probably harmed the results somewhat. Typically you'd want a value much closer to 0 for that. But while a lower, more reasonable temperature setting would probably improve gpt-4's performance, I would still expect a finetuned LM to outperform larger models with just prompting on those kinds of narrow domains, especially once cost is a factor.
It's somewhat surprising to see how bad Bloomberg-GPT was... Even gpt-4 trounced it on every published metric, and it wasn't trained for finance tasks specifically. The bitter scaling lesson, I suppose.
LLM must be a godsend for these companies. All of sudden, low-level tasks like POS can be eliminated. Tasks like NER have only limited use and the companies can enjoy orders more entity types almost for free. Tasks like intention slotting and topic modeling become trivial compared to the pre-LLM-era pipelines.
A. Latency: for some systems, you need near real-time predictions. LLMs (today) are slow for that.
B. Cost: when the low dev. effort (for building and deploying an ML model) and low sample complexity (i.e. zero/few-shot) doesn't translate into proportionate monetary gains over what you pay for LLM usage.
C. Precision: when you want the model to reliably tell you when it doesn't know the correct answer. Hallucination is a part of it - but I think of this requirement as the broader umbrella of good uncertainty quantification. I think there are two reasons why this is worse for LLMs: (1) traditional ML models also suffer from this, but there are some well known ways for mitigation. For LLM's there is still no universal or accepted way to perform this reliably (2) the quality of generated language an LLM produces seems to be more likely to deceive you when it is wrong. I don't know how to scientifically think about this - maybe as LLMs proliferate people would build appropriate mental defenses?
There is also the practical problem of prompt transferability across LLMs: what works best for one LLM might not work well for another, and there is no systematic way to optimally modify the original prompt. This is painful in some setups where you're looking to be not locked-in. But I didn't put it in the list because this seems to be a problem for niche groups - everyone seems to be busy in getting stuff working with one LLM. Maybe this will become a larger issue later.
The nice thing is performance would be different based on the human getting it. It kind of preserves some unique element.
Need to use CRFs for NER? Test on your own validation set Need to use LLMs for NER? Test on your own validation set
The validation set should be made up of carefully curated samples that you've seen errors on, edge cases, just like Test-driven Development, but at a much larger scale. Assessing LLMs by demo is a horrible habit that we've taken to that needs to be changed.
When I met Dan Jurafsky, 10 years later, I definitely felt none of that disappointment that often goes with meeting one of your "childhood heroes" and thanked him for writing the book and the impact it had on my life.
Although much in computer science is reinventing the wheel, LLM so far strikes me as a tool that doesn't have a ton of analogues in history.
For comparison, both Bishop's new deep learning book and Simon Price's Understanding deep learning get straight to transformers and only mention RNN/LSTM in passing - they were written just as state space models and hybrid RNN/transformer models started showing good results.