Semantic Tokenizer for Enhanced Natural Language Processing
arxiv.org
arxiv.org
The line of research here has been going on for 30+ years, from Michael Brent's work, to Linguistica, to Morfessor, and now several approaches to incorporate morphology into tokenizers. The stand-out example is [0]. This paper doesn't seem to acknowledge any of that intellectual legacy. It's not a _research_ paper.
I'm getting a bit tired of people putting their class projects or quick engineering projects on arXiv. I don't know why they're surfacing so high on HN either.
How could a tokenizer do anything about that unless the synonyms actually share substrings? The vector embedding is learned, not part of the tokenizer.
The same problem also exists in the name, "Large Language Model". Sure, the content being modeled contains language, but the model itself is not specific or limited to language patterns. We ought to call them "Large Text Models"; or better yet, "Text Inference Models".
The words we use to describe software are very important: they inform goals and expectations. They define the context that software exists in.
I see our biggest mistake as calling these tools, "Artificial Intelligence". That title began as a goal and a category of work: it doesn't belong in the title or description of software unless that software has actually met the goal.
You're right that what they are doing is morphological, not semantic, but it helps a lot. I would say that
日本語
"Japanese Language" is a good token to apply embedding, attention, etc. to because it has a definite meaning to which the transformer can attach whatever syntax and semantics it learns in terms of activations. If BPE gives up and processes it as UTF-8 bytes e6 97 a5 e6 9c ac e8 aa 9e
there is no clear meaning for any one of those tokens, and the model is going to have to work a lot harder.And yes, what they do helps on their two test tasks. I'm not disputing that. It's the fact that there's no scholarship here.
There are so many thousands of knobs to twiddle with in a model these days, and they went after one that's commonly regarded in the NLP community as the 'defect'—the only part of the model that's not end-to-end trained along with the rest. Which would be great, if they acknowledged it! But there's no citation to any tokenization literature beyond BPE or SentencePiece. The literature review is as superficial as what you could find in a blog.
There are certainly byte-level or character-level tokenizers (think about CANINE or ByT5), and we can argue back and forth about their data-hungriness or slow inference. It would be nice to give more helpful units to a Transformer, so it doesn't have to learn syllables (or even characters) all on its own. Rebracketing/incorrect segmentation is a problem! And these authors have clued into that, but so have several hundred (or thousand?) researchers they don't cite.
What I'm having trouble with is the notion that this paper uncovered some exciting, revelatory fact about tokenization. Yes, "Japanese Language" would be a reasonable semantic unit! But these authors didn't discover that fact. Nobody's questioning whether 'good tokenization is better than bad tokenization'. Tokenization has seen ongoing attention in NLP forever.
These authors tried one variant, compared it against a library default option (and nothing else), evaluated on one task, put a bit of marketing around it, and called it a day. In the NLP course I used to TA, this wouldn't even qualify as a complete final project for the course.
Whenever something becomes a status symbol there will be people willing to exploit it. Perhaps ArXiv should hire some volunteers to check for a minimum of quality before acceptance? (/s, in case it's not clear).
Anecdotally, the second worst paper I've ever read was hosted on ArXiv and presented in an NLP group as a possible breakthrough. Tearing it apart in front of the person presenting it was no fun.
I got downvoted when I expressed similar opinion with regards to MiniGPT4. I guess HN crowd value usefulness more than real contribution.
One more small improvement that will boost future models.
Notably these guys have found semantic tokenization helps with embedding-based search
Before BPE I bailed on a project because the sponsor insisted on using word vectors and I thought "Look, the most important words in our documents will be out-of-dictionary and that's like playing chess down a queen, a rook and two pawns."
Once BPE and similar tokenizers came out now you could say that the model has a chance when it confronts out-of-dictionary situations which will always be important. This was critical to the success of transformers for text.
On the other hand there are many things wrong with tokenization for particular applications. If you want to handle Japanese text you'd think a word like 日本語 "Japanese" should be tokenized as a word or as 日本 + 語 ("japan" + "language")
A multilingual model however is very likely to tokenize those at the unicode character level so you don't even get 日 + 本 + 語 ("sun" + "origin" + "language") but might get underlying UTF-8 bytes like e6 + 97 + a5 + e6 + 9c + ac + e8 + aa + 9e which is just awful.
The trouble is an English language model doesn't want to waste a limited supply of tokens on other languages even though it should be able to handle a few foreign characters. A Japanese language model would clearly make different decisions, a model that supports a large number of languages is going to struggle to allocate tokens between them.
I always wondered if languages with more regular spelling and conjugation rules would converge faster, or if languages like Chinese might be more efficient since they can get more semantic meaning into a pair of bytes than English.
Also, the technique in the paper could be extended to irregular forms with a custom decoder. E.g. encode "mouse" + "##plural", then decode that sequence to "mice".
1. It really hurts the ability to generate creative writing/poetry (e.g. impossible for even ChatGPT-4 to fully understand syllable counts, leading to incorrect haikus and even poor rhyming), see https://gwern.net/gpt-3 and https://paperswithcode.com/paper/most-language-models-can-be...
2. It means that silly stuff like "Glitch Tokens" are a huge issue. i.e. whole tokens dedicated to weird usernames caused by people counting on a counting subreddit so much that their names got their own token. See https://www.youtube.com/watch?v=WO2X3oZEJOA
3. BPE has a lot of terrible vocabulary choices independent of Glitch Tokens. Lots of massive punctuation, garbage sequences, etc.
Glitch tokens are glitchy because they were not present in the training data for the embeddings, so they're essentially uninitialized. The real issue was not the use of BPE, but that they didn't restrict the tokenizer to only output tokens that the model was trained on.
Similarly, garbage sequences are present in the tokenizer vocabulary because those sequences were common in the data that the tokenizer was trained on. The solution is to not include garbage in your training data, though admittedly that's a tad tough when you're starting out with an uncurated dump of random internet content.
Maybe? Wouldn't that require the tokenizer to understand the word's usage in context of the surrounding text, or even worse, dialects?
What do you do with "fire" or "laboratory" or, god help us, "nuclear"?
Any example tokens?