For reference, Let's build the GPT Tokenizer https://www.youtube.com/watch?v=zduSFxRajkE
For reference, Let's build the GPT Tokenizer https://www.youtube.com/watch?v=zduSFxRajkE
The goal of the service is to answer complex queries correctly, not to have a pure LLM that can do it all. I think some engineers feel that if they are leaning on an old school classically programed tool to assist the LLM, it's somehow cheating or impure.
How long before monkey finds tall enough tree to reach moon?
This glosses over a massive issue which is that not everything can be efficiently represented as a vector space via embeddings. So your claim of "general purpose" rings hollow.
Not to mention that there is no feedback mechanism for the supposed "knowledge" advocates claim Transformer-based models have, so things like metacognition are literally impossible with this architecture. As it stands LLM outputs are isomorphic to psychotic stream-of-consciousness babble.
You've managed to find a tall tree, but from your response it seems like you haven't yet gotten to considering rockets.
The problem is, such tricks are sold as if there's superior built-in multi-modal reasoning and intelligence instead of taped up heuristics, exacerbating the already amped up hype cycle in the vacuum left behind by web3.
Most humans also can’t reliably do complex arithmetic without the use of something like a calculator. And that’s no trick. We’ve built the modern world with such tools.
Why should we fault AI for doing what we do? To me, training the AI use a calculator is not just a trick for hype, it’s exciting progress.
Isn't that what it does, when it writes a Python program to compute the answer to the user's question?
And 1000x is just a guess. We have no scaling laws about this kind of thing. It could be a million. It could be 10.
But while we can delegate the math to the calculator, it's essentially sweeping the problem under the rug. It actually tells you your neural net is not very smart. We know for a fact that it was exposed to tons of math during training, and it still can't do even the most basic addition reliably, let alone multiplication or division.
What we want is an actually smart network, not a dumb search engine that knows a billion factoids and quotes, and that hallucinates randomly.
The reason some people have mixed feelings about this because of a historical observation - http://www.incompleteideas.net/IncIdeas/BitterLesson.html - that we humans often feel good about adding lots of hand-coded smarts to our ML systems reflecting our deep and brilliant personal insights. But it turns out just chucking loads of data and compute at the problem often works better.
20 years ago in machine vision you'd have an engineer choosing precisely which RGB values belonged to which segment, deciding if this was a case where a hough transform was appropriate, and insisting on a room with no windows because the sun moves and it's totally throwing off our calibration. In comparison, it turns out you can just give loads of examples to a huge model and it'll do a much better job.
(Obviously there's an element of self-selection here - if you train an ML system for OCR, you compare it to tesseract and you find yours is worse, you probably don't release it. Or if you do, nobody pays attention to you)
That logic doesn’t extend to things we already know how to program computers to do. Arithmetic already works. We don’t need a neural net to also run the calculations or play a game of chess. We have specialized programs that are probably as good as we’re going to get in those specialized domains.
That's actually one of the specific examples from the link I mentioned:-
> In computer chess, the methods that defeated the world champion, Kasparov, in 1997, were based on massive, deep search. At the time, this was looked upon with dismay by the majority of computer-chess researchers who had pursued methods that leveraged human understanding of the special structure of chess. When a simpler, search-based approach with special hardware and software proved vastly more effective, these human-knowledge-based chess researchers were not good losers. They said that ``brute force" search may have won this time, but it was not a general strategy, and anyway it was not how people played chess. These researchers wanted methods based on human input to win and were disappointed when they did not.
While it's true that they didn't use an LLM specifically, it's still an example of chucking loads of compute at the problem instead of something more elegant and human-like.
Of course, I agree that if you're looking for a good game of chess, Stockfish is a better choice than ChatGPT.
Point is, LLM maximalists are wrong. Specialized software is better in many places. LLMs can fill in the gaps, but should hand off when necessary.
You see this type of glitch crop up in tokenizing schemes in large language models. If you attempt working with character level reasoning or output construction, it will often fail. Trying to get ChatGPT 4 to output a sentence, and then that sentence backwards, or every other word spelled backwards, is almost impossible. If you instead prompt the model to produce an answer with a delimiter between every character, like #, also to replace spaces, it can resolve the problems much more often than with standard punctuation and spaces.
The idea applies to abstractions that aren't only individual tokens, but specific concepts and ideas that in turn serve as atomic components of higher abstractions.
In order to use those concepts successfully, the model has to be able to encode the thing and its relationships effectively in the context of whatever else it learns. For a given architecture, you could do the work and manually create the encoding scheme for something like arithmetic, and it could probably be very efficient and effective. What you miss is the potential for fuzzy overlaps in the long tail that only come about through the imperfect, bespoke encodings learned in the context of your chosen optimizer.
I've asked it so many times to count the number of words or letters and it was incredibly bad at it.
Since it is capable of splitting large tokens into smaller tokens, the solution to this problem is to create additional training samples that perform "big token" to "small token" conversion and back, so that the model will learn to dynamically provide the most suitable encoding to itself.
Certain problems are always going to be very algorithmic and computationally expensive to solve. Asking an LLM to multiply each row in a spreadsheet by pi for example would be a total waste.
To handle these kinds of problems, the AI should be able to write and execute its own code for example. Then save the results in a database or other long term storage.
Another thing it would need is access to realtime data sources and reliable databases to draw on data not in the training set. No matter how much you train a model, these will still be useful.
web3 on the other hand have zero use cases other than Ponzi schemes.
Are LLM living up to all the hype? No.
Are they a hugely significant technology? Yes.
Are they web3 style bullshit? Not at all.
No, that's the actual end goal. We want a NN that does everything, trained end-to-end.
As someone applying LLMs to a set of problems in a production application, I just want a tool that solves the problem. Today, that tool is an LLM, tomorrow it could be anything. If there are ~hacks~ elegant techniques that can get me the results I need faster, cheaper, or more accurately, I absolutely will use those until there's a better alternative.
I recognised that the problem, while being beyond what an ANN could do at the time, could be split into two parts each of which was a classic ANN task. For communication between the two I described a very simple electronic circuit - just a few logic gates.
When presenting the design, the professor questioned why this component was not also a neutral network. Thinking it was a trick question, I happily answered that solving it that way would be stupid since this component was so simple and building and training another network to approximate such a simple logical function is just a waste of time and money. He got really upset, saying that is how he would have done it. He ended up giving me a lower score than expected saying I technically had everything right but he didn't like my attitude.
An input analyzer that finds out what kinds of tokens the query contains. A bunch of specialized models which handle each type well: image analysis, OCR, math and formal logic, data lookup,sentiment analysis, etc. Then some synthesis steps that produce a coherent answer in the right format.
The model learns how to split them on its own, and usually splits based not on topic or domain, but on grammatical function or category of symbol (e.g., punctuation, counting words, conjunctions, proper nouns, etc.).
I thought half the point of MoE was to make the training tractable by allowing the different experts to be trained independently?
Eventually there is going to be a model catalog that describes model inputs/outputs in a machine parseable format, all models will use a unified interface (embedding in -> embedding out, with adapters for different latent spaces), and we will have "agent" models designed to be rapidly fine tuned in an online manner that act as glue between all these different models.
This strikes me as the most compute efficient approach.
When I use analysis mode to generate and evaluate code it recently started writing the code, then introspecting it and rewriting the code with an obvious hidden step asking "is this code correct". It made a huge improvement in usability.
Fairly recently it would require manual intervention to fix.
https://paperswithcode.com/paper/most-language-models-can-be...