Understanding large language models: A cross-section of the relevant literature
magazine.sebastianraschka.com
magazine.sebastianraschka.com
Props to the author for giving credit to Bandanau et al (2014), which I believe first proposed the concept of applying a Softmax function over token scores to compute attention, setting the stage for the original transformer by Vaswani et al (2017).
If you're looking for a LLM replacement for regular Google search then Bing and Bard are much more likely to fit the bill.
Another option is to have small re-trainable areas within the big model.
Of course ideally model should be able to work with some sort of "knowledge base", which can be updated daily. As a main, or additional data source.
As for me I'm looking for ways to increase productivity. Have ChatGPT Plus, and in waiting lists for GPT-4 and Bing. Neither solves big problems, but usually there are many small things which can be done faster. I would rather offload them.
But perhaps the LLM can eventually point me towards this to then learn more; somewhat what Bing is already doing but it’s more of a footnote than something upfront (e.g. “you might want to read an excellent write up in The Atlantic about decriminalizing drugs”).
I rarely have to ask it to alter the code it gives me.
These lecture slides from Yann LeCun from 1 month ago. Starting from slide 13, he presents his idea of a roadmap towards autonomous machine intelligence.
https://drive.google.com/file/d/1BU5bV3X5w65DwSMapKcsr0ZvrMR...
[1] https://cyc.com/
You might be able to compose systems that use LLMs for some sub-components and symbolic AI for other functions. So for example if ChatGPT gets asked a question that LLMs are poor at but symbolic systems are good at, it could switch to using a symbolic system, but that's not really combining the actual technologies.
I guess I’m imagining that testing the LLM output might be a problem that some other technique might be suited for, and smells like what we need to mitigate LLM weaknesses.
"Transformer Taxonomy" - https://kipp.ly/blog/transformer-taxonomy/
"Five years of progress in GPTs" - https://finbarrtimbers.substack.com/p/five-years-of-progress...
Do BERTs models learn faster than GPTs?
Assuming the task is predictive, BERT seem to have a stronger/richer signal (obviously in practice you have use existing pre-trained models)
The large corpus of text gives them a general basis of logical patterns, which can then be pruned iteratively in simulated environments.
We already have such a poor idea of how these things seem to understand so much… it’ll be a wild day when cancer is cured and we have absolutely no idea why.
(Just deciding that your LLM has magic powers because you've put it in a category called "AI" and you've decided that category has said magic powers is what that guy Wittgenstein was complaining about in philosophical problems. Besides, intelligence doesn't mean all your thoughts are automatically correct!)
This is the true bitter lesson for HN.