https://grok.com/share/bGVnYWN5LWNvcHk%3D_572b4955-6265-4210...
241 karma · joined January 13, 2023
https://grok.com/share/bGVnYWN5LWNvcHk%3D_572b4955-6265-4210...
https://www.scmp.com/news/hong-kong/society/article/3279201/...
The way we tokenize is just a design choice. Character level models(e.g. karpathy's nanoGPT) exist and are used for educational purpose. You can train it to count number of 'r' in a word.
Imagine the original question is posed in English but it is translated to Chinese and then the LLM has to answer the original question based on the Chinese translation.
It's a flaw of the tokenization we choose. We can train an LLM using letters instead of tokens as the base units but that would be inefficient.
You can feed it with videos of usage of any software or real world footage recorded by a Go Pro mounted on your shoulder(with body motion measured by some sesnors though the action space would be much larger).
Such a "game engine" can potentially be used as a simulation gym environment to train RL agents.
e.g. The minecraft bot Voyager can explore the game environment and extend its skill library(stored as a vector database), is that considered as innovation? There are also systems like leandojo/alphaproof that discover new proofs and use LLM in non-trivial ways(not just naively predict the next token in one shot). Reinforcement learning algorithms like AlphaGo/AlphaZero use self play and use monte carlo tree search to learn to outperform humans. You can similarly use LLMs to generate actions and estimate state values(check the language agent tree search paper).
Most people use LLMs by prompting them with some additional context(chat history and data retrived from database) but there is nothing that stops us from continuously improving a LLM(either by modifying its weight or augmenting it with external database) by asking it to evaluate the task outcome/error message and feeding it back to the LLM. We can also ask it to just keep on generating new tasks to experiment with the environment/internet to get new knowledge.
Reasoning stuff is not useless. They provably(according to benchmarks) improve the performance of coding and math related tasks.
You can try STORM(also from Stanford) and see the prompts it generates automatically, it tries to expand on your topic and simulate the collaboration among several domain experts https://github.com/stanford-oval/storm
An example article I asked it to generate https://storm.genie.stanford.edu/article/how-the-number-of-o...
https://stability.ai/news/stable-diffusion-3-research-paper
Related discussion
https://www.reddit.com/r/StableDiffusion/comments/1c0by2y/wh...
You want planning, you can do monte carlo tree search and use LLM to evaluate which node to explore next. You want verifiable reasoning, you can ask it to generate code(an approach used by recent AI olympiad winner and many previous papers).
What is even "planning", finding desirable/optimal solutions to some constrained satisfaction problems? Is the llm based minecraft bot voyager not doing some kind of planning?
LLMs have their limitations. Then augment them with external data sources, code interpreters, give it ways to interact with real world/simulation environment.
It actually teaches you how to build llama iteratively, test, debug and interpret the training loss rather than just desribing the code.
https://github.com/thomasj02/AiFilter
https://github.com/GoogleChrome/chrome-extensions-samples
https://huggingface.co/docs/transformers/en/tasks/sequence_c...
Another potential source of data is voice acting script of animations. I always thought the storyboards of films/animations can be great annotated training data but it seems there are no open datasets, probably because of copyright issues.
Browser extension that monitors how much time I have wasted on social media/porn
Browser extension that reminds me when I made a comment without reading a linked article(like right now)
Imagine a social media platform with an AI that will spot the most stupid comment and shame it using logical arguments and factual evidences, it can even dig up the user comment history to expose their hypocrisy