543 karma · joined June 22, 2015
The obstacle with LLMs is that they are specifically going at their own uncanny valley pace. I do see how that might be not a problem for some prompts.
I haven't tried a Chinese-trained LLM for learning Mandarin, so I'm not sure which LLM is best..
what
However, I haven't read about it yet. I'm really excited to look into it!
I don't know if this will help for things like understanding code, where the all relevant parts can be the file of 1000 lines that we are analyzing, and where every token is relevant in understanding recursion, loops, function calls, etc.
This sounds like it would be great to do SSA before passing things along to a code model like claude code.
Let me know if I misunderstood
What about only storing the conversation and then recomputing the embeddings in the cache? Does that cost a lot? Doing a lot of matrix multiplication does not cost dollars of compute, especially on specialized hardware, right?
Can you say more?
Each vector has many many dimensions, and when we train the LLMs, their internal understanding of those vectors sees all sorts of dimensions. A simple way to visualize this is a word's vector being <1, 180, 1, 3, ... > which would all mean a certain value at that dimension. In this example say the dimensions are <gender, height in cm, kindness, social title/job, ...> . In this case, our example LLM could have learned that the example I gave is <Woman, 180, 100% kind, politician, ... >. The vector's undergo some transformation so every dimension is not that discretely clear cut.
In this case, elephant and car both semantically look very similar to vehicles. They basically would have most vectors very similar.
See this article. It shows that once you train an LLM, and you assign an embedding vector for each token, then you can see how the LLM can distinguish the difference between king and queen: man and woman.
https://informatics.ed.ac.uk/news-events/news/news-archive/k...