Don't get too knocked back by comments. A) If it works - it works. B) Your learning is as valuable as the outcome.
Oh have a look at https://imagineville.org/software/ for some other things that may be of interest..
Don't get too knocked back by comments. A) If it works - it works. B) Your learning is as valuable as the outcome.
Oh have a look at https://imagineville.org/software/ for some other things that may be of interest..
I appreciate the words of wisdom/motivation :)
Since my last comment the embedding vector is now 16-dimensions, but I'm not quite getting "King - Woman = Queen" from it yet. It actually does find similar words if I score word features, normalize the scores to a min/max spectrum between 0-1, then multiply 2 vectors (dot products) to get similiarity. But in the middle of the experiment I was like "what is this for..?" so I just kinda stopped working on the embedding refactor until I have a real world use case.
I know for example text-embedding-ada-002 uses 1536 length arrays, and others use over 3k, so maybe at some point the usefulness of having very complex embeddings emerges.
For now, the original approach still seems superior for next token prediction, bigram frequency (how often a word follows another) just need to have enough sentences modeled and scored.