I'll love to be wrong though. Please share if anyone has a different experience.
22 karma · joined July 3, 2016
I'll love to be wrong though. Please share if anyone has a different experience.
Do people just figured it out by trial & error like common patterns in x86 / arm / arcade platforms slowly?
I can't really find much discussion on details online.
Also, I found these 2 links pretty good too. 1. http://ai.stanford.edu/blog/understanding-incontext/ 2. http://ai.stanford.edu/blog/in-context-learning/
I'm still not completely convinced. Probably need to dwell on the topic longer.
It still baffles me why such stochastic parrot / next token predictor, will recognize these "Unseen combinations of tokens" and reuse them in response.
I have a feeling it should be a common question, but I just can't find the keyword to search.
PS. If anyone has any links with thoroughly discussion about positional embedding, that would be great. I never got a satisfying answer about the usage of sine / cosine and (multiplication vs addition)
I think my comment was not worded properly. I was thinking "geometry properties = linear properties", what I really should say is:
Why does the latent space has geometry properties where we could use functions like cosine similarity to compare?
So when training, the signal will be mapped to latent space that will minimize the error of the objective function as much as possible.
Many applications already use cosine similarity function at the end the network, it would be obvious why they work. I reviewed other cost functions such as Triplet Loss. They use euclidean distances, so I guess it make sense why the geometry properties exist too.
For "and there I guess the point is it's not maximally information dense, so the geometry exists in the redundancy", what does "maximally information dense" means, I still don't quite get it.
I wasn't able to find a good answer online.