The super effectiveness of Pokémon embeddings using only raw JSON and images
minimaxir.com
minimaxir.com
"base_happiness": 50,
"capture_rate": 190,
"forms_switchable": false,
"gender_rate": 4,
"has_gender_differences": true,
"hatch_counter": 10,
"is_baby": false,
"is_legendary": false,
"is_mythical": false,
Why not treat each of those properties as an extra dimension, and have the embedding model handle only the remaining (non-numeric) fields?Is it because:
A) It's easier to just embed everything, or
B) Treating those numeric fields as separate dimensions would mean their interactions wouldn't be considered (without PCA), or
C) Something else?
I wonder if you might get similar results. Also would be interested in the comperative computation resources it takes. Encoding takes a lot of resources, but I imagine look-up would be a lot less resource intensive (i.e.: time and/or memory).
You almost certainly don't want to use MiniLM-L6-v2.
MiniLM-L6-V2 is for symmetric search: i.e. documents similar to the query text.
MiniLM-L6-V3 is for asymmetric search: i.e. documents that would have answers to the query text.
This is also an amazing lesson in...something: sentence-transformers spells this out, in their docs, over and over. Except never this directly: i.e. it has a doc on how to make a proper search pipeline, and a doc on the correct model for each type of search, but not a doc saying "hey use this"
And yet, I'd wager there's $100+M invested in vector DB startups who would be surprised to hear it.
> It’s super effective!
> minimaxir obtains HN13
I'm not saying that generative AI will crash but if it's indeed at the top of the S-curve there could be issues, notwithstanding the cost and legal issues that are only increasing.
Timeline would be viewed as:
2017: transformers
2018: bert
2018: gpt-1
2019: gpt-2
2020: gpt-3
2022: gpt-3.5 (chatgpt)
For the time, it was large.
Re: that it's not auto regressive, that's correct.
Things built on eachother smoothly. Transformers to BERT to GPT.
Can you compare distances just like that on a 2D space post-UMAP?
I was under the impression that UMAP makes metrics meaningless.
Useless correction, it's king - man, not man - king.
Because it could not tell the difference between word senses I think Word2Vec introduced as many false positives as true positive, BERT was the revolution we needed.
I use similar embedding models for classification and it is great to see improvements in this space.