The comparison i'm making is against a training of encoder/small decoder model so I think we agree on most things. I dont think jev replaces the benefit of embeddings nor think they are mutually exclusive. All part of a handy utility belt.
No comments yet.