New embedding models and API updates
openai.com
openai.com
I've found evidence that the OpenAI 1536D embeddings are unnecessairly big for 99% of use cases (and now there's a 3072D model?!) so the ability to reduce dimensionality directly from the API is appreciated for the reasons given in this post. Just chopping off dimensions to an arbitrary dimensionality is not a typical dimensionality reduction technique so that likely requires a special training/alignment technique that's novel.
EDIT: Tested the API: it does support reducing to an arbitrary number of dimensions other than the ones noted into the post. (even 2D for data viz, but may not be as useful since the embeddings are normalized)
The embeddings aren't "chopped off", the first components of the embedding will change as dimensionality reduces, but not much.
GPT4 is still king but text-embedding-ada-002 is quite bad.
Agree, I tried a few 384-dim models and they perform 95%+ as good, def. not worth the extra space.
Very basic techniques (e.g. SuperBit random projection) have been extremely effective with OpenAI embeddings in the past. E.g. all embeddings on findsight.ai are OpenAI Ada embeddings stored as SuperBit signatures with a code length of 10,000 (i.e. 157 integers each), and there's almost no recall loss compared to the full vectors.
The new GPT-4 Turbo is intended to reduce laziness. I'm updating aider's existing laziness benchmark now.
EDIT: Preliminary results are up.
Overall, the new `gpt-4-0125-preview` model does worse on the lazy coding benchmark as compared to the November `gpt-4-1106-preview` model.
I have a feeling they're purposely reducing the quality of the models, and possibly even relaunching older models as SOTA to show "progress".
Gosh that gets annoying
Btw, of all those I tried so far WhereIsAI/UAE-Large-V1 truly excels and is free/open to use, only downside is a small-ish context size.
However, I see that they also support shortening embeddings. OpenAI says that the text-embedding-3-large w/ shortening to 1536 still outperforms the text-embedding-ada-002 model. So maybe I'll go that route first, and then hope that PGVector begins supporting vectors of size >3072.
That’s still just a few thousand matrices. I’m sure they can handle the training and distribution of that set.
So much information is tied up in tables and images which are so difficult to work with right now. You can hack around it with GPT4V but it’s always going to underperform something that was trained end to end.
I’d also love to see the same support for fine tuning embeddings that they have for their LLM’s. I’m curious how that’d perform over their latest massive model.
Well that's good because the term "lazy" is a literal definition of what I can see GPT4 doing for the last couple of months. It would be a battle to get it actually to complete a task. Let's see how it improves.
The reason I ask is that ChatGPT is helping me debug docker files and a lot of the times I find really hard to Google answers there. Sometimes it waves me away but usually with promptings for more information to go on first.
It does (did?) this all the time to me generating tests.
I know it _couldn't_ know the right answer because I didn't tell it all of the model and serializer schema's and all that.
I don't care, because I know it is capable of generating very good guesses, and it is able to generate comprehensive tests, and I just need to tweak it to actually run. The problem is you have to shove it to even try.
Refusing to make any attempt makes it useless. If it's truly stumped then it'll be pretty obvious (assuming you aren't using it blindly).
It's free, but you're only allowed to use it for checking ChatGPT inputs and outputs. It would be very useful on internet forums, social media, and the like.
That’s some appreciated change from the previous policy, but still it just mentions the API, not the interactive web app.
Or if you're really, really stuck, you could take the source embeddings, some of the sources, get destination new embeddings, then train a small model to update. It's unlikely this will do anything but lower your quality compared to re-embedding everything.
This lock in effect is one reason I've avoided using OpenAI's embedding models; at least if it's open source, you'll be able to embed everything on an open model you have control of. The idea of committing to a large datastore using embedding APIs makes me feel very uncomfortable.
"Next week we are introducing..."
These not only go against the quota, but it hangs for a minute or longer where I have to wait until it realizes that it won't finish the answer.
Right now chat.openai.com won't even load.
I found it to reliably produce JSON correctly, but I've found 3.5 to be a poor performer at things like entity extraction and following directions compared to other fast models such as claude-instant (though that does not have function calling).
How does one solve for this? Wrangling the prompt with "please don't be lazy", or are there inference tricks like running thru the weights differently/multiple times?