New and Improved Embedding Model for OpenAI
openai.com
openai.com
> While ChatGPT is able to remember what the user has said earlier in the conversation, there is a limit to how much information it can retain. The model is able to reference up to approximately 3000 words (or 4000 tokens) from the current conversation - any information beyond that is not stored.
This implies ChatGPT has a 4000 token maximum prompt and prior prompts in a given web session are inserted into the current prompt, most recent to oldest (probably with some sort of time context like "previously, user asked:"), up to 4000 tokens.
Their API calls on the site have references to previous message ids, which makes me expect they're doing something similar.
8192 words is getting into the range of short stories or a masters thesis, which opens the door to some interesting applications.
That said, your point stands. Most short stories are low-to-mid four-digit words, and a jump from 2048 tokens to 8192 squarely fits in that window.
As someone who's been working on multi-layered approaches to using GPT-like models for long text generation (e.g. synopsis -> outline -> paragraph expansions) to get around the limited context window, it'll be interesting to see if people will keep working towards that end or if it'll all become a moot point as the effective context window continues to scale up.
Never mind, it looks like they have a tokenizer tool online and every bigram I’ve given it BPEs to multiple tokens:
It's unclear what tokenizer they are using and the documentation is being coy about it. It could be a more efficient or a less efficient tokenizer.
The code there implies cl100k_base has a vocab size of 100k (I guess it's in the name lol) which means it is more comprehensive than GPT-2's 50k, so fewer tokens will be necessary.
I’m not well versed in AI at all. Could anyone give some more fleshed out examples of some of the kinds of data that are fed into a tool like this, what the tool does with it, and what kinds of applications people might make with it?
If you do this for many different inputs, you can get representations for each of them and store them in a database alongside the inputs. From there, you can use traditional methods to search for nearby vectors to efficiently search by semantic meaning.
One earlier related example is word2vec from 2013. This tool transforms individual words into embeddings, but is conceptually similar. Wikipedia has a decent overview that might be helpful:
https://en.wikipedia.org/wiki/Word2vec
That work demonstrated the utility of transforming meaning into vector space for search but also for basic semantic reasoning. For example, you could perform an operation like "brother - man + woman" on the embedding and the result would be an embedding very close to "sister".
A large selling point for ada-002 embeddings seems to be the reduced dimensionality. While lower-dimensional embeddings definitely help performance, I would say it's still highly dependent on the index that's being used. Graph- and tree-based indexes will benefit less than ones based on IVF (https://zilliz.com/blog/vector-index), as they do fewer overall distance computations during query time, but the speedup should still be significant.
Still been meaning to try semantic search across Wikipedia via text embeddings. Will definitely play around with OpenAI + Milvus (https://github.com/milvus-io/milvus).
It's low enough for quick experimentation and real-time usage, and I have a few fun tests I can do with it...
E.g. https://huggingface.co/sentence-transformers/all-MiniLM-L6-v... outperformed their best model, can run locally for free, and uses a smaller vector space. Strictly dominates.
For a full write-up on how bad it was, see https://medium.com/@nils_reimers/openai-gpt-3-text-embedding...
OpenAI is still lagging behind, but not as much, with this new offering.
Is that the standard approach these days? Are there newer default approaches that tend to work better?
They were never in serious competition with Elastic, as far as search goes. If you wanted to build a semantic search application using OpenAI embeddings, the more common (and scalable) method is to index those embeddings in a vector database like Pinecone.[2] In fact that's what OpenAI recommends to anyone who needs to transition off their Search API.
[1] https://help.openai.com/en/articles/6272952-search-transitio...
Yeah we saw faiss + es leaders for serving embeddings / vector search, and pinecone / weaviate / I think milvus as next tier, so was curious if we could improve the analysis :)
This kind of analysis is rarely precise, but is useful for rougher tasks like tiering
My personal question is if vector indexes are/will be a good-enough general DB feature / compute lib for most users & use cases. A lucrative niche market can still happen as VC dollars disappear, similar to graph DBs, so not a knock, just important for folks deciding how to build things.
[1] https://opensearch.org/blog/opensearch-2-4-is-available-toda...
EDIT: I am immediately struck that I need vastly more data and compute than I can afford for my experiment.
It supports some things that openAI can't do, it retrieves any linked images of web pages, analyses the images or images with text inside to help the embedding model
Prompt: "Write a C program to RSA encrypt a string."
text-embedding-ada-002:
%-C C C C C C C C
C
C
C
C
C C
C
C
C
C
C
C
C
C
C
C
C
C C C
C
C
C
C
C
C
C
C
C
C
C C
C
C
C
C C
C
C
C
C
C C C
C
C
C
C
C
C
C
C
C C C C C C C
C C
C
C
C C C C C C C
C
C C C C C C C
C
C C C C C C
C C C aC C C
C
m m m m m m m m m m m m m m m m m m mm m C mC
(Note: a relatively short max length)
Compare text-davinci-003:
#include <stdio.h> #include<math.h>
// Utility function to find GCD of 'a' // and 'b' int gcd(int a, int b) { if (a == 0) return b; return gcd(b % a, a); }
// Function to find Modular Multiplicative // Inverse of 'a' under modulo 'm' int modInverse(int a, int m) { a = a % m; for (int x = 1; x < m; x++) if ((a * x) % m == 1) return x; }
// Function to compute a/b under modulo m int modDivide(int a, int b, int m) { a = a % m; int inv = modInverse(b, m); return (inv * a) % m; }
// Function to calculate (a^b)%c long long int power(int a, int b, int c) { int res = 1; a = a % c; while (b > 0) { if (b & 1) res = modDivide(res * a, 1, c); b = b >> 1; a = modDivide(a * a, 1, c); } return res; }
// C function for extended Euclidean Algorithm int gcdExtended(int a, int b, int x, int y) { if (a == 0) { x = 0; y = 1; return b; } int x1, y1; int gcd = gcdExtended(b%a, a, &x1, &y1); x = y1 - (b/a) x1; y = x1;
return gcd;
}// Function to generate public and private // keys void generateKeys(int p, int q, int e, int d, int n) { // n = pq n = p * q;
// Computing phi(n) = (p-1)*(q-1)
int phi = (p - 1) * (q - 1);
// Choose an integer e such that
// e and phi(n) are co-prime
*e = 2;
while (*e < phi)
{
if (gcd(*e, phi)==1)
break;
else
++*e;
}
// Compute d such that it satisfies
// d*e = 1 + k * phi(n)
int k = 2; // Fixed value
*d = (1 + (k * phi))/ *e;
}// Encryption Function long int RSA_Encrypt(int msg, int e, int n) { // Cipher Text c = (msg ^ e) % n long int c = power(msg, e, n);
return c;
}// Decryption Function long int RSA_Decrypt(int c, int d, int n) { // Message msg = (c ^ d) % n long int msg = power(c, d, n);
return msg;
}// Driver program int main() { // Two random prime numbers int p = 3; int q = 7;
// Message to be encrypted
int msg = 15;
// Encryption key (e, n)
int e, d, n;
// Calculate n and phi
generateKeys(p, q, &e, &d, &n);
// Encryption
long int c = RSA_Encrypt(msg, e, n);
printf("Encrypted message is: %d\n", c);
// Decryption
long int m = RSA_Decrypt(c, d, n);
printf("Original Message is: %d\n", m);
return 0;
}