My intuition is that as contexts get longer we start hitting the limits of how much comprehension can be embedded in a single point of vector space, and will need better architectures for selecting the relevant portions of the context.
My intuition is that as contexts get longer we start hitting the limits of how much comprehension can be embedded in a single point of vector space, and will need better architectures for selecting the relevant portions of the context.
Multimodality in a model That's between 4-7% the cost per token of OpenAI’s cheapest multimodal model is an important feature when you are talking about production use and not just economically unsustainable demos.
Agree on OpenAI multimodal but it's sort of a stilted example itself, it's because OpenAI has a hole in its lineup - ex. Claude Haiku is multimodal, faster, and significantly cheaper than GPT 3.5.
360 RPM base limit, pricing is posted.
> seriously, try finding info on cost, RPM, or release right now,
I wasn't making up numbers, its on their Gemini API pricing page: https://ai.google.dev/pricing
"Who are the biggest soda or potato chip makers?"
I have tried it for so many use cases in video / audio and it hallucinates an unbelievable amount. More than any other model I've ever used.
So if 1.5 Pro can't even handle simple tasks without hallucination, I imagine this tiny model is even more useless.
I’m not sure it’s public knowledge, but it’s an architecture choice. They choose how big to make the embedding dimension.
My point is just that there’s no limitation in principle, it’s just a matter of how they design it and resource constraints.
So OpenAI's large embedding model has 3072 dimensions, though in practice far fewer are probably used. Clearly you can't compress 1M tokens down to 3072. Yet those 3072 numbers are all you've got for capturing the full meaning of the previous token when predicting the next one; including all 1M tokens of modifying context.
So perhaps human language is simply never complex enough to need more than 3072 numbers to represent a given train of thought, but that doesn't seem clear to me.
Edit: Since Gemini is relevant here, it looks like their text embedding model is 768 dimensions.
Will compute allow that number to go up? Or is that an optimal number?
In general for models we know its trended upward, but for sure it’d be interesting to know what they’re using now.
For example, with Open AI I believe it’s known that the internal dimension for Gpt3 was 12,288.
Mistral uses a 1024 dimension embedding for 8K context. I think the point about trying to capture that rich of a context into a smaller number of dimensions still stands?
For long contexts this is a key consideration along with what self attention optimizations the model chooses to implement.
They don’t make this public, but we can infer they can’t be using full self attention pairs at 1,000,000 tokens because it scales quadratically and would take Terabytes of RAM.
There are different approaches like sparse attention, and the only way to really know how well their choices work is to test it.
Is it possible to explain what this means in a way that somebody only roughly familiar with vectors and vector databases? Or recommend an article or further reading on the topic?
Essentially each token of a text occupies a point in a many-dimensional model that represents meaning, and LLMs predict the next token by modifying the last token with the context of all the tokens before it. Attention heads are basically a way of choosing which prior tokens are most relevant and adjusting the last token's point in vector-space accordingly.
We are dealing with multi-headed attention, therefore we have multiple points per token. You can always increase the number of heads or the size of the key vector.