Gemini Flash
deepmind.google
deepmind.google
pipx install llm # or brew install llm
llm install llm-gemini --upgrade
llm keys set gemini
# paste API key here
llm -m gemini-1.5-flash-latest 'a short poem about otters'
https://github.com/simonw/llm-gemini/releases/tag/0.1a4Not bad compared to rolling your own, but among frontier models the main competitive differentiator was native multimodality. With the release of GPT-4o I'm not clear on why an organization not bound to GCP would pick Gemini. 128k context (4o) is fine unless you're processing whole books/movies at once. Is anyone doing this at scale in a way that can't be filtered down from 1M to 100k?
Gemini's strength isn't in being able to answer logic puzzles, it's strength is in its context length. Studying for an exam? Just put the entire textbook in the chat. Need to use a dead language for an old test system with no information on the internet? Drop the 1300 page reference manual in and ask away.
According to https://ai.google.dev/pricing it's $0.70/million input tokens (for a long context). That will be per-exchange, so every little back and forth will cost around that much (if you're using a substantial portion of the context window).
And while I haven't tested Gemini, most LLMs get increasingly wonky as the context goes up, more likely to fixate, more likely to forget instructions.
That big context window could definitely be great for certain tasks (especially information extraction), but it doesn't feel like a generally useful feature.
It reduces prompt costs by half for those shared prefix tokens, but you have to pay $4.50/million tokens/hour to keep that cache warm - so probably not a useful optimization for most lower traffic applications.
That's on a model with $3.5/1M input token cost, so half price on cached prefix tokens for $4.5/1M/hour breaks even at a little over 2.5 requests/hour using the cached prefix.
Tediously, it would be possible to do this chapter by chapter in order to exceed the output limit building something for future inputs.
Of course, the summary might not fulfill the same functionality as the original source document. YMMV
$20 hosting can serve thousands of users per month. $20 llm sub services just one person. This is fucking impossible.
I'm building a small assistant tool and thought I'm forced to use APIs!
Price for anything, particularly multimodal tasks that with OpenAI GPT-4o is the cheapest model, that doesn't need GPT-4 quality. GPT-3.5-Turbo — which itself is 1/10 the cost of GPT-4o, is $0.5/1M tokens on input, $1.50/1M on output, with a 16K context window. Gemini 1.5 Flash, for prompts up to 128K, is $0.35/1M tokens on input, and $0.53/1M tokens on output.
For tasks that require multimodality but not GPT-4 smarts (which I think includes a lot of document-processing tasks, for which GPT-4 with Vision and now GPT-4 are magical but pricy), Gemini Flash looks like close to a 95% price cut.
It means you can dump context without thinking about it twice and without needing to hack some solutions to deal with context overflow etc.
And given that most use cases most likely deal with text and not multimodal the advantage seems pretty clear imo.
But submitting 1M input tokens instead of 100k input tokens:
- Causes your costs to go up ~10x
- Causes your latency to go up ~10x (or between 1x and 10x)
- Can result in worse answers (especially if the model gets distracted by irrelevant info)
So longer context is great, yes, but it's not a no-brainer like more email storage. It brings costs. And whether those costs are worth it depends on what you're doing.
I tried a half dozen times and gave up, I hope this one is faster and more stable.
I've been trying to work Gemini 1.5 Pro into our workstream for all kinds of stuff and it is so bad. Unbelievable amount of hallucinations, especially when you introduce video or audio.
I'm not sure I can think of a single use case where a high hallucination tiny multimodal model is practical in most businesses. Without reliability it's just a toy.
E.g. I want to send an entire code base in a context. It might not fit into 128k.
Filtering down is a complex task by itself. It's much easier to call a single API.
Regarding quality of responses, I've seen both disappointing and brilliant responses from Gemini. Do maybe worth trying. But it will probably take several iterations until it can be relied upon.
My intuition is that as contexts get longer we start hitting the limits of how much comprehension can be embedded in a single point of vector space, and will need better architectures for selecting the relevant portions of the context.
Multimodality in a model That's between 4-7% the cost per token of OpenAI’s cheapest multimodal model is an important feature when you are talking about production use and not just economically unsustainable demos.
Agree on OpenAI multimodal but it's sort of a stilted example itself, it's because OpenAI has a hole in its lineup - ex. Claude Haiku is multimodal, faster, and significantly cheaper than GPT 3.5.
360 RPM base limit, pricing is posted.
> seriously, try finding info on cost, RPM, or release right now,
I wasn't making up numbers, its on their Gemini API pricing page: https://ai.google.dev/pricing
"Who are the biggest soda or potato chip makers?"
I have tried it for so many use cases in video / audio and it hallucinates an unbelievable amount. More than any other model I've ever used.
So if 1.5 Pro can't even handle simple tasks without hallucination, I imagine this tiny model is even more useless.
I’m not sure it’s public knowledge, but it’s an architecture choice. They choose how big to make the embedding dimension.
My point is just that there’s no limitation in principle, it’s just a matter of how they design it and resource constraints.
So OpenAI's large embedding model has 3072 dimensions, though in practice far fewer are probably used. Clearly you can't compress 1M tokens down to 3072. Yet those 3072 numbers are all you've got for capturing the full meaning of the previous token when predicting the next one; including all 1M tokens of modifying context.
So perhaps human language is simply never complex enough to need more than 3072 numbers to represent a given train of thought, but that doesn't seem clear to me.
Edit: Since Gemini is relevant here, it looks like their text embedding model is 768 dimensions.
Will compute allow that number to go up? Or is that an optimal number?
In general for models we know its trended upward, but for sure it’d be interesting to know what they’re using now.
For example, with Open AI I believe it’s known that the internal dimension for Gpt3 was 12,288.
Mistral uses a 1024 dimension embedding for 8K context. I think the point about trying to capture that rich of a context into a smaller number of dimensions still stands?
For long contexts this is a key consideration along with what self attention optimizations the model chooses to implement.
They don’t make this public, but we can infer they can’t be using full self attention pairs at 1,000,000 tokens because it scales quadratically and would take Terabytes of RAM.
There are different approaches like sparse attention, and the only way to really know how well their choices work is to test it.
Is it possible to explain what this means in a way that somebody only roughly familiar with vectors and vector databases? Or recommend an article or further reading on the topic?
Essentially each token of a text occupies a point in a many-dimensional model that represents meaning, and LLMs predict the next token by modifying the last token with the context of all the tokens before it. Attention heads are basically a way of choosing which prior tokens are most relevant and adjusting the last token's point in vector-space accordingly.
We are dealing with multi-headed attention, therefore we have multiple points per token. You can always increase the number of heads or the size of the key vector.
Google has I think 10 models available (there’s more than ten model names, but several of the models have multiple aliases) through what the Google Cloud console calls the Generative Language API (the documentation calls it the Gemini API) – based on enumerating the model list through the API itself.
Of those, 3 have pricing information on the documentation page for Gemini API pricing, 2 of which are in preview so that pricing applies in the future.
Only one (the same one of the 3 on the documentation page that is not in preview) has pricing listed on the console for the Generative Language API. On the Cloud SKUs list, there is no Generative Language API, but there is for the Gemini API, with, again, the same one model. On the Cloud Price list which the console page links for the “latest pricing” (why are there so many different things?) neither the Generative Language API nor the Gemini API is listed at all.
I think we may get to this point eventually, in the limit we will want multimodal LLMs that understand images and sounds down to the pixel and frequency, and it seems like for text, too, we will eventually want that as well.
Is there a good paper (or talk) how inference looks at scale? (Kinda like ELI-using-single-gpus)
Just make sure to have some big MLPs at the start too, to enrich the "tokens" with the information currently stored in the embedding tables.
1. latency, which would get worse if you have to sequentially generate more output
2. These models very roughly turn tokens -> "average meaning" on the embedding layer, followed by attention layers that combine the meanings, and feed forward layers that match the current meaning combination to some kind of learned archetype/prototype almost. When you move from word parts to characters all of that becomes more confusing (what's the average meaning of a?) and so I don't think there are good enough techniques to learn character-based models yet
Characters are not the semantic components of words—these are syllables. Generally speaking, anyway. I've got to imagine this approach would yield higher quality results than the roman alphabet. I'm curious if this could be tested by just looking at how LLMs handle English vs Chinese.
Besides, considering morphemes as semantic often results in a completely different meaning than we actually intend. We aren't trying to train a chatbot to speak in prefixes and suffixes, we're trying to train a chatbot to speak in natural language, even if it is encoded to latin script before output.
(The point about semantics is also technically wrong. You would first need to specify your view of semantic compositionality before such a point can be evaluated, but the usual views of semantics don't have any such consequence.)
Sure, if you define "morpheme" as a collection of syllables that's meaningful to people using alphabetic script. I don't see any benefit to this compared to working with syllables directly, which is a meaningful concept regardless of the script used to encode them.
Cats, as noted, has two morphemes, despite having only one syllable. Syllables and morphemes are largely orthogonal, morphemes can be less than, equal to, or more than a syllable (and even when more than, may or may not start or end on a syllable boundary.)
(Also, syllables aren’t the minimal semantic units even of spoken speech, those are phonemes – a syllable consists of at least one phoneme, potentially more. But morphemes, even an alphabetic script if it isn’t perfectly phonetic, still don’t necessarily map to one or more phonemes, since is textual semantic unit may have no effect on pronunciation.)
All things that could change but seems late in the game at this point. They certainly had the money to be more creative as they came to market.
"GPT4o"? Seriously?
Even "GPT4 Omni" is easier in conversation, and that's what the "o" stands for!
They severely underestimate the number of casual users they have.
OpenAI could call their model the “[poo emoji] 5000” for all the difference it would make.
Gemini Advanced (“with Ultra 1.0”)
Gemini Ultra
Gemini Pro
Gemini Flash
Gemini Nano-1
Gemini Nano-2GPT-4 turbo (gpt-4-0125-preview) 31.0
GPT-4o 30.7
GPT-4 turbo (gpt-4-turbo-2024-04-09) 29.7
GPT-4 turbo (gpt-4-1106-preview) 28.8
Claude 3 Opus 27.3
GPT-4 (0613) 26.1
Llama 3 Instruct 70B 24.0
Gemini Pro 1.5 19.9
Mistral Large 17.7
-----> Gemini 1.5 Flash 15.3
Mistral Medium 15.0
Gemini Pro 1.0 14.2
Llama 3 Instruct 8B 12.3
Mixtral-8x22B Instruct 12.2
According to https://ai.google.dev/pricing it's priced a bit lower than gpt3.5-turbo but no idea how it compares to it.
I ran Gemini Pro side by side with ChatGPT 4 for a few months on practical coding, systems architecture, and occasional general questions. ChatGPT was more useful at least 80% of the time. Gemini was either wrong or laboriously meandering in reaching a useful answer that it wasn't worth using, in my experience.
Faster isn't what I needed... Maybe it's also "smarter" (more useful) too now?
animals have a lot more intelligence than they typically get attributed
Tool use, names, language, social structure and behavior, even drug use has been shown across many species
Many animals recognize themselves and their species as separate concepts
But... that's not really an excuse any more. Model vendors should understand now that the most natural thing in the world is for people to ask models directly about their own abilities and architecture.
I think models should have a final layer of fine-tuning or even system prompting to help them answer these kinds of questions in a useful way.
Multi-Modal modals running offline on mobile devices with millisecond latencies per token seems the future.
Where is Apple in all of this. Why is Siri still so shit?
Price (output) $0.53 / 1 million tokens (for prompts up to 128K tokens) $1.05 / 1 million tokens (for prompts longer than 128K)
---
Compared to GPT-3.5 Turbo
Input US$0.50 / 1M tokens Output US$1.50 / 1M tokens
I have a huge 500k~ tokens and complex niche codebase that I've worked on for many years, largely alone. There are parts of it I wish to refactor but I struggle because it I've become blind to my own code. It also sometimes feels lonely in a way. If it was a game I could at least show my friends, but this project is too abstract.
Gemini missed the mark a few times, especially when asking about more complex things but overall it was useful. That it got things wrong is sort of ok because I knew the codebase well enough to spot those mistakes.
Gemini 1.5 pro gave me a glimpse into what it was like having "someone" understand your whole codebase, hint at areas to improve, etc. A bit like a true copilot or coworker, but for a dream hobby project.
But I can write code so those can be fixed. At a higher level it's OK, but the most valuable thing is being able to have my codebase in its context. No other public LLM currently as far as I know can do that.
My project "Nattlua" is a typed version of Lua like Typescript is to Javascript.
Some example questions that would give me new insight:
- Asking what the codebase is without supplying the readme. (though it might know because the codebase is public already)
- Asking it to generate complex type code based on existing tests and examples and without.
- Asking for places to refactor, the most fun one. Sometimes the exact solution provided is wrong, but often it's a good start.
> Python code generation. Held out dataset HumanEval-like, not leaked on the web
What I find interesting here is that for this particular benchmark _not_ publishing the benchmark is advertised as a feature (instead of as a sign of 'trust me, bro, we have a great benchmark'), and I can understand why. Still these are strange times we live in.
@
Get blocked by some silly overly sensitive "safety" trigger
You must not work with many companies in the US, especially for online services, at all then.
This suggests names don't stick around for long and can be re-used. Perhaps Google could bring back "Buzz" and "Wave" since enough time has passed!
There's an old saying that if you're selling a commodity, "you can only be as smart as your dumbest competitor."
If we want to be more polite, we could say instead: "you can only price your service as high as your lowest-cost competitor."
It seems that a lot of capital that has been "invested" to train AI models is, ahem, unlikely ever to be recovered.
Fungibility is the defining characteristic of commodities. While these products can be used to accomplish the same task, we're not near real fungibility yet.
This product from Google clearly competes on price/performance ratio, speed and of course, brand.
So people expect to see a return of investment which will create the bottom of pricing (at least as soon as the old money ran out)
I'm also curious if AI is a good example because ai will become fundamental. This means if you don't invest you might be gone therefore it's more like a fee in case the investment would not pan out.
Not so sure about Open AI though…
And if it wasn't for OpenAI, it would still be locked into Google's basement.
It seemingly existed only so in December 2023, Gemini ~= GPT-4. (April 2023 version) (on paper) ("32-shot CoT" vs. 5-shot GPT-4)
You're replying to a comment that points out Gemini Ultra was never released, wasn't mentioned today, and it's the only model Google's benchmarking at GPT-4 level. They didn't say anything about feelings or context window.
What are you even talking about? How do you know it's memory-holed if you haven't used it? The API is not GA, but the model can be used through the chatbot subscription. GP is talking about their lack of trust on Google's claim of 1M context token, not GPT-4 level reasoning. If you're expect GPT-4 level performance with cost-efficient models, that's another problem.
Pretending that they have a model internally that's on par but they're not releasing it is a very "my girlfriend goes to another school" move and makes no sense if they're a business that's actually trying to compete.