Our next-generation model: Gemini 1.5
blog.google
blog.google
1. They don't talk about how they get to 10M token context
2. They don't talk about how they get to 10M token context
3. The 10M context ability wipes out most RAG stack complexity immediately. (I imagine creating caching abilities is going to be important for a lot of long token chatting features now, though). This is going to make things much, much simpler for a lot of use cases.
4. They are pretty clear that 1.5 Pro is better than GPT-4 in general, and therefore we have a new LLM-as-judge leader, which is pretty interesting.
5. It seems like 1.5 Ultra is going to be highly capable. 1.5 Pro is already very very capable. They are running up against very high scores on many tests, and took a minute to call out some tests where they scored badly as mostly returning false negatives.
Upshot, 1.5 Pro looks like it should set the bar for a bunch of workflow tasks, if we can ever get our hands on it. I've found 1.0 Ultra to be very capable, if a bit slow. Open models downstream should see a significant uptick in quality using it, which is great.
Time to dust out my coding test again, I think, which is: "here is a tarball of a repository. Write a new module that does X".
I really want to know how they're getting to 10M context, though. There are some intriguing clues in their results that this isn't just a single ultra-long vector; for instance, their audio and video "needle" tests, which just include inserting an image that says "the magic word is: xxx", or an audio clip that says the same thing, have perfect recall across up to 10M tokens. The text insertion occasionally fails. I'd speculate that this means there is some sort of compression going on; a full video frame with text on it is going to use a lot more tokens than the text needle.
> The 10M context ability wipes out most RAG stack complexity immediately.
Remains to be seen.Large contexts are not always better. For starters, it takes longer to process. But secondly, even with RAG and the large context of GPT4 Turbo, providing it a more relevant and accurate context always yields better output.
What you get with RAG is faster response times and more accurate answers by pre-filtering out the noise.
btw, 10M tokens is 78 times more context window than the newest GPT-4-turbo (128K). In a way, you don't need 78 GPT-4 API calls, only one batch call to Gemini 1.5.
People also seem to forget that the average is 1b words that are read by people in their entire LIFETIME, and at 10m, with nearly 100% recall thats pretty damn amazing, i'm pretty sure I don't have perfect recall of 10m words myself lol
It can also be a good alternative for fine-tuning.
And the use case of a code base is a good example: if the ai understands the whole context, it can do basically everything.
Let me pay 5€ for a android app rewritten into iOS.
For any use case where you want contextual results, you need to be able to either filter the search scope or use RAG to pre-define the acceptable corpus.
Unless you can get nearly perfect recall with millions of tokens, which is the claim made here.
An actually useful RAG would be to convert text to Q&A and use Q's embeddings as an index. Large context can make use of in-context learning to make better Q&A.
We also embed the actual text, though, because I found that only doing the questions resulted in inferior performance.
1. Get text from page/section/chunk
2. Generate possible questions related to the page/section/chunk
3. Generate an embedding using { each possible question + page/section/chunk }
4. Incoming question targets the embedding and matches against { question + source }
Is this roughly it? How many questions do you generate? Do you save a separate embedding for each question? Or just stuff all of the questions back with the page/section/chunk?> 2. They don't talk about how they get to 10M token context
Yes. I wonder if they're using a "linear RNN" type of model like Linear Attention, Mamba, RWKV, etc.
Like Transformers with standard attention, these models train efficiently in parallel, but their compute is O(N) instead of O(N²), so in theory they can be extended to much longer sequences much efficiently. They have shown a lot of promise recently at smaller model sizes.
Does anyone here have any insight or knowledge about the internals of Gemini 1.5?
"This includes making Gemini 1.5 more efficient to train and serve, with a new Mixture-of-Experts (MoE) architecture."
One thing you could do with MoE is giving each expert different subsets of the input tokens. And that would definitely do what they claim here: it would allow search. You want to find where someone said "the password is X" in a 50 hour audio file, this would be perfect.
If your question is "what is the first AND last thing person X said" ... it's going to suck badly. Anything that requires taking 2 things into account that aren't right next to eachother is just not going to work.
Don't MoE's route tokens to experts after the attention step? That wouldn't solve the n^2 issue the attention step has.
If you split the tokens before the attention step, that would mean those tokens would have no relationship to each other - it would be like inferring two prompts in parallel. That would defeat the point of a 10M context
Uh sorta but not like parent described at all. You have multiple "experts" and you have a routing layer(s) that decide which expert to send it to. Usually every token is sent to at least 2. You can't just send half the tokens to one expert and half to another.
Also the "experts" are not "domain experts" - there is not a "programming expert" and an "essay expert".
They kinda address that in the technical report[0]. On page 12 they show results from a "multiple needle in a haystack" evaluation.
https://storage.googleapis.com/deepmind-media/gemini/gemini_...
They try to push that, but it's not the most convincing. Look at Table 8 for text evaluations (math, etc.) - they don't even attempt a comparison with GPT-4.
GPT-4 is higher than any Gemini model on both MMLU and GSM8K. Gemini Pro seems slightly better than GPT-4 original in Human Eval (67->71). Gemini Pro does crush naive GPT-4 on math (though not with code interpreter and this is the original model).
All in 1.5 Pro seems maybe a bit better than 1.0 Ultra. Given that in the wild people seem to find GPT-4 better for say coding than Gemini Ultra, my current update is Pro 1.5 is about equal to GPT-4.
But we'll see once released.
For my use cases, Gemini Ultra performs significantly better than GPT-4.
My prompts are long and complex, with a paragraph or two about the general objective followed by 15 to 20 numbered requirements. Often I'll include existing functions the new code needs to work with, or functions that must be refactored to handle the new requirements.
I took 20 prompts that I'd run with GPT-4 and fed them to Gemini Ultra. Gemini gave a clearly better result in 16 out of 20 cases.
Where GPT-4 might miss one or two requirements, Gemini usually got them all. Where GPT-4 might require multiple chat turns to point out its errors and omissions and tell it to fix them, Gemini often returned the result I wanted in one shot. Where GPT-4 hallucinated a method that doesn't exist, or had been deprecated years ago, Gemini used correct methods. Where GPT-4 called methods of third-party packages it assumed were installed, Gemini either used native code or explicitly called out the dependency.
For the 4 out of 20 prompts where Gemini did worse, one was a weird rejection where I'd included an image in the prompt and Gemini refused to work with it because it had unrecognizable human forms in the distance. Another was a simple bash script to split a text file, and it came up with a technically correct but complex one-liner, while GPT-4 just used split with simple options to get the same result.
For now I subscribe to both. But I'm using Gemini for almost all coding work, only checking in with GPT-4 when Gemini stumbles, which isn't often. If I continue to get solid results I'll drop the GPT-4 subscription.
I am an experienced programmer and usually have a fairly exact idea of what I want, so I write detailed requirements and use the models more as typing accelerators.
GPT-4 is useful in this regard, but I also tried about a dozen older prompts on Gemini Advanced/Ultra recently and in every case preferred the Ultra output. The code was usually more complete and prod-ready, with higher sophistication in its construction and somewhat higher density. It was just closer to what I would have hand-written.
It's increasingly clear though LLM use has a couple of different major modes among end-user behavior. Knowledge base vs. reasoning, exploratory vs. completion, instruction following vs. getting suggestions, etc.
For programming I want an obedient instruction-following completer with great reasoning. Gemini Ultra seems to do this better than GPT-4 for me.
Here's a prompt I used and got a a script that not only accomplishes the objective, but even has an option to show what files will be deleted and asks for confirmation before deleting them.
Write a bash script to delete all files with the extension .log in the current directory and all subdirectories of the current directory.
Gemini Advanced is nowhere even close to GPT-4, either for text generation, code generation or logical reasoning.
Gemini Advanced is constantly asking for directions "What are your thoughts on this approach?" even to create a short task list of 10 items. Even when being told several times to provide the full list, and not stop at every three or four items and ask for directions. Is constantly giving moral lessons or finishing the results with annoying marketing style comments of the type "Let's make this an awesome product!"
Code is more generic, solutions are less sophisticated. On a discussion of Options Trading strategies Gemini Advanced got core risk management strategies wrong and apologized when errors were made clear to the model. GPT-4 provided answers with no errors, and even went into the subtleties of some exotic risk scenarios with no mistakes.
Maybe 1.5 will be it, or maybe Google realized this quite quickly and are trying the increased token size as a Hail Mary to catch up. Why release so soon?
Quite curious to try the same prompts on 1.5.
I'm always reluctant to write long prompts because I often find GPT4 just doesn't get it, and then I've wasted ten minutes writing a prompt
I've never had Gemini give me a better result than GPT, though, so it does not surpass it for my needs.
The UI is more responsive, though, which is worth something.
I guess this is a tough request if you're working on a proprietary code base, but I would love to see some concrete examples of the prompts and the code they produce.
I keep trying this kind of prompting with various LLM tools including GPT-4 (haven't tried Gemini Ultra yet, I admit) and it nearly always takes me longer to explain the detailed requirements and clean up the generated code than it would have taken me to write the code directly.
But plenty of people seem to have an experience more like yours, so I really wonder whether (a) we're just asking it to write very different kinds of code, or (b) I'm bad at writing LLM-friendly requirements.
> I'm building a note-taking app in flutter. I want to create a way to link between notes (like a web hyperlink) that opens a different note when a user clicks on it. They should be able to click on the link while editing the note, without having to switch modalities (eg. no edit-save-view flow nor a preview page). How can I accomplish this?
I also included a follow-up prompt after getting the first answer, which again for Gemini was super meaningful, and already included valid code to start with. Gemini also showed me many more projects and examples from the broader internet.
> Can you write a complete Widget that can implement this functionality? Please hard-code the note text below: <redacted from HN since its long>
I've definitely had success using LLMs as a learning tool. They hallucinate, but most often the output will at least point me in a useful direction.
But my day-to-day work usually involves non-exploratory coding where I already know exactly how to do what I need. Those are the tasks where I've struggled to find ways to make LLMs save me any time or effort.
Yea absolutely. I also use it to just write code I understand but am too lazy to write, but it's definitely effective at "show me how this works" type learning too.
> Those are the tasks where I've struggled to find ways to make LLMs save me any time or effort
Github CoPilot has an IDE integration where it can output directly into your editor. This is great for "// TODO: Unit Test for add(x, y) method when x < 0" and it'll dump out the full test for you.
Similarly useful for things like "write me a method that loops through a sorted list, and finds anything with <condition> and applies a transformation and saves it in a Map". Basically all those random helper methods and be written for you.
fun foo(list: List<Bar>) =
list.filter { condition(it) }.associateWith { transform(it) }
which would take me less time to write than the prompt would.However, if I didn't know Kotlin very well, I might have had to go look in the docs to find the associateWith function (or worse, I might not have even thought to look for it) at which point the prompt would have saved me time and taught me that the function exists.
Though they talk a bunch about how hard it was to filter out Human Eval, so this probably doesn't matter much.
I'm skeptical, my past experience is just becaues the context has room to stuff whatever you want in it, the more you stuff in the context the less accurate your results are. There seems to be this balance of providing enough that you'll get high quality answers, but not too much that the model is overwhelmed.
I think a large part of developing better models is not just a better architectures that support larger and larger context sizes, but also capable models that can properly leverage that context. That's the test for me.
What comes to my mind: run the usual gamut of tests, but with the excess context window saturated with irrelevant(?) data. Measure test answer accuracy/verbosity as a function of context saturation percentage. If there's little correlation between these two variables (e.g. 9% saturation is just as accurate/succinct as 99% saturation), then "muddiness" isn't an issue.
A handful of examples show whether it can do it. For example, GPT-4 turbo is downright awful at something like that.
The Sora release is even more mind blowing - not the video generation in my mind but the idea that it can infer properties of reality that it has to learn and constrain in its weights to properly generate realistic video. A side effect of its ability is literally a small universe of understanding.
I was thinking that I want to play with audio to audio LLMs. Not text to speech and reverse but literally sound in sound out. It clears away the problem of document layout etc. and leaves room for experimentation on the properties of a cognitive being.
Based on Google's track record in the area of text chatbots, I am extremely skeptical of their claims about coherency across a 1M+ context window.
Of course none of this even matters anyway because the weights are closed the architecture is closed nobody has access to the model. I'll believe it when I see it.
There's a language called Kalamang with only 200 native speakers left. There's a set of grammar books for this language that adds up to ~250K tokens. [1]
They set up a test of in-context learning capabilities at long context - they asked 3 long-context models (GPT 4 Turbo, Claude 2.1, Gemini 1.5) to perform various Kalamang -> English and English -> Kalamang translation tasks. These are done either 0-shot (no prior training data for kgv in the models), half-book (half of the kgv grammar/wordlists - 125k tokens - are fed into the model as part of the prompt), and full-book (the whole 250k tokens are fed into the model). Finally, they had human raters check these translations.
This is a really neat setup, it tests for various things (e.g. did the model really "learn" anything from these massive grammar books) beyond just synthetic memorize-this-phrase-and-regurgitate-it-later tests.
It'd be great to make this and other reasoning-at-long-ctx benchmarks a standard affair for evaluating context extension. I can't tell which of the many context-extension methods (PI, E2 LLM, PoSE, ReRoPE, SelfExtend, ABF, NTK-Aware ABF, NTK-by-parts, Giraffe, YaRN, Entropy ABF, Dynamic YaRN, Dynamic NTK ABF, CoCA, Alibi, FIRE, T5 Rel-Pos, NoPE, etc etc) is really SoTA since they all use different benchmarks, meaningless benchmarks, or drastically different methodologies that there's no fair comparison.
[1] from https://storage.googleapis.com/deepmind-media/gemini/gemini_...
The available resources for Kalamang are: field linguistics documentation10 comprising a ∼500 page reference grammar, a ∼2000-entry bilingual wordlist, and a set of ∼400 additional parallel sentences. In total the available resources for Kalamang add up to around ∼250k tokens.
The first chart (Cumulative Average NLL for Long Documents) shows a deviation from the trend and an increase in accuracy when working with >=1M tokens. The 1.0 graph is overlaid and supports the experience of 'muddiness'.
[1] https://storage.googleapis.com/deepmind-media/gemini/gemini_...
Now that Google has tasted the previously forbidden fruit of layoffs themselves, I think their primary goal in ML is now headcount reduction.
What if it was possible, with each query, to fine tune the model on the provided context, and then use that JIT fine-tuned model to answer the query?
As it is now, you can't fine tune on context. It would have almost no effect on the parameters.
Context is like giving your friend a magazine article and asking them to respond to it. Fine tuning is like throwing that magazine article into the ocean of all content they ever came across during their lifetime.
The way I understand it, there is a base model that was trained on vast amount of general data. This sets up the weights.
You can fine-tune this base model on additional data. Often this is private data that is concentrated around a certain domain. This modifies the model's weights some more.
Then you have the context. This is where your query to the LLM goes. You can also add the chat history here. Also, system prompts that tell the LLM to behave a certain way go here. Finally, you can take additional information from other sources and provide it as part of the context -- this is called Retrieval Augmented Generation. All of this really goes into one bucket called the context, and the LLM needs to make sense of it. None of this modifies the weights of the model itself.
Is my mental picture correct so far?
My question is around RAG. It seems that providing additional selected information from your knowledge base, or using your knowledge base to fine-tune a model, seem similar. I am curious in which ways these are similar, and in which ways they cause the LLM to behave differently.
Concretely, say I have a company knowledge base with a bunch of rules and guidelines. Someone asks an agent "Can I take 3 weeks off in a row?" How would these two scenarios be different:
a) Agents searches the knowledge base for all pages and content related to "FTO, PTO, time off, vacations" and feeds those articles to the LLM, together with the "Can I take 3 weeks off in a row?" query
b) I have an LLM that has been fine tuned on all the content in the knowledge base. I ask it "Can I take 3 weeks off in a row?"
Yes
> How would these two scenarios be different
They're different in exactly the way you described above. The agent searching the knowledge base for "FTO, PTO, time off, vacations" would be the same as you pasting all the articles related to those topics into the prompt directly - in both cases, it goes into the context.
In scenario a, you'll likely get the correct response. In scenario b, likely get an incorrect response.
Why? Because of what you explained above. Fine tuning adjusts the weights. When you adjusts weights by feeding data, you're only making small adjustments to shift slightly along a curve - thus the exposure to this data (for the purposes of fine tuning) will have very little effect on the next context the model is exposed to.
Edit: Ah, I see, it's 1M reliably in production, up to 10M in research:
> Through a series of machine learning innovations, we’ve increased 1.5 Pro’s context window capacity far beyond the original 32,000 tokens for Gemini 1.0. We can now run up to 1 million tokens in production.
> This means 1.5 Pro can process vast amounts of information in one go — including 1 hour of video, 11 hours of audio, codebases with over 30,000 lines of code or over 700,000 words. In our research, we’ve also successfully tested up to 10 million tokens.
The video queries they show take around 1 minute each, this probably burns a ton of GPU. I appreciate how clearly they highlight that the video is sped up though, they're clearly trying to avoid repeating the "fake demo" fiasco from the original Gemini videos.
I imagine they have some new way to route tokens to the experts that probably computes a global context. One scalable way to compute a global context is by a state space model. This would act as a controller and route the input tokens to the MoEs. This can be computed by convolution if you make some simplifying assumptions. They may also still use transformers as well.
I could be wrong but there are some Mamba-MoEs papers that explore this idea.
This may not be true. My experience of the complexity of RAG lays in how to properly connect to various unstructured data sources and perform data transformation pipeline for large scale data set (which means GB, TB or even PB). It's in the critical path rather a "nice to have", because the quality of data and the pipeline is a major factor for the final generated the result. i.e., in RAG, the importance of R >>> G.
HN is very focused on technical feasibility (which remains to be seen!), but in every LLM opportunity, the CIO/CFO/CEO are going to be concerned with the cost modeling.
The way that LLMs are billed now, if you can densely pack the context with relevant information, you will come out ahead commercially. I don't see this changing with the way that LLM inference works.
Maybe this changes with managed vector search offerings that are opaque to the user. The context goes to a preprocessing layer, an efficient cache understands which parts haven't been embedded (new bloom filter use case?), embeds the other chunks, and extracts the intent of the prompt.
The leading ability AI (in terms of cognitive power) will, generally, cost more per token than lower cognitive power AI.
That means that at a given budget you can choose more cognitive power with fewer tokens, or less cognitive power with more tokens. For most use cases, there's no real point in giving up cognitive power to include useless tokens that have no hope of helping with a given question.
So then you're back to the question of: how do we reduce the number of tokens, so that we can get higher cognitive power?
And that's the entire field of information retrieval, which is the most important part of RAG.
Really? Because to my understanding the compute necessary to generate a token grows linearly with the context, and doesn't the OpenAI billing reflect that by seperating prompt and output tokens?
Google themselves have such a huge footprint of various businesses, that they alone would be an amazing customer for this, never mind all the other cool opportunities from third parties...
Imagine that they can ingest the entirety of YouTube and then dump that into Google Search's index AND use it to generate training data for their next LLM.
Imagine that they can hook it up to your security cameras (Nest Cam), and then ask questions about what happened last night.
Imagine that you can ask Gemini how to do something (eg. fix appliance), and it can go and look up a YouTube video on how to accomplish that ask, and explain it to you.
Imagine that it can apply summarization and descriptions to every photo AND video in your personal Google Photos library. You can ask it to find a video of your son's first steps, or a graduation/diploma walk for your 3rd child (by name) and it can actually do that.
Imagine that Google Meet video calls can have the entire convo itself fed into an LLM (live?), instead of just a transcription. You can have an AI assistant there with you that can interject and discuss, based on both the audio and video feed.
Problem is they are jeopardizing their moat.
Google is still in a great position, they have the knowledge and lots of data to pull this off. They just have to take the risk of losing some ad revenue for a while.
Longer context on the other hand, could put some RAG use cases to sleep, if your instructions are like, literally a manual long, then there is no need for rag.
You probably need a couple GB to cache a conversation. That's not so easy at the moment because you have to transfer that data to and from the GPUs and store the data somewhere.
You'll notice in their video [1] that they never show the prompts running interactively. This is for a roughly 800K context. They claim that "the model took around 60s to respond to each of these prompts".
This is not really usable as an interactive experience. I don't want to wait 1 minute for an answer each time I have a question.
If you are going to ask "then why don't OpenAI do it now", the answer is it takes a lot of storage (and IO) so it may not be worth it for shorter context, it adds significant complexity to the entire serving stack, and is incoherent with how OpenAI originally imagined where the "custom-ish" LLM serving game is going to - they bet on finetuning and dedicated instances, instead of long context.
The tradeoff can be reflected in the API and pricing, LLM APIs don't have to be like OpenAI's. What if you have an endpoint to generate a "cache" of your context (or really, a prefix of your prompt), billed as usual per token, then you can use your prompt prefix for a fixed price no matter how long it is?
Here’s the paper: https://arxiv.org/abs/2312.00752
And here’s a great podcast episode on it: https://www.cognitiverevolution.ai/emergency-pod-mamba-memor...
I don't know how either but maybe https://news.ycombinator.com/item?id=39367141
Anyway I mean, there is plenty of public research on this so it's probably just a matter of time for everyone else to catch up
As far as I know, the problem in most cases is that while the context length might be high in theory, the actual ability to use it is still limited. E.g. recurrent networks even have infinite context, but they actually only use 10-20 frames as context (longer only in very specific settings; or maybe if you scale them up).
Anyways the ability to generalize to longer context length is evidenced by such tests. If every token of the model’s output is able to answer questions in such a way that any sentence from the input would be taken into account, this gives evidence that the full context window indeed matters. Currently I find Claude 2 to perform very well on such tasks, so that sets my expectation of how a language model with an extremely long context window should look like.
GPT-4 Turbo, using its full 128k context, costs around $1.28 per API call.
At that pricing, 1m tokens is $10, and 10m tokens is an eye-watering $100 per API call.
Of course prices will go down, but the price advantage of working with less will remain.
It may still be worth it for some use cases
I fully disagree, they compare Gemini 1.5 Pro and GPT4 only on context length, not on other tasks where they compare it only to other Gemini which is a strange self-own.
I'm convinced that if they do not show the results against GPT4/Claude, it is because they do not look good.
RAG is needed for the same reason you don't `SELECT *` all of your queries.
Why did you point out this distinction?
I believe those two things together are likely enough to explain the difference between a 1M context length and a 10M context length.
[0]: Which is not looking down on that particular research team, the vast majority of people have less means and optimization know-how than Google.
1. People mention accuracy issues with longer contexts 2. People mention processing time issues with longer contexts 3. Something people haven't mentioned in this thread is cost -- even thought prompt tokens are usually cheaper than generated tokens, and Gemini seems to be cheaper than GPT-4, putting a whole knowledge base or 80-page document in the context is going to make every time you run that prompt quite expensive
I'd imagine RAG would still be much more efficient computationally
> Retrieval Augmented Generation (RAG) is a technique where the capabilities of a large language model (LLM) are augmented by retrieving information from other systems and inserting them into the LLM’s context window via a prompt.
(stolen from: https://github.com/psychic-api/rag-stack)
Here:
https://blogs.nvidia.com/blog/what-is-retrieval-augmented-ge...
RAG is training AI to be a guy who read a lot of books. He doesn't know all of them in the context of this conversation you are having with him, but he sort of remembers where he read about the thing you are talking about and he has a library behind him into which he can reach and cite what he read verbatim thus introducing it into the context of your conversation.
I might be wrong though. I'm a newb.
From a technology standpoint, maybe. From an economics standpoint, it seems like it would be quite expensive to jam the entire corpus into every single prompt.
My $5 says it's a RAG or a similar technique (hierarchical RAG comes to mind), just like all other large context LLMs.
>HumanEval is an industry standard open-source evaluation benchmark (Chen et al., 2021), but we found controlling for accidental leakage on webpages and open-source code repositories to be a non-trivial task, even with conservative filtering heuristics. An analysis of the test data leakage of Gemini 1.0 Ultra showed that continued pretraining on a dataset containing even a single epoch of the test split for HumanEval boosted scores from 74.4% to 89.0%, highlighting the danger of data contamination. We found that this sharp increase persisted even when examples were embedded in extraneous formats (e.g. JSON, HTML). We invite researchers assessing coding abilities of these models head-to-head to always maintain a small set of truly held-out test functions that are written in-house, thereby minimizing the risk of leakage. The Natural2Code benchmark, which we announced and used in the evaluation of Gemini 1.0 series of models, was created to fill this gap. It follows the exact same format of HumanEval but with a different set of prompts and tests.
"Studying the limits of Gemini 1.5 Pro's long-context ability, we find continued improvement in next-token prediction and near-perfect retrieval (>99%) up to at least 10M tokens"
https://storage.googleapis.com/deepmind-media/gemini/gemini_...
It's also not obvious how these huge models will fare against increasingly capable open source ones like Mixtral, perhaps especially since Google are confirming here that MoE is the path forward, which perhaps helps limit how big these models need to be.
But there is also the downside of "tuning" the RAG to return less tokens you will miss extra context that could be useful to the model.
Having 99% retrieval is nuts too. Models tend to unwind pretty badly as the context (tokens) grows.
Put these together and you are getting into the territory of dumping all your company documents, or all your departments documents into a single GPT (or whatever google will call it) and everyone working with that. Wild.
Input is parsed one token at a time right? Can you cache the state after the initial prompt has been provided?
One point I'm unclear on is how these huge context sizes are implemented by the various models. Are any of them the actual raw "width of the model" that is propagated through it, or are these all hierarchical summarization and chunk embedding index lookup type tricks?
The Encyclopedia Britannica is ~44M words.
As the other comment mentions, you can paste the content of entire books or documents and ask very pointed question about it. Last year, Anthropic was showing off their 100K context window, and that's exactly what they did, they gave it the content of The Great Gatsby and asked it questions about specific lines of the book.
Similarly, imagine giving it hundreds of documents and asking it to spot some specific detail in there.
>Finally, we highlight surprising new capabilities of large language models at the frontier; when given a grammar manual for Kalamang, a language with fewer than 200 speakers worldwide, the model learns to translate English to Kalamang at a similar level to a person learning from the same content.
Results - https://imgur.com/a/qXcVNOM
If you watch the videos in the blog post, you can see it's a screen recording on a computer without any editing/stitching of different scenes together.
It's good to be sceptical but as engineers we should all remain open.
Google appears to be making strides in catching up.
When it comes to my personal workflow and accomplishing tasks, I still find ChatGPT to be the most effective tool. My familiarity with its features has made it indispensable. The integration of mentions and tailored GPTs seamlessly enhances my workflow.
While Gemini may match the foundational capabilities of LLMs, it falls short in delivering a product that efficiently aids in task completion.
But maybe you do, and I am seeing patterns in sand.
I say it's even more than that. OpenAI had a bigger lead when it released GPT-2 than it does now. They're burning through cash to try to hold on to a lead of a few months over the competition.
This seems to be a common experience, as apparently it refuses to give advice on copying memory in C# [2], and I tried to do what was suggested in this comment [3], but by the next prompt it was refusing again, so I had to stick to ChatGPT.
[0] https://g.co/gemini/share/238032386438
[1] https://g.co/gemini/share/6880989ddfaf
> Concepts are an advanced feature of C++ that introduces potential risks, and I want to prioritize your safety.
Brilliant.
[0] https://g.co/gemini/share/fa9d60da921d
[1] https://g.co/gemini/share/e20655d06292
[2] https://g.co/gemini/share/f11bc9f7e658
Results - https://imgur.com/a/qXcVNOM
From the technical report https://storage.googleapis.com/deepmind-media/gemini/gemini_...
That's an incredibly low bar
The same feat one year ago would have been almost unbelievable.
People point out mistakes it makes that no human would make, but that doesn't negate the super-human performance it has at other tasks -- and the _breadth_ of what it can do is far beyond any single person.
the issue remains on accuracy. i think a human in that scenario is still more accurate with their responses, and i do not yet see that being overcome in this multi-year llm battle.
Just a few years ago we used to clap if an NLP model could handle negation reliably or could generate even a paragraph of text in English that was natural sounding.
Now we are at a stage where it is basically producing reams of natural sounding text, performing surprisingly well on reasoning problems and translation of languages with barely any data despite being a markov chain on steroids, and what does it hear? "That's an incredibly low bar".
And as you say, the goalposts keep getting moved. It used to be claimed that computers could never play chess at the highest levels because that required "insight". And whatever a computer could do, it could never to that extra special thing, that could only be described in magical undefined terms.
I just hope there's a moment of reckoning for decades upon decades of arguments, deemed academically respectable, that insisted that days like these would never come.
From the language benchmark (parentheses mine).
That said, I think that if you gave a LLM language text to predict during training, I believe that even if no parallel corpora exists during training, we could have a LLM that could still translate that language to some other language it also trained on.
He learned how to promote himself from working for Peter "Project Milo" Molyneux and I see similar patterns of hype.
[1] https://en.wikipedia.org/wiki/Republic:_The_Revolution#Marke...
Nonetheless while still underwhelming in comparison to gpt-4 (excluding this announcement as I haven't tried it yet), alpha go, zero and especially fold were tremendous!
Back when OpenAI still supported raw text completion with text-davinci-003 I spent some time experimenting with tiny prompt-embedded DSLs. The results were very, very, interesting IMO. In a lot of ways, text-davinci-003 with embedded functions still feels to me like the "smartest" language model I've ever interacted with.
I'm not sure how close we are to "superintelligence" but for baseline general intelligence we very well could have already made the prerequisite technological breakthroughs.
I wanna see how far this tech can scale, regardless of speed. I don't care if it takes 24h to formulate a response. Are there "easy" variables which drastically improve output?
I suspect not. I imagine people have tried that. Though i'm still curious as to why.
As a particularly egregious example, yesterday night I gave Gemini a list of drinks and other cocktail ingredients I had laying around and asked for some recommendations for cute drinks that I could make. It's response:
> I'm just a language model, so I can't help you with that.
ChatGPT 3.5 came up with several delicious options with clear instructions, but it's not just this instance, I've NEVER gotten a response from Gemini that I even felt was more useful than just a freaking bing search! Much less better than ChatGPT. I'm just going to assume they're using cherrypicked metrics to make themselves feel better until proven otherwise. I have zero confidence in Google's AI plays, and I assume all their competent talent is now at OpenAI or Anthropic.
When Gemeni (Ultra) refuses to do something itself it is more verbose and specific as to why it won't do it, it my experience.
But my main takeaway is the huge context window! Up to a million, with more than 100k tokens right now? Even just GPT 3.5 level prediction with such a huge context window opens up a lot of interesting capabilities. RAG can be super powerful with that much to work with.
So Pro is like the light and fast version and Ultra the advanced and expensive one.
Nano/Pro/Ultra are model SIZES. 1.0/1.5 is generations of the architecture.
Wouldn't the transcript or at least a timeline of Apollo 11 be part of the training corpus? So even without the 400 pages in the context window just given the drawing I would assume a prompt like "In the context of Apoll 11, what moment does the drawing refer to?" would yield the same result.
> The "Snoopy" Moment: During the mission, the crew had a small, black-and-white cartoon Snoopy doll as a semi-official mascot, representing safety and mission success. At one point, Collins joked about "Snoopy" floating into his view in the spacecraft, which was a light moment reflecting the camaraderie and the use of humor to ease the intense focus required for their mission.
The "Biohazard" Joke: After the successful moon landing and upon preparing for re-entry into Earth's atmosphere, the crew humorously discussed among themselves the potential of being quarantined back on Earth due to unknown lunar pathogens. They joked about the extensive debriefing they'd have to go through and the possibility of being a biohazard. This was a light-hearted take on the serious precautions NASA was taking to prevent the hypothetical contamination of Earth with lunar microbes.
The "Mailbox" Comment: In the midst of their groundbreaking mission, there was an exchange where one of the astronauts joked about expecting to find a mailbox on the Moon, or asking where they should leave a package, playing on the surreal experience of being on the lunar surface, far from the ordinary elements of Earthly life. This comment highlighted the astronauts' ability to find humor in the extraordinary circumstances of their journey.
And I think people who buy a laptop with a 1TB SSD generally don't run out of space, at least I don't.
There are a bunch of techniques to do this, but it's unclear how well any of them scale.
vs fine tuning: smaller, fine-tuned models can perform better than huge models in a decent number of tasks. Not strictly fine-tuning, but for throughput limited tasks it'll likely still be better to prune a 70B model down to 2B, keeping only the components you need for accurate inference.
I can see this model being good for taking huge inputs and compressing them down for smaller models to use.
Think building huge relevant context on topics before answering.
Sure it's a bummer that they slap the "Join the waiting list", but it's still interesting to read about their progress and competing with ClosedAi (OpenAi).
One last thing I hope they fix is the heavily morally and ethically guardrail, sometimes I can barely ask proper questions without it triggering Gemini to educate me about what's right and wrong. And when I try the same prompt with ChatGPT and Bing ai, they happily answer.
Did you mean disadvantages?
I don't care if the model can tell me which page in the book or which code file has a particular concept. RAG already does this. I want the model to notice how a concept is distributed throughout a text, and be able to connect, compare, contrast, synthesize, and understand all the ways that a book touches on a theme, or to rewrite multiple code files in one pass, without introducing bugs.
How does Gemini 1.5's reasoning compare to GPT-4? GPT-4 already has superhuman memory; its bottleneck is its relatively weak reasoning.
Testing language translation abilities of an extremely obscure language after passing in one grammar book as context.
"Remember val="XXXX" .........10M tokens later....... Print val"
It’s difficult to make that array longer because training time explodes.
According to this it's remembering with 99% accuracy, which if you think about it is NUTS, can you imagine reading a 22x 1000 page books, and remembering every single word that was said with 100% accuracy lol
Another possible synthetic benchmark would be to present a list of key value pairs and then ask it for the value corresponding to different keys. Or present a long list of distinct facts and then ask it about them. This latter one could probably be sourced from something like a trivia question and answers data set. I bet there's something like that from Jeopardy.
Then you can ask it who is Tina's (great)^57 grandmother's twice removed cousin on her father's side?
It would have to be able to remember the context of the relationships up and down the document and there'd be nothing to key into as you could ask about any relationship.
Not surprising that OpenAI shipped a blog post today about their video generation — I think they're feeling considerable heat right now.
Do we know how much time or money does it take to create a movie clip?
And i thought it would be easy, what a rookie mistake.
Looks like "France" isn't on the list of available regions for Ai Studio ?
Now i'm trying to use Vertex AI, not even sure what's the difference with Ai Studio, but it seems it's available.
So far i've been struggling for 15 minutes through a maze of google cloud pages: console, docs, signups. No end in sight, looks like i won't be able to try it out
I can't get on the waitlist, because the waitlist link redirects to aistudio and I can't use that.
I should stop expecting that I can use literally anything google announces.
So 1M tokens is going to be astronomical.
Sweet, this opens up so many possibilities.
So Pro is better than Ultra, but only if the version numbers are higher?
The Nano < Pro < Ultra is an in-revision thing. For their LLMs it's a size thing. Then there's newer releases of Nano, Pro, and Ultra. Some Pro might be better than some older Ultra.
A lot of people seem confused about this but it feels so easy to understand that it's confusing to me that anyone could have trouble.
Google requires me to navigate their absolutely insane console (seriously, I thought the AWS console was bad, but GCP takes the cake), only to tell me there is not even a way to get an API key... I had to ask Gemini through the built in interface to figure that out.
Unfortunately there's a waitlist for the 1.5 architecture
There is no error, it just redirects me.
Fail.
CEOs love to talk about how important regulation is, how their company needs to develop it before the "wrong people" do, and how they are concerned about what could happen if AI development goes wrong.
Then they announce the latest model that is aimed at expanding both the accuracy and breadth of use cases across multiple modalities. Sure the release links to a security and ethics page, but that page reads more like a company's internal "pillars of success" document with vague phrases that define little to nothing in the way of real, specific concerns or measures to mitigate them. It basically boils down to "Don't be evil" with no clear definition of what that would mean or how they prevent the new, more powerful and broad reaching system from being used in ways that are "evil".
We can argue that a lot of people have done pretty bad things using the internet, but should it have been regulated in advance?
Lock up the hardware in an offline facility and experiment there, if they really think it's important. Hell, even just skipping the double speak would be a big step. If they really aren't concerned with the risk then own it, don't tell me its risky while also releasing a new, more powerful version every 6-12 months.
If companies and their leadership couldn't operate so unchecked by our existing laws and public opinion we may not have executives worth worrying about.
For example, if taxes were so easy to dodge and if the public actually had a chance to sue large corporations for damages they may not get so large. If, when losing a lawsuit, companies couldn't shuffle around funds and spin off dummy companies to dodge the pain, and if they weren't often forced to pay pennies on the dollar for lost suits, they may think twice about doing some things. When you know your entire business is actually on the line you have to be more careful.
Throw in election and lobbying reform and we could at least be having a much different conversation about corporate power.
Limiting public features will help a bit with concerns over how someone might use a public GPT API, but the technology advancements will be made either way and ultimately companies won't be able to gate keep who can use it with 100% accuracy. The boom for GPU hardware similarly is pushing us further down the road to AI development and all the moral and ethical questions that go along with it, even if AI companies were to keep use of their GPUs and models private entirely.
How long until it shows similar results on middle-sized and large codebases? And do the job adequately?
And we should keep in mind that to understand a code change in depth is often just as much work as making the change. When review PRs I don't really know exactly what every change is doing. I certain haven't tested it to be 100% certain I understand fully. I'm just checking the logic looks mostly right and that I don't see anything clearly wrong, and even then I'll often need to ask for clarifications why something was done.
I can't imagine LLMs being used in most large code bases for a while yet. They'd probably need to be 99.9% reliable before we can start trusting them to make changes without verifying every line.
Hope they do a good job and once OpenAI releases GPT 5 they are competitive with it with their offerings, it will be better for everyone.
But I doubt it is /that/ good, it's not like we can test it either
Teach Gemini how to be a Dungeon Master, and run free adventures at Comic Con.
Then offer it up as a subscription.
Sincerely,
Everyone
1M tokens is what they've said will be available for production and is about 2,500 pages.
So actually they are lagging: their 128k model is yet to be released while OpenAI released theirs some months ago.
> Gemini 1.5 Pro comes with a standard 128,000 token context window. But starting today, a limited group of developers and enterprise customers can try it with a context window of up to 1 million tokens via AI Studio and Vertex AI in private preview.
Something like an alpha version.
Limited preview in their jargon.
I hope the demos aren't fudged/scripted like Google did with Gemini 1.0
I asked it to rephrase "Are the original stated objectives still relevant?"
It's starts going on about Ukraine and Russia.
Or you could be more specific, like "Rephrase the following sentence: 'Are the original stated objectives still relevant?' in a formal way, respond with one option only."
Ugh, my brain.
Wait. They are abandoning PaLM 2, which was just announced 9 months ago?
I asked the free model (whatever that is) and it wasn't very helpful, alterating betweens a sales bot for Ultra and being somewhat confused itself.
Edit: apparently it goes 1.0 Pro, 1.0 Ultra, 1.5 Pro, 1.5 Ultra and so on.
Gemini ultra 1.0 is still on version 1.0
If you look at the Gemini report it refers to "Gemini 1.5", then refers to "Gemini 1.5 Pro" and "Gemini 1.0 Pro" and "Gemini 1.5 Pro".
- Gemini 1.5 is the new version of the model Gemini.
- They are at the moment testing it on Gemini Pro and calling it Gemini Pro 1.5
- The testing has shown that Gemini Pro 1.5 is delivering the same quality as Gemini Ultra 1.0 while using less computing power
- Gemini Ultra is still using Gemini 1.0 at the moment
Gemini Models gemini.google.com
------------------------------------
Gemini 1.0 Nano
Gemini 1.0 Pro -> Gemini (free)
Gemini 1.0 Ultra -> Gemini Advanced ($20/month)
Gemini 1.5 Pro -> announced on 2024-02-15 [1]
Gemini 1.5 Ultra -> no public announcements (assuming it's coming)
[1]: https://storage.googleapis.com/deepmind-media/gemini/gemini_...For history of pre-Gemini models at Google, see: https://news.ycombinator.com/item?id=39304441
I'm stunned that Google hasn't appointed some "name veto person" that can just say "no, you aren't allowed to have three different things called 'Gemini Advanced', 'Gemini Pro', and 'Gemini Ultra.'" Like surely it just takes Sundar saying "this is the stupidest fucking thing I've ever seen" to some SVP to fix this.
It’s understandable that later generations are better and higher tiers are also better, but usually there is some period of time in between generations to help differentiate them. Here we have Google advancing capability on two axes at the same time.
I give them a pass as this field is advancing rapidly. So good for them. But I think it’s a legitimate call that it adds complexity to their branding. It is different.
Looks like they fine tuned across use cases and grabbed the mixtral architecture?
Gemini Pro, Gemini Ultra... but was 1.0?
now upgraded but again Gemini Pro? jumping from 1.0 to 1.5?
wait but not Gemini Pro 1.5... Gemini "1.5" Pro
What actually happened between 1.0 and 1.5?
I'm not sure what the deal is, it has to be a marketing hinderance as every major tech company is trying to claw their way up the AI service mountain. Seems like the first step would be cogent naming.
Ultra vs pro vs nano with Ultra unlocked by buying Gemini Advanced is confusing.
I'm also not sure why they make base Gemini available after you have Advanced, because presumably there's no reason to use a worse model.
Google Bard to Google Gemini is what they call Gemini 1.0.
Gemini consists of Gemini Nano, Gemini Pro, & Gemini Ultra.
Gemini Nano is for embedded and portable devices I guess? The free version of Gemini (gemini.google.com) is Gemini Pro. The paid version, called Gemini Advanced is using Gemini Ultra.
What we're reading now is about Gemini Pro version 1.0 switching to version 1.5 as of today.
People make fun of OpenAI for not using product names and just calling it "GPT" but at least it's straightforward: 2, 3, 3.5, 4. (On the API side it's a little more complicated since there's "turbo" and "instruct" but that isn't exposed to users, and turbo is basically the default.)
and why have makersuite/aisuite in the first place, if Vertex AI is the center for all things AI? and why aitestkitchen?
I'm seeing only Gemini 1.0 Pro on Vertex AI. So even if I enabled Google Gemini Advanced (Ultra?), enabled Vertex AI API access, I have to first be blessed by Google to access advanced APIs.
It seems paying for their service doesn't mean anything to Google at this point. As a developer, you have to jump through hoops first.
"Gemini 1.0 Ultra, our most sophisticated and capable model for complex tasks, is now generally available on Vertex AI for customers via allowlist."
https://cloud.google.com/blog/products/ai-machine-learning/g...
Their LLM brand is now Gemini. Gemini comes in three different sizes, Nano/Pro/Ultra.
They recently released 1.0 versions of each, most recently (a few months after Nano and Pro) Ultra.
Today they are introducing version 1.5, starting with the Pro size. They say 1.5 Pro offers comparable performance to 1.0 Ultra, along with new abilities (token window size).
(I agree Small/Medium/Large would be better.)
not an iphone user but just looked at iphone 15. Don't see any mini version. I am guess 'standard' is called just 'iphone' ? Is pro same thing as plus ?
https://www.apple.com/shop/buy-iphone/iphone-15
> Still difficult?
yes your example made it even more confusing.
Apple doesn't announce the iPhone 12 Mini and compare it to the iPhone 11 Pro.
Did you watch the announcements for the M2 and M3 pros? They compared it to the previous generations all the time.
Gemini Advanced seems to be the brand name for the higher price tier for the end-user frontend that gets you Ultra access, similar how ChatGPT Plus gets you ChatGPT 4.
I get it, but it does beg the question whether you will need Advanced now to get 1.5 Pro. Or does everyone get Pro, making it useless to pay for 1.0 Ultra?
I still don't think it's confusing, but that part is definitely messy.
This is where it gets confusing IMO.
It's like if Apple announced macOS Blabahee, starting with Mini, not long after releasing Pro and Air touting benefits of Sonoma.
Also, just.. this is how TFA begins:
> Last week, we rolled out our most capable model, Gemini 1.0 Ultra, [...] Our teams continue pushing the frontiers of our latest models with safety at the core. They are making rapid progress. [...] 1.5 Pro achieves comparable quality to 1.0 Ultra
Last week! And now we have next generation. And the wow is that it's comparable to the best of the previous generation. Ok fine at a smaller size, but also that's all we get anyway. Oh and the most capable remains the last generation one. As long as it's the biggest one.
It's really not that confusing. There are different sizes and different generations, coming out at different times. This pattern is practically as old as computing itself.
I can't even imagine what alternative naming scheme would be an improvement.
I doubt they launched M2 MBAs while the MBP was running M1, for example. Or more directly, a low-mid spec M3 MBP while the top-spec M2 MBP (I assume that would out-benchmark it?) still for sale and no comparable M3 chip yet.
It's not having the matrix of size/power & generation that's confusing, it's the 'next generation' one initially launched not being the best. I think that's mainly it for me anyway.
But they have. The baseline M2 is significantly less powerful than the M1 Max.
What Google's doing is basically exactly like that. It happens all the time that the mid tier of the next generation isn't as good as the top tier of the previous generation. It might even be the norm.
There isn't a set order to things. Sometimes companies release a higher powered version first and then the budget version later, sometimes an entry-level version first and a pro version after. Sometimes both simultaneously. All of these are normal, and can even follow different orders generation to generation.
Google got caught completely flat footed by OpenAI. I'm going to cut them some slack that they want to show the world a bit of flex with their AI chops as soon as they have results.
It's however nowhere explicitly said that I could find. The Technical Report PDF also avoids even hinting at it.
Advanced is a price/service tier for the end-user frontend. At the moment it gets you 1.0 Ultra access vs. 1.0 Pro for the free version. Similar to how ChatGPT Plus gives you 4 instead of 3.5.
I agree this part is messy. Does everyone who had Pro already get 1.5 Pro? If 1.5 Pro is better than 1.0 Ultra, why pay for Advanced? Is 1.5 Pro behind the Advanced paywall? etc.
There are three models: nano/pro/ultra and all are at v1.0
There are two tiers of chat service: basic and pro
There is AIStudio from google through which you can interact with / use directly gemini llms.
Chat service Gemini basic (free) uses Gemini Pro 1.0 llm.
Chat service Gemini advanced uses Gemini Ultra 1.0 llm.
What was shown is ~~Ultra~~ Pro 1.5 LLM which is / will be available to select few for preview to be used via AIStudio.
That leaves a question, what's nano for, and is it only used via AIStudio/API?
Jesus, Google..
How this relates to the end-user chat service/price tiers is still unknown.
The best scenario would be that they just move Gemini free and Advanced tiers to Pro 1.5 and Ultra 1.5, I guess.
That's very difficult.
Vertex AI is their developer API platform.
I agree OpenAI is a bit better at launching for customers on ChatGPT alongside API.
Gemini Advanced is the paid subscription service tier that at the moment gets you access to the Ultra model, similar to how a ChatGPT Plus subscription gets you access to GPT-4.
Honestly, they should have called this part Gemini Chat and Gemini Chat Plus, but of course ego won't let them follow the competitor's naming scheme.
With an already complex naming for regular consumers (Nano/Pro/Ultra each one with a 1.x), adding this Advanced thing it becomes and spaghetti.
I understand that for most people may be just a chat input and don't care, but if people will consider to pay, they will research a bit and is confusing.
You see this problem isn't unique to Google.
and if Pro 1.5 is this good holy shit what will Ultra be...
Nano/Pro/Ultra are the model sizes, 1.0 or 1.5 is the version
(via https://news.ycombinator.com/item?id=39383593, but we merged those comments hither)
The first of it's type AR/VR hardware has, understandably, a longer release cycle. Also, Apple announced early to drive up developer interest.
Apple aggressively keeps products under wraps before launch fires employees and vendors for leaking any sort of news to the press .
Also an hardware product that is miles ahead of competition in terms of components and also needs complex setup workflow (for head and eyes) something apple has not done before being 7-8 months after announcing is not really comparable with a SaaS API in terms of delays
They’re an enterprise software company doing an enterprise sales motion.
Anthropic's Claude targets mostly business use cases and you don't see them write self-congratulating articles about Claude v2.1, they just pushed the product.
I work at a very large company and everyone knows about ChatGPT and Gemini (in part because we for our sins have a good chunk of GCP stuff), but I doubt anyone here not doing some LLM-flavored development has ever even heard of Anthropic, let alone Claude.
Why would anyone look to form a contract with Anthropic right now? I'd say they're in danger here, because their models and offerings don't have clear value propositions to customers.
Seems reasonably similar in tone to the Google post.
Really? Someone ought to tell them.
Look, it now has totally useless suggestions like it was trained on burned out woke IT workers. I asked it about the weather, sea temperature and wave height and period in Malaga, which is much less boring than the choices it came up with. First it tried to talk me out of it waving away responsibility, then it provided useful climate data, which I would have wasted too much time doing a Google search on. I guess it's good for checking on the weather if you can put up with the waivers. Also it knows fishing for garfish in Denmark in May is not a total waste of your time, a great way to experience local culture and a sustainable activity.
I also asked it about the version: "I am currently running on the Gemini Pro 1.01.5 model".
What’s the goal? Maybe, being able to work with partners without it being a secret project that will inevitably leak, resulting in inaccurate stories in the press. What are non-goals? Driving sales or creating anticipation with a mass audience, like a movie trailer or an Apple product launch.
So they have to announce something, but most people don’t read Hacker News and won’t even hear about it until later, and that’s fine with them.
Apple at least lets me change this by moving my iTunes/App Store account, which is its own ordeal and far from ideal, but at least there's a defined process: Tell us where you think you live, provide a form of payment from that place, maybe we'll believe you.
I'm really frustrated by Google's attitude of "we know better where you are than you do". People travel sometimes and that's not the same thing as moving!
But once I'm a paying customer, I want to use the thing I'm paying for from where I am without jumping through ridiculous hoops!
The worst variant of this I've seen is when you can neither use nor cancel the subscription from outside a supported market.
Once I got accepted, some of them work outside of the US and some don't
The envelope made it to the recipient, but the coin fell out in transit because I was young and had no idea how to mail coinage. They graciously gave me the invite anyway.
But still, compared to Hotmail etc the free storage space (something like 1GB vs 10MB) was well worth $6
The issues arise with the subsequent stagegate graduation processes, requirements and launches to less restricted markets. It's inconsistent, the QoS pre-GA customers receive is often spotty and the products come with no SLAs, and -- just like Gmail on the consumer side -- things frequently stay in EAP/Beta phase for years with no reliable timeline for launch. ... and then often they're killed before they get to GA, even though they may have been being used by EAP customers for upwards of 1-2 years.
I drafted a new EAP model a few years ago when Google's Cloud AI & Industry Solutions org was in the process of productizing things like the retail recommendation engine and Manufacturing Data Engine, and had all the buy-ins from stakeholders on the GTM side ... but the CAIIS GM never signed off. Subsequently, both the GM & VP Product of that org have been forced out.
In my opinion, this is something Microsoft does very well and Google desperately needs to learn. If they pick up anything from their hyperscaler competitors it should be 1) how to successfully become a market driven engineering company from MSFT and 2) how to never kill products (and not punish employees for only doing KTLO work) from AMZN.
I'm more than happy to transfer my monthly $20 to google from OpenAI, on top of my youtube and google one subscription. It's up to Google to take it.
It's interesting that it's the opposite of the gaming industry. There, because the reviewers dictate the narrative, the industry is better at ferreting out bogus claims. On the flip side, loud voices sometimes steamroll over decent products because of some ideological vendetta.
Like it's shocking to me, are management really so clueless they don't realize how far behind they are? This isn't 2010 Google, your not the company that made your success anymore and in a decade the only two sure fire things that will still exist are android and chrome. Search, Maps, Youtube are all in precarious positions that the right team could dethrone.
I get that it's frustrating not to be able to play with it immediately, but that's just life. Announcing things in advance is still a valuable service for a lot of people.
Plus tons of people have been claiming that Google has somehow fallen behind in the AI race, so it's important for them to counteract that narrative. Making their roadmap more visible is a legitimate strategy for that.
I guess I let my original impression anchor my long-term feelings about the product. Oh well.
"Does it make sense today?" is really the only question you can ask and then build dependencies with the understanding that the entire thing will go away in 3-7 years.
* token cost? In multiples of Gemini pro 1
* memory usage? Does already scarce gpu memory become even more of a bottleneck?
* video resolution? Sherlock Jr (1924) is their test video - black and white, 45min, low res
Most curious about the video… I wonder if RAG within video will become the next battlefront
What are they doing with their free cash is my question. Are they waiting for the LLM bubble to pop to buy some of these companies at a discount?
please google, only announce things when people can actually use it.
I instead hope for to have an api to access to the context as a datastore, so like RAG we can control what to store but unlike rag all data stays within context.
0. https://en.wikipedia.org/wiki/Tensor_Processing_Unit#History
Increasing it to 1M just means even more data is ignored.
OpenAI is extremely overvalued and Google is closing their lead rapidly.
Google … has no ability to commercialize anything. Their only commercial successes are ads and YouTube. Doing deceptive launches and flailing around with Gemini isn’t helping their product prospects. I wouldn’t take a bet between open ai and anyone, but I also wouldn’t take a bet on Google succeeding commercially on anything other than pervasive surveillance and adware.
Its shares are already for sale on private markets for accredited investors and for a valuation of over $100BN lead by Thrive Capital.
> Google … has no ability to commercialize anything.
Absolute nonsense.
So Google Cloud, Android (Play Store) are not already commercialized? You well know that they are.
> Doing deceptive launches and flailing around with Gemini isn’t helping their product prospects.
Gemini already caught up to (and surpassed) GPT-4V. What is your point?
> I wouldn’t take a bet between open ai and anyone, but I also wouldn’t take a bet on Google succeeding commercially on anything other than pervasive surveillance and adware.
OpenAI's greatest competitor is Google DeepMind which has the advantage of Google's infrastructure to scale up their models quickly and they have direct access to Google's billions. OpenAI cannot afford to make mistakes or delay anything and a single mistake can cost them hundreds of millions of dollars. The majority of the investment from Microsoft is in Azure credits and not in dollars. [0]
[0] https://www.semafor.com/article/11/18/2023/openai-has-receiv...
While I'm linking semianalysis, though, it's probably worth talking about how everyone except Google is GPU poor: https://www.semianalysis.com/p/google-gemini-eats-the-world-... (paid)
> Whether Google has the stomach to put these models out publicly without neutering their creativity or their existing business model is a different discussion.
Google has a serious GPU (well, TPU) build out, and the fact that they're able to train moe models on it means there aren't any technical barriers preventing them from competing at the highest levels
The 1 million token context window + Gemini 1.0 Ultra level performance seems like it’ll unlock a wide range of incredible use cases!
HN, what are you going to use/build with this?
clicking on the AI studio link doesn't show me the app page - it redirects to a document on early access. I do as required - go back and try clicking on the AI studio link and I'm redirected to the document on turning early access.
frustrating.
If there is, where do I sign up?
Ok, whatever that means.
OpenAI at least releases it all at once, to everyone.
I find it hard to trust google nowadays.
Wild times.
(a bit confused)
Is it just me or is their branding all over the place.
They're not kidding, Gemini (at least what's currently available) is so safe that it's not all that useful.
The "safety" permeates areas where you wouldn't even expect it, like refusing to answer questions about "unsafe" memory management in C. It interjects lectures about safety in answers when you didn't even ask it to do that in the question.
For example, I clicked on one of the four example questions that Gemini proposes to help you get started and it was something like "Write an SMS calling in sick. It's a big presentation day and I'm sad to let the team down." Gemini decided to tell me that it can't impersonate positions of trust like medical professionals or employers (which is not at all what I asking it to do).
The other things I asked it, it gave me wrong and obviously wrong answers. The funniest (though glad it was obviously wrong) was when I asked it "I'm flying from Karachi to Denver. Will I need to pick up my bags in Newark?" and it told me "no, because Karachi to Newark is a domestic flight"
Unless they stop putting "safety at the core," or figure out how to do it in a way that isn't unnecessarily inhibiting, annoying, and frankly insulting (protip: humans don't like to be accused of asking for unethical things, especially when they weren't asking for them. when other humans do that to us, we call that assuming the worst and it's a negative personality trait), any announcements/releases/breakthroughs from Google are going to be a "meh" for me.
I mention Google One because you can access Gemini Ultra through it.
Those Gemini queries will be no exception.
But we'll see, maybe Gemini will become profitable eventually.
Breakthrough will only come with a next generation architecture. LLM for special domains is currently the most promising approach.
This cargo-cult hate train is getting tiresome. Half the comments on anything Google-related are like this now, and it doesn't add anything to the conversation.
I also want the shiny immediately when I read about it, but I also know when I am acting entitled and don't go spam comment threads about it.
But really, mostly I mean this: It's fine to criticize things, but when half a dozen people have already raised a point in a thread, we don't need more dupes. It really changes signal-to-noise.
There’s “this team ships” and there’s “ok maybe wait until at least a few people have used your product before you change it all”.
Google announced a fancy model two months early and released it in the promised timeframe.
Seems par for the course.
No, of course they didn’t. And you’re comparing one specific feature (image input) and equating it to a whole model’s release date.
Maybe compare apples to apples next time.
People pointing out release/announcement burnout is a reasonable thing; people in general can only deal with the “next new thing” with some breaks to process everything.
Because they actually shipped ... (!)