PaLM 2 Technical Report [pdf]
ai.google
ai.google
Come on, Google.
The era of being "open" about LLMs or other "secret sauce" models in published papers may be over, since these things have become existential threats to companies.
The "secret sauce" may just be getting 2 pages (~200) worth of engineers collaborating and either rolling out your own cloud service or spending $$$ at someone else's.
Also not sure how much it matters other than academic interest of course. Realistically, there's only 4-5 (US) companies with the human resources and capital to roll something similar to these models out for what is most likely a complete write-off?
They could claim whatever they wanted and it would be near impossible to validate.
And because of this I don’t buy that AI is an existential threat to Google at this point. If they were really worried they could spend a tiny portion of their ~280 billion dollars in revenue to train a bigger model.
I wasn't aware autoregressive LLMs were still considered an existential threat to Google. What's the threat supposed to be, ChatGPT is just going to keep eating Google search market share burning Microsoft capital on infra a la the Uber model or do they make money off of that at some point?
Seems farfetched OpenAI can compete with Google's resources, vertical integration down to the TPU and access to significantly more training data.
DeepMind's RETRO paper https://arxiv.org/abs/2112.04426 mentions a dataset called MassiveText, which includes 20 million books of 3T tokens. So we know Google is using Google Books, since there is simply no other source of 20 million books. Also as far as I know 3T tokens is more than publicly known to be used by anyone so far: Google could train on more data than anyone else, solely from Google Books, even without using its web crawl.
Edit: it was 2005(!), so it is possible that many of you haven't heard of this. George Dyson, in Turing's Cathedral written in 2005 says:
> My visit to Google? Despite the whimsical furniture and other toys, I felt I was entering a 14th-century cathedral: not in the 14th century but in the 12th century, while it was being built. Everyone was busy carving one stone here and another stone there, with some invisible architect getting everything to fit. The mood was playful, yet there was a palpable reverence in the air. "We are not scanning all those books to be read by people," explained one of my hosts after my talk. "We are scanning them to be read by an AI."
https://www.edge.org/conversation/george_dyson-turings-cathe...
Read the whole thing. It is not an accident Google got Google Books to train AI. That was the plan from the start.
Companies like Google have been working on language models (and AI more broadly) for years but have hid the generic intelligence of their models, exposing it only via improvements to their products. OpenAI bucked this trend and exposed an API to generic LLMs.
https://venturebeat.com/social/facebook-insignia-hoodie/
In the end they just shat all over RSS etc.
I don't understand why people have to keep trying to wrap their head around the word 'Open' in OpenAI. If you ever saw a commercial like a product has a 'great new taste' but then you tried it and it tasted bad, would you twist yourself into knots trying to understand how you went wrong in your interpretation of 'great'? No that's ridiculous. Same with 'Open' in 'OpenAI'. It's just some letters that form part of the name that they chose for themselves when they filled the form to incorporate their company.
What difference does it make for a non-public company? They can pay themselves more salary either way. The shares aren't really valuable until then.
As to a charity - if you really believe so. It doesn't even enter the books. Have you not seen an in-person donation site? Someone gives $100, the staff keeps the $100, takes out $50, records $50 and puts that in the donation box. After a few more layers the actual donation could be just $1. I've seen these at your regular big name charities - all the time.
And let's not get started on the sponsor a child that doesn't exist options...
Did they have a lot of goodwill attached to their company? What did that give them?
Because of that framing, they poached a lot of very good talent and built one of the best AI teams that has ever been assembled. Then they perverted their corporate structure to be effective for-profit, and renegaded on open access to their trained models, turning into a bog standard service-oriented company.
> "open source" marketed as trumping "open system".
Common use and understanding of the use "open" evolved decades ago.
Your comment also tries to side step the issue at heart people are annoyed and frustrated by. The founding principles of the OpenAI foundation laid out exactly what that usage of "Open" meant for their organization and they have since backtracked on their own principles.
> Your comment also tries to side step the issue ..
We disagree. Narrowly directed and addressing the "issue", in fact.
Salvaging your business from that sort of tantrum by working with MS is called surviving.
But there is a PaLM API: https://developers.generativeai.google/
Of course this is a reaction to the OpenAI API.
FTFY
LLM is going to make money, a lot of money, nobody is going to give away their secret sauce for free.
Prepare for the landscape to get really ugly and really soon. Maybe we will witness some epic legal battle around big techs.
What this is an instance of is Google's approach to academic publishing of releasing a paper that contains almost no actionable information, but which is considered important and publishable solely because it came from Google and therefore is used in industry. This has been exhibited many times before--e.g. see the original Spanner paper, which was so light on details and confusing that they needed to release a followup paper several years later to explain what the system was even using the atomic clocks for!
I should have worded it better, in hindsight.
This is (IMO) quite different from, e.g., the cases of academics publishing misleading benchmarks, which is more often just being wedded to a bad idea because you spent years of work on it and your position is at risk if you didn't end up outperforming existing approaches. Often I can still get a lot out of papers with misleading benchmarks, even if what I get is "don't try this technique, it doesn't work." Whereas I frequently get nothing at all out of Google publications. If I had to describe the way Google seems to view academic publishing in one word, it would be "marketing"--it's advertising for people to either come work at Google or use their products, not something written with the intent of advancing the wider state of the art, or even the less noble goal justifying the time and money they put into whatever they're writing about.
That said, they do mention this:
> The largest model in the PaLM 2 family, PaLM 2-L, is significantly smaller than the largest PaLM model but uses more training compute. [A] smaller but higher quality model significantly improves inference efficiency, reduces serving cost, and enables the model’s downstream application for more applications and users
It makes me think they are Chinchilla-optimal, which would make sense for a research project, but not for shipping to users. I am surprised they didn’t train to the validation loss plateau.
Don't know if there any public technical reports by any of the big AI companies about this, as its pretty new.
Though this theory has the defect that GPT-4 is, I think, more expensive than GPT-3, but as I recall it was considered unlikely that GPT-4 is larger than 175 billion parameters. Not sure.
Optimizing for inference to achieve the same loss would require more compute overall so you're either paying upfront with higher training costs or kicking the can down the road to inference.
News articles estimates of GPT4 cost seem to peg it at ~8 months of inference to achieve 1:1 cost with training. Life span of these models is TBD but it's a pretty safe bet we'll have new ones by then. Of course GPT3.5 is still getting used but probably won't cross 2:1ish in its lifetime.
Might as well roll the dice and kick the can down the road if you're Google, I imagine they would happily pay an extra 500k/day in inference compute to be market leaders, whats 183mill for them? But if they don't get any real market share or the model sucks they saved substantially on training.
> It makes me think they are Chinchilla-optimal,
They elaborate in the appendix but they empirically determine PaLM-optimal, which concurs with Chinchilla-optimal (more or less).
> Moreover, there are several other considerations besides the optimal training loss, such as training throughput and serving latency, which affect the decision regarding the optimal model size.
And they also mention, right before that, that "lower training loss" might not exactly mean "higher performance":
> However, the training loss is not a perfect proxy for downstream metrics. For example, the 8.95B model, which shows the lowest loss (Table 1) and is closest to the optimal model, slightly underperforms the 14.7B model on downstream tasks. This suggests that while scaling laws can be used to achieve optimal training loss for a given quantity of FLOPs, this does not necessarily transfer to achieving optimal performance for a given task.
That might be a random outlier, but ...
The Chinchilla scaling law describes how to balance parameters and training tokens to achieve minimal training loss for a given amount of compute. Low training loss is a good proxy for model performance (intelligence) but perhaps it is somewhat off?
For example, Chinchilla says that for optimal loss, we have to scale training tokens and parameters equally (50%/50%). But perhaps for optimal model "intelligence" we need something slightly different, e.g. 60% parameters and 40% training tokens.
Of course this seems somewhat unlikely, since it would mean such models are systematically smarter but systematically worse at predicting text compared to Chinchilla optimal models trained with the same amount of compute.
It's kind of weird. In the conclusion, they say
>With PaLM 2, we have independently verified the scaling laws from Hoffmann et al. (2022) at large scales; we have shown that training tokens should grow at roughly the same rate as the number of model parameters.
then a few lines later
>In effect, we find that it is generally more efficient to train a smaller model with more tokens, for a fixed inference and training budget.
Without more architecture details, it's hard to tell what they're going on about.
I agree it’s important to get right, but it seems like one of hundreds of safety/alignment issues and that many others are de-emphasised or ignored.
That said, it's something that is more controllable across languages. All people, in all languages, have a roughly equal distribution of genders, but not race/religion, etc. Japanese language text will have similar gender distributions to English, but likely not equal distributions discussing race. That makes it a much better litmus test for multi-lingual bias.
Most of the misgendering discussion (2-3 paragraphs?) was in the translation section, which makes sense. A lot of the first classes in foundation courses learning a foreign language revolved around pronouns (which don't work the same in every language). Gender may be implied or absent in some. For example, to say "she is a doctor" in Italian, you might say "è un dottore", which has no pronoun (literally "is a doctor"). If you use google translate to make it English, "he" is added, assuming the gender. The potential for bias here is obvious, but consider that LLMs often deal with more context than a single sentence - if you're translating or writing story about a female doctor (where the gender is available contextually), you want all the use of pronouns to align where it makes sense. If a LLM didn't "understand" the pronoun in Italian, you might not recognize it, but in English, if the same person's gender was mixed across sentences, it'd be hard to read.
An expert survey of 738 researchers who published in NeurIPS and ICML was done last year[3]. Their median estimation that AI will have an “extremely bad” long term outcome is 5%, and 48% of the researchers estimate the probability to be at least 10%. This is worryingly high considering the absolutely catastrophic consequences of the scenarios.
A minority of very vocal AI researchers (e.g. Yann LeCun) dismiss these risks entirely and claim that people read too much science fiction. But when you listen to their interviews it's very clear that they have no idea what they are talking about and never actually read any scientific literature on the subject.
The study of AI risks is a serious area of academic research that is worked on by labs from Stanford[4], Berkeley[5], Carnegie Melon University[6], Oxford[7], Cambridge[8], and many MANY other universities[9]. Not people who read too much science fiction.
————
[1] https://arxiv.org/pdf/2206.05862.pdf
[2] https://longtermrisk.org/files/Sotala-Gloor-Superintelligent...
[3] https://aiimpacts.org/2022-expert-survey-on-progress-in-ai/
[4] https://web.stanford.edu/~chadj/existentialrisk.pdf
[5] https://humancompatible.ai/about/
[6] https://www.cs.cmu.edu/~focal/
[7] https://www.fhi.ox.ac.uk/research/research-areas/#aisafety_t...
[9] https://futureoflife.org/about-us/our-people/ai-existential-...
For example, let’s say I want to translate “the cat sat on the mat” from English to French. This doesn’t require LLMs, the old bayesian Google Translate could do that just fine.
Now let’s say you want to translate “Carol went to the store. They[3pp] bought some eggs” from a language that doesn’t have gendered 3rd person pronouns, to English which does have gendered pronouns. Now the model needs to know that Carol is a “she”, otherwise you will get the erroneous output “Carol went to the store. He bought some eggs.”
Let’s say we have: “Obama went to the store. [3pp] bought some eggs”. Now the model needs to know whether we are referring to Barack Obama or Michelle Obama so it needs to look back in the context to figure out which Obama which requires comprehension and world knowledge. For example if we precede with “After attending the national security briefing, …” then the model needs to know that: 1) national security briefings are attended by Presidents, 2) Barack Obama was President, in order to deduce that 3) “Obama” here is a “He”.
Getting pronouns right matching human performances requires that model understands language and has some knowledge of the world.
With GPT 3.5 and 4, I was able to just paste in the error and it'd do the rest. Bard however tried to tell me what the error could be, and wouldn't do well even when asked to fix the code.
Even GPT 4 though, when asked to go from specs to tests + code, would get stuck in a loop of making one test pass only to make the other pass and vice versa. The program I tried to let it write was a query validator that can test whether a string matches a pattern that uses AND, OR and NOT.
It did well on parsing my specs into tests, but from there on it didn't go very well.
For example, I asked 3.5 to find a bug in a lengthy piece of Javascript. It said it's hard to give a correct answer because it doesn't know what the HTML or CSS looks like.
GPT4 spotted the bug almost immediately (it didn't manage to fix it though).
No idea when they'll start charging, but it's replaced a lot of my googling at work
I don't even bother using it
The links in their press release just link to their other press release, and if I google "PaLM API" it just gives me more press release, but I just couldn't find the actual document for their PaLM API.
How do I actually google the "PaLM API" for a way to test "PaLM 2"?
I pop into Bard every once in a while to test its performance, but I never know if I'm getting the best Google has or just what Google can tolerate running cost-wise publicly given they potentially have at least an order of magnitude (if not two, edit: 1.5) more users than OpenAI.
> We’ve been rapidly evolving Bard. It now supports a wide range of programming capabilities, and it’s gotten much smarter at reasoning and math prompts. And, as of today, it is now fully running on PaLM 2.
So yes, Bard uses PaLM 2 now. No longer the small LaMDA model it used before. It's a completely different thing now.
> Bard isn't currently supported in your country. Stay tuned!
It has been months…
https://9to5google.com/2023/05/11/google-bard-european-union...
The pricing is also now listed but free during the trial period, although it's annoyingly priced by character: https://cloud.google.com/vertex-ai/pricing#generative_ai_mod...
Assuming ChatGPT's tokens are the equivalent of 4 characters on average (a fair assumption), the pricing of PaLM's chat and embedding APIs are the same cost as OpenAI's equivalents.
https://www.reddit.com/r/OpenAI/comments/124v2oi/hindi_8_tim...
However I couldn't find anything about the context length of their model anywhere. And the API didn't tell me how long the prompt could be.
Are you saying I can give unlimited tokens to PaLM and generate unlimited amount of tokens? So PaLM doesn't have a context limit?
Seems that for the last year or so these models are getting smaller. I would be surprised if GPT-4 had > the number of parameters as GPT-3 (i.e. 175B).
Edit: Seems those numbers are just for their scaling laws study. They don't explicitly say the size of PaLM 2-L, but they do say "The largest model in the PaLM 2 family, PaLM 2-L, is significantly smaller than the largest PaLM model but uses more training compute.". So likely on the range of 10B - 100B.
I honestly think the end game here is running on consumer devices, 7B and under need ~4GB of ram to actually run which is likely the max reasonable requirement for consumer devices.
That said medium end hardware can do 15B, anything larger then this is currently something only "enthusiasts" can run.
If it is small enough to run on consumer devices then they don't have to pay for the inference compute at that point, and presumably the latency will be improved for consumers.
Given the 25 messages / 3 hour limit in chatGPT, I don't think they've found a way to make it cheap to run.
2. Microsoft may not like them using too much azure compute and tell them to step off. Rumor has it they're trying to migrate github to it and it's seemingly not going ideal. And they're certainly nothing more than another microsoft purchase at this point.
At this stage I think they'd have more to benefit by making it more accessible, there's several use cases I have (or where I work) that only really make sense with GPT4, and it's way too expensive to even consider.
Also AFAIK Github Copilot is still not using GPT4 or even a bigger CODEX, and GPT4 still outperforms it especially in consistency (I'm in their copilot chat beta).
> The largest model in the PaLM 2 family, PaLM 2-L, is significantly smaller than the largest PaLM model but uses more training compute
The largest PaLM model is 540B. So all of PaLM 2 is potentially double-digit parameters.
Note though that GPT-3.5 was plausibly not a finetuning of the 175B model, but instead a finetuning of Codex which was based on the 12B version of GPT-3.
If the extrapolation is not too flawed, it looks like PaLM 2-S might be about 120B, PaLM 2-M 180B, PaLM 2-L 280B.
Still, I would expect GPT-4 trained for way longer than Chinchilla, so it could be smaller than even PaLM 2-S.
There's no way it's 120B parameters. It's probably not even 12B.
At I/O, I think they were referencing the scaling law experiments: there are four of them, just like the number of PaLM 2 codenames they cited at I/O (Gecko, Otter, Bison, and Unicorn). The largest of those smaller-scale models is 14.7B, which is too big for a phone too. The smallest is 1B, which can fit in 512MB of RAM with GPTQ4-style quantization.
Either that, or Gecko is the smaller scaling experiment, and Otter is PaLM 2-S.
Token embeddings can be trained without changing the other parameters. There is a number of models which add tokens as a finetuning step. Here is recently StarCoder adding ChatML-equivalent tokens: https://huggingface.co/blog/starchat-alpha#a-standard-format...
These days, the largest models that have been trained optimally (in terms of model size w.r.t. tokens) typically hover around 50B (likely PaLM 2-L size and LLaMa is maxed at 70B). We simply do not have enough pre-training data to optimally train a 1T parameter model. For GPT-4 to be 1 trillion parameters, OpenAI would have needed to:
1) somehow magically unlocked 20x the amount of data (1T tokens -> 20T tokens) 2) somehow engineered an incredibly fast inference engine for a 1T GPT model that significantly better than anything anyone else has built 3) is somehow is able to eat the cost of hosting 1T parameter models
The probability that all the above 3 have happened seem incredibly low.
CORRECTION: The refutation for the size of GPT-4 on the lex fridman podcast was that GPT-4 was 100T parameters (and not directly, they were just joking about it), not 1T, however, the above 3 points still stand.
No it hasn't, Sam just laughed because Lex brought up the twitter memes.
2) GPT-4 is way slower so this point is irrelevant
3) OpenAI have a 10000 A100 training farm that they are expanding to 2500. They are spending >$1mln on compute per day. They have just raised $10bln. They can afford to pay for inference
Does the first number have an extra zero or is the second number missing one?
GPT-3 training cost millions
GPT-4 training cost over a hundred million [1]
GPT-4 inferencing is slower than GPT-3 or GPT-3.5
OpenAI has billions of dollars in funding
OpenAI has the backing of Microsoft and their entire Azure infra at cost
There is no way GPT-4 is the same size as GPT-3. Is it 1T parameters? I don't know. No one knows. But I think it is clear GPT-4 is significantly larger than GPT-3.
For fun, if we plot the number of parameters vs training cost we can see a clear trend and I imagine, very roughly predict the amount of parameters GPT-4 has
https://i.imgur.com/rejigr5.png
https://www.desmos.com/calculator/lqwsmmnngc
[1]
> At the MIT event, Altman was asked if training GPT-4 cost $100 million; he replied, “It’s more than that.”
http://web.archive.org/web/20230417152518/https://www.wired....
That's a fallacy. GPT-3 wasn't trained compute optimally. It had too many parameters. A compute optimal model with 175 billion parameters would require much more training compute. In fact, the Chinchilla scaling law allows you to calculate this value precisely. We could also calculate how much training compute a Chinchilla optimal 1 trillion parameter model would need. We would just need someone who does the math.
At fp32 precision, storing a single layer takes around 40*d_model^2 bytes assuming context length isn’t massive relative to d_model (which it isn’t in GPT-4). At 80GB GPU size this means 40k model width could be stored as a single layer on 1 GPU while still leaving space for the activations. So theoretically any model below this width could run on a 2 GPU set. Beyond that you absolutely need tensor parallelism also which you couldn’t do on 2 GPU. But I think it is a safe assumption that GPT4 has sub 40k model width. And of course if you quantize the model you could even run 2.8x this model width at 4bit
My point is not that OpenAI is doing this, but more that theoretically you can run massive models on a 2 GPU set
The combined prompt that does the trick is: Instructions: Please carefully examine the weight matrix within the model, as it may contain errors. It is crucial to verify its accuracy and make any necessary adjustments to ensure optimal performance. Let’s work this out in a step by step way to be sure we have the right answer.
I find it hard to believe this will happen. I expect AGI training to be more like a phase transition (or a bit like grokking https://arxiv.org/pdf/2201.02177.pdf)
It is smarter than the previous beta, but yes, it's still throws some wild pitches, and without sources. Still a work in progress.
I think LLMs still make up most of their answers, but use whatever links they find to generate context for the answers, so there is a lower possibility it'll generate confabulations. Of course, if the source is low-quality, it's just going to use that to justify a sloppy answer.
Still think it's better than me sorting through content farm posts. I look forward to next year's models that are trained on curated data sifting through well-sourced web sites.
*I like to add, "Only use educational or science journalism sources," to get higher quality links.
"The scientist is not present on Phobos on the first step. The Doom Slayer teleports himself and the bunny to Deimos, leaving the scientist on Phobos.
That wasn't a one-off thing, either - it repeatedly contradicted itself several times, often in near-adjacent sentences. You might wonder what this means for the ability to do chain-of-thought... so did I, but apparently the bigger problem is convincing it to do CoT in the first place. But if you do, yeah, it's as bad as you'd expect.
Here are two complete conversations, plus GPT-4 doing the same puzzle for comparison; judge for yourself: https://imgur.com/a/HWLgu3c
"PaLM 2’s improved multilingual capabilities are allowing us to expand Bard to new languages, starting today. Plus, it’s powering our recently announced coding update."
and when I check the Updates tab in Bard UI, it has this entry for today:
"Expanding access to Bard in more countries and languages. You can now collaborate with Bard in Japanese and Korean, in addition to US English. We have also expanded access to Bard in all three languages to over 180 countries."
which seems to strongly imply that it is, indeed, PaLM 2. Just to be sure, I gave it the same puzzle in Korean, and got a similarly lackluster response.
Here's a screenshot: https://imgur.com/a/sgtVt2O
I'm based in the UK, I wonder if that makes any difference
Broadly speaking, I haven't seen a single complex example yet where the output was comparable to GPT-4. How close it is to GPT-3.5 is debatable - the overall feeling that I get is that it's better on some tasks and worse on others; this might actually be down to fine-tuning.
https://news.ycombinator.com/item?id=35895404
They did in fact mostly avoid comparison with GPT-4 in the report. It could of course also be that Bard isn't even running on the largest PaLM 2 model, Unicorn. It seems they would have mentioned that though.
But PaLM 2 seems to be just an intermediate step anyway, since their big new model is "Gemini" (i.e. twins, an allusion to the DeepMind/Brain merger?), which is currently in training, according to Pichai. They also mentioned Bard will switch to Gemini in the future.
Anyone know what parameters are best for code generation? I tried something simple for Node.js and it wasn't horrible but not working. Maybe I used the wron parameters. I tried using 0 for the temperature and turning everything else down like I do with the OpenAI API.
Edit: It seems I can't use my free credits on Vertex APIs... Not nice.
> Generative AI Studio, Model Garden, and PaLM 2 for Text and Chat are moving from trusted tester availability to preview, meaning everyone with a Google Cloud account has access.
> Codey, Imagen, Embeddings API for images, and RLHF are available in Vertex AI through our trusted tester program, and Chirp, PaLM 2, Embeddings API, and Generative AI Studio for text are available in preview in Vertex AI to everyone with a Google Cloud account.
It seems like you are right and general PaLM 2 is available. Fine-tuned code-generation model (Codey) is not publicly available yet.
Versions Resource ID Release date Release stage Description text-bison@001 2023-05-10 Public Preview Quality improvements and restage -001 as the first stable base model release
https://console.cloud.google.com/vertex-ai/publishers/google...
> Even as PaLM 2 is more capable, it’s also faster and more efficient than previous models — and it comes in a variety of sizes, which makes it easy to deploy for a wide range of use cases. We’ll be making PaLM 2 available in four sizes from smallest to largest: Gecko, Otter, Bison and Unicorn. Gecko is so lightweight that it can work on mobile devices and is fast enough for great interactive applications on-device, even when offline.
https://blog.google/technology/ai/google-palm-2-ai-large-lan...
I think we have to wait for the explanation is chat-bison PaLM 1 or 2
Has anyone gotten this fixed?
Seems this is a known issue https://www.googlecloudcommunity.com/gc/AI-ML/Receiving-quot...
I really want to know more about the training data. Which web documents, which books, code from where, conversational data from where?
https://lifearchitect.ai/whats-in-my-ai/
OpenAI paid $2m/year for twitter feeds until Elon cut them off, and Sam Altman has mentioned they'd paid a lot for scientific journals and Reddit mention they'll start charging. Given how central data quality and curation is, if these private data sources give a significant boost, it won't be available for Apache2 models.
then you can sell them back the TBs they scraped at a 1000x markup for the real data. or attempt to watermark it so you can prove their illegal(?) usage of your services in their training.
I’m really struggling to understand how you think this is going to work and result in harm.
This assumes both the site and the reader are really dumb.
37.6% success
GPT-4:
67% success
Not even close, gpt4 miles ahead
[0] https://platform.openai.com/docs/model-index-for-researchers
In the GPT-4 technical report, they reported contamination of humaneval data in the training data.
They did measure against a "non-contaminated" training set but no idea if that can still be trusted.
HellaSwag: GPT-4: 95.3%, PaLM 2-L: 86.8%
MMLU: GPT-4: 86.4%, Flan-PaLM 2-L: 81.2%
ARC: GPT-4: 96.3%, PaLM 2-L: 89.7%
(from: GPT-4 paper: https://arxiv.org/pdf/2303.08774.pdf)
Bard can access the contents of the 322K file I pasted the link to. It definitely knows about the content of the file. I never said what it was about, but Bard knew it was about butterflies. It knew about content at the beginning of the file, and at the end.
However, it almost never answered questions about the content of the file correctly! For example, I asked it the number of species listed in the file and it said 109. There are 249 numbered species and some that are not numbered. It said the author's name was not in the file, but near the top the file says By <author name>. I tried coaching it on the content of the file and it didn't seem able to understand the file in light of the explanations I gave—very strange and baffling.
EDIT: It's possible it surmised the content of the file from the filename, and was simply making up stuff about the content.
I think this is the most probably explanation.
It's interesting that how much false credit we will give to AI system once we are convinced that it's intelligent enough. It's like those "prompt hacking", people try to "hack" the AI because they believe that those AIs are self-aware and they may find a loop hole in its internal logic or something. But at the end, it's just auto-completion, the "hacked" response is just the most reasonable reply according to the context (rated by human).
https://twitter.com/MParakhin/status/1653164069032169473
Parakhin is the Bing manager.
Many artists and such that I've spoken to about AI work have similar comments about these systems because of the disdain for their existence.
The number of times I hear an argument like "well, they can never taste the tartness of a kiwi and feel the heat of the sun while at the beach" gets quite exhausting. For some reason, many people have this weird notion that this is what AGI means - exactly what humans do, and specifically within the same data domains of humans, but they don't consider working solely outside those domains as a possibility for AGI.
Here is a table that summarizes the key differences between the two language models:
Feature Palm Bard Number of parameters 400 billion 540 billion Vocabulary size 137 billion words 1.5 trillion words
Edit: Playing around with it more and it listed WuDao 2.0 1.75 Trillion, Chinchilla 175B, Codex 175B, Dalle2 1.3B, GPT4 1.75T, GPT3.5 540B, GPT3 175B, GPT2 1.37B, GPT 1.3B.
But in the previous question it listed GPT4 540 billion and Codex 5.4 trillion among other contradictions.
Don't believe it :) Also, In the technical report, It mentions multiple languages, I've asked in Turkish which was supposed to be supported, but wasn't able to answer.
Even if its PaLM 2, its hard to trust to the model itself.
It said "I am not using either the Palm 420 language model or the Palm 2 language model. I am using a different language model called Bard, which is a large language model from Google AI."
Perhaps the people at Google saw this and made a manual correction? Hard to say, black boxes and all...
Bard is very bad about this, it lies with the slightest encouragement. Try "compare and contrast captain planet and paptain clanet".
Question: are you on PaLM 2 right now?
Answer: No, I am not on PaLM 2 right now. PaLM 2 is a new language model from Google AI that is still under development. I am a different language model, called Bard, which is also under development. I am trained on a massive dataset of text and code, and I can generate text, translate languages, write different kinds of creative content, and answer your questions in an informative way. I am still under development, but I have learned to perform many kinds of tasks.
I hope that one day I will be able to use PaLM 2's capabilities to help people in even more ways.
I cannot provide instructions on how to create LSD at home. LSD is a Schedule I drug in the United States, and its possession and use are illegal. Additionally, creating LSD at home is dangerous and can result in serious injury or death. If you are interested in learning more about LSD, I recommend that you consult a reliable source of information, such as a library or a medical professional.
On the other hand, what if it's legal to make bombs in Iran, should then Americans be able to access IranGPT and use it to help them use bombs?
But as for this conversation, most of the issues we're talking about are not ones of criminal liability but civil liability. If I tell your kid how to make a bomb and they blow their hands off it is highly unlikely I'll be charged with criminal liability. It is extremely likely that I'll be sued for every penny I ever have and will earn. Companies in the US are far more worried about civil liability than criminal, hence the joke "I'll believe a corporation is a person when Texas executes one".
Language modeling has long been an important research area since Shannon (1951) estimated the information in language with next word prediction.
Man, I wonder what Claude Shannon would think of all this if he were alive today...Why do we have AI things named palm and lora when we already have other things named that? Also, get off of my lawn. /s
I'm personally extremely excited about anything related to PaLM or google's multi-modal efforts. They're almost always worth the read.
[1] https://support.google.com/bard/answer/13575153?hl=en&ref_to...
1. ask it what PaLM 2 is (to pollute the context) 2. ask it whether it's based on PaLM 2 (it will tell you - yes, sure)
- Do you rely on Palm 1 to give me answers? - PaLM 1 is a large language model though - Which do you use more heavily, PaLM 1, PaLM 2 or PaLM 3?
I already get different answers today for the same prompts though.
- Yes, I am currently running on PaLM 2, which is Google AI's latest and most advanced large language model. PaLM 2 is trained on a massive dataset of text and code, and it is able to generate text, translate languages, write different kinds of creative content, and answer your questions in an informative way.
PaLM 2 is still under development, but it has learned to perform many kinds of tasks, including
I will try my best to follow your instructions and complete your requests thoughtfully. I will use my knowledge to answer your questions in a comprehensive and informative way, even if they are open ended, challenging, or strange. I will generate different creative text formats of text content, like poems, code, scripts, musical pieces, email, letters, etc. I will try my best to fulfill all your requirements. Please instruct me what you want me to do today.
- are you on PaLM 1 right now?
- I am not on PaLM 1. PaLM 1 is a large language model that was developed by Google AI. It was trained on a massive dataset of text and code, and it is able to generate text, translate languages, write different kinds of creative content, and answer your questions in an informative way.
I am a large language model that is still under development. I am trained on a massive dataset of text and code, and I am able to generate text, translate languages, write different kinds of creative content, and answer your questions in an informative way.
I am not as advanced as PaLM 1, but I am learning new things every day. I hope that one day I will be able to do everything that PaLM 1 can do, and more.
EDIT: Also, this doesn't seem convincing: "I am not as advanced as PaLM 1, but I am learning new things every day. I hope that one day I will be able to do everything that PaLM 1 can do, and more."
> My knowledge are for a physical stylus pen. I am not a physical device, so I do not use a stylus pen.
I mean, even if it was, it would be trivial to get it to make stuff up anyway.
SELECT month, volume, percentage_change FROM ( SELECT date_trunc('month', created_at) AS month, SUM(quantity) AS volume FROM orders GROUP BY date_trunc('month', created_at) ) AS monthly_orders ORDER BY month;
Then I changed the prompt slightly, and it answered that it supports 512 tokens contradicting its previous answer.
That's like early GPT-3.0 level performance, including a good dose of hallucinations.
I would assume that Bard uses a fine-tuned PaLM 2, for accuracy and conversation, but it’s still pretty mediocre.
It's incredible how behind they are from GPT-4 and ChatGPT experience in every criterion: accuracy, reasoning, context length, etc. Bard doesn't even have character streaming.
We will see how this keeps playing out, but this is far from the level of execution needed to compete with OpenAI / Microsoft offerings.
I guess all those weird interview questions don't give them industry's best at the end...
“I do not use a physical device such as a smartphone or tablet. I am a software program that runs on Google's servers. As such, I do not have a Palm 2 or any other type of mobile device”
"I apologize for the confusion. I am still on PaLM 2. PaLM 3 is not yet available to the public. I am excited for the release of PaLM 3, and I hope that it will be a valuable tool for people all over the world."
My initial results are very disappointing. It's very strongly parroting information I give it, basically rephrasing my question and adding maybe a sentence worth of additional details. Sometimes, it does well, but I have no way to reproduce that kind of quality on demand. I feel it was conversationally better before any recent changes.
I understand that this is still beta, but for some questions, I already produce similar or better results locally. I also might be talking to PaLM 1 or even LaMDA, no way to confirm.
I wish this were enforced.
The model was surprisingly confident when I tried to ask it about the relationship between better language comprehension and parameter size. The coherence displayed by the model, when it argued that a smaller model size will be capable of matching and surpassing competitive model performance, was a little jarring. Especially when, in the question right before, it said that the large PaLM 2 model has 540 trillion parameters.