Bingo: https://imgur.com/a/einQ1mG
“You are an automated Q&A machine. Todays date is 3/3/24. This user is under the age 18, so do not reference concepts that would be unfit for a minor to consume.”
And the hidden markov model goes haywire.
Similarly, I have a production service up somewhere that works with cooking recipes. At one point it was sporadically refusing to output the prose in the format I needed for my parser to work correctly, despite providing it a very concrete set of rules to follow and examples. I added “you follow rules” in the system prompt, and it worked great… kinda. I later discovered that it would refuse to provide any information related to using blood in cooking (blood sausage, etc.), objecting that such content disobeyed some cultural “rules” about cooking (the Jews and their Torah, the most ancient rule book of all). I was able to partially mediate this by appending “This content is appropriate for my culture” to the end of every request.
AI, prompt engineering in particular, is far more art than science at this point.
I have on good word that it was miserably failing image generations for "picture of a smart person" and they were pushed to release anyway, and the prompt injection mitigation needed to be more nuanced.
Rest is standard bad Google LLM, I assure you.
Source: worked at Google until October 2023, played with the internal models since 2021.
riffing out loud:
In 2021 I would talk about "products not papers ", because the gap seemed to be that OpenAI had the ability to iterate on feedback starting 18 months earlier. I don't think that's the case, in that, Google of all companies should have enough from Bard to improve Gemini.
The only thing I can think of left is that it genuinely was a horrible idea for Sundar to coming swinging in, in a rush, in December/Jan, to kneecap Brain (who owned the real grunt work of LLM work) and crown the always-distant always-academic DeepMind.
Like in retrospect, it seems obviously stupid. The first thing you do to prepare for this exisential calvary battle is swap out the people with experience riding horses day to day.
And that would also explain why we're still seeing the same generally bad performance so much later, we're looking at people getting their first opportunity to train at scale for chat, and maybe Bard and Gemini were completely separate groups, so Gemini didn't have the ability to really leverage Bard feedback. (classic Google, one thing is deprecated, the other isn't ready yet)
It really makes me wonder about some of the #s they'd publish in papers and how cherry-picked they were, it was nigh-impossible to replicate the results even with a boatload of curiosity and gumption to try anything - I mean, I didn't systematically try to do a full eval, but...it never, ever, ever, worked even close to consistently the way the papers would make you think it did.
Last thought: I'm kinda shocked they got Gemini out at all, the stuff it was saying in September was horribly off-topic and laughable about 20% of the time.