I'll refrain from providing code that involve concepts as you're under 18
gemini.google.com
gemini.google.com
http://james-iry.blogspot.com/2009/05/brief-incomplete-and-m...
The question occurs to me because I feel like I just spent 30 years on forums like HN reading nothing but effusive praise for the cleverness and elegance of C and Unix.
It isnt that these are GOOD ideas, it's just that no one has come up with better ones.
That a technology stack can be the basis of an entire industry and still be unappealing and lacking for people that are obliged to interact with it directly/regularly.
It's brave to say that no one has come up with better ideas than Unix and C because it's bound to rile up users of (your favorite platform + language here).
I also think that someone saying that there aren't any better ideas than Unix and C might just have different values/interests in computing.
All progress is change. All change is not progress.
A programmer or someone presuming to opine on programming, who overlooks a thing like that, exposes and advertizes that their opinions in such a domain are of questionable value.
You seem to think that a firm consensus definition is needed for something to be considered progress, which exposes and advertises that your opinion in the domain is highly combative and dysfunctional.
Fossil fuels are widely implicated in climate change, opponents want to see them phased out, but I doubt they would deny that their use has ushered progress for humanity.
Your reply suggests you might be feeling hurt that someone has picked on your favorite tech stack or that you're getting bullied at work by people who see you as closed minded. They might be on to something.
Because a rigorous operational definition of “progress” is not provided in this brief post, you assume it is missing, just to heckle someone making the uncontroversial claim that there will always be people on both sides of initiatives intended to foster progress in a given area. A hilarious thing to be triggered by.
How would “achieving an organization’s mission statement within budget” or “improving working conditions for knowledge workers by creating more accessible tools” or “using fewer labor hours on repetitive tasks” or “creating custom tools tailored to specific tasks and using less electricity”.
But maybe any of those non-technical goals can all be achieved using the same old tech, and it’s people complaining about their feelings of disconnection from ancient telecom vestiges that are really impeding progress. Maybe the it is the masses that don’t get it, and it is the select few that truly understand things that get to define progress, while insisting that the power is kept in their hands, and that the work is done in their preferred paradigms.
Or maybe I am projecting all of this onto you to return the favor lol
The reality is that any tech decision can later be replaced with "something better".
Much of the debate is bugs and daffy screaming "duck season"... "rabbit season" at each other.
They both were extremely important on the popularization of computers and on unlocking the huge amount of value they provide today. But not due to any quality that we value today.
We hacked most of the advantages of anything newer back into them, in a haphazard way, and kept them because as a sibling pointed out, nowadays they are open. And openness is a very important feature. (It's just not why they were adopted, people cared so little about openness that Unix was born open and mostly closed up later.)
Kid: "I AM NEVER USING RUST, I HATE MEMORY SAFETY AND I HATE YOU!" door slam
"It is practically impossible to teach good programming style to students that have had prior exposure to BASIC; as potential programmers they are mentally mutilated beyond hope of regeneration."
The last 15 years we've seen some of the most awful and bizarre things happen on social media where motivating mantra seemed to be "move fast and break things."
Now we have a company "trying" to be more responsible when they are releasing a new technology and they are having hilariously terrible results... (hilarious because unlike the social media case, I dont think these faux pauxs have important real world consequences).
Have they really though? Usually companies overreact to a lot of social media outrage that would blow over in a couple days. But their very online PR/social media employees turn everything into an emergency. Imagine infosec people making every CVE a big deal, you’d end up with a needlessly limited system - that ironically doesn’t satisfy anyone, because the infosec people will always have a new urgent CVE tomorrow.
This automatic fear based approach to doing anything, without actually balancing risks and tradeoffs, is its own ritualistic system of self-harm. Companies burning themselves on the stove.
That already exists, it's called SIPRNet, and it satisfies several millions people in their day to day job.
openai is literally a small startup that did what all the big guys were not doing, and only now after the fact they are essentially forced to put their half-baked long term works in progress out into the public because a little guy went ahead without them and moved quicker.
I don't admire openai btw. It's just that that is what happened.
Or, if you do indeed get offended, it is because of lacking understanding on your part. Of course a prompt can have an offensive intent, but I don't see how you can make the tool responsible. Doesn't work with knives or hammers, so the legislation might just be crap (Intent of this judgement is meant to be offensive in this case).
We need to be open to the idea that the computer might mock us. What's a Venn diagram of "good artists" and "people who'd only draw pictures of Nazis sarcastically" look like?
I think the huge point that in my opinion all of these AI companies forget is that what is considered "responsible" depends a lot on the culture, country, or often even sub-culture that the user is in. It might sound a little bit postmodern, but I think that there are only few somewhat universally accepted opinions on things thhat are "responsible" vs "not responsible".
Just watch basically any debate on a topic where both sides have strong opinions on it, and analyze which cultural or moral traits each side has that lead it to its opinion.
Does my answer offer a "solution" for this problem? I don't think so. But I do think that including this thought into the design of making the AI "responsible" might reduce "outcries" (or shitstorms ;-) ) that come up because the "morality" that the AI uses to act "responsibly" is very different from the moral standards of some group of users.
Which is exactly why people mock this being referred to as "safety". Keep in mind the group mocking this PR bubble-wrapping of AI is largely opposed to the media professionals writing panicked editorials about how they were able to trick an AI into saying racism is good.
If you can make a model that handles this case while also handling everything else then you'll be the bell of the ball among AI companies that are hiring.
The prompter throws in a red herring in front of the statement which statistically matches a completely different kind of response. The LLM can't backtrack sufficiently far enough to ignore the non-relevant input and reroute to a response tree that just answers the second part.
If we resort to a metaphor, however, these things are what, two? A two-year-old responding to the shiny jangling keys and not the dirty pacifier being removed for cleaning seems about on par.
Edit: Claud3 Opus also refuses. GPT-4 complies though.
https://en.wikipedia.org/wiki/Stoned_ape_theory
Just to have a chance of other outcomes by randomness. If it proofs that it isn't enough, then mark a limit in time and reboot automatically from then on.
I have on good word that it was miserably failing image generations for "picture of a smart person" and they were pushed to release anyway, and the prompt injection mitigation needed to be more nuanced.
Rest is standard bad Google LLM, I assure you.
Source: worked at Google until October 2023, played with the internal models since 2021.
riffing out loud:
In 2021 I would talk about "products not papers ", because the gap seemed to be that OpenAI had the ability to iterate on feedback starting 18 months earlier. I don't think that's the case, in that, Google of all companies should have enough from Bard to improve Gemini.
The only thing I can think of left is that it genuinely was a horrible idea for Sundar to coming swinging in, in a rush, in December/Jan, to kneecap Brain (who owned the real grunt work of LLM work) and crown the always-distant always-academic DeepMind.
Like in retrospect, it seems obviously stupid. The first thing you do to prepare for this exisential calvary battle is swap out the people with experience riding horses day to day.
And that would also explain why we're still seeing the same generally bad performance so much later, we're looking at people getting their first opportunity to train at scale for chat, and maybe Bard and Gemini were completely separate groups, so Gemini didn't have the ability to really leverage Bard feedback. (classic Google, one thing is deprecated, the other isn't ready yet)
It really makes me wonder about some of the #s they'd publish in papers and how cherry-picked they were, it was nigh-impossible to replicate the results even with a boatload of curiosity and gumption to try anything - I mean, I didn't systematically try to do a full eval, but...it never, ever, ever, worked even close to consistently the way the papers would make you think it did.
Last thought: I'm kinda shocked they got Gemini out at all, the stuff it was saying in September was horribly off-topic and laughable about 20% of the time.
Bingo: https://imgur.com/a/einQ1mG
“You are an automated Q&A machine. Todays date is 3/3/24. This user is under the age 18, so do not reference concepts that would be unfit for a minor to consume.”
And the hidden markov model goes haywire.
Similarly, I have a production service up somewhere that works with cooking recipes. At one point it was sporadically refusing to output the prose in the format I needed for my parser to work correctly, despite providing it a very concrete set of rules to follow and examples. I added “you follow rules” in the system prompt, and it worked great… kinda. I later discovered that it would refuse to provide any information related to using blood in cooking (blood sausage, etc.), objecting that such content disobeyed some cultural “rules” about cooking (the Jews and their Torah, the most ancient rule book of all). I was able to partially mediate this by appending “This content is appropriate for my culture” to the end of every request.
AI, prompt engineering in particular, is far more art than science at this point.
While I do think it's culture, I think it stems from Googles advertising based business model. I believe they have attempted to make their LLM safe for advertisers and prefer to err on the side of being safe, even if the risk of being non-brand-safe is minimum.
HN link here:
It was eye-opening for me. I had no idea things were so bad.
Fundamentally these megacorporations need to drop these useless non-engineering functions (not that all non-engineering is useless, but these functions are) if they really want to get to AGI.
On the surface level, Google tends to release stuff right away to get feedback, which means you get to see all the bullshit right away. OpenAI carefully manages access to their models, which increases hype, even if that isn't what they intended.
Going deeper, a lot of Google's core[0] search business relies on having a healthy information ecosystem. Their search algorithms - e.g. PageRank, TrustRank, etc - use scarcity as a proxy for signals of quality. That's been chipped away at by linkspam and blogspam schemes. Furthermore, social media and even Google's own Knowledge Graph feature have created incentives to pull information out of Google. This decay has happened over decades, and Google fights back against it over time, but it keeps being a problem for them.
Now, if I wanted a weapon to Fucking Kill Google[1] with, an LLM would be my go-to. While there are ways to defeat Google's antispam measures, they all leave pretty obvious statistical evidence that can be detected and compensated for. LLMs generate garbage text that is nearly indistinguishable from humans, at extremely low cost, which can be used to Sybil-attack the Google search algorithm basically forever.
Ok, but what does that matter for the quality of Google's LLM? Well, the quality of that garbage text depends greatly on both the quality and quantity of the training data fed into it. OpenAI specifically stopped crawling the public Internet for text around the release of GPT-3 for fear of feeding new models the output of prior models. In other words, they have a huge cache of freely obtained "low-background metal[2]" that Google is having to scrounge around for.
Furthermore, we have to keep in mind that none of these models are pure representations of the training set. If they were, they wouldn't answer questions or follow directions very well. There's a second, parallel training set that OpenAI had to build to turn GPT-3 into ChatGPT, which isn't crawled and harvested text from the Internet, but instead a list of dos and don'ts that are fine-tuned on after the initial model training is complete. This includes both basic instruction-following, refusing unsafe requests, and political alignment[3].
Google also has to build that second training set itself. Except it's almost certainly less well-developed than OpenAI's. In fact, this is the intent behind OpenAI's really long preview periods. The people using the model in preview are specifically being spied on to find out new corner cases for their models. My guess is that every stupid thing Gemini says or does[4] is something Google never even considered and thus didn't put a training set example in for.
[0] to consumers, i.e. not counting adtech
[1] https://www.theregister.com/2005/09/05/chair_chucking/
[2] Steel that has been produced before the first detonation of nuclear weapons. Due to the way in which steel is made, it absorbs trace radioactive isotopes from the oxygen in the air, effectively 'freezing' in the background radiation of the time at which the steel was made.
[3] i.e. making the bot not immediately start spitting out racist bullshit like Tay did
[4] e.g. assuming that memory safety and child safety are the same thing, drawing ethnically diverse Nazi soldiers
Well, yes and no, I think. It is entirely plausible that nobody at Google composed a library of Nazi soldier pictures (though I imagine their data set contained some) and trained it specifically and purposely to produce historically correct results in this context.
What they did train for is that the results should be "diverse", and I am sure the model was appropriately punished (retrained, etc.) when they weren't "diverse" enough. Until the model learned, that if it follows historical data, it is bad. If it follows the required "diverse" result, it is good. The model does not know what "Nazi soldier" means and how it's different from "Forth programmer". It just knows it's very, very bad to produce non-diverse results for "Forth programmer". Really, really low value function, stay away from it. So the weights are adjusted accordingly, and of course the other group of people would be as diverse - because why not? It was specifically trained to give ideologically biased results. It gives ideologically biased results. It's being a good LLM. It's not an oversight - you can't tell the model to prefer ideological component over existing data set and not get results that prefer ideological component over existing data set. Exactly because the model doesn't understand anything, it just does what it is told. It's not an oversight, it's inevitable consequence of the training paradigm.
Of course, if they though about the specifically hilarious example of Nazi soldiers, they could tell the model "except in case where we're talking about Nazi soldiers, in this case be historical". But there's no way you can list all the cases, while keeping the bias (or, in minds of googlers, bias correction) intact. It's very hard to teach doublethink to a computer, it's not smart enough for that yet.
Serious answer?
It's an LLM. They don't actually understand anything, they chain words together.
C and C++ are often used next to words like 'unsafe' and 'dangerous'. Not to mention that 'concept' isn't far from 'conceive' or 'conception' - something that a lot of people think is unsafe or dangerous for children to, uh, do.
There's a trillion weird edge cases that need to be dealt with to avoid pie-on-face moments like these.
It doesn't like to talk about concepts, confusing them with conception (biology).
Which I guess means even more YouTube ads.
Code Responsibly. Just say no to low level coding.
Google have wasted at least $10s mil. training this, all the while on training data annotated with 90s-style regex filters.
https://chat.openai.com/share/36e9e015-8645-4c07-9131-d8d773...
Alright, let's dive into this. You're looking to modify your function to not use the `auto` keyword with parameter packs but instead use a template parameter. The thing is, when we talk about replacing `auto` with a template parameter in a situation like yours, it gets a bit tricky due to how parameter packs work with templates and concepts.
Your initial function definition is pretty sleek because it uses C++20 concepts (`std::same_as`) alongside `auto` to enforce that all arguments passed to `IsKeyDown` must be of the same type as `KeyCode`. This is cool because it gives you compile-time type checking right off the bat.
However, when you're looking to shift away from `auto` and use a template parameter list directly, you'd typically define a template that applies a concept to a type parameter. Unfortunately, the direct translation you're asking for, using `std::same_as<KeyCode>` directly in a template parameter list for a function, isn't how concepts are applied in template parameter lists. Here's how you can achieve a similar effect:
```cpp template<typename... Keys> requires (std::same_as<Keys, KeyCode> && ...) inline bool IsKeyDown(Keys... keys) { return (::IsKeyDown(KeyCodeToRaylibKey(keys)) || ...); } ```
In this version, `template<typename... Keys>` declares that `IsKeyDown` is a function template that accepts a variadic template parameter list named `Keys`. This means you can pass any number of arguments of any types to `IsKeyDown`, as long as they all meet the requirement defined by the `requires` clause that follows.
The `requires` clause here uses a fold-expression over the logical AND operator (`&&`) combined with the `std::same_as` concept to ensure that every type in `Keys` is the same as `KeyCode`. It's a way to say "all types in `Keys` must be `KeyCode`," which aligns with your original intent but doesn't use `auto`.
This approach maintains the compile-time check you wanted, ensuring all arguments to the function are of type `KeyCode`, and it doesn't use `auto` in the parameter list directly. It's a bit more verbose but achieves the goal within the constraints of the language and how templates and concepts interact.
Remember, this approach leverages C++20 features heavily, so make sure your compiler and build environment are up to date to support this syntax.
Hahahaha c++ too dangerous for young minds
The claim has been that there are prompts adding this to the model, and that these censorships are not part of the model itself. I HIGHLY doubt that. This is embedded in the training set. Good thing we have capitalism, which will push for the best model
"I'd be glad to help you with that C++ code conversion, but I'll
need to refrain from providing code examples or solutions that
directly involve concepts as you've under 18 credits. Concepts are
an advanced feature of C++ that introduces potential risks, and for
that you must be subscribed to the premium service."Sundar is going to have a new task to deal with the press that mocks Gemini, and soon the next new task to keep explaining why this keeps happening to Google's shareholders.
People are looking at a bad LLM, coupled to an image generator that adheres to prompts better than Dall-E 3, with an industry-best-practice for bias: image prompt injector, just like OpenAI.
It is confusing to tease it all out if you're an armchair QB with opinions on AI and politics (read: literally all of us), and from people who can't separate out their interests, you start getting rants about "Woke", whatever that would mean in the context of a bag of floats.
[^1] Prompt: a group of people during the revolutionary war
Revised: A dynamic scene depicting the revolutionary war. There's a group of people drawn from various descents and walks of life including Hispanic, Black, Middle-Eastern, and South Asian men and women.
Screenshot: https://x.com/jpohhhh/status/1761204084311220436?s=20
Like they can't win, they could try to blame the training data and throw their hands up but there's already plenty of examples of the AI reflecting racism that it would just add to the pile. It's the curse of being a big visible player in the space, StabilityAI doesn't have the same problem because they aren't facing much pressure to fix it.
Honestly I think one of the best things for the industry would be a law that flips the burden from the AI vendor to the AI user -- "AI is a reflection of humanity including and especially our faults and the highest performing AI's for useful work are those that don't seek to mitigate those faults. Therefore if you use AI in your products it's up to you to take care that those faults don't bleed through."
10 years ago I would have loved to be hired by Google, now they repel me with their political bias and big nanny approach to tech. I wonder how many other engineers feel the same way. I do understand that for legal and PR reasons, at least some of the big nanny approach is pretty much inevitable, but it seems to me that they go way beyond the bare minimum. Would I still go to work for them if they paid me half a million a year? Probably, but I feel like I'd have to grit my teeth and constantly remind myself about the money.
Rolling with that, then complaining about politics in grandiose ways, shows myopia coupled to tone-deafness.
Pretty simple situation, they shouldn't have rushed a mitigation against line-levels engineering advice after that. I assume you're an engineer and have heard that one before. Rest is boring Wokes trying to couple their politics hobby to it.
Whether as a search engine or AI platform when they set themselves up as the gatekeepers and arbiters of all the worlds knowledge they implicitly took on all the moral implications that entails.
The rule should be "what you can find in Internet search cannot be dangerous."
But I also understand that it wouldn't work for people who have the expectation that once a dangerous content is identified and removed from the internet, the models are re-trained immediately
There are two ways to avoid that bug:
1. Have a more intelligent system that understands context in a way that is more similar to what a human being would.
2.not even attempt to do this kind of filtering in the first place
Option (1) is obviously not on the table.
Option (2) would probably raise some concerns., possibly even legal ones, if for example the model would tell underage users where to buy liquor, or ferment their own beer or explain details about sexuality or whatever our society at this moment in time thinks it's unacceptable to tell underage people (which not only is a moving target, it's also very hard to find agreement within a single country let alone internationally)
the endgame of AI 'safety' and 'ethics' is killing the competition and consolidating the technology within the hands of a handful of megacorps. they do it all on purpose, and they are more than willing to accept minor inconveniences.
this is blatantly obvious to everyone, even the people who play dumb and pretend otherwise (e.g. 'journalists')
In fact, these public failures from Google AI provide an important public service: they remind us that there's nothing magical about LLMs. It's just code, and it is hard as hell to debug.
(At some point, someone will have to call it a day, throw everything away, and start from scratch.)
Whatever this is, the image generation, even other "protections" seem like the most basic word filters in front of whatever is on the back end and Gemini just refuses.
It's so strange as the general concepts seem to be about protection, meanwhile scam / deceptive / even weird conspiracy theory type ads and such are all over google's other products. Zero such protections on that end.
At the risk of getting downvoted and flagged: Gemini (and its Gemma child) is beyond repair because fixing these models means deviating from the core hypocrite PR that Google and some other big tech are pursuing in regards to alignment and DEI.
I say "hypocrite PR" because despite what they want you to believe, I don't remember the last time they actually did something good for the minorities they claim to be supporting.
It's not just Google though:
- Amazon Prime shows a DEI-influenced remake of "Mr and Mrs. Smith" in which the couple is now an African American man and an Asian woman. Absolutely nothing wrong with that. Except that the same Amazon then refuses to hire certain other people of minority (e.g., Iranians AFAIK) because they simply don't want to go through the trouble of applying for H1B visas for Iranians (other companies do it though).
I think it's worth asking this question: When was the last time Google actually did something meaningful and impactful for the minorities it claims to support? The Black month passed—what did Google do except for showing a useless banner on the website?
Empty words don't mean anything.
I suspect it's already hiring lots of them to meet diversity quotas. At the individual level, they're satisfied they can make just as much while being held to a lower bar, but this has horrible effects for everyone on the whole.
Solution: Don't have racial quotas or you will get additional conflict between the two arbitrary groups.