CriticGPT: Finding GPT-4's mistakes with GPT-4
openai.com
openai.com
According to his website he previously ran the language model alignment team at OpenAI. https://paulfchristiano.com/
It's wonderful to see his idea coming to fruition! I'm honestly a bit skeptical of the idea myself (it's like proposing to stabilize the stack of "turtles all the way down" by adding more turtles - as is insightfully pointed out in this other comment https://news.ycombinator.com/item?id=40817017) but every innovative idea is worth a try, in a field as time-critical and urgent as AGI safety.
For a good summary of technical approaches to AGI safety, start with the Future of Life Institute AI Alignment Podcast, especially these two episodes which serve as an overview of the field:
- https://futureoflife.org/podcast/an-overview-of-technical-ai...
- https://futureoflife.org/podcast/an-overview-of-technical-ai...
In both of those episodes, Cristiano's publication series on Iterated Amplification is link #3 in the list of recommended reading.
That's completely fine. Say each layer uses the same amount of turtles, but half the size of the layer above. Even allowing for arbitrarily small turtles, the total height of the infinite stack will be just 2x the height of the first layer.
Point being, some series converge to a finite result, including some defined recursively. And in practice, we can usually cut the recursion after first couple steps, as the infinitely long remainder has negligible impact on the final result.
I genuinely laughed. Oh no somebody please save me from a chatbot that's hallucinating half the time!
Joke aside, of course OpenAI is gonna play up how "intelligent" its models are. But it's evident that there's only so much data and compute that you can throw at a machine to make it smart.
Other people using half-baked AI can still kill you, and that doesn't have to be a chatbot as we have current examples from self-driving cars that drive themselves dangerously, and historical examples of the NATO early warning radars giving a false alarm from the moon and the soviet early warning satellites giving false alarms from reflected sunlight, but it can also be a chatbot — there are many ways that this can be deadly if you don't know better: https://news.ycombinator.com/item?id=40724283
Every software bug is an example of a computer doing exactly what it was told to do, instead of what we meant.
AI safety is about bridging the gap between optimising for what we said vs. what we meant, in a less risky manner than if covid — and while I think it doesn't matter much if covid did or didn't come from a lab leak (the potential that it did means there's an opportunity to improve bio safety there as well as in wet markets), every AI you can use is essentially a continuous supply of the mystery magic box before we know what the word "safe" even means in this context.
That is only as long as the person describing the behaviour as a bug is aligned with the programmer. Most of the time this is the case, but not always. For example a malicious programmer intentionally inserting a bug does in fact mean for the program to have that behavior.
No real way to mathematically prove this, considering there is also no way to know if the training data also had this “hallucination” inside of it.
Investigate it with the tools of psychologically, as suited for use on a new non-human creature we've never encountered before.
That being said, people new to the field tend to believe that these models are fact machines. In fact, they are the complete opposite.
There isn't really such thing as a "hallucination" and honestly I think people should be using the word less. Whether an LLM tells you the sky is blue or the sky is purple, it's not doing anything different. It's just spitting out a sequence of characters it was trained be hopefully what a user wants. There is no definable failure state you can call a "hallucination," it's operating as correctly as any other output. But sometimes we can tell either immediately or through fact checking it spat out a string of text that claims something incorrect.
If you start asking an LLM for political takes, you'll get very different answers from humans about which ones are "hallucinations"
There's definitely room for a better label, though. "Empirical mismatch" doesn't quite have the same ring as "hallucination," but it's probably a more accurate place to start from.
Sure, but that would require semantic mechanisms rather than statistical ones.
If someone wants info to make their model to be more reliable for a specific domain, it's in the existing papers on model training.
Is is possible for a chess engine to compute the next move and be absolutely sure it is the best one? It's not, it is a statistical approximation, but still very useful.
People say it's "anthropomorphizing" but honestly I can't see it. The I in AI stands for intelligence, is this anthropomorphizing? L in ML? Reading and writing are clearly human activities, so is using read/write instead of input/output anthropomorphizing? How about "computer", a word once meant a human who does computing? Is there a word we can use safely without anthropomorphizing?
[1]: And please don't argue what's "wrong".
You will be told that linear algebra is just a model and the fact that epistemology has never turned up a decent result for what knowledge is will be ignored.
We are meant to believe that we are somehow special magical creatures and that the behaviour of our minds cannot be modelled by linear algebra.
If a company does a thing that's bad, it doesn't matter much if the work itself was performed by a blacksmith or by a robot arm in a lights-off factory.
> We are meant to believe that we are somehow special magical creatures and that the behaviour of our minds cannot be modelled by linear algebra
I only hear this from people who say AI will never reach human level; of AI developers that get press time, only LeCun seems so dismissive (though I've not actually noticed him making this specific statement, I can believe he might have).
No, it’s more specific than just wrong.
Hallucination is when a model creates a bit of fictitious knowledge, and uses that knowledge to answer a question.
You can argue if it matters how a wrong answer came about ofc but there is a difference
You can still get that with zero bad labels in a supervised training set.
Multiple causes for the same behaviour makes progress easier, but knowing if it's fully solved harder.
Context is "don't call it hallicination" picked up meme energy since https://link.springer.com/article/10.1007/s10676-024-09775-5 on the thesis that "Calling their mistakes ‘hallucinations’ isn’t harmless: it lends itself to the confusion that the machines are in some way misperceiving but are nonetheless trying to convey something that they believe or have perceived."
Which is meta-bullshit because it doesn't matter. We want LLMs to behave more factually, whatever the non-factuality is called. And calling that non-factuality something else isn't going to really change how we approach making them behave more factually.
If they could predict facts, then these would be gods, not machines. It would be saying that in all the written content we have, there exists a pattern that allows us to predict all answers to questions we may have.
It logics its way to it.
By predicting the next word in a sequence of words.
Sure? It kinda sounds plausible? But man, if it’s that straight forward, what have we been doing as a species for so many years ?
TLDR: Sure. A rose by any other name would be just as sweet. It’s when I use the name of the rose and imply aspects that are not present, that we create confusion and busy work.
Hey, calling it a narrative is to move it to PR speak. I know people have argued this term was incorrect since the first times it was ever shared on HN.
It was unpopular to say this when ChatGPT launched, because chatGPT was just that. freaking. cool.
It is still cool.
But it is not AGI. It does not “think”.
Hell - I understand that we will be doing multiple columns of turtles all the way down. I have a different name for this approach - statistical committees.
Because we couched its work in terms of “thinking”, “logic”, “creativity”, we have dumped countless man hours and money into avenues which are not fruitful. And this isnt just me saying it - even Ilya commented during some event that many people can create PoCs, but there are very few production grade tools.
Regarding the L in ML, and the I in AI ->
1) ML and AI were never quite as believable as ChatGPT. Calling it learning and intelligence doesnt result in the same level of ambiguity.
2) A little bit of anthropomorphizing was going on.
Terms matter, especially at the start. New things get understood over time, as we progress we do move to better terms. Let’s use hallucinations for when a digital system really starts hallucinating.
To me the real danger comes from when the models get things wrong but also correct at the same time. Not so much in software engineering, I doubt your average programmer without LLM tools will write “better” code without getting some bad answers. What consents me is more how non-technical departments implement LLMs into their decision making or analysis systems.
Done right, it’ll enhance your capabilities. We had a major AI project in cancer detection, and while it actually works it also doesn’t really work on its own. Obviously it was meant to enhance the regular human detection and anyone involved with the project screamed this loudly at any chance they got. Naturally it was seen as an automation process by the upper management and all the humans parts of the process were basically replaced… until a few years later when we had a huge scandal about how the AI worked as it was meant to do, which wasn’t to be on its own. Today it works along side the human detection systems and their quality is up. It took people literally dying to get that point through.
Maybe it would’ve happened this way anyway if the mistakes weren’t sort of written into this technical issue we call hallucinations. Maybe it wouldn’t. From personal experience with getting projects to be approved, I think abstractions are always a great way to hide the things you don’t want your decision makers to know.
We won't, and we'll see this constant distraction.
Well, parent is lamenting the lack of lowerbound/upperbound for "hallucinations", something that cannot realistically exist as "hallucinations" don't exist. LLMs aren't fact-outputting machines, so when it outputs something a human would consider "wrong" like "the sky is purple", it isn't true/false/correct/incorrect/hallucination/fact, it's just the most probable character after the next.
That's why it isn't useful to ask "but how much it hallucinates?" when in reality what you're out after is something more like "does it only output facts?". Which, if it did, LLMs would be a lot less useful.
LLM don't need to be perfect fact machines at all to be honest, and non-hallucinating. They simply need to ground statements in other grounded statements and identify the parts which are speculative or non-grounded.
Otherwise, how do you prove the grounding isn't "hallucinated"?
That's simply not true. You're confusing how they're trained and what they do. They don't have some store of exactly how likely each word is (and it's worth stopping to think about what that would even mean) for every possible sentence.
It's a simplification. Temperature also influences it to not always be the most probable character, as an example.
It is somewhat humorous when humans have ontological objections to the neologisms used to describe a system whose entire function is to relate the meanings of words. It is almost as if the complaint is itself a repressed philosophical rejection of the underlying LLM process, only being wrapped in the apparent misalignment of the term hallucination.
The complaint may as well be a defensive clinging "nuh uh, you can't decide what words mean, only I can"
Perhaps the term "gas lighting" is also an appropriate replacement of "hallucination," one which is not predicated on some form of truthiness standard, but rather THIS neologism focuses on the manipulative expression of the lie.
Hallucination might not be the best word, but I don't think it's a bad word. If a weather model predicted a storm when there isn't a cloud in the sky, I wouldn't have a problem with saying "the weather model had a hallucination." 50 years ago, weather models made incorrect predictions quite frequently. That's not because they weren't modeling correct weather, it's because we simply didn't yet have good models and clean data.
Fundamentally, we could fix most LLM hallucinations with better model implementations and cleaner data. In the future we will probably be able to model factuality outside of the context of human language, and that will probably be the ultimate solution for correctness in AI, but I don't think that's a fundamental requirement.
People still want it to be used for thinking.
This isnt going to happen with better data. Better data means it will be better at predicting the next token.
For questions or interactions where you need to process, consider, decompose a problem into multiple steps, solve those steps etc - you need to have a goal, tools, and the ability to split your thinking and govern the outcome.
That isnt predicting the next token. I think it’s easier to think of LLMs as doing decompression.
They take an initial set of tokens and decompress them into the most likely final set of tokens.
What we want is processing.
We would have to set up the reaction to somehow perfectly result in the next set of tokens to then set up the next set of tokens etc - till the system has an answer.
Or in other words, we have to figure out how to phrase an initial set of tokens so that each subsequent set looks similar enough to “logic” in the training data, that the LLM expands correctly.
Humans also confabulate but not as a result of "hallucinations". They usually do it because that's actually what brains like to do, whether it's making up stories about how the world was created or, more infamously, in the case of neural disorders where the machinery's penchant for it becomes totally unmoderated and a person just spits out false information that they themselves can't realize is false. https://en.m.wikipedia.org/wiki/Confabulation
This is a very "closed world" view of the phenomenon which looks at an LLM as a software component on its own.
But "hallucination" is a user experience problem, and it describes the experience very well. If you are using a code assistant and it suggests using APIs that don't exist then the word "hallucination" is entirely appropriate.
A vaguely similar analogy is the addition of the `let` and `const` keywords in JS ES6. While the behavior of `var` was "correct" as-per spec the user experience was horrible: bug prone and confusing.
What am I suppose to call that?
https://www.nytimes.com/2023/11/06/technology/chatbots-hallu...
For us, we treat hallucinations as the ability to accurately respond in an "open book" format for retrieval augmented generation (RAG) applications specifically. That is, given a set of information retrieved (X), does the LLM-produced summary:
1. Include any "real" information not contained in X? If "yes," it's a hallucination, even if that information is general knowledge. We see this as an important way to classify hallucinations in a RAG+summary context because enterprises have told us they don't want the LLMs "reading between the lines" to infer things. To pick an absurd/extreme case to show a point, the case of a genetic research firm, say, using CRISPR and finding they can create a purple zebra, if the retrieval system in the RAG bits says "zebras can be purple" due to their latest research, we don't want the LLM to override that knowledge with its knowledge that zebras are only ever black/white/brown. We'd treat that as a hallucination.
2. On the extreme opposite end, an easy way to avoid hallucinating would be for the LLM to say "I don't know" for everything thereby avoiding hallucinating by avoiding answering all questions. That has other obvious negative effects, so we also evaluate LLMs for their ability to answer.
We look at the factual consistency, answer rate, summary length, and some other metrics internally to focus prompt engineering, model selection, and model training: https://github.com/vectara/hallucination-leaderboard
Some people are surprised by smaller models having the ability to outperform bigger models, but it's something we've been able to exploit: if you fine tune a small model for a specific task (e.g. reduced hallucinations on a summarization task) as Intel has done, you can achieve great performance economically.
So, slightly better than random chance. I guess a win is a win but I would have thought this would be higher. I'd kind have assume that just asking GPT itself if it's sure would be this kind of lift.
You can see the plots if you prefer, or think of it this way: out of a total of 100 trials, one team gets 40 and the other gets 60 = 40 + 40 * 50%
If you want to think of a 75% win rate as a more extreme example: you could say 25% above random or you could say one team wins 3 times as many cases as the other. Both are equivalent but I think that the second way conveys the strength of the difference much better.
The results in this work are statistically significant and substantial.
I'm sure ChatGPT would outperform me, and I could only aid it in very limited ways.
That doesn't mean an expert iOS programmer wouldn't run circles around it.
Improving alignment of dialogue agents via targeted human judgements - https://arxiv.org/abs/2209.14375
Teaching language models to support answers with verified quotes - https://arxiv.org/abs/2203.11147
The EU seems to be very much toward a regulate early, safety-first approach. Where USA is very much toward unregulated, move fast, break things, assess the damage, regulate later.
I don't know which is better or worse.
There really is comparatively little reward for shooting for the moon though – the fragmented stock markets don't provide great exit opportunities, so less money goes into funding ambitious companies. Then, scaling throughout all of Europe is notably hard, with dozens of languages, cultures, and legal frameworks to navigate. Some of these cultures are more risk-averse, and that's not easy to change. Not to mention English being the _lingua franca_ of business and tech.
I would love Europe to reach the States' level of tech strength, but these are all really hard problems.
Because of policy choices over the past several decades, as well as the ongoing war with Russia, the EU is already struggling to provide enough energy for its existing industry. There just isn’t any slack left for a newcomer.
This is a relatively recent (2022) comparison of the Industrial electricity prices including taxes:
https://www.gov.uk/government/statistical-data-sets/internat...
Roughly, electricity for industrial uses is 50% more expensive in France that in the USA.
In Germany, it is over 120% more expensive.
In the UK, over 150%.
The most valuable tech companies have long histories, incredibly broad, and deep technological portfolios that go much farther than 'we use the latest frameworks' and/or 'we own the most eyeballs at the moment' and dubious business models that are a combination of spending free VC money to grow and having exploitative business models.
Such companies are certainly represented, in the top 100 list, but I for one am not sad that Europe missed out on them. It's more of a problem that we don't have or NVIDIA, Apple, Intel, Samsung imo.
As for regulation, there are a bunch of US companies that got to where they are by abusing their monopolistic reach by locking out and disadvantaging potential competitors (think of the smartphone, OS and social media spaces), where the swift kick in the butt from European regulators could've come sooner.
Does Anthropic do something like this as well, or is there another reason Claude Sonnet 3.5 is so much better at coding than GPT-4o?
It's impossible to say because these models are proprietary.
"Which data specifically? Gerstenhaber wouldn’t disclose, but he implied that Claude 3.5 Sonnet draws much of its strength from these training sets."[0]
[0]https://techcrunch.com/2024/06/20/anthropic-claims-its-lates...
I remember when I first started using activation maps when building image classification models and it was like what on earth was I doing before this... just blindly trusting the loss.
How do you discover biases and issues with training data without interpretability?
This reminds me of the passage found in the description of the fuckitpy module:
"This module is like violence: if it doesn't work, you just need more of it."
If I ask it a question, I try not to trust it immediately, and I independently look the answer up and I argue with it. In turn, it actually is one of my favorite learning tools, because it kind of forces me to figure out why it's wrong and explain it.
Fighting with the AI's wrongness out of spike is an unexpectedly good motivator.
I guess at some level this is almost what "prompt engineering" is (though I really hate that term), but I use it as a learning tool and I do think it's been really good at helping me cement concepts in my brain.
Interesting, that's the basic process I follow myself when learning without ChatGPT. Comparing my mental representation of the thing I'm learning to existing literature/results, finding the disconnects between the two, reworking my understanding, wash rinse repeat.
It can be hard for me to directly figure out when my mental model is wrong on something. I'm sure it happens all the time, but a lot of the time I will think I know something until I feel compelled to prove it to someone, and I'll often find out that I'm wrong.
That's actually happened a bunch of times with ChatGPT, where I think it's wrong until I actually interrogate it, look up a credible source, and realize that my understanding was incorrect.
Finally, LLMs teach us the good habits.
Arcane language is actually kind of a pet peeve of mine in theoretical CS and mathematics. Sometimes it feels like academics really obfuscate relatively simple concepts but using a bunch of weird math terms. I don't think it's malicious, I just think that there's value in having more approachable language and metaphors in the process of explaining thing.
ChatGPT (Gemini/Anthropic/etc) have the advantage of never getting sick of arguing with me. I can go back and forth and argue about any weird topic that I want for as long as I want at any time of day and keep learning until I'm bored of it.
Obviously it depends on the person but I really like it.
Beyond just subject-wise, finding people who argue in good faith seems to be an issue too. There are people I'm friends with almost specifically because we're able to consistently have good-faith arguments about our strongly opposing views. It doesn't seem to be a common skill, but perhaps that has something to do with my sample set or my own behaviors in arguments.
Politically? Yeah, nearly impossible to find anyone who argues in good faith.
Even with that, it took a process that had taken multiple hours before down to about 30-45 minutes. It was super cool.
[1] Just to be clear, I always did the homework assignments myself beforehand to make sure that a solution was solvable and fair before I assigned it.
Usually for it to get a correct answer, I have to provide it a bit of context.
We've even been getting "prompt engineering" meeting invites of 3+ hours to get an introduction into their usage. 100-150 participants each time I joined
It's amazing how much they're valuing it. from my experience it's usually a negative productivity multiplier (x0.7 vs x1 without either)
If they're not doing their job, then why do they still have one?
You handle it by teaching them how to write good code.
And if they refuse to learn, then they get bad performance reviews and get let go.
I've had junior devs come in with all sorts of bad habits, from only using single-letter variable names and zero commenting, to thinking global variables should be used for everything, to writing object-oriented monstrosities with seven layers of unnecessary abstractions instead of a simple function.
Bad LLM-generated code? It's just one more category of bad code, and you treat it the same as all the rest. Explain why it's wrong and how to redo it.
Or if you want to fix it at scale, identify the common bad patterns and make avoiding them part of your company's onboarding/orientation/first-week-training for new devs.
Because if it's bad, at least it's simple. Meaning simple to review, quickly correct and move on.
And if you know a solution should be 50 lines and they've given you 500, it's not like you have to read it all -- you can quickly figure out what approach they're using and discuss the approach they should be using instead.
(Practical nuclear fusion has 10 years away in 1954, and it's still 10 years away now. I suspect in practice LLMs are in a similar space; everyone seems to be fixated on the near, supposedly inevitable, future where they are actually useful.)
> CriticGPT’s suggestions are not always correct
Now you have two problems...
What do you mean? The model is for improving their RLHF trainers performance. RLHF does get applied "at the source" so to speak. It's a modification on the model behind the API.
Perhaps if you were to say what you think this thing is for and then share why you think it's not "at the source".
You'd like to get the "correct" answer straight away, not watch a discussion between two bots.
It's plausible that there are potential avenues for improving language models through adversarial learning. GANs and Actor-Critic models have done a good job in narrow-domain generative applications and task learning, and I can make a strong theoretical argument that you can do something that looks like priority learning via adversarial equilibria
But why in the world are you trying to present this as a human-in-the-loop system? This makes no sense to me. You take an error-prone generative language model and then present another instance of an error-prone generative language model to "critique" it for the benefit of... a human observer? The very best case here is that this wastes a bunch of heat and time for what can only be a pretty nebulous potential gain to the human's understanding
Is this some weird gambit to get people to trust these models more? Is it OpenAI losing the plot completely because they're unwilling to go back to open-sourcing their models but addicted to the publicity of releasing public-facing interfaces to them? This doesn't make sense to me as a research angle or as a product
I can really see the Microsoft influence here
Any system that helps you more accurately label data with good critiques should help the model. I'm not sure how you come to your conclusion. Do you have some data to indicate that even with improved accuracy that some LLM bias would lead to a worse trained model? I haven't seen that data or assertion elsewhere, but that's the only thing I can gather you might be referring.
The idea of RLHF as a mechanism for tuning models based on the principle that humans might have some hard-to-capture insight that could steer them independent of the way they're normally trained is the very best steelman for its value I could come up with. This aim is directly subverted by trying to use another language model to influence the human rater, so from my perspective it really brings us back to square one on what the fuck RLHF is supposed to be doing
Really, a lot of this comes down to what these models do versus how they are being advertised. A generative language model produces plausible prose that follows from the prompt it receives. From this, the claim that it should write working code is actually quite a bit stronger than the claim that it should write true facts, because plausibile autocompletion will learn to mimic syntactic constraints but actually has very little to do with whether something is true, or whatever proxy or heuristic we may apply in place of "true" when assessing information (supported by evidence, perhaps. Logically sound, perhaps. The distinction between "plausible" and "true" is in many ways the whole point of every human epistemology). Like if you ask something trained on all human writing whether the Axis or the Allies won WWII, the answer will depend on whether you phrased the question in a way that sounds like Phillip K Dick would write it. This isn't even incorrect behavior by the standards of the model, but people want to use these things like some kind of oracle or to replace google search or whatever, which is a misconception about what the thing does, and one that's very profitable for the people selling it
I tried to understand the paper and I can't really make sense of it for "code".
It seems like this would inherit a subtler version of all the problems from expert systems.
A press release of this does feel rather AI bubbly. Not quite Why The Future Doesn't Need Us level but I think we are getting close.
But, as per the first graphic, CriticGPT alone has better comprehensiveness than CriticGPT+Human? Is that right?
[1] - https://en.wikipedia.org/wiki/Generative_adversarial_network
All these humans make up too much stuff, I don't see how that can be fixed.
Reddit, while it has some niche communities with tribal info and knowledge, is FULL of spam, bots, companies masquerading as users, etc etc etc. If people are truly relying on reddit as a source of truth (which OpenAI is now being influenced by), then the world is just going to be amplify all the spam that already exists
I hope LLMs will offer a "-reddit" model to switch to when needed.
The biggest issue is how confidently wrong GPT enjoys being. You can press GPT in either right or wrong direction and it will concede with minimal effort, which is also an issue. It’s just really bad russian roulette nerdspining until someone gets tired.
How do you think leadership at OpenAI would respond to that?
2. While I agree that it's a stretch to call ChatGPT agentic, it's nonetheless "motivated" in the sense that it's learned based on an objective function, which we can model as a causal factor behind its behavior, which might improve our understanding of that behavior. I think it's relatively intuitive and not deeply incorrect to say that that a learned objective of generating plausible prose can be a causal factor which has led to a tendency to generate prose which often deceives people, and I see little value in getting nitpicky about agentic assumptions in colloquial language when a vast swath of the lexicon and grammar of human languages writ large does so essentially by default. "The rain got me wet!" doesn't assume that the rain has agency
> deliberately cause (someone) to believe something that is not true, especially for personal gain.
Emphasis on the personal gain part. It seems like you have a different definition.
There's no point in arguing about definitions, but I'm a big believer in that if you can identify a difference in the definitions people use early into a conversation, you can settle the argument at that.
When it comes to technical details, current LLM's have a bias towards sycophancy and bullshitting that humans only show when especially desperate to impress or totally fearful.
Humans make mistakes too, but the distribution of those mistakes is wildly different and generally much easier to calibrate for and work around.
# Post 1
> The problems of epistemology and informational quality control are complicated, but humanity has developed a decent amount of social and procedural technology to do these, some of which has defined the organization of various institutions.
Very fluffy, creating very uncertain parsing for reader.
Should cut down, then could add specificity:
ex. "Dealing with misinformation is complicated. But we have things like dictionaries and the internet, there's even specialization in fact-checking, like Snopes.com"
(I assume the specifics I added aren't what you meant, just wanted to give an example)
> The mere presence of LLMs doesn't fundamentally change how we should calibrate our beliefs or verify information. However, the mythology/marketing that LLMs are "outperforming humans"
They do, or are clearly at par, at many tasks.
Where is the quote from?
Is bringing this up relevant to the discussion?
Would us quibbling over that be relevant to this discussion?
> combined with the fact that the most popular ones are black boxes to the overwhelming majority of their users means that a lot of people aren't applying those tools to their outputs.
Are there unpopular ones aren't black boxes?
What tools? (this may just indicate the benefit of a clearer intro)
> As a technology, they're much more useful if you treat them with what is roughly the appropriate level of skepticism for a human stranger you're talking to on the street
This is a sort of obvious conclusion compared to the complicated language leading into it, and doesn't add to the posts before it. Is there a stronger claim here?
# Post 2
> I wonder what ChatGPT would have to say if I ran this text through with a specialized prompt.
Why do you wonder that?
What does "specialized" mean in this context?
My guess is there's a prompt you have in mind, which then would clarify A) what you're wondering about B) what you meant by specialized prompt. But a prompt is a question, so it may be better to just ask the question?
> Your choice of words is interesting, almost like you are optimizing for persuasion,
What language optimizes for persuasion? I'm guessing the fluffy advanced verbiage indicates that?
Does this boil down to "Your word choice creates persuasive writing"?
> but simultaneously, I get a strong vibe of intention of optimizing for truth.
Is there a distinction here? What would "optimizing for truth" vs. "optimizing for persuasion" look like?
Do people usually write not-truthful things, to the point it's worth noting that when you think people are writing with the intention of truth?
Like I said, I'm very interested!
Maybe it doesn't mean anything other than what it says on the tin? You think people should treat an LLM like a stranger making claims? Makes sense!
It's just unclear what a lot of it means and the word choice makes it seem like there's something grander going on, coughs as our compatriots in this intricately weaved thread on the international network known as the world wide web have also explicated, and imparted via the written word, as their scrivening also remarks on the lexicographical phenomenae. coughs
My only other guess is you are doing some form of performance art to teach us a broader lesson?
There's something very "off" here, and I'm not the only to note it. Like, my instinct is it's iterated writing using an LLM asked to make it more graduate-school level.
I can't tell you exactly what you find "off" about my prose, because while you have advocated precision your objection is impossibly vague. I talk funny. Okay. Cool. Thanks.
Anyway, most benchmarks are garbage, and even if we take the validity of these benchmarks for granted, these AI companies don't release their datasets or even weights, so we have no idea what's out of distribution. To be clear, this means the claims can't be verified even by the standards of ML benchmarks, and thus should be taken as marketing, because companies lying about their tech has both a clearly defined motivation and a constant stream of unrelenting precedent
You mean on this planet?
If not, what do you think of that idea? Does something not seem....weird?
Particularly:
> Also, in general it seems unlikely humans function as optimizers natively, because optimization tends to require drastically narrowing and quantifying your objectives. I would guess that if they're describable and consistent, most human utility functions look more like noisy prioritized sets of satisfaction criteria than the kind of objectives we can train a neural network against
Considering this, what do you think us humans are actually up to, here on HN and in general? It seems clear that we are up to something, but what might it be?
In general? Slowly dying mostly. Talking. Eating. Fucking. Staring at microbes under a microscope. Feeding cats. Planting trees. Doing cartwheels. Really depends on the human
> Talking.
Have you ever noticed any talking that ~"projects seriousness &/or authority about important matters" around here?
Yes, absolutely. I view this as one of the criteria by which I assess emotional maturity, and despite societal pressures to never do so, many manage to, even though most don't
I'm not a sociologist, but I think the degree to which people can't turn it off maps fairly well onto the "low-high trust society" continuum, with lower trust implying less willingness or even sometimes ability to stop trying to do this on average, though of course variation will exist within societies as well
I have this intuition because I think the question of whether to present vulnerability and openness versus authority and strength is essentially shaped like a prisoner's dilemma, with all that that implies
We're not fully aligned here....I'm thinking more like: stop (or ~isolate/manage) non-intentional cognition, simulated truth formation, etc.....not perfectly in a constant, never ending state of course, but for short periods of time, near flawlessly.
I think we're not talking about exactly the same thing though, which I'd say is my fault. I would like to modify this:
> stop (or ~isolate/manage) non-intentional cognition, simulated truth formation, etc.....not perfectly in a constant, never ending state of course, but for short periods of time, near flawlessly.
...to this (only change is what I appended to the end):
> stop (or ~isolate/manage) non-intentional cognition, simulated truth formation, etc.....not perfectly in a constant, never ending state of course, but for short periods of time, near flawlessly, without stopping cognition altogether (such as during "no mind" meditation or "ego death" using psychedelics). Think more like a highly optimized piece of engineering, where we have ~full (comparable to standard engineering or programming) access to the code, stack, state, etc.
1. I explicitly acknowledged I misspoke and wanted to clarify: "I think we're not talking about exactly the same thing though, which I'd say is my fault. I would like to modify this:"
2. What is circuitous about my question? Is my refined question non-valid?
> At this point I'm not sure what point, if any, you're trying to get at
I encourage you to interpret my question literally, or ask for clarification.
> ...and it's hard not to form the impression that you're being deliberately obtuse here...
obtuse: ": lacking sharpness or quickness of sensibility or intellect : insensitive, stupid. He is too obtuse to take a hint. b. : difficult to comprehend : not clear or precise in thought or expression".
I'd like to see you make the case for that accusation, considering the text of our conversation is persisted above.
Rhetoric is popular, and it will work on most people here, but it will not work on me. I will simply call it out explicitly, and then observe what technique you try next. You do realize that you people can be observed, and studied, don't you?
> ...though it also could just be the brainrot that comes of overabstraction
Perhaps. Alternatively, my question could be valid, challenging to your beliefs (which I suspect are perceived as knowledge), and you lack the self-confidence to defend those beliefs.
You are welcome to:
1. genuinely address my words
2. engage in more rhetoric
3. stay silent (which may be interpreted as you not seeing this message, regardless of whether that is true)
4. something else of your choosing
Why would I want to ask your average person a physics question? Of course, their answer will probably be wrong and partly made up. Why should that be the bar?
I want it to answer at the level of a physics expert. And a physics expert is far less likely to make basic mistakes.
We have human institutions dedicated at least nominally to finding and publishing truth (I hate having to qualify this, but Hacker News is so cynical and post-modernist at this point that I don't know what else to do). These include, for instance, court systems. These include a notion of evidentiary standards. Eyewitnesses are treated as more reliable than hearsay. Written or taped recordings are more reliable than both. Multiple witnesses who agree are more reliable than one. Another example is science. Science utilizes peer review, along with its own notion of hierarchy of evidence, similar to but separate from the court's. Interventional trials are better evidence than observational studies. Randomization and statistical testing is used to try and tease out effects from noise. Results that replicate are more reliable than a single study. Journalism is yet another example. This is probably the arena in which Hacker News is most cynical and will declare all of it is useless trash, but nonetheless reputable news organizations do have methods they use to try and be correct more often than they are not. They employ their own fact checkers. They seek out multiple expert sources. They send journalists directly to a scene to bear witness themselves to events as they unfold.
You're free to think this isn't sufficient, but this is how we deal with humans making up stuff and it's gotten us modern civilization at least, full of warts but also full of wonders, seemingly because we're actually right about a lot of stuff.
At some point, something analogous will presumably be the answer for how LLMs deal with this, too. The training will have to be changed to make the system aware of quality of evidence. Place greater trust in direct sensor output versus reading something online. Place greater trust in what you read from a reputable academic journal versus a Tweet. Etc. As it stands now, unlike human learners, the objective function of an LLM is just to produce a string in which each piece is in some reasonably high-density region of the probability distribution of possible next pieces as observed from historical recorded text. Luckily, producing strings in this way happens to generate a whole lot of true statements, but it does not have truth as an explicit goal and, until it does, we shouldn't forget that. Treat it with the treatment it deserves, as if some human savant with perfect recall had never left a dark room to experience the outside world, but had read everything ever written, unfortunately without any understanding of the difference between reading a textbook and reading 4chan.
I tried recently to have ChatGPT an .htaccess RewriteCond/Rule for me and it was extremely confident you couldn't do something I needed to do. When I told it that it just needed to add a flag to the end of the rule (I was curious and was purposely non-specific about what flag it needed), it suddenly knew exactly what to do. Thankfully I knew what it needed but otherwise I might have walked away thinking it couldn't be accomplished.
If you are still seeing the issue with memory and GC, you can submit it to https://github.com/dotnet/runtime/issues especially if you are doing something that is expected to just work(tm).
* difficult as in retrieving data detailed enough to trace individual allocations, otherwise `GC.GetGCMemoryInfo()` and adjacent methods can give you high-level overview. There are more advanced tools but I always had the option to either use remote debugging in Windows Server days and dotnet-dump and dotnet-trace for containerized applications to diagnose the issues, so haven't really explored what is needed for the more locked down environments.
Mostly it's good at structure and syntax, so I'll often find the library/spec I want, paste in the relevant documentation and ask it to write my function for me.
This may seem like a waste of time because once you've got the documentation you can just write the code yourself, but A: that takes 5 times as long and B: I think people underestimate how much general domain knowledge is buried in chatgpt so it's pretty good at inferring the details of what you're looking for or what you should have asked about.
In general, I think the more your interaction with chatgpt is framed as a dialogue and less as a 'fill in the blanks' exercise, the more you'll get out of it.
Just like I do with my own code.
Both AI and I "hallucinate" sometimes, but with good tests you make things work.
If you are knowledgeable on a subject matter you're asking for help with, the LLM can be guided to value. This means you do have to throw out bad or flat out wrong output regularly.
This becomes a problem when you have no prior experience in a domain. For example reviewing legal contracts about a real estate transaction. If you aren't familiar enough with the workflow and details of steps you can't provide critique and follow-on guidance.
However, the response still stands before you, and it can be tempting to glom onto it.
This is not all that different from the current experience with search engines, though. Where if you're trying to get an answer to a question, you may wade through and even initially accept answers from websites that are completely wrong.
For example, products to apply to the foundation of an old basement. Some sites will recommend products that are not good at all, but do so because the content owners get associate compensation for it.
The difference is that LLM responses appear less biased (no associate links, no SEO keyword targeting), but are still wrong.
All that said, sometimes LLMs just crush it when details don't matter. For example, building a simple cross-platform pyqt-based application. Search engine results can not do this. Wheras, at least for rapid prototyping, GPT is very, very good.
Iteration also is when your brain meets the external world and corrects. This is a closed system.
Probably just evaluation on benchmarks.
In the particular case of spotting mistakes made by ChatGPT, a mistake is spotted if it is spotted by the human reviewer or by the critic, so even a critic that makes many mistakes itself can still increase the number of spotted errors. (But it might decrease the spotting rate per unit time, so there are still trade-offs to be made.)
From the article:
"In our experiments a second random trainer preferred critiques from the Human+CriticGPT team over those from an unassisted person more than 60% of the time."
Of course the second trainer could be wrong, but when the outcome tilts 60% to 40% in favour of the *combination of a human + CriticGPT that's pretty significant.
From experience doing contract work in this space, it's common to use multiple layers of reviewers to generate additional data for RLHF, and if you can improve the output from the first layer that much it'll have a fairly massive effect on the amount of training data you can produce at the same cost.
I've done work on what (at least to my belief) is the very high end of that scale (not for OpenAI) to fill gaps, so I know firsthand that it's available, and sometimes the work is complex enough that a single response can take over an hour to evaluate because the requirements often include not just reading and reviewing the code, but ensuring it works, including fixing bugs. Most of the responses then pass through at least one more round of reviews of the fixed/updated responses. One project I did work on involved 3 reviewers (none of whom were on salaries anywhere close to the Kenyan workers you referred to) reviewing my work and providing feedback and a second pass of adjustments. So four high-paid workers altogether to process every response.
Of course, I'm sure plenty lower-level/simpler work had been filtered out to be addressed with cheaper labour, but I wouldn't be so sure their costs for things like code is particularly low.
A human reviewer might have trouble catching a mistake, but they are generally pretty good at discerning a report about a mistake is valid or not. For example, finding a bug in a codebase is hard. But if a junior sends you a code snippet and says "I think this is a bug for xyz reason", do you agree? It's much easier to confidently say yes or no. So basically it changes the problem from finding a needle in a haystack to discerning if a statement is a hallucination or not.