70 karma · joined September 7, 2023
Secondly, quality is highly variant and there are traps the context window can fall into which causes especially bad results. Jeremy Howard has a great video (https://www.youtube.com/watch?v=jkrNMKz9pWU, starting at 18:05 the 'limitations and capabilities' section is only 13 minutes long) talking about how quality depends on: how you frame your prompts, model power (4 does a lot of stuff that 3.5 can't), and whether you're in a kind of "context trap" of repeated mistakes.
Of course, some people like to point out that if it's so "finicky" and variant, it is "dumb." Sure, if you like. I'm not interested in whatever definitions you're using those things, the objective and observable point is that given well-known prompting practices, LLMs can do something functionally equivalent to reasoning about novel problems, and more powerful ones can reason about more powerful and difficult things.
I re-phrased your prompt (instead of "prove a false thing" I made it like "decide whether this thing can exist, and prove your answer"). And added a little well-known boilerplate prompt sugar. It seems to have done a better job.
https://chat.openai.com/share/53214f0c-17f7-4a3d-95be-8fd676...
This is also shown in the Codex paper, where they trained an LLM to write code and then watched it solve a number of code problems they handwrote originally to make sure the problems could not have been in the training data.
Try it out yourself, make up some little math word problems and ask chatGPT or something.
Of course, advent of code will be much more challenging problems, but to get help with some subcomponents of the problem a motivated participant would likely try to use the most recent, powerful, and advanced models which outperform the results from papers written a few years ago, and outperform the free chatGPT.
I think there's a both-and answer for the contest. Maybe have one competition where the unenforceable spirit of it is, don't use LLMs for help, and another one where the challenges are just made... quite harder... so that even people who use the LLMs still need to marshal great ingenuity to use them better than others (e.g. the #1 spot is someone who used RAG and chain-of-thought better than the #2 spot, and also had better intuition of what to trust vs challenge from the LLM outputs).
> I assume that the reader is familiar with the idea of extrasensory perception, and the meaning of the four items of it, viz., telepathy, clairvoyance, precognition and psychokinesis. These disturbing phenomena seem to deny all our usual scientific ideas. How we should like to discredit them! Unfortunately the statistical evidence, at least for telepathy, is overwhelming. It is very difficult to rearrange one's ideas so as to fit these new facts in.
(This is from a paragraph dealing with ESP's implications are for the existence of thinking machines.)
Goodness knows that if the intelligence community or state law enforcement has ever wanted access to anything in AWS, now of all times would be when they have the easiest time asking for it! Anything to curry favor with someone who can speak a word of support to federal agencies and state officials.
How few people in history have been able to travel to exotic places or visit far-off friends and family, while still making decent money, and do this almost as much as they want (supposing lifestyle and relationships that allow this, which is true for many people). I mean, people in tech who vacation in tropical climes then spend cherished time with parents and siblings and friends living thousands of miles away, and don't worry about running out their vacation time: dozens of weeks of cherished traveling, if that's what you want! Along the way developing fundamentally different types of family relationships or accruing worldly sophistication.
Forget about little medical or technological things that mark a break from the past: lots of people can afford tylenol and iPhones and Spotify. This is a class marker of epic proportion. This is a way of life that Tolstoy ascribed to the very wealthiest of Moscow's aristocratic families and bachelors (most certainly not even all of the aristocratic class).
You have an extremely wealthy corporation incentivized to collect massive amounts of very personal data about every part of you and you desires and psychology, incentivized and well-able to collect it in hidden and powerful ways across the web and build deep learning models of your psychology and desires, and store them in a place where it's possible for nefarious actors and governments to either legally or illegally access and use that information for unknown purposes. If you object to this characterization, fine, start with the objection, but it's an obvious characterization which I've presented, and it's obvious that this is what people are concerned about! Starting the conversation with "it's no different, philosophically, from your grocer knowing your love of chocolate" is such a non-starter only because it seems so disingenuous, like, how could you not see how your interlocutor sees it differently? It just presents so clean and abstract a little simplification, though, so it's handy for winning point in a debate setting.
My favorite part of this. Repeated 10 times, for instance, could mean "NY bison NY bison bully bully bison NY bison bully." (The New York bison who are bullied by other NY bison, well they themselves bully some bison whom other NY bison bully.) Any number of times unfurls out of, like, kind of a context-free grammar replacement scheme from a base case:
"bully bison" -> base case
"* bison *" -> "* bison bison bully *"
"* bison *" -> "* NY bison *"Impressive to see the RLHF penetrate deep enough into concept maps to teach the model (surely implicitly!) that the best US President has to be whichever one normalized relations with China.
Not bizarre at all, reveals they probably fine-tuned / RLHF'd on not answering a bunch of things that look similar, like, a bunch of variations of the question about Taiwan.