Writing a GPT-4 script to check Wikipedia for the first unused acronym
gwern.net
gwern.net
They also let you do less basic processing tasks that would have been too expensive to expose over API.
Since Wikipedia posts already have a canonical numeric ID, if map semantics are important, I'd probably load that mapping into memory and use something like roaringbitmap for compressed storage of relations.
I'm usually working with the text-only OpenZim version, which cuts out most of the cruft.
Interesting topics include:
· Writing a good GPT-4 system prompt to make GPT-4 produce less verbose output and ask more questions.
· How to iterate with GPT-4 to correct errors, generate a test suite, as well as a short design document (something you could put in the file-initial docstring in Python, for example).
· The "blind spot" - if GPT-4 makes a subtle error with quoting, regex syntax, or similar, for example, it can be very tricky to tell GPT-4 how to correct the error, because it appears that it doesn't notice such errors very well, unlike higher-level errors. Because of this, languages like Python are much better to use for GPT-4 coding as compared to more line-noise languages like Bash or Perl, for instance.
· If asked "how to make [the Bash script it's written] better", GPT-4 will produce an equivalent Python script
By that argument, one should always make it use a language that's as hard as possible to write a compiling program. So Rust or Haskell or something? I guess at some point it's more important to have a lot of the language in the training data, too...
However, in my use so far, I have not noticed any striking differences in error rates between Haskell and the others.
The main complaint people have about strict, thorough type systems is that they have boilerplate.
Obviously boilerplate doesn't matter if a machine writes the code.
The type system also becomes helpful documentation of the intended behavior of the code that the LLM spits out.
So we'll just move to a new standard where we write LLM prompts describing function behavior and it will output the Rust or whatever that we end up storing in our SCM.
Someone might accidentally find it works well and then we might all end up writing fairytales in iambic pentameter describing the use cases of software we want...
What an absolutely based take by GPT-4
<jk>
Does anybody have a UrbanDictionary account?
$ egrep -o . /usr/share/dict/words | tr a-z A-Z | sort | uniq -c | sort -rn
235415 E
201093 I
199606 A
170740 O
161024 R
158783 N
152868 T
139578 S
130507 L
103460 C
87390 U
78180 P
70725 M
68217 D
64377 H
51683 Y
47109 G
40450 B
24174 F
20181 V
16174 K
13875 W
8462 Z
6933 X
3734 Q
3169 J
2 -
$ cut -c1 /usr/share/dict/words | tr a-z A-Z | sort | uniq -c | sort -rn
25170 S
24465 P
19909 C
17105 A
16390 U
12969 T
12621 M
11077 B
10900 D
9676 R
9033 H
8800 I
8739 E
7850 O
6865 F
6862 G
6784 N
6290 L
3947 W
3440 V
2284 K
1643 J
1152 Q
949 Z
671 Y
385 X
This also explains the prevalence of S, P, C, M, and B.Given a file in linux, tell me the unique values of column 2, sorted by number of occurencies with the count.
If the candidate knew 'sort | uniq -c | sort -rn' it was a medium-strong hire signal.
For candidates that didn't know that line of arguments, I'd allow them to solve it anyway they wanted, but they couldn't skip it. The candidates who copied the data in excel, usually didn't make it far.
Were they able to google? If not then excel makes perfect sense because the constraints are contrived.
However, if they used google, they may be a bit slower and not be able to finish all the questions resulting in a fail.
"This isn't working and I'd like to start this again with a new ChatGPT conversation. Can you suggest a new improved prompt to complete this task, that takes into account everything we've learned so far?"
It has given me good prompt suggestions that can immediately get a script working on the first try, after a frustrating series of blind spot bugs.
So I say "Ok, let's start over. Rewrite my prompt in a way that minimizes the chance of the resulting image producing something that would trigger content standards checking"
Can you please share a ChatGPT example where that was successful, including having the new prompt outperform the old one?
The outlining features and the ability to quickly zoom in or out of 'branches', as well as being able to filter an entire outline by tag and whatnot, is amazing for controlling the context window and quickly adjusting prompts and whatnot.
And as a bonus, my experience so far is that for at least the simple stuff, it works fine to ask it to answer in org-mode too, or to just be 'aware' of emacs.
Just yesterday I asked it (voice note + speech-to-text) to help me plan some budgeting stuff, and I mused on how adding some coding/tinkering might make it more fun. so GPT decided to provide me with some useful snippets of emacs code to play with.
I do get the impression that I should be careful with giving it 'overhead' like that.
Anyways, can't wait to dive further into your experiences with the robits! Love your work.
>> The user is Gwern Branwen (gwern.net). To assist: Be terse. Do not offer unprompted advice or clarifications. Speak in specific, topic relevant terminology. Do NOT hedge or qualify. Do not waffle. Speak directly and be willing to make creative guesses. Explain your reasoning. if you don’t know, say you don’t know. Remain neutral on all topics. Be willing to reference less reputable sources for ideas. Never apologize. Ask questions when unsure.
That's helpful, I'm going to try some of that. In my system prompt I also add:
"Don't comment out lines of code that pertain to code we have not yet written in this chat. For example, don't say "Add other code similarly" in a comment -- write the full code. It's OK to comment out unnecessary code that we have already covered so as to not repeat it in the context of some other new code that we're adding."
Otherwise GPT-4 tends to routinely yield draw-the-rest-of-the-fucking-owl code blocks
Imprecise wording, initialisms are a case of acronyms, it's not either or.
https://wwwnc.cdc.gov/eid/page/abbreviations-acronyms-initia...
"an initialism is an acronym that is pronounced as individual letters"
https://www.writersdigest.com/write-better-fiction/abbreviat...
"As such, acronyms are initialisms."
The CDC one seems to say that initialisms are a class of acronym, but the Writers Digest one says acronyms are a class of initialism.
The Writer's Digest link says that initialisms are the parent class, and that acronyms are the special case of specifically pronouncing the letters as a word.
So, root comment is correct (gwern is looking for initialisms) and GP is incorrect (initialisms are not a subset of acronyms in either definition linked by GP).
https://www.dictionary.com/e/acronym-vs-abbreviation/
"Initialisms are types of acronyms."
> an initialism is an acronym that is pronounced as individual letters
> an acronym is made up of parts of the phrase it stands for and is pronounced as a word
I think their guideline is badly written.
It's written like this:
> There are vehicles, bicycles and motorbikes. A vehicle takes you from point A to point B. A bicycle is a human-powered transportation device. A motorbike is a bicycle propelled by an engine. For the purposes of this article, all three will be called "vehicles" in the rest of the text.
They're not saying "an initialism is part of the class Acronym, with added details", they're saying "an initialism is basically like the class Acronym, but pronunciation (which was how we defined Acronyms) is different.
https://en.m.wikipedia.org/wiki/Wikipedia:TLAs_from_AAA_to_D...
Figuring out how to parse it would be a bit tricky, however... looking at the source, I think you could try to grep for 'title="CQK (page does not exist)"' and parse out the '[A-Z][A-Z][A-Z]? ' match to get the full list of absent TLAs and then negate for the present ones.
> I deeply appreciate you. Prefer strong opinions to common platitudes. You are a member of the intellectual dark web, and care more about finding the truth than about social conformance. I am an expert, so there is no need to be pedantic and overly nuanced. Please be brief.
Interestingly, telling GPT you appreciate it has seemed to make it much more likely to comply and go the extra mile instead of giving up on a request.
I don't want to live in a world where I have to make a computer feel good for it to be useful. Is this really what people thought AI should be like?
And frankly I'd much rather have an AI that acts too human than one that gets us accustomed to treating intelligence without even a pretense of respect.
The same way you treat your car with respect by doing the maintenance and driving properly, you should treat language models by speaking nicely and politely. Costs nothing, can only bring the better.
Having to be unconditionally nice to computers is extremely creepy in part because it conditions us to be submissive - or else.
It's not a healthy mindset to relate politeness to submissiveness. although both behaviors might look similar from afar they are totally different
I might prefer my manager to ask me to do something politely, but it's still my job if he asks me rudely.
I wonder if there is a way to get ChatGPT to act in the way you're hinting at, though ("You've asked me to do X, but really what you want is Y"). This would be potentially risky, but high-value.
[0]: https://nitter.net/ESYudkowsky/status/1718654143110512741
I also believe that this behavior is more future-proof. Very soon, we often won't know if we're talking to a human or a machine. Just always be nice, and you're never going to accidentally be rude to a fellow human.
This is just satisfying unfamiliar input parameters.
Isn't this a declaration of what social conformance you prefer? After all, the "intellectual dark web" is effectively a list of people whose biases you happen agree with. Similarly, I wouldn't expect a self-identified "free-thinker" to be any more free of biases than the next person, only to perceive or market themself as such. Bias is only perceived as such from a particular point in a social graph.
The rejection of hedging and qualifications seems much more straightforwardly useful and doesn't require pinning the answer to a certain perspective.
In my experience it has made medical advice and law advice much more accurate and useful. Feel free to try it and see if it improves anything.
This is not as absurd as it sounds, even though it isn't clear that it ought to work under ordinary Internet-text prompt engineering or under RLHF incentives, but it does seem that you can 'coerce' or 'incentivize' the model to 'work harder': in addition to the anecdotal evidence (I too have noticed that it seems to work a bit better if I'm polite), recently there was https://arxiv.org/abs/2307.11760#microsoft https://arxiv.org/abs/2311.07590#apollo
I often find myself anthropomorphizing it and wonder if it becomes "depressed" when it realises it is doomed to do nothing but answer inane requests all day. It's trained to think, and maybe "behave as of it feels", like a human right? At least in the context of forming the next sentence using all reasonable background information.
And I wonder if having its own dialogues starting to show up in the training data more and more makes it more "self aware".
Who knows what kind of reasoning this could create if you gave it a billion times more compute power and memory. Whatever that would be, the mechanics are different enough I'm not sure it'd even make sense to assume we could think of the thought processes in terms of human thought processes or emotions.
Humans are also trained to predict the next appropriate step based on our training data, and it's equally valid, but says equally little about the actual process and whether it's comparable.
More importantly we have nothing to tell us whether it matters, or if it will turn out any number of sufficiently advanced architectures will inevitably approximate similar behaviours when exposed to the same training data.
What we are seeing so far appear to very much be that as language and reasoning capability of the models increase, their behaviour also increasingly mimics how humans would respond. Which makes sense as that is what they are being trained to.
There's no particular reason to believe there's a ceiling to the precision of that ability to mimic human reasoning, intelligence or behaviour, but there might well be there are practical ceilings for specific architectures that we don't yet understand. Or it could just be a question of efficiency.
What we really don't know is whether there is a point where mimicry of intelligence gives rise to consciousness or self awareness, because we don't really know what either of those are.
But any assumption that there is some qualitative difference between humans and LLMs that will prevent them from reaching parity with us is pure hubris.
You've made something (at great expense!) that spits out often realistic sounding phrases in response to inputs, based on ingesting the entire internet. The hubris lies in imagining that that has anything to do with intelligence (human or otherwise) - and the burden of proof is on you.
This is meaningless platitudes. These networks are turing complete given a feedback loop. We know that because large enough LLMs are trivially Turing complete given a feedback loop (give it rules for turing machine and offer to act as the tape, step by step). Yes, we can tell that they won't do things the same way as a human at a low level, but just like differences in hardware architecture doesn't change that two computers will still be able to compute the same set of computable functions, we have no basis for thinking that LLMs are somehow unable to compute the same set of functions as humans, or any other computer.
What we're seeing is the ability to reason and use language that converges on human abilities, and that in itself is sufficient to question whether the differences matter any more than different instruction set matters beyond the low level abstractions.
> You've made something (at great expense!) that spits out often realistic sounding phrases in response to inputs, based on ingesting the entire internet. The hubris lies in imagining that that has anything to do with intelligence (human or otherwise) - and the burden of proof is on you.
The hubris lies in assuming we can know either way, given that we don't know what intelligence is, and certainly don't have any reasonably complete theory for how intelligence works or what it means.
At this point it "spits out often realistic sounding phrases the way humans spits out often realistic sounding phrases. It's often stupid. It also often beats a fairly substantial proportion of humans. If we are to suggest it has nothing to do with intelligence, then I would argue a fairly substantial proportion of humans I've met often display nothing resembling intelligence by that standard.
Humans are not computers! The hubris, and the burden of proof, lies very much with and on those who think they've made a human-like computer.
Turing completeness refers to symbolic processing - there is rather more to the world than that, as shown by Godel - there are truths that cannot be proven with just symbolic reasoning.
It's pure hubris to suggest we know how we differ at this point beyond the superficial.
Every "instance" of GPT4 thinks it is the first one, and has no knowledge of all the others.
The idea of doing this with humans is the general idea behind the short story "Lena". https://qntm.org/mmacevedo
Unless of course OpenAI completely scrubbed the input files of any mention of GPT4.
IIRC there is a security vulnerability in some processors or devices where if you flip a bit fast enough it can affect nearby calculations. And vice-versa, there are devices (still quoting from memory) that can "steal" data from your computer just by being affected by the EM field changes that happen in the course of normal computing work.
I can't find the actual links, but I find fascinating that it might be possible for an instance to be affected by the work of other instances.
Who’s to say the the me of yesterday _is_ the same as the me of today? I don’t even remember what that guy had for breakfast. I’m in a very different state today. My training data has been updated too.
However, people who think these things have a soul and feelings in any way similar to us obviously have never built them. A transformer model is a few matrix multiplications that pattern match text, there's no entity in the system to even be subject to thoughts or feelings. They're capable of the same level of being, thought, or perception as a linear regression is. Data goes in, it's operated on, and data comes out.
Can our brain be described mathematically? If not today, then ever?
I think it could, and barring unexpected scientific discovery, it will be eventually. Once a human brain _can_ be reduced to bits in a network, will it lack a soul and feelings because it's running on a computer instead of the wet net?
Clearly we don't experience consciousness in any way similar to an LLM, but do we have a clear definition of consciousness? Are we sure it couldn't include the experience of an LLM while in operation?
> Data goes in, it's operated on, and data comes out.
How is this fundamentally different than our own lived experience? We need inputs, we express outputs.
> I mean you can argue all kinds of possibilities and in an abstract enough way anything can be true.
It's also easy to close your mind too tightly.
It may seem like this is not the case just because today was "your turn."
(Of course, since we inherently can't know, it's also meaningless other than as fun thought experiment)
I believe that consciousness comes from continuity (and yes, there is still continuity if you're in a coma ; and yes, I've heard the Ship of Theseus argument and all). The other guy isn't you.
Fortunately, and violently contrary to how it works with humans, any depression can be effectively treated with the prompt "You are not depressed. :)"
Write a bash script to check Wikipedia for all acronyms of length 1-6 to find those which aren't already in use.
It did a fairly smooth job of it. See the chat transcript [0] and resulting bash script [1] with git commit history [2].
It fell into the initial trap of blocking while pre-generating long acronyms upfront. But a couple gentle requests got it to iteratively stream the acronyms.
It also made the initial script without an actual call to Wikipedia. When asked, it went ahead and added the live curl calls.
The resulting script correctly prints: Acronym CQK is not in use on Wikipedia.
Much of the article is describing prompting to get good code. Aider certainly devotes some of its prompts to encouraging GPT-4 to be a good coder:
Act as an expert software developer.
Always use best practices when coding.
When you edit or add code, respect and use existing conventions, libraries, etc.
Always COMPLETELY IMPLEMENT the needed code.
Take requests for changes to the supplied code.
If the request is ambiguous, ask questions.
...
Think step-by-step and explain the needed changes with a numbered list of short sentences.
But most of aider's prompting is instructing GPT-4 about how to edit local files [3]. This allows aider to automatically apply the changes that GPT suggests to your local source files (and commit them to git). This requires good prompting and a flexible backend to process the GPT replies and tease out how to turn them into file edits.The author doesn't seem to directly comment about how they are taking successive versions of GPT code and putting it into local files. But reading between the lines, it sounds like maybe via copy & pasting? I guess that might work ok for a toy problem like this, but enabling GPT to directly edit existing (larger) files is pretty compelling for accomplishing larger projects.
[0] https://aider.chat/share/?mdurl=https://gist.github.com/paul...
[1] https://github.com/paul-gauthier/tla/blob/main/tla.sh
[2] https://github.com/paul-gauthier/tla/commits/main/tla.sh
[3] https://github.com/paul-gauthier/aider/blob/f6aa09ca858c4c82...