New Capabilities for GPT-3: Edit and Insert
openai.com
openai.com
When Jane saw mister Bingley,
She knew that he was fine,
But it took her a while to realize
He wanted to be more than friends,
Then she met Fitzwilliam Darcy,
And she thought he had a lot of money,
But it took her a while to realize
He didn't want to talk about money,
She was just a young girl,
But she knew what was going on,
She knew who she wanted to be with,
But she also had her doubts,
She knew he was great,
But she also thought he was a snob,
She was afraid that,
He was going to break her heart.
She loved him,
But didn't know if she could trust him,
So she decided to wait,
Until his feelings became clear.
Then all of a sudden,
She realized,
That she had been waiting all this time,
And that she was in love.
So she decided to go after him,
And she got to him just in time,
And they lived happily ever after.
It's a terrible rap but a fairly impressive summary of the book (from someone that's admittedly never read it)Odd that it took "rewrite (the opening paragraph) as a rap" to mean "summarise the whole book".
It's beyond me that Facebook squandered its data value by half in not including even a privately tracked 'dislike' button.
The value in what people both like and dislike is going to be massive in the next few years, as it can help train the evaluation of content in both general and personalized scales.
I'd wager with where the tech is today we Tinder could generate synthetic matches that would be far preferred over any natural matches, even at equal or lower physical attractiveness scores in generated photos.
Pandora might be relevant again for music soon as their like/dislike data could allow for the generation of new music in line with preferred tastes and not just selection.
Transistor models like GPT-3 are only part of the equation, and as we continue to improve not only generation of new data, but evaluation of data too -- the application of the technology is likely going to move faster than anything we've seen before in the next few years.
Codex generated 35% of the code written by people with access to it, and Microsoft and OpenAI have the data on what was and wasn't selected. That data alone could produce an AI linter that could have significant value in catching sophisticated natural programmer errors - but in combination with Codex too?
(Keep in mind it was only summer 2020 when the tweet showcasing normal GPT-3 writing HTML was blowing minds, and only last summer Copilot was announced.)
We really haven't seen anything yet.
Things are moving so fast that anyone building products relying on 3rd party AI at a foundational level right now should absolutely be concerned with how they are planning around obsolescence within ~3 years for whatever they are building on top of in their product design and implementation.
(Not 'Transistor' - what I get for typing it out on my phone and not proofreading)
This is why that rap looks so weird: so high-quality in general, yet so totally lacking a critical element we expect from rap. GPT-3 has a pretty good idea how to write rap, but it just doesn't know what words sound like. It is blind to that. It can't see that difference between rap and poems. To it, rap is just a sort of poetry, and some words are chosen for reasons it can't understand, no matter how many rap lyrics or poems it looks at.
The better the models get at writing, the more this difference throws us for a loop: "how can it be so good at ABC, and yet fail so totally at D?"
Doesn't matter how many parameters it has; and as a matter of fact, the GPT models have not improved their rhyming noticeably even as they go from a hundred million parameters to over a hundred billion parameters.
(And I have extensive evidence consistent with my interpretation at the link, ranging from several years of failure by many people to prove me wrong by eliciting rhymes even a fraction as good as its non-rhyme poetry to fixing GPT-3 performance on tasks by working around BPEs to different models by different groups but the same data encoding also showing the same utter failure to rhyme.)
Now if you ran an experiment comparing MLM (or any LM) on rhyming tasks with different encodings, then you could certainly make a statement like "BPE is worse at rhyming than other encodings" and it would be scientifically supportable. And that very well might be true. But your extreme conclusion is not supportable.
A rhyme dictionary would still not replicate human rhyming capabilities. Think about neologisms or misspellings. A model can memorize every single entry in the rhyming dictionary (and let's say this somehow cashes out as apparent rhyming proficiency in being able to for every word recall an entry in the rhyme dictionary of valid other words), but it would not be able to write something like "Jabberwocky" inventing a bunch of new words or phrases or names which rhyme. (How would it know to rhyme "wabe" and "outgrabe" when they appear in no dictionaries - because they were just invented?) A model which has "actually" learned to rhyme would be able to take new words (not necessarily invented by it, but possibly invented by humans after it was trained, or invented on the spot for a prompt, or part of a new fictional work like worldbuilding) and rhyme them appropriately. A model which has memorized a rhyming dictionary would not.
At least that's what I am getting from Gwern's write-up. I might have misunderstood Gwern, or Gwern might be wrong, of course.
And eg it would probably perform worse at translating from English to French than a naive system. (Unless you preprocess your corpus extensively to figure out when you get English and when you get a snippet of French.) GPT-3 is surprisingly good at translating English to French.
Another problem is that English has not just one text-to-phonetic transcription: different accents pronounce words differently. For just one example:
> n most non-rhotic accents, if a word ending in written "r" is followed immediately by a word beginning with a vowel, the /r/ is pronounced, as in water ice. That phenomenon is referred to as "linking R". Many non-rhotic speakers also insert an epenthetic /r/ between vowels when the first vowel is one that can occur before syllable-final r (drawring for drawing). The so-called "intrusive R" has been stigmatized, but many speakers of Received Pronunciation (RP) now frequently "intrude" an epenthetic /r/ at word boundaries, especially if one or both vowels is schwa. For example, the idea of it becomes the idea-r-of it, Australia and New Zealand becomes Australia-r-and New Zealand, the formerly well-known India-r-Office and "Laura Norder" (Law and Order). The typical alternative used by RP speakers (and some rhotic speakers as well) is to insert a glottal stop wherever an intrusive R would otherwise have been placed.
https://en.wikipedia.org/wiki/Rhoticity_in_English
Btw, even without any explicit pronunciation data, I would expect the system to get a pretty good idea of how pronunciation works by essentially using the same technique human linguists use to reconstruct old pronunciations:
They observe what kinds of mistakes people make.
Eg when you see people mixing up "there" and "their" and "they're" in the corpus, that tells you that in modern English these three are pronounced almost the same.
From spelling mistakes in words like carburetor you can figure out that unstressed vowels in English are pretty much all pronounced the same: as a schwa. https://en.wikipedia.org/wiki/Schwa
You can also learn a lot from observing rhymes.
Your example of their/they're/there has some data because the whole words can be mistaken for each other, but even if you take a billion pages of prose, you won't get data to deduce that 'there' rhymes with 'pear' but not 'here', that 'great' rhymes with 'straight' but 'cough', 'dough' and 'through' do not rhyme. A model can't learn something that's simply not represented in the training data.
So you either have to bring in external information (i.e. pronunciation models, or dictionary data with pronunciation guides, or audio recordings) or you have to have sufficient rhyming text inside your training data - i.e. train it on corpora of poetry instead of prose.
Also, I'm not sure if the variation of pronunciation is critical - any accent variations that affect every instance of some sound similarly would preserve rhyming. There are certain differences e.g. I recall a discussion about some parts of Shakespeare which rhyme perfectly in pronunciation of that period but do not in modern English, but I think that the variations of modern English accents should be mostly OK from that perspective.
Yes, the latter. Just mix some rhyming text into your corpus. If you take eg Wikipedia, there's already plenty of rhyming stuff in there. Nursery rhymes, song lyrics, etc. Similar for other corpora.
Because training is largely running inference, and then correcting any mistakes. Over and over again.
(I say largely, because inference also does stuff like sample top-k completions etc.)
In the end, aside from boilerplate, I spend most of my time in Qt Creator (which doesn't have a Copilot plugin) rather than VS Code, so I mostly stopped using Copilot anyway.
When I'm problem solving or not sure how to code something though, its constant suggestions are just noise, and then I turn it off.
If only we had programming languages that didn't force us to write boring code in the first place ...
(I probably remember that story all wrong. But I like it anyway!)
Overall, a tremendous time saver.
It's a little like pair programming with an incredibly eager junior developer who has read a lot of the documentation of every popular API in the world. I need to review the code it produces, but it's very fast, and its suggestions are usually great.
It's annoying when I know exactly what I want to write, and most helpful when I'm unsure (either because I'm trying things out, or if I'm using a new API or a language I'm rusty at).
I can write a comment outlining what I want a function to do, and 90% of the time it will generate the code I need (or very close with a couple of small tweaks needed).
Coincidentally, my Stack Overflow visits have decreased by approximately 90%.
PS - it is also AMAZING with CSS and saves so much time on easy things that I just haven't memorized. "Style the LI so there are no bullet points" and boom...'list-style: none;'
It’s also quite impressive how Copilot will learn from any patterns you’ve typed on previous lines — so when I’m using a certain color name in a variable, it knows that I likely want to use it in subsequent code.
Or, more concretely, when using tokens for colors (e.g. blue.50 for light blue, blue.900 for dark blue) it can figure out that I probably want my background to be an x.50 colour and my text to be, say, x.700. So cool!
I would have expected the demonstration that your work is easily replaced by an AI to have the opposite effect. :)
It's really good for boilerplate. Things like TDD tests where I'm just modifying a few parameters. You can get it to write functions like "parse this DateTime object into a format like Tuesday, 15 May 2020".
It's a useful lookup too. Like often I just want to extract a variable from a List and would spend 15 mins looking up the docs or sifting Stack Overflow. Codex is more faster and accurate.
With GPT-3 it's garbage in, garbage out. You have to invest a few days in learning the prompts that work.
It's possible to set up an API to it, but it's not yet at the point where it's worth doing. I think it's possible to get Codex to write the script, now that I think about it.
It prompted the (joke) thought that perhaps it is making me less productive because of how often I end up sitting back and marveling at how amazing it is. I really can’t believe how good it is.
It is dumb but very good at doing repeating things, or being a very smart auto-complete. For writing tests, it is actually quite decent.
I find them to be an annoying Clippy-like companion that interrupts my train of thought and introduces a whole separate class of programming bugs: "autocomplete errors" which are like copy & paste errors, but written by another programmer.
I like my auto-complete functionality built in to Jetbrains and Visual Studio for hinting at variable names, function parameters, classes, imports, and so forth. But the boiler-plate code that "assistant AI programmers" provide are not worth the effort at this time and do not live up to the hype. I think they help less experienced developers out, but the kind of code I am writing won't usually be found in Copilot or Tabnine. I could see the appeal if all I was doing was churning out boilerplate CRUD app code all day long, which honestly, I would outsource to Craigslist for thirty bucks an hour. Or just hand off to my assistant.
It's quick to scan and ignore things that aren't right, and it's either completely right or close enough that it definitely feels like a timesaver.
The best parts are where it's doing something long-winded but fairly straight forward (e.g. assigning variables). But it has moments of shocking ability with more complex things.
Now I can't live without it.
But you need to set the right expectations that it is not magic, which requires some tuning.
Both the edit endpoint (used in those tweets) and the insert endpoint have mixed performance and tend to go off-the-rails often, especially compared to the new "Instruct" models which do a much better job, although expensive while these new endpoints are in a free beta.
The coding endpoints are slightly better but a more narrow domain.
In all cases I recommend looking at the docs for examples.
Also, the edit/insert endpoints specifically should hypothetically help a lot with plot divergence, which has been a huge problem trying to generate long-form works with standard completions, even with top-down outline expansion strategies, scene transition metadata, etc.
Excited to see what the millions of "AI word processors" that've sprung up over the past year actually do with it, besides the obvious.
I'm excited about the potential of things like this, I just don't like this wish-wash way of presenting a technical tool.
the U.S. military most likely having an interest in suppressing information to avoid someone else creating AI first.
someone else creating ai first? You make it sound like they are trying to create general ai skynet... there are strong incentives for secrecy (trade secrets)
they do not have the right to call themselves "open" if they are focusing on secrecy. They are trying to paint themselves in a positive opensource light when really they are no different then any other corporate overlord hiding in the shadows. Open implies developed in the open, free for use + modification, and/or key details NOT being hidden from the public.
Simply sharing details of things created or research done is no different than what researchers at any other ABC corporation/organization do.But it definitely saved me time. It's won't write your software for you. It's more like autocomplete but much, much better.
As larger requests (ex: spellcheck or translate 1000 words) can cost 20c+ each and the model often requires to be ran 2 or 3 times before giving an acceptable answer on more complex tasks (ex: translation), that makes for a very expensive tool if you want to edit a significant amount of text.
I wonder if the free usage only apply to the Codex engine and not the GPT-3.
(I realise this sounds like I’m making it up, but I promise this is a real story. It was quite fun.)