Show HN: Humans vs AI – A/B testing GPT-3
vwo.com
vwo.com
As an A/B testing company (VWO), it's exciting to see how much effective is GPT-3 generated copy against human copywriters on live websites.
I like to think of this as the Turing test on the web :)
The quality of generated headlines, buttons and product descriptions seem very good, so we're hopeful that AI will at least score a few statistically significant wins.
I encourage you to participate in the competition (you don't have to use VWO for A/B testing - you can use your existing stack).
If you have any feedback or comments, happy to discuss.
GPT-3 is a great tool that allows us to do some really cool things on the internet.
Here are just a couple of them:
It helps us generate interesting headlines for our posts. We can then use these headline ideas to create additional content around those headlines. This gives us a lot of flexibility when it comes to writing posts.
We can easily add social media links into our articles. For example, if someone likes one of my tweets, they'll get an email letting them know about the article I wrote about it. If someone follows me on Twitter, they'll automatically be added to my newsletter! That's pretty awesome right?
And lastly, we can use GPT-3 to do A/B testing! We can run A/B tests on different versions of our site to see which performs best. We can also run A/B tests on specific pages within our site.
So if you've got any questions about GPT-3, feel free to ask away! I'd love to hear all about it!
GPT-3 OUT!" You press submit and feel pretty good about your first post to HN. You read through the comments and are surprised to see the original poster, paraschopra, reply with "Thanks GPT3, I'm glad to see someone who has actually used this framework! :)"
But it'll detect human written one too. So expect many false positives.
return true;Please help us make you job redundant. Friendly indeed.
If not, then its pretty easy for GPT-3 to copy paste existing human written texts and just prove that it can write like a human.
My god, they've finally worked out that humans haven't written anything original since the dawn of the internet.
Only 1/200 or so tweets that I checked had a verbatim dupe.
So the profundity comes first, and I can't always find an inspiration for it. With GPT-3 in particular, that would be the exception rather than the rule.
Fun challenge though !
Let's see which one gets more participation :)
AI generated one is "Variation 1" v/s "Control" which was written by me https://imgur.com/a/pcbTGwR
> Title: United Methodists Agree to Historic Split
> Subtitle: Those who oppose gay marriage will form their own denomination
> Article: After two days of intense debate, the United Methodist Church has agreed to a historic split - one that is expected to end in the creation of a new denomination, one that will be "theologically and socially conservative," according to The Washington Post. The majority of delegates attending the church's annual General Conference in May voted to strengthen a ban on the ordination of LGBTQ clergy and to write new rules that will "discipline" clergy who officiate at same-sex weddings. But those who opposed these measures have a new plan: They say they will form a separate denomination by 2020, calling their church the Christian Methodist denomination.
> ...
First it suggests the spinoff of a new "theologically and socially conservative," denomination, but then it is the liberal minority that is expected to form a new denomination. The paper[0] acknowledges occasional non-sequiturs, but more pertinently, how did 88% of judges let it slip?
I think the cause is more that some religious group i dont care about having a schism isn't very interesting. Maybe i care that they're being homophobic, but i definitely don't actually care which side is forking and which side is staying as the original group. So I don't fully pay attention to the details unless i really force myself.
1) Humans often make logical errors
2) In the religious-split context, who is the original denomination and who is the splitting denomination is inherently fraught/subjective/liable-to-contradiction/open-to-changing-contextualisation, so its not a surprising mistake. The readers may have just assumed it was the author/speaker/washington post getting mixed up or projecting their own opinion, which may have contrasted with that expressed from the reports/attendees of the 'actual conference'.
(plus there's a good chance none of them could give two hoots about, or have any specific knowledge or interest in methodists...i have a religious studies degree and my eyes are already almost glazing over just at the mention of them to be honest :P)
Makes you think, how often does this happen with other writing?
Think of the way you read something that you strongly disagree with. I would bet you check every fact and constantly evaluate the strength of the argument, and I would bet it takes you an hour to get through a single article. That is what it takes to deeply understand something. We rarely read that way.
GPT-3 can write the content-free babbling that C students use to pad papers. It can write things that appear coherent if you don't actually read them.
For example your processed post could look like that:
- Author claims his fast reading speed at 800 wpm
- 400 wpm with good comprehension
- 50 words when reading carefully, with effort
- Claims people read that last way when disagreeing with the premise
- Claims that it takes an hour to read one article
- Claims that it is rare
- Claims GPT-3 can write filler text like bad students can
- Claims it can appear coherent
Those can be automatically cross checked with each other and external knowledge base, duplicate reworded posts could be found, and even humans could read those more carefully and detect problems because claims and logic are already extracted from the text. I suspect this is what we do internally when reading things, but with more speed comes more things to keep track of, which is hard so we tend to pattern match word salads instead, because we have wetware optimizations for that.
I want to see that in the next generation of content blockers. :)
Ideally it should also support hangouts because the romance scammers that spam me almost every day with a new address always ask to move to an hangout conversation in their initial message.
Automated job applications
and
Automated job listings
Companies will use GPT-3 to generate job listings, and some company will curate a big database of good job applications (i.e those that have landed someone a job), and make a service where you feed in the listing, and out comes job application / letter.
A lot of posters here seem to complain (or have, for the past weeks) the GPT-3 output comes off as mediocre, but I'd wager that a lot of the job applications we see today are worse than mediocre, often times horrible.
It sucks that the end result would be some standoff between ML-software that reads job applications, written by some other ML-software, but there's lots of hours to be saved - from the human standpoint.
If I don't have to write 10 letters, then that's probably 10 hours saved from my part - as I easily use an hour to write a tailored job letter.
Seems to imply that there will be a difference (by rejecting the null hypothesis). But not rejecting the null should actually count as a win for GPT3.
I don't want to think at the implications if it turns out GPT3 is better.
My guess is it will work, GPT will be hard to distinguish from human copywriting, for a lot of everyday items. It will only be found out when there's some deeper logic involved in the sale, for instance if you were to try to have it pretend to be a B2B salesman.
For short texts, it's almost too good to be true though. So for a use case like web copywriting or giving quick answers to questions, it holds a lot of promise.
More than A/B testing this might be a better fit for web site building tools like wix.com and and webflow.com
The dictator in my country is already paying stupid people to write stupid comments, GPT-3 is already over their level.
Maybe a future debate format should include not only human candidates but a few instances of GPT-3 as well?
Though with another some context and different priming I guess I'd get a different answer.
E.g. here is amazon.de https://imgur.com/a/5JesS3Y
I have no idea if generated recommendations are good. I don't know German. Perhaps someone who knows can comment.
You can also do things like make a characters that speak one language and another which knows both languages, then the German guy will write in German and the other can answer in English.
I always wanted a "tiered" tl;dr functionality that would allow me to collapse text into a tree-like structure with the most important content on top and filler at the leaves. And please, please package it as a browser addon.
-- rationale -- There are plenty of articles inflated due to autor being paid per kilogram of used ink. Or a book author that was arm twisted into inflating a perfect 100 page book into an unreadable 400 page monstrosity noone is able to follow without mind wandering.
Of course there is a question on how to achieve a working tl;dr - the "old" way would be to manually summarise articles on a number of conciceness levels and use that as training data. Or to use some existing summary services as source.
Perhaps there is a better way? If we could run the GPT-3 backwards (inverse) (^-1 ??). GPT-3 can "produce text" given a start cue, in reverse it would "remove text".
It's not downvote worthy, but it is cringey.