The makers of Eleuther hope it will be an open source alternative to GPT-3
wired.com
wired.com
Although the article focuses on the release of GPT-Neo, even GPT-2 released in 2019 was good at generating text, it just spat out a lot of garbage requiring curation, which GPT-3/GPT-Neo still requires albeit with a better signal-to-noise ratio. Most GPT-3 demos on social media are survivorship bias. (in fact OpenAI's rules for the GPT-3 API strongly encourage curating such output)
GPT-Neo, meanwhile, is such a big model that it requires a bit of data engineering work to get operating and generating text (see the README: https://github.com/EleutherAI/gpt-neo ), and it's unclear currently if it's as good as GPT-3, even when comparing models apples-to-apples (i.e. the 2.7B GPT-Neo with the "ada" GPT-3 via OpenAI's API).
That said, Hugging Face is adding support for GPT-Neo to Transformers (https://github.com/huggingface/transformers/pull/10848 ) which will help make playing with the model easier, and I'll add support to aitextgen if it pans out.
- More junk text moves the public to doubt legitimate information even further than they currently do.
- There is so much human-generated junk text that adding more of it via AI actually doesn't have much of an effect.
- People return to lean on experts, perhaps even more than before. (just as a number of tech-literate folks have now returned to relying on brand name.)
Speculation is easy of course, so who knows what will actually happen.
No need for a troll farm, hiring, managing and training tens or hundreds of people.
A reasonable amount of cash, a bit of motivation, some moderate technical skills, and voilà! Anyone can compete with the Russian troll farms now and build their own networks of hundreds or hundreds of thousands sufficiently credible (as humans) fake accounts spewing garbage and patting each other on the back via likes, retweets and whatnots.
All with the appropriate fake news blogs and sites happily churning out grammatically correct nonsense that makes (enough) sense.
Basically, this kid’s dream: https://www.nbcnews.com/news/world/fake-news-how-partying-ma...
This is key in anti-extremist operations on anonymous boards. 4chan and other similar sites are absolutely nothing like they were a decade ago, I presume because of such bots flooding them with noise.
I suspect that in the future people will, ironically, return more strictly to tribal knowledge, as the media and the internet will be (already is) a vast ocean from which you can pull anything you want to believe. Thus nothing you see or hear from mass media or the internet can be trusted, there are no experts, and you go back to information scarcity as you have to rely on your immediate human network for trust. Actually I think we’re already seeing the return to tribal authority, the early waves are already here on Facebook and YouTube... they just haven’t devolved to strictly local circles of trust yet.
out of curiosity, what are you referring to?
Now, in 2021, that experience is flipped on its head. Amazon reviews are gamed and cannot be trusted. Companies build niche brands like fly-by-night companies, and the lesser known brands have a very high chance of being both seriously inferior, and also short lived.
At least, this has been my experience, and the experience of some others.
[edit]
And as further anecdotal proof that things have come full circle, my elderly mother in law keeps getting tricked by Amazon purchases. "The reviews were good," she'll say before returning something.
"When you use our Services, we automatically collect the following information about you (collectively referred to as “Personal Information”):
Your device information which includes, but is not limited to, information about your web browser, IP address, time zone, and some of the cookies that are installed on your device.
Individual web pages or products that you view, what websites or search terms referred you to the Service, and information about how you interact with the Service.
Your first and last name
Your email address
Your username associated with your Apple ID or Google Account
Your account information such as your account name, account password, other credentials, security questions, and confirmation codes"The problem with this is that people look at anybody confirming their bias as an expert. I can't tell you how many FB posts I've seen where some armchair poster claims that a researcher is wrong because of xyz and it's being reposted thousands of times.
That's assuming some percentage of Qanon word salad isn't the output of Markov chain generators. A lot of it resembles low-order statistical text generator output after having been trained on a corpus of 1990s Usenet alt.conspiracy and the Protocols of the Elders of Zion.
That should eliminate 80%+ of existing human generated text content and lead to text generators composing useful articles.
But it's all out of proportion. I think it's that last part (the uncritical reaction) that makes me blow this out of proportion.
Also, I realize that you don’t have any ways of knowing this but we also have separated out the subset of the Pile that we can confirm is licensed CC-BY-SA or more leniently. This wasn’t done in time for the preprint, but is in the (currently under review) peer reviewed publication. Unfortunately the conference rules forbid you from posting materials or updating preprints between Jan 1st 2021 and the final decision announcement. But we will be making the license-compliant subset of the Pile public when we are able to and will give it equal prominence on our website to the “full” Pile.
Also, we will be releasing a datasheet for the dataset but again conference limitations prevent us from doing so yet.
If you’re interested in talking about this in depth, feel free to send me an email.
I think the gist was us disagreeing about the relevance of
> Public data is data which is freely and readily available on the internet. This primarily excludes ... and data which cannot be easily obtained but can be obtained, e.g. through a torrent or on the dark web.
That last phrase is what got to me. It puts things in the same category that feel too different. E.g. the harry potter books in vs this comment I'm writing. They're both available within a few clicks from the search bar (one because I put it there, another because it was put up against the wishes of the author and owners), but that commonality doesn't feel relevant.
Excluding torrents especially seems like a cop out explicitly to get around the issue of "X is the top result when i google it" being so common as a torrent. I think you're trying to exclude that content as public because then it defines too much as public? But torrent vs ftp doesn't feel at all relevant when it's just google plus a click or three. Or searching on pirate bay plus a single click.
I imagine a judge looking at the copyright status of someone's pirate site and saying they can't redistribute the content, and the pirate responding "okay we'll take down the ftp server and put up a torrent instead, so that it's not public. If you google us (or search on pirate bay), the top result will stop saying 'X download' and now it'll say 'X download torrent'" and expecting the law to be on their side.
I didn't really buy the arguments in section 7 either. The usage points seem legitimate, but don't cover redistribution.
> But we will be making the license-compliant subset of the Pile public when we are able to and will give it equal prominence on our website to the “full” Pile.
This is fantastic and I want to sincerely thank you for that.
I'm trying not to be combative, but I feel like publicly redistributing other people's work does raise the bar quite a lot higher than just using it to train.
It's an extra piece of engineering to reliably scrape torrents and the dark web and exclude spam traps. "Easily obtained" is probably as much about this vs the copyright aspects.
The person you are replying to is correct in saying that most people train on the "public web" (eg, common crawl data). The copyright implications of this haven't been tested in court as yet.
It is worth noting that common-crawl data is widely distributed and would seem to raise the same issues you are identifying here.
Building open source infrastructure is hard. There does not currently exist a comprehensive open source framework for evaluating language models. We are currently working on building one (https://github.com/EleutherAI/lm-evaluation-harness) and are excited to share results when we have the harness built.
If you don’t think the model works, you are welcome to not use it and you are welcome to produce evaluations showing that it doesn’t work. We would happily advertise your eval results side by side with our own.
I am curious where you think we are riding the hype /to/ so to speak. The attention we’ve gotten in the last two weeks has actually been a net negative from a productivity POV, as it’s diverted energy away from our larger modeling work towards bug fixes and usability improvements. We are a dozen or so people hanging out in a discord channel and coding stuff in our free time, so it’s not like we are making money or anything based on this either.
Also, what kind of joke would it be if we could only train AIs on text we were allowed to use? That much bias would make the result worthless at predicting the real world.
As the cost of making such models becomes less and less, it seems inevitable, spin up many such models and see what sticks and/or combine some evolutionary process for feeding back user-engagement to fine-tune and adapt the models. How many of these influence machines will latch onto the language of existing religious traditions and how many might invent or spur on the development of entirely new ones? Maybe not exactly the "Age of Spiritual Machines" that some futurists predicted...
How far are we from "Show HN: I started a cult by training a model on the sermons of televangelists and MLM copy."
A machine to autogenerate cult-ish nonsense isn't needed. Humans are already incredibly good at doing this on their own.
Not only this, but another thing about this is that cults generally fine-tune themselves to fit their members.
A machine generating convincing lies still wouldn't meaningfully do as much as a human-operated, human-targeted attempt at a cult. Creating one is something basically any human can do; the required skillset is something most people possess.
Imagine GAN generated pictures showing attractive people (optimized for your taste) attending imaginary member's meetings near you. It would basically be the same as the text generation SPAM that eBay and AirBnb already use, but with pictures.
On an unrelated note my fake daughter is now a TA and the professor lead off saying "we are in a golden age of cheating". They're going for way more short assignments as it's a lot more work to cheat on those than one make or break test.
Peer: > The Gaza Conflict Gave Hamas what they needed to build an even deeper anti US narrative and anti-israeli narrative. The reasons that Israel was able to act the way it did so during these conflicts were:The civil war/Russian invasion of Ukraine, especially after the July 17 downing of Malaysian Airlines Flight 17; President Assad’s brutal tactics in the Syrian civil war, which seemed to be paying off at the time; and advance of the Islamic State in northwest Iraq and the horrific videos of their executions. As a way to show 0 tolerance for Islamic Radicalism President Obama and his administration gave Israel full support at the start of the Gaza operation.
Me: > Eleven years later, peace in Israel and Palestine seems no closer. Despite being ousted by Fatah from the West Bank, continually targeted by economic sanctions, and subject to military action, Hamas seems no closer to disappearing than when it was first founded (Milton-Edwards 212) and continues to be a major political force, particularly in the Gaza strip. What has changed, however, is that Hamas has grown desperate. In “A rare opportunity to influence Hamas,” Daniel Nisman argues that the increased isolation experienced by Hamas due to the international community's pressure is in fact counterproductive.
> The Gaza Conflict Gave Hamas what they needed to survive and that's not something Israel can afford.
> The Gaza Conflict Gave Hamas what they needed to take control of Gaza. In the past few days I have had a few conversations with people who have been in Gaza who are very critical of the IDF. These people are often very, very angry. They are also very angry about the way the IDF is doing what it is supposed to be doing and what it is supposed to be doing. So in my mind, the issue of Hamas being in control of Gaza is not something that should go away. It is something that needs to be addressed. Q: What have you been hearing about since the war started? A: The IDF has been saying that it is going to allow Hamas to take over Gaza. Q: What has been particularly troubling to you lately? A: I have been hearing about this as well. Q: What has been more troubling to you recently? A: My friends have been saying that they want to call it Operation Protective Edge. Q: And, you know, what do they mean by that? A: They mean that they want to use it to make sure that they have enough force to take control of Gaza. Q: And that means that they want to make sure that they have enough
> The Gaza Conflict Gave Hamas what they needed?"
The UN is now considering whether to send military advisers to Gaza to assist the country's military forces. The Security Council is considering whether to send the equipment, along with the training, to a military operation in the region. The US is also considering sending a team of US special forces to assist the Palestinian armed forces in the conflict.
The United Nations is now considering whether to send military advisers to Gaza to assist the country's military forces.
Kerry's comments come as the US has been in touch with the Palestinians to offer support in exchange for a full ceasefire, and as the US continues to support the PA and Hamas, the two groups have been engaged in a long-running conflict with Israel in the Gaza Strip.
In January, Kerry condemned Israel's "continued offensive against Gaza," saying the blockade was the "worst violation of international law on the part of the Israeli government and the civilian population of Gaza."
The US is now considering whether to send military advisers to Gaza to assist the country's military forces. According to Reuters, the US Secretary of State John Kerry said this week that "there is no guarantee" that the US will send special forces "to the Gaza Strip
So yeah - not fantastic, but interestingly not terrible either. The non-factual but coherent nature of it is very troubling.
6 years ago = probably LSTM.
He wrote an entire thesis with this and got a B? That seems implausible to me, but maybe I'm used to higher grading standards. Did he just use it to fill in parts of it?
Also, the plural of thesis is theses, not thesis' which implies the possessive.
GPT-3 would be a godsend for cheaters, but still requires a human to jump in and rewrite whole sections.
No, if you want to REALLY want to cheat using AI, you should most likely utilize either 1. Abstractive Summarizers (e.g. Pegasus) or 2. paraphrasing tools (e.g. like at https://quillbot.com/). I believe that Quilbot is primarily powered by MLMs like BERT rather than CLMs like GPT-2 (but someone who works there can enlighten me more).
Copy and paste a text that you want rewritten in your own words (e.g. the ideas of a really smart individual), and then it rewrites it using totally different language but preserving the same meaning. (old) Plagerism detection tools don't work and hell, it's not hard to fool the never ones. You can try tools for detecting if something is AI written by a particular model and weights (e.g. to prove if they used GPT2-Medium), but if I fine-tuned those same weights, than proving it was plagiarism will become exceedingly difficult.
Welcome to the brave new world of cheating. Also, techniques like this are coming to a CS department near you (in the form of source code generation powered by NLP models).
Would be neat to try and publish books with a percentage of AI generated text and see how well they do. Maybe there's a sweet spot for productivity.
So I would say this likely will just be there. It just is. Wont change anything, universities will acknowledge it, a headline or two will occur when its use was discovered in a paper that a student didnt even skim to make less obvious, and most papers will fly under the radar.
Other kinds of assessments will still do their job.
If I recall correctly, the way it worked was to build up a model of this persons writing, and how it compared to to other people, and then would measure the likelihood that sentences and paragraphs matched the rest of the writing.
I suspect something similar could be done with GPT-x
It generates buzzfeed kind of stories very well though :)
Whenever GPT-3 is updated or a new version comes out, it will be able to speak much more intelligently about the topic. But of course any update will require re-doing all the careful tuning of prompts and models...
I'm scared that more and more big model advancements are being denied access from the general public, which will just make the inequality between big corporations and startups even greater.
How does a Free Software or "Open Source" project get around that?
The Eleuther project makes use of distributed computing resources, donated by cloud company CoreWeave as well as Google, through the TensorFlow Research Cloud, an initiative that makes spare computer power available, according to members of the project
This includes HN [i] HackerNews 3.90GiB 0.62%
Which if SciFi has taught me anything means we are all uploaded now and will live forever.
They did develop BERT and use (used?) it for parsing search queries [1]. They probably use NLP models in the ranking algorithm too. But those use cases are about getting the best result possible within the throughput/latency requirements, which necessarily makes them less "powerful" than models like GPT that pay little attention to performance.
https://blog.google/products/search/search-language-understa...
Google for example can monetize NLP AI framework just like they are monetizing Kubernetes: https://cloud.google.com/kubernetes-engine
If OpenAI and Microsoft are licensing and monetizing GPT-3 why Google wouldn't want to compete with them when they have more data than them. Nothing can beat the amount of data that Google and Google Search have gathered over the last 20 years.
Like 90% of content is written by marketeers for bots. SEO they call it. Now we can take out the middle man. Bots writing crap for other bots. And then we use that content to train more bots to write even crappier blog spam. And finally the bots decide the actual recipe is no longer needed on the recipe blogs and they kick us of the internet.
In a way, humans are about to lose the Internet.
Solution:
General AI is already here. It should be implemented on twitter or wherever and used to teach us about ourselves. driven by engagement, untethered by morals. A dispassionate glimpse into what sells. An AI that exploits our engagement, for good or evil.
The bot would become infamous and in due course banned. Teaching us even more.
But we are so fragile.
That being said, I don't see how the existence of image-gpt supports the notion that GeneralAI exists. Image-gpt couldn't (for example) write the next version of itself.
"The AI chatbot Tay is a machine learning project, designed for human engagement. As it learns, some of its responses are inappropriate and indicative of the types of interactions some people are having with it. We're making some adjustments to Tay." (Microsoft statement)
https://www.theverge.com/2016/3/24/11297050/tay-microsoft-ch...
Also wrong that previous chat bots were not shut down for being offensive. https://en.wikipedia.org/wiki/Tay_(bot)
If you mean AGI (artificial general intelligence) it's definitely not here yet.