OpenAI is Using Reddit to Teach An Artificial Intelligence How to Speak
futurism.com
futurism.com
Source: http://spraktidningen.se/artiklar/2007/11/darfor-ar-din-mobi...
Translated by Google: https://translate.google.se/translate?sl=sv&tl=en&js=y&prev=...
very easy to text even from your pocket etc - something you cant do on a smartphone - notice texting accidents became more common with the widespread adaptation of smartphones ;)
I think this feature was actually key to being able to memorize how to perform operations quickly. There's no way you can do that with modern smartphone, where every other interaction lags for anything between 10 and 1000 ms, and inputs are sometimes dropped at random. It's this non-determinism of smartphone UI that makes me look at the screen all the time when using it.
You can't do that when you have to actually wait for the thing you want to select to show, otherwise you'll be selecting something completely different.
I guess I'll need to wait until shape changing tactile surfaces come to phone screens, which might be a while / never.
Anyways, i could see someone use that term in frustration about a car having been parked across two spaces.
Then I remembered this [0].
Still wonder about the plurality/majority thing, though.
https://english.stackexchange.com/questions/55920/is-most-eq...
We used it to train a syntax-enriched word2vec model. Write up and demo: https://explosion.ai/blog/sense2vec-with-spacy
Btw, the above was run on CPU in a couple of days, because spaCy doesn't use GPUs yet. I've applied for a grant from NVidia so I can fix that. If anyone from NVidia is reading, email me? :)
"it was gone" meaning the association between Carrot Top and Kate Mara? So after better training, who is now most_similar(['Carrot_Top|PERSON'])?
EDIT: RTFA, used the interactive demo. most_similar() is now a category I would describe as "actors/comedians popular in the 90's: "Bill Murray, Gary Busey, David Spade, Charlie Sheen, Ashton Kutcher, Chris Farley".
Weirdly there's a bug that's dropped PERSON from the sense list. Fixing.
Edit: Fixed.
Edit2: Ah this is super misleading atm. I'll have a think about how to do this better. auto is case insensitive, but if you specify a sense, it's case sensitive. So you need to do "Carrot Top" and set PERSON. Btw, contrast with carrot top "NOUN".
> I've applied for a grant from NVidia so I can fix that.
A g2.xlarge is 65c/hour on AWS, FWIW.I'd use a spot instance and stop it whenever possible.
Unless you're replicating someone else's thing exactly, You can't really get by with one training process. You want to be trying different things, and running a few samples of each configuration to account for random variation. I'm not even talking about decadent hyper-parameter sweeps to fine-tune. I'm talking like, how wide do I need my layers to be, what optimizers are good, how deep should I make the network, etc.
I want to be training 5-10 models at a time minimum. 20-30 would be much more productive. If I can only train one model at once, it's not really worth the effort — it's better to work on one of the other tickets for the library.
The full table (up to end of 2015) is available on BigQuery, with separate tables for each month thereafter: https://bigquery.cloud.google.com/table/fh-bigquery:reddit_p... (there is a similar table for comments)
And here's a year-old post I wrote on how to use that Reddit dataset with BigQuery: http://minimaxir.com/2015/10/reddit-bigquery/
Sometimes I can remember the comments on an article, but not the article itself. Unfortunately I can't search google using comment text, because Reddit doesn't allow google to index its comments.
Disallow: /*/comments/*?*sort=
Disallow: /r/*/comments/*/*/c*
Disallow: /comments/*/*/c*"site:reddit.com don't quote me"
I use this all the time to search specific subreddits for phrases in comments I remember, but I don't want to give examples or subreddits I visit.
Look at this maverick!
Google was getting overzealous and indexing every page hundreds of times because it would follow every link, which included every "context" link and every sort.
Try adding site:reddit.com to your search.
Even training models on that is possible in realistic times on normal systems with that.
However, uncompressed, it hits over 1+TB, if the BigQuery sizes are indication.
Even then, that’s easily doable on a consumer system.
I’ll download it in the night between friday and saturday, after I install my new HDD, and just run queries over it for fun. (far slower, but also far cheaper than BigQuery. Even at German electricity prices).
magnet:?xt=urn:btih:UGFLA4QNEXGEFKYYY5ZU37JIHWEEYY5R&dn=reddit_data&tr=udp%3a%2f%2ftracker.openbittorrent.com%3a80&tr=udp%3a%2f%2fopen.demonii.com%3a1337&tr=udp%3a%2f%2ftracker.coppersurfer.tk%3a6969&tr=udp%3a%2f%2ftracker.leechers-paradise.org%3a6969
which I have put together. It contains the data up to April 2016.If you want to work with this dataset on your workstation, there are some code examples in https://github.com/dewarim/reddit-data-tools
Imagine you have a bot that convincingly passes the Turing test - what would you do with it?
Build a chatbot business? B2C or B2B?
Sell it to one of the big companies and if yes then how much do you think it would go for?
Give it to OpenAI? Open source it? If you answer yes to any of this questions, then why?
Edit: let me qualify - this would not be AGI, just a much more advanced bot than whatever is currently on the market.
The gulf between the smartest human being ever and the dumbest human being ever is not really that wide. It's just a small blip in the long line of intelligence progression. There is no reason why this point should be the limiting point of AI intelligence. It is very likely that the stable point of AI intelligence is far beyond human intelligence, from just the sheer quality of processing hardware that exists that is better than, or could become better than human hardware.
A chat bot that passes the Turing test is limited by whatever programming got it to that point.
Also, possibly medical applications? Eliza (the first known chatbot) was built to simulate a Rogerian psychologist and was quite convincing for its time (1960s)...
What else would you use it for?
The question then is, "When everyone has access to free chatbots that can pass the turing test, what will they be used for?" The answer is "tons of stuff", and lots of people will try it at once. I think many applications will be niche.
Also, people will argue about what constitutes a Turing Test. For instance: https://twitter.com/mattdpearce/status/784162089397092352
However, we are not there yet. If someone had this today, how valuable would this be to Apple, Google, Facebook, Microsoft and Amazon?
As for the definition of what the Turing Test is, it's definitely a fuzzy subject. My own arbitrary definition is "ability to convince a human that he's talking to another human after a sustained (based on time or length of) conversation whereas human is aware that there's a possibility that his interlocutor maybe a machine".
So, it's more of convincing a judge in Loebner Prize competition than a random troll on Twitter.
Obviously such programs, while being works of art, are not interesting from AI/Machine Learning point of view.
Ok, lets get this straight: ~6 million real human men paid real money that they earned through their labor or whatever to talk to bots and then paid more real money to do it again. Admittedly, they are 'cheaters', but 6 million men must have an IQ distribution nearly identical to that of the general population, i.e. they represent heterosexual human males in general. And yes, they were trying to get laid, these conversations are likely pretty brief, and mammalian males are not generally known for using their neocortex during mating.
Still, I think that 'counts' as far as passing the Turing Test. Yes, now we can move the goal posts to say that the bot has to teach me something, or guess what I was thinking, or generally be better than a man on tinder. But as a first pass of the TT, I think we have been here for a few years now.
Online lectures are great, but a personalized tutor could change things. If I restate back my understanding of a subject, and it clearly tells me why I'm mistaken, that's useful.
Reddit does this, kind of, today... but it's not really from an informed position. It's mostly uninformed people arguing with equally uninformed people. There are gems occasionally, but it's rare. That's why /r/depthhub was created
While it can answer your questions and check your understanding (it's a First Order Logic application), I haven't thought of the educational applications but I don't see why not... Thank you for bringing this up.
Edit: Corrected.
The main thing stopping me is NLP. Ideally, i want this offline only as i am unsure how much of my life i want to leave my home network.
I haven't started yet, I'd be grateful if you would be open to discussing on this topic :-)
So far I didn't find any useful stuff on the internet
For me, functionality of the assistant was fundamentally difficult. I could think of a dozen things i would like a home assistant to do, but most of them i don't want to program. Things like playing a song of spotify, changing the device spotify is playing on, and etc.
For me, what i was willing to program a bot to do was manage my personal server, take a load off of me. Check for updates, notify me about them, ensure backups are being triggered, etcetc.
At the moment i don't even have a bot though. I've taken many iterations on the literal programmatic API, and still don't have it quite right. I started in Go, and am now sitting in Rust (though not actively working on it). The difficulty, is i want to find a nice way to write a handler for an event. Eg, a web page visit is a single handler for a single event. However a bot response is not a single event. It's a conversation - so i'm trying to figure how to manage the state. This is my biggest point of internal struggle.
Anyway, if given what i've said above you're still interested, feel free to let me know if you'd like to talk more :)
Easy peasy. (Not really, but for an in-home piece of software, it doesn't seem too bad.)
If I can say predict with a 30-40% accuracy things like riots in 3rd world nations based on collective analysis of data provided from thousands of sources (just by looking at places and sentiment), broken up by groups and affiliations, and correlated an analysis of a country's monetary and political situations then I could probably sell it for a nice chunk of change.
Lots of work, but probably huge pay off. Then again I'm not a "Data scientist" so I'll leave this up to those experts.
PS: you could definitely use this + gender detection for finding information about products and services and correlate that to corporate success of advertisements. Technology like this is applicable to many industries. Just looking for different correlations of the same sets of data.
Science fiction author and scientist/science writer Isaac Asimov popularized the term in his famous Foundation series of novels, though in his works the term psychohistory is used fictionally for a mathematical discipline that can be used to predict the general course of future history.
If you have data like payroll and bank account info (so, state-level hacking), it'd be interesting to see how economic pressures turn to uprisings. CIA/NSA probably has access to logs of all the world's telecom companies, I wonder if they can see trends of how e.g. riots grow organically (many mobile phones registering at particular antenna = huge crowd = riot (or a concert, or a football game...)), and to see who the instigators are.
2. Privately invite representatives from each company to sign up for the site, and invite them to a blind auction on each brand pool their brand is involved with (Toyota exec sees car pool with top bid of $0.87/comment, decides to bid over it to put Toyota at the top of the pool)
3. Maintain a million Turing-beating chatbots that trawl reddit/facebook/twitter/quora/g+/HN/etc looking for brands, looks up the pools that brand is in, and then leaves a good PR comment for whichever brand has the highest bid across related pools. Swarm properly to distribute these comments evenly across the internet instead of clumping together.
4. ????
5. Profit (and/or become a real illuminati)
Also, I would use it to troll bureaucrats when I give up in frustration. It should try to force the bureaucrat to admit flat out that their logic is fundamentally flawed, and then ask them to propose a solution }:)
Spambot. Contact someone over chat services, start an interesting conversation, then subtly promote a product.
More seriously, interview bots. Talk to people and ask them questions, turn them into a coherent whole. Let the elderly talk about their lives and record their stories, let people who have some problem they need solved talk about it so it can be turned into a succinct description, and so on.
Of course, it depends on whether hypothetical Turing test passing computers can do that. Let's just assume we'll ask contestants in a Turing test to do those things, then we know the winners can.
In the brief time between humanity's destruction and the Internet going down, Twitter and other social media will be just spambots re-tweeting each other. If we get that far, "AI" will be able to keep the Internet and the infrastructure it needs (power generation, power grid, actual cables) alive, and when aliens discover our planet, it will just find bots recycling the trending topics ad-infinitum.
The anti-spam bots are at a disadvantage because at some point in the arms race you have to make the bots actually fully intelligent, and then exposing them to spam 24/7 is torture and illegal. Spammers don't care about legal issues.
My second reaction was, "at least they're not using 4chan."
- SYSTEMD IS AN ABERRATION AGAINST ALL THAT IS UNIX!
- But A.I, you use systemd to boot...
Thinking about it now, this is deeper... There's a fear that AI will take over the world, use weapons in unethical ways, say one thing and do another, etc... If we use news channel and politicians debates to teach AI, I'm afraid that this is exactly what we're going to get!
https://twitter.com/dorfsmay/status/785907475480350720
"The same way adults stop swearing once they have kids, we might become more honest and ethical by fear that AI will learn from us."
The core learning algorithms are not changing.
Correct me if I'm wrong, but I think it's basically a couple hundred NVIDIA 10-series cards strapped together with a full custom NVIDIA software stack.
It uses their P100 HPC cards instead of consumer grade cards (8x P100s), plus two Xeon E5v4 chips, half a TB of RAM and 7.5 TBs of SSD storage - all wrapped up nicely configured for you with their CPU-GPU speed up stack.
I believe the only way to get P100s right now is in the DGX-1, so there's that.
Sure, there's been improvements to the computational performance.
But the big deal (to me at least) is the unified memory model between the GPUs and the Xeon host processors. This makes a lot of things easier to code for on a single system, and it makes multi-system applications easier to scale. This is because you're streaming data in over the network (10G Ethernet) and then the GPUs can operate on it without an extra copy step. The copy step also implies more management and shuffling around of the data you're operating on.
NVIDIA gimped half-precision on the consumer cards to drive datacenters, hedge funds, machine learning companies, etc. towards the "professional" cards (and their huge markup).
After that, it's going to be mostly about memory size and bandwidth.
We really need some more Frameworks that work with OpenCL, so that we can have some competition from AMD, who's consumer cards are not gimped.
I don't see the issue with a company making a very high-end product, adding stuff that doesn't have good use for consumers, and asking extra money for their effort.
AMD doesn't have double speed FP16 on its current FPUs either. The latest version has FP16 at the same speed as FP32, but if you're doing that you might as well use FP32 always.
And let's not forget: the Nvidia consumer GPU have deep learning quad int8 operations enabled at all time. They didn't need to do that and could have reserved it for their Tesla product line only.
Reddit can get vitriolic and rude, insightful at times too, but once the system learns the syntax hopefully they'll be able to use sentiment analysis to weigh more strongly the polite conversion that occurs.
Also interested to see how many memes this AI picks up.
I also hope they are able to follow links through to sources when a comment cites another page -- not only can this bot learn syntax but also data extraction by comparing what is said to the source material.
If they're taking the whole of reddit it could start to identify enough context to know when to be smart, sarcastic or simply helpful.
With some of the subs there are long discussions that stay mainly civilised. Same for the support subs it could learn the context and how of sympathy and empathy. Things that end up on front page, filled with snap sarcasm, will be a tiny fraction.
I think it's going to be very interesting see what comes out.
They should limit it to top comments only, and for training, you might as well assume 90% of top comments are sarcastic/tongue in cheek. Or let a user dial the sarcasm/wittiness/seriousness as they want it, kind of like TARS from 'Interstellar'.
I have at least one paper about sarcasm in my Zotero (link to PDF): http://www.aclweb.org/anthology/P/P11/P11-2.pdf#page=621
If you don't want to click on that nondiscript looking link, the title is: "Identifying Sarcasm in Twitter: A Closer Look"
In seriousness, between all of the garbage there is a ton of knowledge and intelligent conversation uploaded to Reddit every day. And, it's all hierarchically organized and scored by domain semi-experts. It really would be wonderful if someone could mine that knowledge IBM Watson style. For example, I'd love to ask the /r/BuildAPC collective AI for PC building advice.
Everything on Reddit is on a bell curve, with a fat mediocre middle and trailing awesome and superbad ends.
And that includes the quality of the scoring process.
----
"Siri, get me dinner date reservations."
. . . DID YOU MEAN 'false rape accusations' ?
One problem for data geeks to solve: Reddit data fits nicely into a graph structure and not so nicely in table form. It would be fantastic if someone put the Reddit data set into a graphdb and made it open.
[1]https://wisdomofreddit.com and https://github.com/qxf2/wisdomofreddit
[2]For now, my search engine currently just uses Whoosh's (out of the box) BM25F.
Anyone have any good defensive technology ideas?
Social network camouflage.
Does anyone know of any projects in this direction?
I've also seen some that register you for sites by feeding in random demographic data.
[0]: http://www.theverge.com/2016/3/24/11297050/tay-microsoft-cha...
I am currently doing a double degree in communication studies and information science. They are both interdisciplinary. Communication sciences integrates aspects of both social sciences and the humanities (both "soft"), and so far when doing research both of these fields were taken into account and no students have problems with combining these fields.
Information sciences integrates aspects of formal sciences ("hard") and social sciences ("soft"). When the course is about analysing communication data, the methodology of social sciences is also important - for instance questioning the validity of your data. That's the thing what you're mentioning: the majority of reddit speaks like and holds the views of college-aged white males, so the data does not represent everyone, and is not valid if you truly want to develop an AI for everyone.
Whenever the "soft" science comes around, like writing an assignment analysing the validity of data, many of my fellow students struggle with the concept of data not being neutral. This is where the two fields collide, and usually it just ends up with students scoffing at that "illegitimate" scientific field. Many teachers also don't spend much time discussing that field during the lectures. I admit, I have written some lazy essays which probably had been given a negative grade if they were written for communication studies, but easily passed in information science.
Of course information science is not AI, but they're both sciences that have parts of formal sciences and social sciences (I know AI has many more fields). I am afraid many talents within AI research miss essential knowledge about social sciences or deliberately ignore it, because it's not "hard" science. Case in point: your comment is now at the bottom of this thread. And then you get nasty surprises, like Google Photos categorising pictures of black people as monkeys.
For Reddit, I like to imagine that it's basically the training data for all of the emotional and societal nuances that a human goes through.
Think about all of those stories that people post in ask Reddit that explain western norms and no nos. how to treat people with respect, when to call the police, how to communicate properly, etc.
Obviously we're far away from using the data to its full potential but one day I could see Reddit data to make our AIs more relatable and human like.
> "Just post it to the ship's 4chan, and check after a few hours to see if anything was modded up to +5 Insightful."
For interest, how many HN comments are there? Miles fewer, no doubt, but perhaps far more erudite and less likely to offend.
Time to read the article.
Ignorant person speaking here: this still doesn't sound like AI, you're just making something follow patterns and regurgitating them. Is that AI? Maybe that's what I do a tech parrot. Ahh well time will tell.
Of course we imitates our parents/others to learn how to speak.
I was interested in parsing vocal sound bytes and learning how sound was created/formed letters/words.
Alright ignorant person out.
Human speech is produced from the conscious experience of being a human being. If your dataset contains just the speech, without the experience, there's simply not enough there. Any machine trained on this data is doomed to talk hollow rubbish.
A virtual assistant that has the personality of a smug know-it-all, know-nothing 20 year-old with little motivation to do anything but regurgitate surface knowledge and sarcasm in an attempt to look intelligent without expressing genuine interest in helping anyone.
Humans have opposite problem. We understand what we talk about, but have little idea how our brains create language.