DeepText: Facebook's text understanding engine
code.facebook.com
code.facebook.com
I have no mistakes that these features are probably loved by some, maybe even most, but more and more-while I don't want to disconnect FULLY from social media (while I much love the laconic, brevity inherent design of twitter, the fact that the 140 character limit is going away has me sighing heavily), I do sometimes finding myself wishing I could opt out and take a bit more control over the content I'm ostensibly subscribing to.
Has anyone else felt similarly, or could maybe phrase the phenomenon better than I have?
I'm starting to feel like some good old brains would serve us well sometimes. Take Google Maps' traffic again, it can tell me the traffic now. It can tell me expected traffic tomorrow. But it doesn't understand that if I depart at 8am and I arrive at 10am, then at 9am that traffic jam is worse than when I departed. I'm better at planning this route myself this way. Kudos to the people who built the whole model (almost certainly using some sophisticated algorithms) that automatically detects traffic when there's enough data, probably correlating it with their map data and stuff, but this is an elementary feature apparently nobody there thought to implement. It's not all just training sets and enough CPU power, some human intelligence to provide it with the concept of time would be smart, too.
Now this might seem like one silly example, but in many cases it's like this. People promote machine learning and neural nets and all sorts of automated training in situations where some research and an if statement or five would have worked as well.
Another simple example: we had security monitoring class and one of the topic was machine learning. Let your intrusion detection system learn normal traffic patterns and alert on, say, traffic spikes on Sunday nights when there should be few if any people working. I feel like looking at traffic data and some if statements would be more useful, especially since you can then integrate it with holidays and events and whatnot. Indeed, one of the warnings in the course was that "during events this will go off as well". Yeah, that's why we shouldn't use it here in the first place. This just a hammer and screw issue.
As a Machine Learning Engineer / Data Scientist I would always evaluate my model against a baseline (in this case the if statements) and only put the model in production if it outperforms the baseline.
[1] https://www.google.com/search?q=%22google+now%22+cards+causi...
1. The weather
2. Stocks that have the same acronyms as things I've searched for
3. Flight information for any flight with the same code as mine
4. A recommendation of a film I might like to see because I searched for actors in them (which might be good except the reason I was looking up those actors was that I had just watched that film)
5. Sports scores for teams I don't follow
6. Possibly updates to sites I visit regularly (equivalent to a highly unreliable RSS client)
This is what annoys me the most in Google Now. It's a damn black box. I'm happy that they provide "updates to sites I visit regularly", but even more than those updates I'd like to see the list of those sites. Without that, and any kind of information about how reliable this reporting is, I have no way of trusting I won't miss an "important update". I just don't know, and Google Now does nothing to reassure me.
The only time they annoyed me was when picking up a false article about Breaking bad being renewed. I shared the headline with several people before realizing it was a fake article.
http://www.businessinsider.com/breaking-bad-season-6-hoax-20...
Parking information would be useful if it were both accurate (it isn't, due to taking a park and ride bus) and consistent. Like many Google Now cards it seems difficult to predict whether it will turn up or not.
Another trick it used to do was show me cards with navigation information to somewhere I'd searched for, usually the night before a trip. But, when I got into my car at 6am the next day such information was nowhere to be seen and I'd have to search for it again. Now I save the location in the calendar so at least there's a link to click, causing Google Now to remind me to leave on time after I've already left, sometimes when I'm nearly there.
Overall, it's not a useful product and relies upon data being collected which I'd rather wasn't. The only reason I leave it on is so that I can dictate reminders to my watch, which I find a very useful feature. The requirement to use Google Now for this last feature seems rather arbitrary to me.
Then on the way out of work I try an "OK Google. Directions home"[0] to get a time estimate or best route estimate and it tells me I have to turn on browser history to continue or some other such bs. I wonder when Google Now will understand "Ok google. F?!k you Google."
[0] correction: that's more likely "ok Google. OK Google! OK GOOGLE! Damnit (click google now, click microphone). Directions home."
Me: "OK Google, navigate home" Google: "You need to turn on web and app activity." Me: "OK Google, navigate to $NUMBER $STREET, $TOWN." Google: "No problem".
But, they then decided that web and app activity was essential in order to give me commute information, and the only way I found around this was to turn it back on and start doing almost all my searches via Duckduckgo.
Yes, because they're not doing it for us, but for themselves, investors and customers (advertisers) fueled by our data.
That includes the people you mentioned yes, but also users who get to use a valuable tool for free.
The whole "if it's free, you are the product" party line may resonate in HN circles, but you have to realize the public as a whole doesn't think it's a bad deal at all.
Try asking someone "hey, you know Facebook use a lot of state-of-the-art technology to algorithmically understand the meaning of everything you post and write, feeding this massive data set into their huge artificial intelligence program?" I don't think it's likely that many people will say "cool, that sounds like a great deal, because I love it when I get very specifically targetted ads."
Still there is nothing which could totally replace Facebook so old users are staying and baiting new users. It still has unique offer for large user base and very useful features so even if it's getting worse it can stand. People might be clicking their feed more but that's because they are so bored with content that they try if they can find something more interesting behind the links. Or they are scrolling down since trying to find something interesting, something different which is hard since algorithms just push same boring content all over again. But statistics look good.
What's the magic of Facebook? If you could answer that and build the service people will flock out from Facebook.
But neither do you really, right ?
Sure, you may have a better idea than them of what's technically feasible, but the general HN consensus on the motives, values and internal decision-making process at these companies ("always assume the worst from them, they're optimizing for money!") is in my experience (of working at FB a few years ago) a simplistic view that's way, way off the mark.
>"hey, you know Facebook use a lot of state-of-the-art technology to algorithmically understand the meaning of everything you post and write, feeding this massive data set into their huge artificial intelligence program?" I don't think it's likely that many people will say "cool, that sounds like a great deal, because I love it when I get very specifically targetted ads."
I think this is again where nerd-bias is misleading you : 1) Many people would answer your question with "yeah ? is this the thing that enables translations, picture captions for the blind and suggests me events I might like ? Please, bring it on!" 2) people actually do prefer very specifically targeted ads, rather than annoying+irrelevant pop-up banners of the past. This is why FB is making a killing.
I don't think it's nerd-bias... it's more that many of my friends and acquaintances who aren't nerds are still instinctively skeptical of large corporations and the power structures they create with respect to privacy.
As for what people in general think, we really should ask around. I heard about a consumer trust survey where Facebook's trust was very low, and based on how I see people relate to Facebook, I'm not surprised at all.
I think we ought to be suspicious and critical of corporations that collect large amounts of private and personal data, just as we ought to be so of government surveillance (of course they're inseparable).
http://techcrunch.com/2013/01/27/facebooks-categorial-impera...
But you add human intelligence to this kind of thing and it starts getting creepily accurate much of the time if you're looking to promote certain behaviors, expose others, or what have you.
It doesn't have a specific profile of the opponent, though, instead it thinks they're another AlphaGo. That's an area they could work on.
Here's a link: http://www.nature.com/nature/journal/v529/n7587/full/nature1...
I think you search on the title, you can find probably find a non paywalled version.
Minimax and MCTS are just different algorithms.
MCTS simulates games randomly and creates a distribution of expected value for each one based on some cool math. It estimates the value of each state based on a sample. Minimax actually calculates them all (or all of them except for some pruned ones).
That's the whole reason MCTS can work on Go and minimax doesn't. Minimax actually needs to look at all possible branches all the way down. MCTS can just say "well its been 100ms, let's use the estimates I have".
No, I assure you that stockfish definitely does not look "all the way down (ie to states where the game is over)" in most positions, even with pruning - that would take too long. It uses a static evaluation function to evaluate positions without searching down the tree further.
> MCTS simulates games randomly and creates a distribution of expected value for each one based on some cool math.
You described the "Monte Carlo" part of MCTS but not the "Tree Search" part. It's right in the name! See for example the first paragraph on page 20 of https://gogameguru.com/i/2016/03/deepmind-mastering-go.pdf, which clearly describes AG constructing a search tree.
Minimax obviously will produce perfect play in both chess and Go, but we can't use it because it takes too long. Hence we prune and truncate the search tree. When we truncate it we use an approximation of the value of the node (the static evaluation function); when we prune we sometimes use the static evaluation function as well. This is true in both MCTS and in Stockfish.
Ehh alright, it calculates the value of a node exactly (or exactly assuming its evaluation function). MCTS does not.
To be clear, what I meant, that you pulled out of context is that Minimax requires an exact calculation of the value of each child before choosing the best and returning it, this requires either simulating every possible game, or an evaluation function (depth-pruning). MCTS uses monte-carlo methods to estimate the best paths instead of trying all of them naively. At t=infinity, MCTS is equivalent to Minimax, because it will try all nodes. But it can run in time constrained systems better because it picks a much smarter order to run in.
>You described the "Monte Carlo" part of MCTS but not the "Tree Search" part. It's right in the name! See for example the first paragraph on page 20 of https://gogameguru.com/i/2016/03/deepmind-mastering-go.pdf, which clearly describes AG constructing a search tree.
yes. That has nothing to do with minimax though.
Which is what I said. AlphaGo uses pruning. AlphaGo doesn't use minimax with or without pruning, it uses MCTS (with pruning).
Edit: Also, goddamn dickheads wasting human upload/duplication....
Well, I figured out today that when Facebook has these, it doesn't necessarily send them to you right away. Instead it waits until it looks like you've bailed off the site to do something else. Then you get an email notification -- drawing you back into Facebook.
Yep. Pretty tired of it. Thousands of really smart people using the latest in tech to try to play me like a musical instrument.
Then, once I'm away from the computer for 30-60 seconds or so (I'm sure they have a system training for this), I get the email notification -- even if the original back-and-forth occurred an hour ago. The timing of when I get the email has nothing at all to do with when the event occurred. It's completely dependent on my interaction with Facebook. Many times, it takes so long, I'm usually thinking "Wow! I wonder what else X had to say?" but I click over and it's the same damned thing I already interacted with.
I've never got a later email about something that happened live while I was browsing the site, and was already notified about.
So you're saying that there's a bug. And the bug waits for me to stop using Facebook, then emails me things from Facebook.
I find this quite difficult to believe. At the least, it's a very convenient bug.
The bug would be to get emailed about things that happened while you were using the site and were notified in real time about already.
Absolutely! At least give me an option (even if it is buried nineteen pages beneath about:config) to switch to a simplified reverse chronological stream. Until this happens, I've in followed everyone on Facebook. If you didn't send something specifically to me, then no I definitely didn't see it. (:
[1] https://en.wikipedia.org/wiki/Filter_bubble
(even mentions Facebook's news-stream)
And Facebook just can't seem to remember that I never am interested in whatever it considers "top stories", but always, every time I log on, switch it to "most recent". Not that I know that I actually get all the most recent stuff, but on top of that Facebook is clearly telling me it would prefer me to see its own curation of it, even though I obviously have no interest in that. That this option doesn't stick seems like the intended behaviour, and I consider that yet another hostile aspect of it. So I'll assume whatever they're now doing instead of a UI that deserves the title, it will understand text just fine when it suits advertisers, and will be useless otherwise. Not a claim, but a guess.
Further reading:
https://aeon.co/essays/how-the-internet-flips-elections-and-...
So I try to embrace the bubble when it is useful for me. I appreciate that when I type "something something python" into google the top results are relavant to my work and not about snakes or comedians. But when I'm want to form a politcal opinion I use duckduckgo and activly look for contrary positions.
I have a hard time feeling like it's helping me when I get filter bubbled; I feel like I'm more productive and on point when I know where I am in relation to the rest of the world for whatever reason. The feeling when you learn what to search to get the desired specialized results instead of other general usages by adding keywords or going to sites that are more specialized themselves (e.g. stackoverflow where you can search [r] to search posts tagged with "r" in particular) is wonderful. You opt in to a bubble of a particular sort versus not knowing what else there is out there. I tend to feel a bit... disgusted with Google deciding what's good for me during a given search.
Earth provide opportunities for success for both men. In the long run, one of them will dominate the other.
Note - I don't just mean, "content that disagrees with you", because 95% of that is shit (Sturgeon's Law), but, content that effectively disagrees with you. That actually has a chance of changing your mind.
Information Quotient, as I understand it, is the idea of "What information will teach me the most?"
Count me in the camp that loves them.
I don't want to spend hours a day keeping up with social activity. Facebook does a great job of giving me a quick summary of everything going on in my friends' lives by sorting through all the drivel algorithmically.
The only way you can actually know that is to compare the summary with the raw data. The only way to remain sure of it is to keep doing that.
i'm not saying your experience is untrue... just that facebook haven't done a great job of this for everyone at all.
Sounds like you really care a lot [0].
By the way, I believe that some filtering might be necessary. It seems that when I joined Facebook there were far fewer memes and 'folklore' posts (we have a nice word for that in Dutch: 'tegeltjeswijsheid'). But that would be a user-trained filter, which is not interesting to Facebook.
So, I agree with both camps: we need filtering, but Facebook's filtering is rarely good, because it's not aligned with my interests.
The problem to me comes only when the big corporate AI systems no longer are fully on your side, like spam filters which let the "right" spam through.
Everyone in the field has read that paper. It was good work! But there are lots of intriguing things mentioned in the post which deserve further details.
The most interesting thing to me is "more than 20 languages"!! That's pretty nice - the paper had some early results for Chinese, but if it can perform similarly to the English results across 19 other languages that is probably the state-of-the-art for many of them.
edit - I meant to say even if they had end-to-end encryption then your text would still need to be decrypted so that other user can read it.
I mean, my message contents are personal data. The topics "buy a bike" and "sell a bike" are not (I mean, not by themselves), so those you can transmit to ad companies.
Now of course you need to make the detection lightweight enough that you don't need to talk to the server anymore, but that's just a matter of time. The person you're responding to, who's got downvoted, has a point (if this was his point).
Edit: I sounded like I agreed with it, which is not the case. Even if it's stored "private and securely" at Facebook, they would still be able to connect my topics of interest with my identity and contacts, which would "technically" comply with the concept of end to end encryption (only metadata leaks because topics alone are metadata, which always happens because you need to route messages)... but which still doesn't comply with my own definition of proper privacy.
This is just how marketing would sell these features to the world (end to end encryption, and yet targeted ads), but disclosing each topic I ever discuss with a given person is still private in my opinion.
I've edited my post's original text slightly, adding a bigger edit at the bottom.
OP has shown that and how FB does in fact use text analysis of messages to show adverts. You habe been a re-targeting target by correllating your browsing habits and other factors. That has been going on for ages.
It's a little tough to figure out what Facebook's goal with the post was though since it's not very technical.
You mean show us more "relevant" ads ? Last time I checked, Facebook's *product" was advertising. Thanks, this is exactly what I miss in my life - more ads for products and services I don't need.
So excited that Facebook is going to understand everything I'm talking about privately with my "friends" and sell my identity to more ad buyers, who'll design more subliminal ads to squeeze the last millisecond of what's left of my attention span.
Here is example code for text classification from character-level using convolutional networks in Torch 7 https://github.com/zhangxiangxiao/Crepe
Is there a way for us to play with it? Or are they just bragging?
I guess those things would be fine to share without revealing the inner workings.
When these things get just a little bit better, the Five Eyes agencies will suddenly not have a staffing problem any more.
> Facebook is strip-mining human society. Watching everyone share everything in their social lives and instrumenting the web to surveil everything they read outside the system is inherently unethical.
> But we need no more from Facebook than truth in labelling. We need no rules, no punishments, no guidelines. We need nothing but the truth. Facebook should lean in and tell its users what it does.
> It should say: "We watch you every minute that you're here. We watch every detail of what you do. We have wired the web with 'like' buttons that inform on your reading automatically."
> To every parent Facebook should say: "Your children spend hours every day with us. We spy upon them much more efficiently than you will ever be able to. And we won't tell you what we know about them."
> Only that, just the truth. That will be enough. But the crowd that runs Facebook, that small bunch of rich and powerful people, will never lean in close enough to tell you the truth.
I was hoping for some open source code to read, or more detail on their models. Facebook is pretty good at open sourcing things, so hopefully more papers will be released and open source software as it makes sense for FB to do that.
EDIT: typos
They still have their business advantage intact: they have the data and the awesome infrastructure.
Fun that you guys can train a neural net to deduce what I like better than me, but as they said at the CCC conference, everyone has to decide for themselves how close they are with their machines (though that was in the context of taking your laptop to the toilet, in context of leaving it alone unattended).
[1]: This past weekend, a close friend was telling me about his friend who works at Google and the following morning, this Google person came up on my "Friends suggestions". It was the most bizarre thing... I also confirmed with my friend if he looked this Google guy up later on Facebook, which may have prompted his Google friend to show up on mine, and he said he didn't touch Facebook after our conversation.
EDIT: For further clarity.
Having said all this, it was actually in the news[1] yesterday that Facebook will soon be providing end to end encryption, though I doubt it'll apply to the listening in on people's conversations chatting.
[1]: https://www.theguardian.com/technology/2016/may/31/facebook-...
I'd expect a text understanding engine to do more than classify, I'd expect it to read a bunch of sentences in context and then be able to answer questions about it:
John was walking down the stairs. John tripped and fell. John lays on the floor.
Describe John's status: Is he hurt or in pain? Is he standing?
Daniel picked up the football.
Daniel drops the football.
Daniel got the milk.
Daniel took the apple.
How many objects is Daniel holding?
There have been plenty other papers published by them and others tackling that and similar but harder problems eg[2]Seems like a PR stunt to attract talent.