Introducing FBLearner Flow: Facebook's AI Backbone
code.facebook.com
code.facebook.com
So, everything that i hate about FB is related to the AI.
Ranking and Personalizing: Horrible ranking results and the "personalized" results that they push into my feed have nothing to do with me, my interests and is not something i would ever want in my social feed. Its almost as bad as their advertising results (which im guessing is based on the same AI).
Offensive Content: I could not give 2 f@@@@ about this, but i guess some people don't like strong language or pictures (only point that i see as acceptable use of the AI for now).
Trending Topics: FB, you are not Twitter, please stop trying to copy them. If i want trending topics I'll go to Twitter. My Facebook is a way to contact people i have not talked to in months.
Search results etc: What search??? Facebook as a search engine now? Or are you talking about the BS do you know John? You should know john! - FB: If I have not added John, it might mean i dont want to talk to John cause he is an a@@. Until your AI learns to detect aholes then please switch off this crap.
I know its a rant, but Im so tired of services trying to fill my feeds with crap, and also the crap BS about AI saving / ending the world. It might create nice cool articles for the news, and buzzwords in the bubble out west. We can only hope that Winter will not come this time, and we can advance far enough that AI becomes an actual useful product for more then spam filters, and detecting if John went to the same school as me.
* John is an invented person based on people i know. * Swearing is blocked as I'm not sure what people prefer on here. * Yes, I know AI is advancing a lot, and I'm a big supporter. I just don't like how its used, and how the PR makes it look like its the 7th wonder of the world.
You probably use FB for one hour per month because it's the most convenient way to catch up with your friends and family and they spend 4 hours per day because for them the internet is Facebook. FB uses statistics to become better for them, not for you.
Or you could just install Social Fixer, and control your feed that way. However, in my experience, it's quite hard to produce a better feed than Facebook's AI...
You can also game this system by getting others to interact with your posts more often, which will push your posts into their feed at a higher rate.
This argument seems remarkably similar to the sort of argument once used to justify the dictatorial rule of colonial strongmen. "You don't understand these people, they don't value the thing you value, for them a sense of security is much more important than choice." etc.
Which is to say that its pretty patronizing assumption.
I use FB for hours a day, Facebook is probably my main Internet outlet. I think it's problematic that FB substitutes their version of how they expect you to filter things for giving their users control over how their news feed is filtered. Lots of "unsophisticated" users complain about this and there are even cargo-cult methods used by the unsophisticated to get control over their feed (such the "disclaimer post" purporting to give the average Facebook user some extra legal rights over their posts).
I like what you did there :P
The optimist in me says that the reason this hasn't been open sourced is because a lot of the distributed code is tied specifically to FB's infrastructure - but I guess we'll wait and see.
Dismissing a blog post simply because it describes closed source software is silly.
seems to be mostly about the post
> Machine Learning Automation
Tuning models is a pain, and if you don't understand the model and all its parameters (like most software engineers who's job is the write software and not build models!) it takes time and can be immensely frustrating. Randal Olson[1] just announced TPOT[2], a Python tool that "automatically creates and optimized machine learning pipelines using genetic programming". This is going to be a huge lever for engineers wanting to experiment/implement with ML algorithms.
[1] http://www.randalolson.com/2016/05/08/tpot-a-python-tool-for... [2] https://github.com/rhiever/tpot
I know that folks are having some success with Gaussian process optimisation for hyper parameter tuning (eg https://github.com/Yelp/MOE), but I would be sceptical about the application of genetic programming to this area, particularly as broadly defined a search space as TPOT seems to set up (genetic programming either doesn't make assumptions about the objective function being optimised which can be used to search more efficiently, like Gaussian optimisation does - or uses implicit assumptions that may be unsuitable, depending on how you want to think about it); has anyone seen benchmarks against other search methods? The workflow does look great though - integrating it into scikit looks really clever.
https://github.com/calvinschmdt/EasyTensorflow
"This package provides users with methods for the automated building, training, and testing of complex neural networks using Google's Tensorflow module. "
http://docs.cascading.org/cascading/2.0/javadoc/cascading/fl...
http://submitteddenied.github.io/azkaban2/documents/2.1/crea...
- Lua/Torch appears to not be designed for distributed systems as much, and looks less mature in terms of the scientific computing side. - PHP is clearly the wrong tool for this job - it's only really the right tool for websites (debatable). - Could do Node.js, but has little/no data-science/scientific computing side.
C++ would be the only reasonable competitor to Python here I think. Python is a pretty good combination of great for science, good for servers and easy to pick up and use for all the engineers who will need to interact with it.
> The body of the workflow looks like a normal Python function with calls to several operators, which do the real machine learning work. Despite its normal appearances, FBLearner Flow employs a system of futures to provide parallelization within the workflow, allowing steps that do not share a data dependency to run simultaneously.
"Futures" did they implement this with concurrent.futures (which also has a backport prior to Ptyhon 3.2)?
> Operators: Operators are the building blocks of workflows. Conceptually, you can think of an operator like a function within a program. In FBLearner Flow, operators are the smallest unit of execution and run on a single machine.
> Channels: Channels represent inputs and outputs, which flow between operators within a workflow. All channels are typed using a custom type system that we have defined.
I think this is related to the thoughts in the following essay: https://colah.github.io/posts/2015-09-NN-Types-FP/
When you think about it. What do NNs fundamentally do (once trained)? Map inputs to outputs, right? So they behave like functions. In type theory functions are intensional and we explicitly construct them. With NNs we have our ML algorithm and construct them based on training data and feedback. With regular functions we have explicitly constructed A -> B. Let's represent NNs as A ~> B. Let's represent channels as A||B. Say we have two operators: E ~> F and G ~> H then to join them,
So E ~> F||G ~> H we'd have to construct a channel F||G.
At least this is how I imagine they mean things! Corrections from FB most welcome!
I could be wrong, but the Facebook team has built an Python internal DSL for building DAGs of operators in a procedural manner, helping maintain type-safeness between the input channels and the input-and-outputs of each operator.
Really cool idea.
Because all the channels and operators have typed interfaces, it means they can then materialize UI for configuring any collection of workflows. React would be ideal for this, mapping types to React components. Would be awesome to know if they built up a GraphQL layer on top of the workflow config parameters to get really elegant frontend-backend binding.