Real World Recommendation System
blog.fennel.ai
blog.fennel.ai
We had an api layer where another team runs inference on their model as new user data comes in, then streams it to our api which inboards the data.
On top of this, you have extensive A/B testing systems
In practice, good old Matrix Factorization works really well. Can you beat it with a huge team and tons of GPU hours to train fancy neural nets? Probably. Can you set up a nightly MF job on a single big machine and serve results quickly? Sure can.
The point here is that you don't typically need a huge amount of computation power to serve recommendations, even if the underlying model is sophisticated and required a lot of computation to train.
Likewise for data access, the online recommendation system typically does not need full access to the databases that the researchers need access to.
https://www.linkedin.com/pulse/personalized-recommendations-...
If you want cross entropy loss, use word2vec.
Follow it up with a lambdarank ranker.
I’d love to see a lightweight, flexible recommendation system at a low level, specifically the scoring portion. There are a few flexible ones (Apache has one) but none are lightweight and require massive servers (or often clusters). It also can’t be bundled into frontend applications which makes it difficult for privacy-centric, own-your-data applications to compete with paid, we-own-your-data-and-will-exploit-it applications.
It has everything you need at a platform level to build a production recommendation system given that it’s the engine that powered a lot of yahoo product’s search and recommendation capabilities. I have been experimenting with it, the number of capabilities are immense. It’s really an untapped resource.
Take a look at the features: https://vespa.ai/features and the ranking syntax: https://docs.vespa.ai/en/ranking.html
Really cool stuff, I haven’t even scratched the surface of what it can do.
I mean it does. As far as I'm aware Facebook's ad platform is mostly backed by hundreds of thousands of Mysql instances.
But more importantly this post really doesn't describe issues of scale.
Sure it has the stages of recommendation, that might or might not be correct, but it doesn't describe how all of those processes are scheduled, coordinated and communicate.
Stuff at scale is normally a result of tradeoffs, sure you can use a ML model to increase a retention metric by 5% but it costs an extra 350ms to generate and will quadruple the load on the backend during certain events.
What about the message passing, like is that one monolith making the recommendation (cuts down on latency kids!) or micro services, what happens if the message doesn't arrive, do you have a retry? what have you done to stop retry storms?
did you bound your queue properly?
none of this is covered, and my friends, that is 90% of the "architecture at scale" that matters.
Normally stuff at scale is "no clever shit" followed by "fine you can have that clever shit, just document it clearly, oh you've left" which descends into "god this is scary and exotic" finally leading to "lets spend half a billion making a new one with all the same mistakes."
[1] http://people.csail.mit.edu/matei/courses/2015/6.S897/readin...
Kind of. It's part of the recipe but one you find at these large tech companies (I've worked at FB and GOOG) is they have the resources to bend even large/standard projects like MySQL to their will, while ideally preserving the good ideas that made them popular in the first place. There are wrappers/layers/modifications/etc that eventually evolve to subsume the original software, such that is acting more like a library than a standalone service/application. So, for example, while your data might eventually sit in a MySQL table, you'll never know, and likely didn't write anything specific to MySQL (or even SQL) to get there.
What you're describing sounds like you mean something on the level of Cockroach, talking the Postgres wire protocol but implemented entirely independently underneath (which came indirectly out of Google). Facebook's MySQL deployment sounds more like a heavily-patched-but-basically-MySQL installation. I think Facebook is overanalogised to Google sometimes, as an engineering org.
(Admittedly I haven't worked at either whereas you have - though I have at another FAANG fwliw - but am basing this impression partly on what I hear from friends & partly on plain old stuff I read on the internet.)
Same for YouTube itself https://www.mysql.com/customers/view/?id=750 and they use Vitess for horizontal scaling: https://vitess.io/
Talking w/a friend who works at Netflix, it sounds like this is a warranted assumption. The way he told it, they were tearing their hair out at one point b/c users wouldn't put much into it.
You just keep your simple interface, but allow the power users to, say, click through to a particular menu and change their setting – the setting in this case being ~"let me provide feedback / configure how recommendations work". For that kind of user, finding a 'cheat code' is actually a gratifying product experience anyway.
I believe it can also have performance implications especially for things like recommender systems where you are depending a lot on caching, pre computation and training.
TikTok, on the other hand, has way more data. Things like time-to-swipe, shares, comments presumably form the basis of some sentiment metric.
However, a lot of usecases are time insensitive rankings. Like recommending content on netflix, spotify etc. (spotifys discover weekly even has a one week! request time :D).
In which case you can just run your ranking and store the recs in your DB and its much much easier.
https://blog.youtube/news-and-events/youtube-now-why-we-focu...
https://www.eugenewei.com/blog/2020/9/18/seeing-like-an-algo...
There’s a more technical recent paper from bytedance as well: https://arxiv.org/pdf/2007.07203.pdf
And another recent one on bytedance user profile system, this paper gives the deepest understanding of their recommendation system: https://www.cs.princeton.edu/courses/archive/spring21/cos598...
There’s a good breakdown of the recent nytimes article on TikTok’s internals here: https://read.deeplearning.ai/the-batch/issue-122/
If anyone has found anything better let me know.
If you try searching yourself you will want to try switching between the keywords “douyin” “toutiao” “bytedance” and “tiktok.”
For example, the HN front page is a recommendation system if you literally mean system-that-recommends-web-pages-to-look-at. But it's not personalized; every visitor sees the same front page. This fundamentally makes it a different sort of thing.
I don't think they're really paying much attention to the dimension you're splitting it along, i.e. whether the recommendations are personalised for each user. The huge important idea they have in their head is that recommendations can apply to user-generated social media content too.
* HN and classic Reddit sort their items on a single dimension ("hotness"), calculated using a few input variables and producing a single output variable. This is about as cheap to calculate as recommendation systems get. The XKCD comment recommender is a bit more complex, but still in the same complexity class. Since the whole point of an algorithm like this is to be timely, the naive approach is to compute it on-the-fly, which it's perfectly simple enough to manage.
* At a somewhat more complex level, you get stuff like a basic, uncustomized Similar Items list. If YouTube has no data on you, this is what you get from their sidebar recommender (and their front page would be analogous to Reddit and HN, but sharded by region and language). It's also pretty close to what AdWords used to be, before they started doing user profiling. The thing with this method is, even though it involves some level of AI, it's presenting the same thing to everyone and it's expensive, so the natural solution is to precompute it.
* Personalized recommenders are the worst of both worlds. You can't naively compute it on-the-fly, because it's too slow, but you also can't naively precompute it, because there's a combinatorial explosion of users and items. You actually have to be clever about it.
I know that they probably optimize for ads etc. but if they actually showed me videos I would like to see, then I would spend more time on the platform.
Not directly, most people believe they optimize for session time. It tries to serve you a result that will keep you on the platform watching videos as opposed to needing to keep scrolling or leaving the platform. They truely do want to serve the best videos that they can to you that they think you will be interested in watching. Thankfully for YouTube you watching videos and YouTube getting ad rev is correlated.
>where will you go and watch medium to long videos created by "normal" people?
No one forces you to watch that format of video. You can go on TikTok, Twitch, Netflix, etc and be entertained.
Imagine if FB were to pay a fraction% of how yur data was used and paid you for it...
It may be a small amount, but in super 4th world countries, it could affect change in their lives...
Now imagine that this becomes big... and it works well.
Now imagine that the populous is aware of the hand of god above them just pressing keys to affect land masses (yes I am referring to the game from the 80s)
but this cauterizes them into union building...
So when the people realize their metrics are the product to feed consumerism for capitalistic profits, and decide to organize, what happens?
Is FB going to need a military force to protect their DCs?
---
With "Zuck Bucks" (I still am not sure if true)
This makes this ultimate "company store"
Tokens?
So how get?
How EARN? (What service on FB GENERATES '$ZB'?)
How spend?
WHAT GET? (NFTs?, Goods? Services?)?
The entire fucking model of EVERYTHING FB DOES is to MAP SENTIMENT!
Sentiment is the tie btwn INTENT and SENTIMENTAL VALUE
The idea is to map interest with emotional drivers which make someone buy (spend resources their time and effort went into building up a store-of)...
---
So map out your emotinal response over N topics and forums.. Eval your documented Online comments, NLP the fuck out of that, see what your demos are and build this profile to you....
THEN THEN THEN THEN
Offer an "earnable" (i.e. Grindable by farms and bots alike) -- "Zuck Buck" which is a TOKEN (etymology that fucking word for yourself)
of value...
Meaning, zero INTRINSIC value, Zero accountability (managed by a central Zuck Bank) <-- Yeah fuck that)
And the vaule both determined AND available to you via not INTRINSIC CONTROL, nor VALUE.
---
FB Bots Galore.
Carry On with Dangs Blessing.
This doesn't make much sense to me since a recommendation is rarely needed instantly. Why not spend, say, 10 s constructing a better recommendation while the user is doing something else, during which the recommendation can simply be blank. Obviously if the user requests a recommendation on first visit, you're out of luck, but I'm thinking the typical use case is for a recommendation after the primary reason for visiting has been completed.
"the Recommender Systems research community is facing a crisis where a significant number of papers present results that contribute little to collective knowledge […] often because the research lacks the […] evaluation to be properly judged and, hence, to provide meaningful contributions"
https://doi.org/10.1145%2F2532508.2532513
More here... https://en.wikipedia.org/wiki/Recommender_system#Reproducibi...
Recommender systems is one of the few areas in ML where almost all of the knowledge is contained in industry, not academia.
And the reason it happens despite the 'invisible hand' etc is because it still works, it just happens to be horrendously inefficient. I think that's the main area of inefficiency in the industry: not in getting the job done, nor even arguably in accuracy - at least not severely - but in overcomplicating the solution[0] because we've formed a cargo cult around one particular method of optimisation, beyond all nuance.
[0] I mean 'overcomplicating' in absolute terms. Of course the very crux of my point is that, from the data scientist's perspective, it's not overcomplicated - it's less complicated than using e.g. ILP precisely because we have made libraries like TensorFlow so incredibly easy and tempting to use.
This part was the one I was interested in. As most of the rest are obvious.
Good feedback, noted. Will get the next post focused on training within the next couple of days.
2. For any single request, there are thousands of things to recommend from. As a result, a single request is not scoring a single ML model but thousands of models - one (or often more, see value modeling the post) for each candidate.
The devil's in the details, which are surely domain specific and hopefully not too morally questionable.
Millions of cores of compute, exabyte scale custom data stores. Good recommendations are expensive. If you try to build a similar system on AWS, you will spend a fortune.
Most recommender models just use co-occurrence as a seed, this can actually work pretty well on it’s own. If you want to get fancy then build up a vectorized form of the document with something like an an autoencoder, then use some approximate nearest neighbors to find documents close by. 95% of the compute and storage is just spent on calculating co-occurrence though.
And then it will be gamed, and become as useless as every other recommendation system already going.
For my money - and, for what little it's worth, I work in this field – I think most of the impressive feats of data science attributed to 'machine learning' are really just a function of now having hardware capacity so insanely great that we're able to 'make the map the size of the territory', so to speak. These models are essentially overfitting machines, but that's OK when (a) it's an interpolation problem and (b) your model can just memorise the entire input space (and deal with any inaccuracies by regularisation, oversampling, tweaking parameters till you get the right answers on the validation set, then talking about how 'double descent' is a miracle of mathematics, etc).
Don't get me wrong, neural nets are obviously not rubbish. They are a very good method for non-convex, non-differentiable optimisation problems, especially interpolation. (And I'm grateful for the hype cycle that's let me buy up cheap TPUs from Google and hack on their instruction set to code up linear algebraic ops, but for way more efficient optimisation methods, and also in Rust, lol.) It's just a far more nuanced story than "this method we discovered and hyped up for a decade in the 80s suddenly became the key to AGI".
These algorithms are not benign. They make choices about what information you consume, whose opinions you read, what movies you watch, what products you are exposed to, even which politicians messages you hear.
When people complain about the takeover of algorithms, they don't mean databases or web interfaces. They mean this: content selection or preference algorithms.
We should be deeply suspicious. We should demand greater accountability. We should require that the algorithms explain themselves and offer alternatives. We should implement better. Give control back to the users in meaningful ways
If software engineering is indeed a profession, our professional responsibilities include tempering the damaging effects of content selection algorithms.
Do you know how a TV channel decides to schedule stories?
Humans, its all humans. Looking at the metrics, and steering stuff that feeds that metric.
Content filters are dumb and easy to understand. seriously, open up a fresh account at FB, instagram, twitter or tiktok.
First it'll try and get a list of people you already know. Don't give it that.
Then it'll give you a bunch of super popular but click baity influencers to follow. why? because they are the things that drive attention.
if you follow those defaults, you'll get a view of whats shallow and popular: spam, tits, dicks and money.
If you find a subject leader, for example a independent tool maker, cook, pattern maker, builder, then most of your feed will be full of those subjects, save for about 10% random shit thats there to expand your subject range (mostly tits, dicks, spam or money)
What you'll see is stuff related to what you like and stare at.
And thats the problem, they are dumb mirrors. Thats why you don't let kids play with them. Thats why you don't let people with eating disorders go on them, thats why mental health needs to be more accessible, because some times holding up a mirror to your dark desires is corrosive.
Could filter designers do more? fuck yeah, be we also have to be aware that filters are a great whipping boy for other more powerful things.
I was smart enough to see what collaborative filtering (CF) could be early on, and to file a patent that issued. I wasn't smart enough to make it a complicated patent, or to choose the right partners so I could have success with it.
But the patent makes a good way to learn how to get from "what are your desert island 5 favorite music recordings?" over to "here is a list of other music you might like". Basic CF, which is at the core of a lot of this stuff. Enjoy!:
If you call that garbage, what on earth - or, for that matter, off it - is not garbage!?
MANTAMAN
-TFA
I'll top it off with an interview at Twitter with the Eng MGR ~2009-ish?
--
Him: So tell me how you would do things differnetly here at twitter based n your experience?
ME: "Well, I have no idea what your internal processes are, or architecture, or problems, so my previous experience wouldn't be relevant."
I'd go for the best option that suits goals.
[This was my literal response to the question, which I thought was a trap but responded honestly -- as a previous mgr of teams, the "well, we did it at my last company as such"]
Dont reply this way. <--
Here was his statement:
This is a literal quote from a hiring manager for DevOps/Engineering at Twitter:
"Thank god!, We have hired so many people from FB, where that was there only job out of school, and no other experience, and the biggest thing they told me was "well - the way we did this at FB was... X"
--
His biggest concern was engineering-culture-creep...
Wait, I got it, I would rewrite everything as AWS Lambdas. That's the right answer! Screw your (almost certainly SQL) DB, let's move it all to DynamoDB too.
I was stating that the eng mgr was relieved to NOT hear an answer of "the way we did it at company X, and it was successful for them, so I assume that the same approach maps to your company"
--
Are we talking about the same thing?
Have you literally never come across the "you're not Google!!!" trope before now, during the whole ~decade leading up to this very day? Gosh I envy you.
(Also, I am reaaally struggling to understand that story. Who is speaking? It sounds like a story within a story within a story. I can just about piece together the gist, but I'm very confused by all the formatting and nested quotes.)
"Fennel AI: Building and deploying real world recommendation systems in production Launched 18 hours ago"
Caveat reader.
Of course I didn't write the original comment and there's something to say for flag-and-move-on or whatever, and other people did enjoy it. I'm just saying I understand the impulse to short-circuit the entire tedious conversation!
I'd disagree with this pretty strongly since there are many examples of content marketing that are also very useful pieces of content.
Eh, the new parent company names aren't really what people know them as still. I don't think most people are even aware that Google has a parent company.
I have a friend that works at Google, and that's what we say. I don't think him or anyone would ever say he works at Alphabet.
But Google is still Google, and probably always will be. Just like Youtube is still Google, and Waymo is still Google.
But for every other one of the FAA(N)G companies, I can barely work a day as a developer without touching every one of their technologies. Yeah, Netflix got into ML years before most, but the netflix prize exists as a distant cautionary memory, and as an ML professional, I'd literally never heard of metaflow before. Just sayin'.
Nowhere was the argument made that somehow Netflix was more influential than Twitter/Uber/AirBnB, but your counter-argument that somehow it's less influential because you haven't heard of/used some projects directly holds no ground.
Oh come on, they are indisputably right that Microsoft, Twitter, Uber, Airbnb, hell, even Cloudflare are more technically influential than Netflix is.
Apple and Google would make anyone's top 5, that's his point. No argument about it. Their products collectively dominate anyone's life, along with MSFT. Netflix is maybe in your top 10, top 20 for sure, but it's not up there as one of the few 'platform that everyone's lives are built on' techcos.
(Like, Netflix vs Microsoft? Seriously? For that matter, Amazon probably wouldn't be in my top 5 either, and not only because it's not mainly a tech company. I s'pose it depends how you define 'Amazon', and if you include AWS. But for Netflix there's just no argument that they win a spot there.)
It's now been taken over by the tech industry to be shorthand for places that are highly selective in their hiring and tend to work on cutting edge tech at scale.
That being said, the impact of Netflix on tech is pretty big. They pioneered using the cloud to run at massive scale.
Which is to say they were AWS's biggest early customer? Doesn't really seem like Netflix should get the credit for that one.
Netflix tech even spawned a company to sell their open source tools:
And they codified the entire practice of Chaos Engineering:
That, and FAAG had less of a ring to it.
Edit: Dammit, the GP made the same observation. Oh well, I'm keeping it.
In Europe you nearly always see "GNAFAM", which includes Microsoft too. It's certainly weird to exclude MSFT, worth at times more than Amazon+Meta+Netflix combined.
I find it oddly poetic, but, this is my last day of magic.