Show HN: Seldon – Open Predictive AI released on GitHub under Apache 2.0
seldon.io
seldon.io
Sign up for early access to betas here: http://eepurl.com/6X6n1
Thanks!
Seldon is available on Github: https://github.com/SeldonIO/seldon-server Technical docs: http://docs.seldon.io
# User Clusters Improve relevance of recommendations in high churn media services. - Cluster users based on historical activity. -- configurable taxonomy (category, price range, brand, visit referrer, .. -- unsupervised (fuzzy k-means) - Apache Spark to handle large historical data sizes. - Load user clusters into front-end servers periodically and count content hits for users in same cluster - Decay counts to provide activity dynamics as new content is published. - Recommend by combining counts for content based on cluster membership of user. - Real-time stream processing for adding short-term dynamics to recommendations.
# Item Activity Correlation Built for static slowly changing historical inventory - Similar to Amazon’s “people who bought this also bought…” - Use historical user activity to find items that share similar user activity. - Apache Spark scalable offline implementation. - Upload for each item: top-N similar items. - For each user: item recommendations based on their historical activity.
# Topic Models Built for sites needing long tail recommendation - Assume activity is associated with a set of topics. - Users individuals tastes are covered by a subset of topics. - Describe users by the set of keywords for the items they have interacted with. - Built with Apache Spark and Vowpal Wabbit implementation of Latent Dirichlet Allocation. - Online serving layer scores user association with items in real time.
# Latent Factor Models Best for e-commerce sites lower churn sites - Netflix Prize-winning solution. - Use Matrix Factorization to reduce activity matrix to two low dimension user and item factor matrices. - Load factors into API servers and score users and items in real time. - Fold-in new users and items until next batch update of model. - Utilize Apache Spark mllib and streaming modules.
# Content Similarity Built for services with rich metadata and high sparsity - Requirement – fast content based technique to match user history to similar content based on text/tags of content. - Utilize random vectors technique. Each word/tag is assigned a random high-dimensional vector. - Open-source Semantic-Vectors and word2vec implementations. - Periodically process recent content into vectors and update servers. - Servers load vectors into memory. - Recommendation on recent user activity to find similar content in real-time.