Seldon – Serves predictions via a REST API
github.com
github.com
Our plan is to build out and support the strongest ecosystem of data scientists in the predictive machine learning space. We adopted the permissive Apache 2.0 license, so you are free to build domain-specific algorithms and functionality to either contribute back into the core platform or release under a separate license.
We would love to hear your feedback, both first impressions and thoughts as you start using the platform. And, of course, how you would like to see the project evolve in the long term.
Suggestions on the functionality that you would find useful would also be appreciated - we're hooked on customer and developer community feedback! :)
As a developer while we can try VM setup and as well as Amazon instance, can you also host it on one of your server and provide a demo access to developers to get acquainted with system and architecture before they can try on VM or Amazon AMI?
We were originally a SaaS recommendation engine and will continue to provide a fully-managed cloud-based solution. It’s now serving recommendations on hundreds of millions of page views per month.
There are a few pretty significant technical tasks we need to complete to make the API provisioning streamlined enough to support trial accounts since there are some manual steps required to configure the metadata properly, etc.
Since we decided to go open-source last September, all of our R&D focus has been in preparing this open-source release instead of the self-serve. Streamlining our cloud service is in the roadmap, and it's now powered by the same codebase as the open-source project.
The feedback we received from developers who tried the AMI and VM are that it's super quick to get started. And as you can see we've spent a lot of time recently on documentation.
If you're interested in the hosted API, please email support@seldon.io.
Some thoughts on how to make the webpage more informative:
- Improve the schematic overview[1], e.g. what's the purpose of a logging cluster? General logging? The MySQL database, what's it's role? My initial reaction is that logging/sql-storage could be handled outside of Seldon. I may be wrong, but you should help people here.
- Same with load-balancers, seems like it should be outside the scope of a prediction service. If not, an explanation would help.
- Clarify which parts are open-source, and which ones are closed. E.g. with ElasticSearch, it's obvious on their front-page which components are open-source and which-ones are paid.
- Add a comparison to prediction.io (RethinkDB's comparison with MongoDB[2] is a good example). It's an obvious question people will have, answering it upfront is helpful. Keep it fair/unbiased.
[1] http://www.seldon.io/open-source/ [2] http://rethinkdb.com/docs/comparison-tables/
You're right about the load balancers - this doesn't form part of the project codebase, but it's important for people to know that the API servers are stateless and can scale horizontally. However, the other infrastructure components are required as Seldon provides the full end-to-end platform not just collection of predictive algorithms.
We will review your comments and for now we linked to the infrastructure page in our docs: http://docs.seldon.io/tech.html
We are planning to release some functionality under a separate commercial license, just as anyone is free to do. Everything you currently see on the website and Github repos is open-source and licensed under Apache 2.0.
Try this:
I was looking at using Prediction.IO and this looks suitable too, can you elaborate on the high level differences between this project and prediction.io?
I'm also curious about horizontal scaling.
With regards to horizontal scaling, there are two parts to consider. Creating the models and serving the recommendations (the Seldon server project).
Model creation is done in a variety of ways, but but can be managed with scalable technologies such as Spark.
The Seldon Server project can be deployed on as many machines are you require and they will work together to provide recommendations behind a load balancer. We have experience working with some very large news websites so this part of our technology is well developed.