PredictionIO – A machine learning server
github.com
github.com
https://github.com/PredictionIO/PredictionIO/blob/develop/process/engines/itemsim/evaluations/scala/topkitems/src/main/scala/io/prediction/evaluations/itemsim/topkitems/TopKItems.scala
Why is this structured like this? These directories seem ridiculous to me, it would be great if someone could explain.sigh
/develop/process/engines/itemrec/algorithms/hadoop/...
cascading/popularrank/src/main/java/io/prediction/algorithms/...
cascading/itemrec/popularrank/PopularRankAlgo.java
Most of these contain a single folder. io/prediction/evaluations/itemsim/topkitems/
Is the folder structure required by the package definition.EDIT: mostly it seems to be the result of the project being constructed as dozens of separate modules each with it's own build process...yikes...
* process/engines/itemsim/evaluations/scala/topkitems
They're using sbt's awesome multi-project feature here (http://www.scala-sbt.org/release/docs/Getting-Started/Multi-...) so basically every "sub-project" that makes up the whole can have its own dependencies, options, versions, etc while also maintaining which projects depend on each other. This really helps keep all of the logic separated and sbt deals with all the compilation-order madness that ensues when you have a tangled nest of inter-dependencies.
Note: this isn't necessarily reflected on what is published as most projects that do this will still publish it as one single jar file or project; it just helps with development and really helps with compilation speed (in my experience).
Again not really condoning what they're doing here as they're really taking it to an extreme; for my "big" project I basically just have a top-level "modules" folder and each sub project is one below that. You can see how the hierarchy is defined here: https://github.com/PredictionIO/PredictionIO/blob/develop/bu... which I find to be quite human-readable but I've been using sbt for years so ymmv. The customized settings for that particular project are here: https://github.com/PredictionIO/PredictionIO/blob/develop/pr...
* src/main/scala
This is the basic structure of an sbt project. By default, you put all of your code/resources in the `src` directory. Then you have two directories, main and test(optional) which is how you seperate code/resources that belong in the final project and which is just used for testing. The last level there are three (default) directories that are processed: java, scala, and resources. The first two should be pretty self explanatory and the last is where you put any files that you need to be packaged/available to your project. So if you have main/resources/aDir/logback.xml then you can reference that (via class resources which is a java thing) with "aDir/logback.xml" (I didn't include a leading slash because it's ~complicated).
Example layout:
src
- main
- java
- resources
- scala
- test
- java
- resources
- scala
* io/prediction/evaluations/itemsim/topkitems/TopKItems.scalaIn Java it is mandatory that your package name be reflected in your directory structure. So here we can see that the TopKItems class is in the package "io.prediction.evaluations.itemsim.topkitems" if they followed that convention. As hinted at, scala does not mandate this silly requirement but it's considered best practice to follow along as it keeps things separated and easy to follow. Scala projects mostly used short package names so it isn't as nested as this.
This might all make it seem that development would be a nightmare trying to manage everything but all of this integrates beautifully with a good IDE such as IntelliJ (which is the recommended one for scala -- eclipse is just way too slow and freezes constantly, even on beefy machines). You just run a quick gen-idea command and the entire thing is recognized by intellij, sub-projects and all. You never even see the crazy nesting of folders!
P.S. I mostly just lurk here so I'm sure I butchered the markdown. Sorry.
[1] https://dl.dropboxusercontent.com/u/2938195/predictionio-ema...
"You starred Play so come look at our thing" is not. It's an email blast. It's spam.
Note: Please be patient. It may take a long time to train the data model the first time even for very small dataset. It is normal because PredictionIO implements an distributed algorithm by default, which is not optimized for small dataset. You can change that later.
Sums up my experience with the Mahout/Hadoop world nicely. Not a good fit for small-medium projects -- too complex, too cumbersome, too slow. By the time you really need the scale (=often never, save for using "Big Data" for marketing), you're big enough and know enough about your domain to roll a custom, efficient, domain-optimized solution.Bringing machine learning to the masses is an honourable goal though, so thumbs up for PredictionIO.
I wonder if this title was given in jest.
i think letting the community upvote and downvote the titles themselves would give clear indication to mods for changing them. instead, they choose to justify doing nothing by saying they dont have the resources to read all articles and evaluate each of them.
just seems like intentional friction and reluctance to fix a recurring and aggravating issue that has so many viable solutions.
I'm asking, because for us (dawanda.com, one of the biggest ecommerce platforms in germany) most of the development effort on our soon-to-be-opensourced recommendation engine was spent on scaling the CF up from a few thousand test records to a 150 million record production data set.
In the first iteration we also built it completely in scala, but as we were putting more and more data into it, memory usage was exploding. We realized that boxed types had too much overhead and that we had to implement the whole sparse rating/similarity matrix in C [1]. Also we decided to go for a hybrid memory/disk approach which allowed us to process 80GB datasets on a machine with only 64GB main memory.
How did you manage to solve the memory consumption issue for prediction.io in scala? Did you use java raw memory access or did you also swap out data to disk/ssd?
Computation time and resource requirement depend on the choice of technology. If a non-distributed implementation is chosen using the framework, the rule of thumb from Apache [2] is a good guideline. For distributed implementations based on Hadoop, the 10M MovieLens data set [3] finish training on a single m1.large AWS instance (7.5GB RAM) within 30 minutes. Although we do not have an accurate account of how much computation time and resource will be required for your production data set's scale, a user has reported using his own production data set of similar size with 2M users, and finished training in about an hour using Amazon EMR.
That said, PredictionIO does not do anything special on memory consumption or has a special memory access model. It really depends on the underlying libraries that do the actual work.
We imagine your project requires a much faster turnaround time according to your spec, which is an interesting application to us as well.
PS. The work you posted is pretty cool. :)
[1] http://mahout.apache.org/ [2] https://cwiki.apache.org/confluence/display/MAHOUT/Recommend... [3] http://grouplens.org/datasets/movielens/
It was pretty straight-forward to implement and has a nifty backend to it as well for managing the algorithms.
I don't have the full engine source code or benchmarks to share right now (hopefully in the next two weeks as it is approved by our legal department), but I can tell you all the hard lifting is done by a library called "libsmatrix", which implements a fast, memory-efficient sparse matrix data structure that is used at the core of the CF algo. libsmatrix is 100% threadsafe and persists to disk (this way you can also work with datasets larger than the available main memory). this library is already open sourced:
https://github.com/paulasmuth/libsmatrix
Using libsmatrix, you can build the rest of the "engine" in a matter of a hundred lines of C or so... #include "smatrix.h"
smatrix_t* my_smatrix;
// libsmatrix example: simple CF based recommendation engine
int main(int argc, char **argv) {
my_smatrix = smatrix_open(NULL);
// one preference set = list of items in one session
// e.g. list of viewed items by the same user
// e.g. list of bought items in the same checkout
uint32_t input_ids[5] = {12,52,63,76,43};
import_preference_set(input_ids, 5);
// generate recommendations (similar items) for item #76
void neighbors_for_item(76);
smatrix_close(my_smatrix);
return 0;
}
// train / add a preference set (list of items in one session)
void import_preference_set(uint32_t* ids, uint32_t num_ids) {
uint32_t i, n;
for (n = 0; n < num_ids; n++) {
smatrix_incr(my_smatrix, ids[n], 0, 1);
for (i = 0; i < pset->len; i++) {
if (i != n) {
smatrix_incr(my_smatrix, ids[n], ids[i], 1);
}
}
}
}
// get recommendations for item with id "item_id"
void neighbors_for_item(uint32_t item_id)
uint32_t neighbors, *row, total;
total = smatrix_get(my_smatrix, item_id, 0);
neighbors = smatrix_getrow(my_smatrix, item_id, row, 8192);
for (pos = 0; pos < neighbors; pos++) {
uint32_t cur_id = row[pos * 2];
printf("found neighbor for item %u: item %u with distance %f\n",
item_id, cf_cosine(smatrix, cur_id, row[pos * 2 + 1], total));
}
free(row);
}
// calculates the cosine vector distance between two items
double cf_cosine(smatrix_t* smatrix, uint32_t b_id, uint32_t cc_count, uint32_t a_total) {
uint32_t b_total;
double num, den;
b_total = smatrix_get(smatrix, b_id, 0);
if (b_total == 0)
b_total = 1;
num = cc_count;
den = sqrt((double) a_total) * sqrt((double) b_total);
if (den == 0.0)
return 0.0;
if (num > den)
return 0.0;
return (num / den);
}> PredictionIO, a machine learning server for data engineers and software developers.
It strikes me as a lot of overengineering with no real meat at the core.
edit
Why downvote? This is a good question. smh
It. Solves. Everything.