Well, thinking more about it leads me to tf-idf and naive bayes (of course), at which point you pretty much already have a classifier. So it seems feature selection is learning in itself and defines the maximum accuracy you'll be able to reach ? This is border philosophical but I'd love to read more about these matters. Pointers welcome !