Ludwig, a code-free deep learning toolbox
eng.uber.com
eng.uber.com
This should work well with a Feature Store, where features are already pre-processed and ready for input to model training. With a feature store, this could be like the Tableau/Qlik/PowerBI tool for Data Science.
Regarding the code-free vs declarative, I mean, it really depends where you draw the line :) Is calling a command line program with some parameters declarative programming? What if there are a lot of parameters? And what if those parameters have a nested structure and are contained in a YAML file? And what if the parameters become so many and so detailed that you can write almost each single operation than the command line program will perform (like in Caffe configurations files for instance)? Do you see what I mean? Anyway, if it's not code-free you may agree that it's the closest open source tool to be code-free :)
* The field is evolving quite rapidly, and so most options will require a lot of configuration, which IMO is not suitable for a declarative approach.
* It's hard to debug, eventually you will need to dive into the code at some point.
* Extending the library would mean touching the code anyways.
One major difference here, is that Ludwig tries to expose one interface, that for instance users without python knowledge can use, and inside a company, some machine learning engineers can extend the tool and do support by debugging other users use-cases.
I think at some point such approach can be useful for a very limited mature use-cases.
[1]: https://github.com/polyaxon/polyaxon-lib/blob/master/example...
* That may be true, but I think it also depends on the level of abstraction. Fon instance in Ludwig the encoder, although configurable, are pretty monolitic. It's a trade off with flexibility, in Caffe or in your configuration file you are super flexible in specifying each single operation, while in Ludwig that is abstracted from you, but the advantage is that, as long as you trust the encoder implementation, those encoders mimic papers / state of the art models and require much less configuration to be written in order to run (in many cases if you are happy with the default parameters you don't have to configure them at all). So if the field moves and a new text encoder comes in, one can easily add a new encoder to Ludwig too. It's a dangerous game to play catch up, but hopefully releasing it as open source and making it easy to extend may help spontaneous contributions from the community.
* debugging is an issue, that is true, I answered to another post about that, but again, it's a matter of tradeoffs. Debugging is kind of a nightmare in SQL too for instance, or is in TensorFlow (even if the tfdbg improved things a bit).
* extending requires coding, that is also true, but if you have an idea for a new encoder for instance, all you have to do is implement a function that takes a tensor of rank k ad input and output a tensor of rank z as output, and all the rest (preprocessing, training loop, distributed training, etc.) comes for free, which kind of a nice value proposition imho and let you focus on the model rather than all the rest.
Thanks for the interesting discussion!
It's nice to see other, (much more) reputable engineering organizations taking a similar approach and treating the construction of different predictive models as an exercise in configuration.
Although in my solution I haven't looked at any DL models, and typically default to feeding everything through XGBoost and performing a grid search for the best hyperparameter config there. My product is basically focused entirely around taking a raw dataset and a configuration file and producing an analytic dataset off of which algorithms can be tested.
I'd be really interested in hearing others' experiences with this type of stuff.
So for example I'll have a section that says "data" and with variables "binned = ['var1', 'var2']" and likewise for log/power-transformed variables, and have some Python code turn that into column transformers. In other examples there's even more custom logic hidden (some stuff from my master's thesis that I'm not at liberty to discuss) so it's not just a matter of reusing scikit-learn.
I use TOML for little languages that are really flexible configuration files, and YAML for situations where there may be multiple similarly-specified objects (because the hierarchical syntax in TOML is somewhat obscure).
strategies = [
{
'name': 'column1',
'kind': 'categorical',
'strategy': 'ohe'
},
{
'name': 'column2',
'kind': 'continuous',
'strategy': 'center'
}
...
]
It has worked very well thus far. I received a request from a stakeholder a few weeks back about building a new model using a slightly different target. I told him it'd take a couple of days at least (it was a very similar target), but I had it completed (at least from a re-engineering perspective) in about 5 minutes. I simply changed the target variable config and removed any leaky features after changing the target.I'm convinced that this approach, coupled with configurations for tree-based models like minimum samples per leaf and max depth, is the most efficient way of building predictive models. Those configurations specific to tree-based model software help to skirt things like the rule of five, etc., IMO.
I'm looking into more general parsing that would allow me to define semi-verbose little languages. In my heart Stata is still the gold standard for rapid fire usability, even at the cost of idiosyncrasy; I'd like to further make the engineering of, say, REST APIs more and more code/language independent and more logically specified, since a lot of the new "data science" crowd coming from stats and applied maths can't code their way out of a Tequila Sunrise.