End-to-end implementation of a machine learning pipeline (2017)
spandan-madan.github.io
spandan-madan.github.io
A non-academic observation - the 'real-world' challenge of ML pipelines is what I call the 'last-mile' problem of ML - operationalizing your model. You begin to run into problems of:
1. How often do you 'score' live data? How will this affect latency, data ingestion etc?
2. How often do you have to update your weights, if you want your model's performance to be consistent?
3. Integration with source systems
4. If you build your final scoring model on library-dependent languages like Python, how do you ensure no breakages? (Docker solves this to a large extent though)
[1] https://ai.google/research/pubs/pub46555
[2] https://developers.google.com/machine-learning/rules-of-ml/
- https://eng.uber.com/ - https://code.fb.com/
and many more. Google also publishes papers on various engineering practices obviously, some ML-related, but I can't find a blog where they focus on that specifically.
Also it's not "to keep up to date", but there's a great paper (from Google) that's often cited:
Machine Learning: The High Interest Credit Card of Technical Debt https://ai.google/research/pubs/pub43146
It talks about issues you face over the long run (I've experienced some of those). It also provides interesting pointers for further reading, e.g. about "pipeline jungles".
If others have pointers, I'm curious to hear about them as well.
I see a similar pattern with Pandas: some people use Pandas not because it's the right tool for the job (Pandas has many strengths), but because they're scared of writing comprehension loops and basic data structures. To avoid the CS-y stuff. But without the CS-y stuff, the result ends up a mess of lambdas, weird reindexing and buggy copy/view semantics.
And then "the next guy", the one who's job it is to clean up and productionalize the maverick's output, ends up having to reinvent and fix the entire solution. Basically doing both jobs.
I had a "data scientist" submit notebooks to us as if we could ship any of that in production. (We fired him.) It's for hacking and blogging, not for production work.
For future tutorial suggestions, mail me at smadan@mit.edu. A new one on NLP is coming soon!
Genre_ID_to_name=dict([(g['id'], g['name']) for g in list_of_genres])
In other places, you would benefit a lot from the enumerate(..) function, which returns (index, item) tuples when called on a list.
It is a suspect proposition that anything is gained by turning 3 lines of code into one line of code. Unless it is javascript for the Google homepage or somesuch where the bytes matter. Moving code from a bad data model to a good one usually correlates with a big reduction in line count, but the gain is in choosing more appropriate data structures and not in the number of lines removed.
Every reader of code, including the author after 3 months, is going to have to read and understand the code from scratch. One line doing a multidimensional transform of the data is going to scan for a small fraction of people. That one liner would take about 3 times as long to understand as any one line of the tutorial code. The data model hasn't changed either. If anything, I'd argue that the nature of the transform being done is clearer in 3 lines.
> It is a suspect proposition that anything is gained by turning 3 lines of code into one line of code.
Some code is way more verbose than its description would be. A named function signals intent:
function multidimensional_transform(the_data)
And for Javascript ES2015 there is new syntax (like spread operators) that improves code the same way.In the end, choose code style that matches level of your audience. If you write a one-off python script that automizes some build process in mostly non-python codebase, it should probably be very verbose and easy to understand. If, on the other hand, you're writing code in a decently advanced codebase and most of your colleagues are fluent in the language (or at least, supposed to be), it makes sense to use as much condensed syntax sugar as possible.
id_to_name = {g['id']: g['name'] for g in list_of_genres}
And for i in range(len(list_of_genres))
is really a dangerous antipattern better replaced with for genre in list_of_genres: for idx, genre in enumerate (list_of_genres)Also, looks like that notebook needs an update.
For me personally, I find this off-putting. Let your content speak for itself. No added credibility when the affiliation is advertised like this.
Content looks good, though! :)
[0] - https://github.com/Spandan-Madan/NLP-Intuition-and-Applicati...
- some values in the command outputs dont match the author's comments (or maybe I misunderstood some?)
- there are some big red blocks of errors in the outputs
- the outputs of the trainings are way too verbose for mobile reading
I guess they are issues on Jupyter's framework side. It would be nice if mobile were treated as first-class viewer.