Useful Python decorators for data scientists
bytepawn.com
bytepawn.com
1. @dataclass
https://docs.python.org/3/library/dataclasses.html
2. @cuda.jit
I use decorators quite a bit. The feature is incredibly powerful. Once you find @lru_cache, you start realizing how many ways you can take advantage of the feature.
from fastapi import FastAPI
from pydantic import BaseModel, constr, conlist
from typing import List
from transformers import pipeline
classifier = pipeline("zero-shot-classification",
model="models/distilbert-base-uncased-mnli")
app = FastAPI()
class UserRequestIn(BaseModel):
text: constr(min_length=1)
labels: conlist(str, min_items=1)
class ScoredLabelsOut(BaseModel):
labels: List[str]
scores: List[float]
@app.post("/classification", response_model=ScoredLabelsOut)
def read_classification(user_request_in: UserRequestIn):
return classifier(user_request_in.text, user_request_in.labels)
*: Production grade if used in combination with workers, A python quirk I felt is not relevant to the topic of decorators.Which means your dev and production environments are quite different, increasing risks of letting stupid mistakes slip into prod.
Besides switching between something like production and test you can catch a case where an envar isn't set or is an unexpected value.
@reloading to hot-reload a function from source before every invocation (https://github.com/julvo/reloading)
%load_ext autoreload %autoreload 2
now all your imports will be automatically reloaded.
So good that I have it in ipython startup scripts
Great for prod if an error is unnacceptable
(or more realistically, web scraping where you just don't care about one off errors)
It has some ast stuff built in.
@fuckit
def this_is_fine(garbage_in):
x = 1 / 0
return "garbage_out"
this_is_fine("") ### Returns "garbage out"I just wouldn't use them as decorators in proper code.
Applying these functions as decorators, i.e. with @, means you can't run the non parallel version, or the test not in production, etc.
In the end, decorators, though nice on the first day of usage, reduce composability by restricting usage to whatever you wanted on that same day.
(this is not a general remark, it doesn't apply to DSLs that use decorators, e.g. flask)
You could even write a higher order function, callable as a decorator, that would transform a decorator that didn't do this into one that was otherwise identical that did.
def my_func:
...
my_decorated_func = my_decorator(my_func)The decorator pattern is a well known one, where one "decorates" a function by passing it into another function. GP expresses that they would avoid the pattern with these decorators.
The decorator operator is essentially prefix notation of the form `f1 = f2(f1) = @ f2 f1` which is what the GP alluded to, i.e. that f2 is a higher order function since it takes a function and produces another function. In-fact, the @ operator is a higher order function as well since it takes 2 higher order functions.
I've found I like a decorator better than some envar testing custom logging/telemetry. If I have a custom_log function I use everywhere, I have to use it everywhere I might want it. With a decorator I can add it to only a few functions and get far less noise.
Those decorators are exactly what data scientists would do, while software engineers would be terrified.
At least surprised. There are solutions to these problems already.
Author seemed to have learned decorators and is enthusiastic about abusing them, instead of learning the stdlib.
The capturing of print statements. Why not use the Logging machinery, instead?
Or the @stacktrace, it seems what they really want/need is a debugger.
But anyway, if these solutions fit their programming style better, so be it.
When I create my own solutions, I tend to underestimate how hard it will be to get it working bug-free and to maintain it.
We use a library that is very...chatty (some function calls send a screenful of info/progress to the screen), and I think I'm going to steal this to make it quieter.
It would be more interesting to point out the parts you feel are so terrible.
Not everything has to be designed for a super critical prod environment with >10 coders working non stop on it.
You don't need a super critical prod environment to have decent code. Half of these are hardcoding environment configuration, others have hidden side effects that the caller of the function cannot control at all, and others are badly reimplementing things that already exist (@redirect -> you want logging for this, @stacktrace -> use a debugger)
That said, I'm not really defending these specific decorators.
I don't believe I've ever encountered this. Can you elaborate on what you mean?
Some folks _really_ insist on changing everything to be "modern" and follow "best practices" using "up to date tooling" (invariably for a non-consensus-but-very-cool definition of "modern" and "up to date"). Often, it's switching to something that's only been around for 6 months over tooling that's _incredibly_ well supported and has been around for decades. I'm not opposed to using something new, but give me a reason beyond "it's new and everyone uses it now". That's doubly true when the new approach has the old tooling as a dependency and is basically a different interface to the same things (i.e. adding a dependency without taking one way).
There's lots of things that need to be updated to be more modern, sure. But there's also a trap of lots of tempting-but-relatively-low-value "best practice" updates that some folks will insist on spending 100% of their time on.
Another common example is some variant of this situation:
"Yes, X looks like a wart and is for many common use cases. It's there because of functionality Y needed by projects A,B,C. Downstream projects D,E,F,G already have workarounds in place where it matters. If you remove the wart, it breaks key functionality for projects A,B,C and means that D,E,F,G have to change the way they use this. Sure, you could handle this in a different way that could be a bit cleaner, but is it worth changing? Changing it is non-trivial and means a bunch of other people suddenly need to do extra work for no clear benefit. Oh, you really think it is, and want to devote the next 6 months to doing that and only that..."
Sometimes things really need some love and attention to get up to date. However, it's also important to avoid work that's temping to do, but low-impact and high-risk (in terms of unintended consequences).
But as I'm recently trying to improve my software skills, I notice that while those are indeed useful in the short term, in the long term they are not worth the price. The @production one, seems like a disaster waiting to happen.
And even when it does, cargo-culting rules-of-thumb is generally the wrong way to do that. Best practices are better treated as the Pirate Code than the Divine Writ.
Actually, when I read the post I'd guessed this is what an ex Software Engineer who is now a Data Scientist would do. And looking at the author's LinkedIn confirmed it.
You have to have a software engineering background to come up with this stuff in the first place.
It sweeps complexity under the rug with a deceivingly simple facade.
If your code links against C/C++ extensions that have internal printf calls, then you need something lower-level, which can intercept the data at the file handle.
There's a great solution for that on StackOverflow: https://stackoverflow.com/a/22434262/162094
I looked into this because I love structured logging, and I wasn't able to figure out how to make that work easily with the default logging module.
But I have to say, the native python logging module is pure evil.
In this case its much more convenient to design it as a higher order function.
for i, line in enumerate(lines):
i += 1> Let's assume I write a really inefficient way to find primes
let me stop you right there
- @parallel: Doesn't let the caller configure the actual amount of parallelism, and no, the only overhead is not "having to define a merge function". Creating the processes has quite a bit of overhead and calls might end up being far slower than just a regular serial execution. Also, you might not want to eat all CPUs on a single function (not to mention that cpu_count tells you the amount of CPUs in the system, not the amount that are available to your process).
- @production/@deployable: Functions that sometimes return values and sometimes don't depending on the hostname? That's just a bug waiting to happen. Also, at least use an environment flag and not a hardcoded list of servers.
- @redirect: Don't reinvent the wheel, use the logging module. Log messages will be more useful, you'll be able to redirect them wherever you want, enable, disable, increase the level, even use different logging modules and enable/disable them at runtime. There are even handlers for pretty IPython output.
- @stacktrace/@traceable Somewhat more useful but, again, the logging module usually does this. Also, debuggers exist.
If anyone is using this, be very aware of where you're using them, how and possible side effects and maintenance problems. Personally, I wouldn't let anything like this run in a somewhat serious codebase.
Long story short, if you want to have a cache, have a cache. If you want to make something parallel, make it parallel. This stuff just gets in the way. The only real exceptions I make for decorators by and large are for pydantic things and similar. It's very noticeable that if you're doing stuff in pytorch lightning you don't tend to fiddle with decorators much at all but in tensorflow it's common. Cute code isn't good code, end of story
This suggestion (logging) doesn't help if it's e.g. other code you don't have control over, but this is also available as a context manager in the standard library as "redirect_stdout": https://docs.python.org/3/library/contextlib.html#contextlib...
At least the article doesn't advocate inserting them into __builtins__ (which I have seen people do with their own special custom functions before).
Most of the time, other code uses the logging module so you do have control over that output. I haven't seen code not controlled by me that uses print() without me wanting it to.
If you are a traditional software engineer, the other people's code you deal with as an upstream depency probably looks different than the other people's code data (or many other) scientists work with.
While logging is good for production code, the tracing helpers look very useful for developing or debugging. One of the first things I do when troubleshooting why something is failing (in tests, or when reproducing things locally where I can run things with a visible console) is to print out _what was passed in_, or what the mid-function state is, so that I don't have to constantly query them in the debugger. Many many lines that look like
print(f">>> foo: {foo}, bar: {bar}")
Being able to automate this in a more convenient way seems very useful. The icecream library [0] does this, probably in a more robust way, but I've never used it because I always forget about it.Interleaving configuration management in the business logic code is a massive technical debt waiting to engulf future generations in a world of maintainance pain.
That seems obtuse to me in the extreme. Maybe I'm just not used to that style however.
Not being facetious! I’m genuinely curious and would love to learn more. Thanks
This is incredibly popular in single cell analysis
Data scientists need to be trained with software engineering skills, software engineers need to be trained with data science skills. This is the only way we do not end-up with non-sense like this.
I straddle both worlds - I’m much more on the tech side, but sometimes interact with scientists / MATLAB codebases.
What data science skills / methdologies would be useful for me to learn?
P.S. And what did you mean by MATLAB to C++? That was a specific time frame (I suppose in the early 2000s ish) when C++ was taught to scientists in the hope they’d be able to productionize their MATLAB code? With not great results (i.e. C++ learning curve + lack of software engineering skills…?) Thanks!
My take (non-exhaustive) with the current ecosystem is to apply Agile and DevOps methodologies: - Use Git everywhere all the time, always - Use Jupyter early one: great for quick prototypes & demos, keynotes, training material, articles - Once the initial prototype is approved, archive Jupyter notebooks as snapshots - Write functional tests (ideally in a TDD fashion) - Build and/or integrate the work into a real software product, be it in Python, C++, Java, etc. - Use tools for deterministic behavior (package manager, Docker, etc.) - Use CI/CD and GitOps methodologies - Deliver iteratively, fast and reliably
And by "MATLAB to C++" was a reference to a time (2010's) when corporation were deeply involved with MATLAB licenses, could not afford to switch easily to Python and lots of SWE with applied math background had to deal with MATLAB code written by pure math guys without any consideration for SWE best practices and target product constraints. Nowadays, if the target product is also in Python, there is way less friction, hopefully :)
This is where standard SWE guidelines could be of help (function interfaces, contracts definition, etc)
class mymodel(): ... def feature_engineering(self): ... def impute_missing(self): ... def fit(self) ... def predict_proba(self) ...
Then it is pretty trivial to test with new data. And you can parameterize the things you want, e.g. init fit method as random forest or xgboost, or stack on different feature engineering, etc. And for debugging you can extract/step through the individual methods.
It's a nice idea, but it turns out to be a pretty big ask. Particularly at my employer, where a large proportion of new hires come straight out of college. Data scientists have usually studied something like economics, math, statistics, or physics; most of them haven't been introduced to software engineering at all. We try to bring new hires up to speed, but there's only so much we can do with a series of relatively short sessions on Python and git.
Similarly, software engineers don't necessarily have the requisite background to understand the kind of work data scientists do. They'll have had a few semesters of calculus but it's likely they won't have had much if any exposure to data analysis or machine learning. They might not have even had a stats course in college. Further, in my experience they have had little inclination to understand how data scientists work, nor how their software products may or may not fit data scientists' needs.
Opining for a moment here ...
I've had the privilege of working for a few years at a position that kind of straddles the line between data scientist and software engineer (though I was technically a data scientist), and part of that job was mentorship and training. Getting good code out of data scientists and software engineers can be tough to do. I've seen nearly as much messy, uncommented, unformatted, unoptimized code from engineers as I have from data scientists, it's just that when I make recommendations to data scientists they'll actually listen to me.
I'm just lucky engineers started finally using the internal libraries I maintain rather than their own questionable alternatives (though if I never have the "why aren't you pinning exact dependencies for your library? My code broke!" discussion again it'll be too soon.)
That's kind of like saying the solution to liability issues arising in the practice of medicine is for physicians to learn lawyer skills and lawyers to learn physician skills.
It's a great idea, if you ignore the costs to get people trained and the narrowing of the pool for each affected profession.
Heck, it's hard enough to get software engineers who work almost entirely on systems where a major part of the lifting is done by an RDBMS to learn database skills.
But other then that, thought it was fine. The use of type hints is pretty awesome, in particular. Do you guys just not like higher-order functions?
Higher order functions are ok, but the problem with this what they're using it for. Code that behaves different based on environment without any explicit warning for the caller? That's fairly dangerous.
True enough. For the reader of these comments, it's possible (but non-obvious) to properly respect type hints in a decorator. See here: https://mypy.readthedocs.io/en/stable/generics.html#declarin...
I'm reminded of that joke:
Business logic (n): your stuff, which is shit, as opposed to my code, which is beautiful
I say this as a machine learning engineer
primes: set[int] = set()
candidate: int = randint(4, domain)
It's just so ugly and redundant.My feeling is that redundant stuff like type hints make programmers create more bugs because they don't see the forest for the trees anymore.
As if you wrapped the driver of a car in so many seat belts that they cannot see the road anymore.
Not ones that ship to production. Some developers have a terrible time trying to get through the mypy errors. But from my observations, I haven't seen any typing related errors in a very long time.
Sometimes adding a type to a single object (e.g. `df: pd.DataFrame = ...`) only once at the top of a 20-50 line function, is usually sufficient for Pycharm to infer all other types downstream without ambiguity.
I've been using type hints heavily and it's been such a great help in avoiding bugs and also faster completion in IDEs. Most of the time, with correctly typed functions you don't have to write explicit types for variables. It's a little bit more code, yes, but it's definitely worth it for any codebase that grows more than a handful of files. Even for smaller ones I have found it very useful.
Agreed. I've been using type hints for the past few years now even though most of my projects are fairly small (< 10,000 LOC). I wouldn't say they've prevented a huge amount of bugs, but they've definitely prevented a few, and I think they've helped in library design and documentation.