Collaborating with nontechnical people is oddly my favorite part of doing MLE work right now. It wasn't the case when I did basic web/db stuff. They see me as a magician. I see them as voodoo priests and priestesses. When we get something trained up and forecasting that we both like, it's super fulfilling. I think for both sides.
Most of my modeling is healthcare related. I tease insights out of a monstrous data lake of claims, Rx, doctor notes, vital signs, diagnostic imagery, etc. What is also monstrous is how accessible this information is. HIPAA my left foot.
Since you seemed to be asking about the temporal realities, it's about 3 hours of meetings a week, probably another 3 doing task grooming/preparatory stuff, fixing some ETL problem, or doing a one-off query for the business, the rest is swimming around in the data trying to find a slight edge to forecast something that surprised us for a $million or two using our historical snapshots. It's like playing wheres waldo with math. And the waldo scene ends up being about 50TB or so in size. :D
As we grow, I'm ever watchful for our metamorphosis into a big-dumb-company, but no symptoms yet. :)
That was my experience as well - training documentation for fresh college grads (i.e. me) directed new engineers to just... send SQL queries to production to learn. There was a process for gaining permissions, there were audit logs, but the only sign-off you needed was your manager, permission lasted 12 months, and the managers just rubber-stamped everyone.
That was ten years ago. Every time I think about it I find myself hoping that things have gotten better and knowing they haven't.
That was also the company where prod pushes happened once a month, over the weekend, and were all hands on deck in case of hiccups. It was an extraordinarily strong lesson in how not to organize software development.
(edit: if what you're really asking is "did every engineer have write access to production", the answer was, I believe, that only managers did, and they were at least not totally careless with it. not, like, actually responsible, no "formal post-mortem for why we had to use break-glass access", but it generally only got brought out to unbreak prod pushes. Still miserable.)
Is it surprising that Engineers in healthcare dont read the actual HIPAA documentation?
Use of health data is permitted so long as it’s for payment, treatment or operations. Disclosures and patient consent are not required.
There are helpful summaries on the US Department of Health and Humans Services website of the various rules (Security, Privacy & Notification)
Source: https://www.hhs.gov/hipaa/for-professionals/privacy/guidance...
This allowance is permitted to covered entities and by extension their vendors (business associates) by HIPAA.
If it wasn’t then, theoretically the US healthcare industry would grind to a halt considering the number of intermediaries for a single transaction.
Example:
Doctor writes script -> EHR -> Pharmacy -> Switch -> Clearinghouse —> PA Processing -> PBM/Plan makes determination
Along this flow there are other possible branches and vendors.
It’s beyond complex.
It just occurred to me that cleaning up our country's data privacy / data ownership mess might have extraordinarily positive second-order effects on our Kafkaesque and criminally expensive healthcare "system".
Maybe making it functionally impossible for there to be hundreds of middlemen between me and my doctor would be a... good thing?
A national identity service, like a normal mature economy, would be a massive public policy and public health win.
But it'd slightly decrease profits. So of course is quite impossible to implement.
Pre-HIPPA, the hospital would tell a news reporter about your status. Now, drug marketing companies know about your prescriptions before your pharmacy does.
In attendance were marketing, biz dev, execs. Ours and theres. But no legal. Hmmm.
And it would have worked if it weren't for those meddling teenagers.
From your comment, I'm inferring they're now allowed to use health records for marketing.
On the would-be due date, a box arrived via FedEx of Enfamil samples and a chirpy “welcome baby” message. Of course there was no baby.
It turns out that using prescription data, admission data and other information that is aggregated as part of subrogation and other processes, you can partially de-anonymize/rebuild / health record and make an inference. Enfamil told me what list they used and I bought it, it contained a bunch of information including my wife’s. I also know everyone in my zip code who had diabetes in 2015.
There’s even more intrusive stuff around opioid surveillance.
> Emfamil samples...
Insult to injury. They're all just ghouls. Profitting from other people's pain.
> what list they used
Who watches the watchers? Nobody.
We had an instance of Lexis/Nexus' (formerly Seisent) demographic database to play with. Most people just can't wrap their heads around how bad things were in the mid-aughts. I was called a "sweaty paranoid kook" by our local paper for simply describing how things worked (no opinion or predictions).
Oh well. Maybe society has started to recognize what's happened.
Thanks for reading my rant. Peace.
Can you please explain more about this or point me towards a source?
Worked on EMRs (during the aughts). Had yearly HIPAA and other security theater training. Not optional for any one in the field.
Of course we had root access. I forget the precise language, but HIPAA exempts "intermediaries" such as ourselves. How else could we build and verify the system(s). And that's why HIPAA is a joke.
Yes, our systems had consent and permissions and audit logs cooked in. So theoretically peeking at PII could be caught. Alas, it was just CYA. None of our customers reviewed their access logs.
--
I worked very hard to figure out how to protect patient (and voter) privacy. Eventually conceded that deanon always beats anon, because of Big Data.
I eventually found the book Translucent Databases. Shows how to design schemas to protect PII for common use cases. Its One Weird Trick is applying proper protection of passwords (salt + hash) to all other PII. Super clever. And obvious once you're clued in.
That's just 1/2 of the solution.
The other 1/2 is policy (non technical). All data about a person is owned by that person. Applying property law to PII transmutes 3rd party retention of PII from an asset to a liability. And once legal, finance, and accounting people get involved, orgs won't hoard PII unless it's absolutely necessary.
(The legal term "privacy" means something like "sovereignty over oneself", not just the folk understanding of "keeping my secrets private.)
please, do tell me, where "estimating better pricing models to extract more profit" fit into those?
> Doctor writes script -> EHR -> Pharmacy -> Switch -> Clearinghouse —> PA Processing -> PBM/Plan makes determination
All the OP mentions happens way after the fact of those paths you described. People have already been charged. Treatment was already decided.
I'm curious to know the tech stack behind converting unstructured to structured data(for reporting and analysis)
If you want to learn machine learning for healthcare in general, it may help to start problems with tabular data like CSVs instead of images. Image processing is a lot harder, and takes a lot of time and computational power. But it's best to learn what you're interested in the most.
Anyway, first you need to be familiar with basic; Python, machine learning, and popular libraries like scikit-learn, matplotlib, numpy, and pandas. Those are tons of articles, textbooks, and videos to help you learn them.
If you grasp the basics, I think it's better to learn from actual code to train/evaluate models rather than more theories. Kaggle may be a good starting point. They host a lot of competitions for machine learning problems. There are easy competitions for beginners, and competitions and datasets in the medical field.
You can view notebooks (actual code to solve those problems well written by experts) and popular ones are very educational. You can learn a lot by reading those code, understanding concepts and how to use libraries, and modifying some code to see how it changes the result. ChatGPT is also helpful.
If you want to learn image classification, the technology used to detect objects from images and videos is called image classification and object detection. It uses CNN, one of the deep neural networks. You also need to learn basic image processing, how to train a deep neural network, how to evaluate, and libraries like OpenCV/Pillow/PyTorch/TorchVision. There are a lot of image classification competitions in the medical field on Kaggle too[0][1].
To run those notebooks, I recommend Google Colab. Image processing often uses a lot of GPUs, and you may not have GPUs, or even if you have it's difficult to set up the right environment. It's easier to use those dedicated cloud services and it doesn't cost much.
It's hard to learn, but sometimes fun, so enjoy your journey!
[0] https://www.kaggle.com/datasets/paultimothymooney/chest-xray... [1] https://www.kaggle.com/code/arkapravagupta/endoscopy-multicl...
I like this so much I'm definitely stealing it!
The people who have lasted in those roles have built up a large degree of intuition on how their domains work (or they would've done something else).
I am guessing every datascientist that works for a BlueCross BlueShield at individual states deals with processes that touch multiple-million dollars of claims.
We even now have various dueling systems -- one company has a model to tack on more diagnoses to push up the bill, another has a process to identify that upcoding. One company has models to auto accept/deny claims, another has an automatic prior authorization process to try to usurp that later denial stage, etc.
I do not know the chances that any individual provider uses a system to identify what codes to tack onto your bill.
Collaborating with them is great, and has been a great exercise in learning how to explain complex ideas to non-technical business people. Which has the side effect of helping me get better at what I do (because you need a good understanding of a topic to be able to explain it both succinctly and accurately to others). It has also taught me to appreciate the business context and reasoning that can drive decisions about how a business uses or develops data/software.
The other half is spent doing tech support for the bunch of recently hired "AI scientists" who can barely code, and who spend their days copy/pasting stuff into various chatbot services. Stuff like telling them how to install python packages and use git. They have no plan for how their work is going to fit into any sort of project we're doing, but assert that transformer models will solve all our data handling problems.
I'm considering quitting with nothing new lined up until this hype cycle blows over.
But yes. My work is kind of similar… I do some data curation / coding, and help 2 engineers who report to me. I enjoy it.
It means I'm marginalized in terms of planning. The company has long term goals that involve making good use of data. Right now, the plan is that "AI" will get us there, with no plan B is it doesn't work. When it inevitably fails to live up to the hype, we're going to have a bunch of clobbered together systems that are expensive to run, rather than something that we can keep iterating on.
It means I'm marginalized in terms of getting resources for projects. There's a lot of good my team could be doing if we had the extra budget for more engineers and computing. Instead that budget is being sent off to AI services, and expensive engineer time is being spent on tech support for people that slapped "LLM" all over their resume.
95% of the job is data cleaning, joining datasets together and feature engineering. 5% is fitting and testing models.
There is a lot of toil and unnecessary toil in the whole data field, but if you define away all of the "yucky" parts, you might find that all of those "someone elses" will end up eating your lunch.
See: the use of "devops" to encapsulate "everything besides feature development"
Should your reseacher have to manage nvidia drivers and infiniband networking? Should your operations engineer need to understand the math behind transformers? Does your researcher really gain any value from understanding the intricacies of docker layer caching?
I've seen what it looks like when a company hires mostly researchers and ignores other expertise, versus what happens when a company hires diverse talent sets to build a cross domain team. The second option works way better.
If other peoples work is reliant on yours then you should know how their part of the system transforms your inputs
Similarly you should fully understand how all the inputs to your part of the system are generated
No matter your coupling pattern, if you have more than 1 person product, knowing at least one level above and below your stack is a baseline expectation
This is true with personnel leadership too, I should be able to troubleshoot one level above and below me to some level of capacity.
I've seen these too, and you aren't wrong. Division into specializations can work "way better" (i.e. the overall potential is higher), but in practice the differentiating factors that matter will come down to organizational and ultimately human-factors. The anecdotal cases I draw my observations from organizations operating at the scale of 1-10 people, as well as 1,000s working in this field.
> Should your reseacher have to manage nvidia drivers and infiniband networking? Should your operations engineer need to understand the math behind transformers? Does your researcher really gain any value from understanding the intricacies of docker layer caching?
To realize the higher potential mentioned above, what they need to be doing is appreciating the value of what those things are and those who do those things beyond: these are the people that do the things I don't want to do or don't want to understand. That appreciation usually comes from having done and understanding that work.
When specializations are used, they tend to also manifest into organizational structures and dynamics which are ultimately comprised of humans. Conway's Law is worth mentioning here because the interfaces between these specializations become the bottleneck of your system in realizing that "higher potential."
As another commenter mentions, the effectiveness of these interfaces, corresponding bottlenecking effects, and ultimately the entire people-driven system is very much driven by how the parties on each side understand each other's work/methods/priorities/needs/constraints/etc, and having an appreciation for how they affect (i.e. complement) each other and the larger system.
Personally I think we calcified the roles around data a little too soon but that's probably because there was such demand and the space is wide.
If you outsource that to somebody else, you'll miss out on all the pattern-matching eureka moments, and will never know the answers to questions you never think to ask.
Yeah, but also knowing which features to get right. Right?
However, having a quality and diverse dataset is more important now than ever.
At the staff/principal level it’s all about maintaining “data impedance” between the product features that rely on inference models and the data capture
This is to ensure that as the product or features change it doesn’t break the instrumentation and data granularity that feed your data stores and training corpus
For RL problems however it’s about making sure you have the right variables captured for state and action space tuple and then finding how to adjust the interfaces or environment models for reward feedback
That's my fun part. The discovery process is a joy especially if it means ingesting a whole new domain and meeting people.
Environment broken
Spend 4 hours fixing python environment
pip install Pillow
Something something incorrect cpu architecture for your Macbook
Spend another 4 hours reinstalling everything from scratch after nuking every single mention of python
pip install … oh time to go home!
All of a sudden you have people with C problems, who have no idea they're even using compiled dependencies.
It's not really much of a counterargument to say that Python is good enough that you don't have to care what's under the hood, except when it breaks because C sucks so badly.
I never experienced CPython itself segfault, it's always due to some package.
The point is the same - we had it simpler and now, with all capabilities for automation, we have it more complex.
Frankly, I suspect most of the efforts now are spent fighting non-essential complexities, like compatibilities, instead of solving the problem at hand. That means we create problems for ourselves faster than removing them.
(I don't do most of my work locally, but for smaller models its pretty convenient to work on my mbp).
Funnily, the only real competitor for Nvidias' GPUs are Macbooks with 128GB of RAM.
I'm an Applied Scientist vs. ML Engineer, if that matters.
Conda.. it interfere with OS setup and has not always the best utils. Like ffmpeg is compiled with limited options, probably due to licensing.
It causes so many entirely unnecessary issues. The conda developers are directly responsible for maybe a month of my wasted debugging time. At my last job one of our questions for helping debug client library issues was “are you using conda”. And if so we just would say we can’t help you. Luckily it was rare, but if conda was involved it was 100% conda fault somehow, and it was always a stupid decision they made that flew in the face of the rest of the python packaging community.
Data scientist python issues are often caused by them not taking the 1-3 days it takes to fully understand their tool chain. It’s genuinely quite difficult to fuck up if you take the time once to learn how it all works, where your putbon binaries are on your system etc. Maybe not the case 5 years ago. But today it’s pretty simple.
The interminable solves were awful. Mamba made it better, but can still be slow.
Plenty of more esoteric packages are on PyPI but not Conda. (Yes, you can install pip packages in a conda env file.)
Many packages have a default version and a conda-forge version; it's not always clear which you should use.
In Github CI, it takes extra time to install.
Upon installation it (by default) wants to mess around with your .bashrc and start every shell in a "base" environment.
It's operated by another company instead of the Python Software Foundation.
idk, none of these are deal-breakers, but I switched to venv and have not considered going back.
This works way better than pip, as it does more checks/dependency checking, so it does not break as easily as pip, though this makes it definitely way slower when installing something. It also supports updating your environment to the newest versions of all packages.
It's no silver bullet and mixing it with pip leads to even more breakages, but there is pixi [0] which aims to support interop between pypi and conda packages
- If they're so good at dependency management, why is Conda installed through a magical shell script?
- It's slow as molasses.
- Choosing between Anaconda/Miniconda...
When forced to use Python, I prefer Poetry, or just pip with freezing the dependencies.
The Python people probably can't even imagine how great dependency management is in all the other languages...
> The Python people probably can't even imagine how great dependency management is in all the other languages...
Yep, I wish I could use another language at work.
> Choosing between Anaconda/Miniconda...
I went straight with mamba/micromamba as anaconda isn't open source.
This happened about a week ago
Your comment makes me feel a little better that this is not merely some personal failing of focus, but happens in a professional setting too.
For hobby stuff at home though I don't tend to hit those types of issues because my projects are pretty frozen dependency-wise. Do you really have OS updates break stuff for you often? I'm not sure I recall that happening on a home project in quite a while.
It’s why I still use react, their backcompat is amazing
or pyenv at least
>Something something incorrect cpu architecture for your Macbook
I’m glad I have something in common with the smart people around here.
Surely work can provide that?
[0] https://bugs.launchpad.net/ubuntu/+source/ubuntu-meta/+bug/1...
Although I think the UX of poetry is stupid and I do not agree with some design decisions, I have not had any dependency conflicts since I used it.
To avoid spending hours on fixing broken environments after a single "pip install", I would make it easy to rollback to a known state. For example, recreate virtualenv from a lock requirements file stored in git: `pipenv sync` (or corresponding command using your preferred tool).
Half of the problems I've helped some people solve stem from python devs insisting on shuffling std libraries around between minor versions.
Some libraries have a compatibility grid with different python minor versions, because how often they break things.
Since then the whole python ecosystem has gotten worse.
We are building towers on quicksand.
It's not about python, it's about people who don't care about dependencies.
13 years ago when I was trying to explore the field R seemed to be the most popular, but looks like not anymore. (I didn't get into the field, and do just a regular SWE, so I'm not aware of the trends).
There is also a lot of development in Elixir ecosystem around the subject [1].
[1](https://dashbit.co/blog/elixir-ml-s1-2024-mlir-arrow-instruc...)
There are Microsoft-backed F# bindings for Spark and Torch, but no one seems interested. And this is despite a pretty compelling proposition of lightweight syntax, strong typing, script-ability and great performance.
The answer will probably be JavaScript.
Everyone already knows the language - all that's missing is operator overloading and a few key bindings.
For exactly the reasons your mentioned, i feel like F# would have been the perfect match for both MLE/ETL(spark pipeline) work and some of the deep learning/graph modelisation such as pytorch. Saddly, even from MSFT, the investment in F# as dried up
Why is python dependency management so cancerously bad? Why are there so many “solutions” to this problem that seem to be out of date as soon as they exist?
Are python engineers just bad, or?
(Background: I never used python except for one time when I took a coursera ML course and was immediately assaulted with conda/miniconda/venv/pip/etc etc and immediately came away with a terrible impression of the ecosystem.)
- You can’t have two versions of the same package in the namespace at the same time. - The Python ecosystem is very bad at backwards compatibility
This means that you might require one package that requires foo below version 1.2 and another package that requires foo version 2 and above.
There is no good solution to the above problem.
This problem is amplified when lots of the packages were written by academics 10 years ago and are no longer maintained.
The bad solutions are: 1) Have 2 venvs - not always possible and if you keep making venvs you’ll have loads of them. 2) Rewrite your code to only use one library 3) Update one of the libraries 4) Don’t care about the mismatch and cross your fingers that the old one will work with the newer library.
Most of the tooling follows approach 1 or 4
I threw it online at https://github.com/fragmede/python-wool/
Now your 15 venvs deep, and have over 3000 different package version combinations installed. Your job is to upgrade them right now because of a level 8 CVE
for venv in $(find ~/projects/ -type d -name 'venv'); do
(
source ${venv}/bin/activate
pip install --upgrade pip
pip install --upgrade package_with_cve
deactivate
) &
done
should do the trick. Sucks that we're in that world, but let's not work any harder than we have to.pyenv for installing and managing python versions.
direnv for managing environments and environment variables (a highly underrated package imo).
With those two installed I just include a .envrc file in every project. It looks like this:
layout python ~/.pyenv/versions/3.11.0/bin/python3
export VARIABLE1=variable1
export VARIABLE2=variable2
etc...
PYTHONPATH="`pwd`/venv/lib/python-3.11/site-packages"
or whatever in .envrc and Bob's your uncle.The ML system is a whole another stack of problems. The elephant in the room is Nvidia who is not known for playing well with others. Aside from that, the state of the art in ML is churning rapidly as new improvements are identified.
I can only imagine it seemed like an oasis to R which is bottom tier.
So when you combine data scientists, academics, mathematicians, juniors, grifters selling courses…
things like bad package management, horrible developer experience, absolutely no drive in the ecosystem to improve anything beyond the “pinnacle” of wrapping C libraries are all both inevitable and symptoms of a poorly designed ecosystem.
> Something something incorrect cpu architecture for your Macbook
I wonder how “real” ML people deal with the stochastic/gradient results and people’s expectations.
If I do ordinary software work the thing either works or it doesn’t, and if it doesn’t I can explain why and hopefully fix it.
Now with ML I get asked “why did this text classifier not classify this text correctly?” and all I can say is “it was 0.004 points away to meet the threshold”, and “it didn’t meet it because of the particular choice of words or even their order” which seems to leave everyone dissatisfied.
At least with symbolic regression you can treat the model as an analyzable entity from first principles theories. But that's not really particularly relevant to most failure modes in practice, which usually boil down to either missing some qualitative change such as a bifurcation or else just parameters being off by a bit. Or a little bit of A and a little bit of B.
I build the systems to support ML systems in production. As others have mentioned, this includes mostly data transformation, model training, and model serving.
Our job is also to support scientists to do their job, either by building tools or modifying existing systems.
However, looking outside, I think my company is an outlier. It seems in the industry the expectations for a ML Engineer are more aligned to what a data/applied scientist does (e.g. building and testing models). That introduces a lot of ambiguity into the expectations for each role in each company.
I gave a talk at the Open Source Summit on MLOps in April, and one of the big points I try to drive home is that it's 80% software development and 20% ML.
Highly paid motherboard troubleshooter, because those all those H100's really get hot, even with watercooling, and we have no dedicated HW guy.
Fighting misbehaving third-party deps, as everyone else.
$variable = something() if sanity_check()
And do_something() unless $dont_do_thatfoo = something() if sanity_check else None
Can replace None with foo (or any other expression), if desired.
variable = something() if sanity_check() else None
do_something() if not dont_do_that else Noneif sanity_check(): variable = something()
variable = sanity_check() and something
(but the standard python two line way is more readable unless it's a throwaway single use code, then whatever)
- Collaboration with stakeholders & TPMs and analyzing data to develop hypotheses to solve business problems with high priority
- Framing business problems as ML problems and creating suitable metrics for ML models and business problems
- Building PoCs and prototypes to validate the technical feasibility of the new features and ideas
- Creating design docs for architecture and technical decisions
- Collaborating with the platform teams to set up and maintain the data pipelines based on the needs of new and exiting ML projects
- Building, deploying, and maintaining ML microservices for inference
- Writing design docs for running A/B tests and performing post-test analyses
- Setting up pipelines for retraining of ML models
Not my main work, but spending a lot of time gluing things together. Tweaking existing open source. Figuring out how to optimize resources, retraining models on different data sets. Trying to run poorly put together python code. Adding missing requirements files. Cleaning up data. Wondering what could in fact really be useful to solve with ML that hasn't been done years ago already. Browsing the prices of the newest GPUs and calculating whether that would be worth it to get one rather than renting overpriced hours off hosting providers. Reading papers until my head hurt, that is just 1 by 1, by the time I finish the abstract and glanced over a few diagrams in the middle.
Dedicated groups on social. e.g: https://www.linkedin.com/newsletters/top-ml-papers-of-the-we...
Sometimes it also involves building internal tooling for our team (we are a mixed team of researchers/MLEs), to visualize the data and the inferences as again, it's a pretty niche sector and that means having to build that ourselves. That allowed me to have a lot of impact in my org as we basically have complete freedom w.r.t tooling and internal software design, and one of the tools that I built basically on a whim is now on the way to be shipped in our main products too.
I take responsibility for the end to end experience of said API, so I will do whatever gives the best value per time spent. This often has nothing to do with the ML models.
I'd argue that if you are not spending >50% of your time in model development and research then it is not a machine learning role.
I'd also say that nothing necessitates the vast majority of an ML role being about data cleaning, etc. I'd suggest that indicates that the role is de facto not a machine learning role, although it may say so on paper.
In my case, there are established processes and designated teams for cleaning & collecting data, but you still do a part of it yourself to provide guidelines. So, even though data is always a perpetual problem, I can shed off most of that boring stuff.
Ah, and of course you're not a real engineer if you don't spend at least 1-2% of your time explaining to other people (surprisingly often to a technical staff, but not ML-oriented) why doing X is a really bad idea. Or, just explaining how ML systems work with ill-fitted metaphors.
* 15% of my time in technical discussion meetings or 1:1's. Usually discussing ideas around a model, planning, or ML product support
* 40% ML development. In the early phase of the project, I'm understanding product requirements. I discuss an ML model or algorithm that might be helpful to achieve product/business goals with my team. Then I gather existing datasets from analysts and data scientists. I use those datasets to create a pipeline that results in a training and validation dataset. While I wait for the train/validation datasets to populate (could take several days or up to two weeks), I'm concurrently working on another project that's earlier or further along in its development. I'm also working on the new model (written in PyTorch), testing it out with small amounts of data to gauge its offline performance, to assess whether or not it does what I expect it to do. I sanity check it by running some manual tests using the model to populate product information. This part is more art than science because without a large scale experiment, I can only really go by the gut feel of myself and my teammates. Once the train/valid datasets have been populated, I train a model on large amounts of data, check the offline results, and tune the model or change the architecture if something doesn't look right. After offline results look decent or good, I then deploy the model to production for an experiment. Concurrently, I may be making changes to the product/infra code to prepare for the test of the new model I've built. I run the experiment and ramp up traffic slowly, and once it's at 1-5% allocation, I let it run for weeks or a month. Meanwhile, I'm observing the results and have put in alerts to monitor all relevant pipelines to ensure that the model is being trained appropriately so that my experiment results aren't altered by unexpected infra/bug/product factors that should be within my control. If the results look as expected and match my initial hypothesis, I then discuss with my team whether or not we should roll it out and if so, we launch! (Note: model development includes feature authoring, dataset preparation, analysis, creating the ML model itself, implementing product/infra code changes)
* 20% maintenance – Just because I'm developing new models doesn't mean I'm ignoring existing ones. I'm checking in on those daily to make sure they haven't degraded and resulted in unexpected performance in any way. I'm also fixing pipelines and making them more efficient.
* 15% research papers and skills – With the world of AI/ML moving so fast, I'm continually reading new research papers and testing out new technologies at home to keep up to date. It's fun for me so I don't mind it. I don't view it as a chore to keep me up-to-date.
* 10% internal research – I use this time to learn more about other products within the team or the company to see how my team can help or what technology/techniques we can borrow from them. I also use this time to write down the insights I've gained as I look back on my past 6 months/1 year of work.
The research becomes relevant immediately because my team is always looking to incorporate it into our production models right away. Of course it does take some planning (3-6 months) before it's fully rolled out in production.
the users are researchers and have deep technical knowledge of their use case. it is still a challenge to map their needs into design decisions of what they want in the end. thanks to open-source efforts, the model creation is rather straightforward. but everything around making that happen and shaping it like a tool is a ride.
especially love the ever-changing technical stack of "AI" services by major cloud providers rn. it makes mlops nothing more than a demo imho.
I worked on other ML projects as well. A system that analyzed the syslogs of dozens of computers to look for anomolies. I wrote the SQL (256 fields in the query! Most complex SQL I've ever written) that prefiltered the log data to present it to the ML algorithm. And built a server that sniffed encrypted log data we broadcast on the local network in order to gather the data in one place continuously. Another system used heart rate variability to infer stress. I helped design a smartwatch and implemented the drivers that took in HRV data from a Bluetooth chest strap. We tested the system on ourselves. None of our ML projects involved writing new ML algorithms, we just used already implemented ones off the shelf. The main work was getting the data, cleaning the data, and fine tuning or implementing new feature extractors. The CS people weren't familiar with the biological aspects (digging into Grey's Anatomy), sensors, wireless, or electronics, so I handled a lot of that. I could have done all the work, it's not hard to run ML algorithms (we often had to run them on computing clusters to get results in a reasonable amount of time, which we automated) or figure out which features are important, but then the students wouldn't have graduated :-) Getting labeled data for supervised learning is the most time consuming part. If you're doing unsupervised learning then you're at the mercy of what is in the data and you hope what you want is in there with the amount of detail you need. Interfacing with domain experts is important. Depending on the project a wide variety of different skills can be required, many not related to coding at all. You may not be responsible for them all, but it will help if you at least understand them.