35 karma · joined December 6, 2017
Bay -> LA (Go Bruins!) -> NYC
- going to therapy
- taking PTO
- going outside
- hanging out with friends
- cooking a meal
- going on a date
- going to LA and getting sun
- reconnecting with old friends
I'm not out yet but I feel like with some time and effort I can be. How did I do it? I think there are really two parts of this - helping others and asking for help. I've always enjoyed the first but never done the second and that was really holding me back.
As you mentioned, ML training can be parallelized but this requires either model/data parallelism.
Data parallelism means spreading the data over many different compute units and then synchronizing gradients somehow. The heterogeneous nature of @home computing makes this particularly challenging, as you will be limited by the smallest compute unit. I've personally only ever seen data (and model) parallel done on a homogenous compute cluster (i.e. 8x GPUS)
For model parallelism, we split the model across different compute units. However, this means that you need to synchronize the different parts of the model together, which can get very expensive when you do it across the internet. If you have 8xGPUS on one machine, your latency is limited by PCIe instead of TCP/IP in a distributed @home cluster.
But I would say it's not impossible, someone clever could definitely figure it out.
I don't regret it much, as I was tired of being a broke student. Also I think the problems you run into in industry are more interesting. However, I do think my skills stagnated a bit and I learned more when doing research.
There's so much emphasis on writing "clean" code (rightly so) that it's nice to hear an opposing viewpoint. I think it's a good reminder to not be dogmatic and that there are many ways to solve a problem, each with their own pros/cons. It's our job to find the best way.
https://jcaip.github.io/Startup-Equity-TLDR/
You negotiate around dilution by asking for additional equity grants, like you would for a raise / promotion. But in general dilution is a good problem to have because it means you are raising more money and growing.
In my experience I've found that trying to develop deep expertise requires a solid understanding of many different fundamentals. For example, when I was doing NLP research, I had to learn about distributed systems in order to diagnose problems when training with multiple GPUs, or about dynamic libraries for fixing CUDA errors.
If you want to buy a GPU for ML this is a good resource: https://timdettmers.com/2023/01/30/which-gpu-for-deep-learni...
Basically buy the NVIDIA GPU with the largest amount of VRAM that's in your budget.
Also I'm well compensated and surrounded by smart and kind people.
I would get him the gaming console, but otherwise maybe nice headphones?
FWIW at the uni I attended, projects were a huge part of some classes (usually ~50%) and there were some SWE specific classes where you learn best practices / tooling. You would need to be able to program in order to graduate. I feel like this is the norm and not the exception, but I may be wrong.
There were a lot of CS clubs at my school that hosted workshops and tried to teach people some of this stuff. Maybe you can find a club (or start one) at your school.
I feel it's kind of like writing tests in SWE - something you do because it's beneficial even if it's not enjoyable.
A second bit of advice: Programming (and execution) skills are IMO heavily undervalued by people looking to get into ML. The faster you can write code, debug, and implement new things, the easier it is to produce good research.
Some books I liked: PR & ML (Bishop), Deep Learning Book (Goodfellow), AI: A Modern Approach (Norving), Elements of Statistical Learning (Friedman)
I know you mentioned working at a research lab but have you tried more academic research? You seem to value learning and mentorship a lot, so it could be a lot of fun to pick a professor and work with them for a summer. Given your background, you could probably convince any STEM professor to work with you, and I would heavily recommend exploring stuff outside your area of expertise.
The example I gave of multi-modal learning was really just highlighting a dichotomy in the techniques that we use in machine learning today. FWIW I am a couple of years removed from working heavily with tabular data, so do take this with a grain of salt. But there are essentially two different modeling approaches for two different types of datasets. On the one hand, you have deep learning (BERT, language models, CV models) which does well on raw data like text or images. These usually work by mapping the raw data to dense embeddings, which are the output of neural models. On the other hand, you have decision trees / forests (think XG boost) that work great on tabular data - spreadsheets or other data of that nature.
But what do you do if you have a spreadsheet of data and one of the columns is raw text data but the other columns are say sparse boolean features? How can you incorporation the extra information from the spreadsheet into your language model? I think this is a common problem in industry that there's not a clear solution for right now.
For industry: Feels like “making BERT / some other language model do things” is a common job nowadays. On the more engineering side - I think we’ll see more tools to quickly and efficiently fine-tune language models, especially tools that allow a human in the loop.
Overall it feels like we’re getting to a point where there’s a pretty standardized approach to simple NLP problems like text classification - no more real feature engineering, just throw BERT at the problem. I expect for this trend to continue - with more and more of a focus on dataset creation and validation and less of an emphasis on model architecture.
I also think there will be a rise in multi-modal language models - combination of language and vision models for example. But I think the more interesting application will be combining dense language model representations with sparser tabular data. Think of trying to predict a users likelihood to buy a product given a review of another product (dense embedding of text), but also their clicks over the last 2 hours. (sparser tabular data) - this feels like a much more common problem people have.
To stay updated: read papers (arxiv-sanity.com is a lifesaver) and watch talks (usually just on youtube or a lot of uni reading groups are public on zoom nowadays).
What is success?
To laugh often and much;
to win the respect of intelligent people and the affection of children;
to earn the appreciation of honest critics and endure the betrayal of false friends;
to appreciate the beauty; to find the best in others;
to leave the world a bit better, whether by a healthy child, a garden patch Or a redeemed social condition;
to know even one life has breathed easier because you have lived.
-- Ralph Waldo EmersonI think because it expressed something that I've felt but could never quite articulate - that there isn't one big well-defined reason for living but a bunch of little reasons that are all worthwhile.
It has worked quite well for us. The snorkel public package is a bit out of date now, as I think they're building a SaaS solution and focusing more on that. But aside from that it's quite easy to use. Other downside is that lot's of cool ideas are present in the papers but not fully implemented (not complaining though!). Also thinking of a diverse set of heuristics can be hard.
We use snorkel a lot for bootstrapping text classifiers. Our classification models don't require much domain expertise, as it's pretty easy to tell if a text sample is classified correctly, so the main advantage is just avoiding labeling costs and quicker prototyping. We find that we can usually use embedding similarity as a good heuristic. I wrote up a little bit about this approach here if you're curious: https://cultivate.com/why-cultivate-uses-embeddings-for-rapi...
Happy to answer any additional questions you have too :)
Hope it's helpful!