Even so: reading ingredients is honestly not that hard.
204 karma · joined March 1, 2013
Even so: reading ingredients is honestly not that hard.
I meant that my linguistics background helps me understand & solve problems: studying linguistic field work has helped me design crowd labeling jobs, knowing about morphology helps me understand why BPE tokenizers work so well (and when they might not), knowing about syntax/dominant word order makes me think that multilingual Bert should probably do something more intelligent with positional embeddings, methods from psycholinguistics are useful for understanding entropy/surprisal wrt LM next-word probabilities... just a few examples but the list could go on.
Yes, you can easily use AutoModel.from_pretrained('bert-base-uncased') to convert some text into a vector of floats. What then?
What are the properties of downstream (aka actually useful) datasets that might make few-shot transfer difficult or easy? How much data do your users need to provide to get a useful classifier/tagger/etc. for their problem domain?
Why do seemingly-minor perturbations like typos or concating a few numbers result in major differences in representations, and how do you detect/test/mitigate this to ensure model behavior doesn't result weird downstream system behavior?
How do you train a dialog system to map 'I'm good, thanks' to 'no'? How do you train a sentiment classifier learn from contextual/pragmatic cues rather than purely lexical ones (example: 'I hate to say it but this product solves all my problems.' - positive or negative sentiment?)
How bad is the user experience of your Arabic-speaking customers compared to that of your English-speaking customers, and what can you do to measure this and fix it?
My linguistics background really helps me think through a lot of these 'applied' NLP problems. Knowing how to make matmuls fast on GPUs and knowing exactly how multihead self-attention works is definitely useful too, but that's only one piece of building systems with NLP components.
Also if your research/biz needs are satisfied by sklearn, then why not just use sklearn? But for a lot of NLP systems, BERT is actually really really useful. And if you don't want to use a pretrained BERT, you can easily initialize their BERT implementation randomly and train it on your own data.
And they amortize the cost of hiring their own NLP engineers by developing a few models/model-based services that lots of businesses would be willing to pay for. E.g. 'foundation models' for different verticals like healthcare etc. Then it'll also be a lot easier to either fully automate or at least scale up work that's specific to each paying customer (because fine-tuning should go much more quickly, just essentially be a hyperparameter tuning cycle in as many cases as they can get away with).
Yes the former changes every 1-5 years. Doing the latter well is much harder, no single tool can solve these problems, and I think years of experience really does help.
One thing that works great for me at home is, compromising on a temperature with my partner such that we're both physically comfortable throughout the day.
Like, who cares??
* What I mean is, text gen models are big enough. We need controllable text generation; like, so it can talk about a specific THING sensibly. Rather than spew statistically plausible nonsense.
NYT on 24 August 2017: "The Flatiron School in New York may have discovered one path. Founded in 2012, Flatiron has a single campus in downtown Manhattan and its main offering is a 15-week immersive coding program with a $15,000 price tag. More than 95 percent of its 1,000 graduates there have landed coding jobs."
two months later
NY AG, 17 October 2017: "However, Flatiron did not disclose clearly and conspicuously that the 98.5% employment rate included not only full time salaried employees but also apprentices, contract employees and self-employed freelance workers, some who were employed for less than twelve weeks. Similarly, Flatiron failed to clearly and conspicuously disclose that its $74,447 average salary claim included full time employed graduates only, which represent only 58% of classroom graduates and 39% of online graduates."
https://docs.google.com/document/d/1504Sy29t1uUw_B8zTbeLIPCH...
Paid for once: Threes, Clear, Convert (for all of my unit conversion needs)
I use but don't pay for: Dropbox, IntellijIDEA, Sublime Text 2
I've also spent a shameful amount of money on Candy Crush...