Goodbye, data science
ryxcommar.com
ryxcommar.com
In my experience it's even a little bit worse than that. Approaches that are wrong from a statistics point of view are more likely to generate impressive seeming results. But the flaws are often subtle.
A common one I've seen quite many times is people using a flawed validation strategy (e.g. one which rewards the model for using data "leaked" from the future), or to rely on in-sample results too much in other ways.
Because these issues are subtle, management will often not pick up on them or not be aware that this kind of thing can go wrong. With a short-term focus they also won't really care, because they can still put these results in marketing materials and impress most outsiders as well.
I had one candidate who was in charge of a multi-armed-bandit project at their current company. I asked them how it worked, and how they settled on that. Their response was "you know, I'm not really sure, the code was set up when I got there". He had been there for over a year, and could tell me nothing!
> A common one I've seen quite many times is people using a flawed validation strategy (e.g. one which rewards the model for using data "leaked" from the future), or to rely on in-sample results too much in other ways.
It's funny you mention this, we have a direct competitor who does this and advertises flawed metrics to clients. Often times our clients will come back to us saying "XYZ says they can get better performance", the performance in this case being something which is simply impossible without data leakage or some flawed validation strategy.
A simple stats question. If I double the number of samples, how much will the confidence interval change? Most FAANG ML engineers can't answer this question.
The definition of Standard deviation is in chapter 1 of Stats 101. https://www.google.com/search?q=standard+deviation&tbm=isch Apparently, asking Stats 101 chapter 1 question of a so called "Data Scientist" is too much of an irrelevant question!
> expect people to have very high Stats skills
Or as you have made apparent, expect people to have ZERO stats skills!
Some of the innumerate activities I have observed in "expert" data scientists and ML engineers who have years of experience without once thinking about sample sizes
1. Using A/B tests to accept the Null hypothesis instead of rejecting it
2. Squandering away 30M $ in annual revenue because they wanted to avoid a situation/meeting in which they might look like they don't understand statistics. This is hilarious because they simply nodded their head as if they understand all the calculations and then simply dropped any other meetings or followups and left 30M $ on the table
3. Not refreshing a key revenue generating model for 18 months because the were "trying to figure out" why the AUC was improving when the performance on "golden set data" was dropping
4. Using thresholding and aggregation to produce poor quality distorted training data of rich perfectly sampled data
5. Trying to use A/B tests to estimate impact even when the control and variant are not independent
All of the above at FAANGS! My coworkers in a non FAANG company were much more sophisticated. These are the kind of candidates a "build recommendations for youtube" interview selects. Template appliers.
The list of stupidities goes on and on! But yeah, none of them think that a basic understanding of statistics is necessary for work. The good thing about Javascript engineers is that they don't have an understanding of Statistics and are aware of it. However the DS/MLEs are unskilled and unaware of it.
Oh yes, good old marketing.
Along with buying off "Industry Awards" – hey, we're objectively the "Best cybersecurity company of 2022!" With a matching "platinum/gold badge" to go on our website! Or buying a place in the "10 Best Products for X" and "Independent X-vs-Y Comparison", another classic.
Because it works. Are your customers not sophisticated? Are they unable (or unwilling) to follow up on defects and outright lies? Or reality simply doesn't matter all that much to them? Humans LOVE a good story more than reality, after all.
Then your contribution as an engineer to your company's success, and hence its longevity and your job security, is strictly inferior to that of marketing. Not everything is the work of evil marketers – a lot of the supplied BS is in response to an existing demand for BS.
You would probably be depressed if you knew who our customers were, and how technologically unsophisticated they are.
Lies and deceptions.
Can you do your analysis both ways? Give your customers both, then tell them you method is more modern, but if they want outdated methods you have those too.
If someone tells you that the data says their work is good, the only real way to know if they’re right or wrong is to look at what the data says yourself. If 99% of the work is building and 1% is checking something like latency, then you’re likely to have more than one set of eyeballs on that 1%. But if 99% of the work is putting the data together and doing the analysis, then you’re unlikely to have more than one person ever look at that part.
So incompetence goes unchecked (or worse, it is rewarded).
The output of Data Science is harder for non-specialists to evaluate.
Not as management. You just have to see that other people's similar sites are not slow with the same resources, therefore it is possible for your site not to be slow. You don't have to know why you're failing to know that the totality of the people you hired were not as good as the people those others hired.
This is of course barring management failure; but if you're failing at management, that's about the same as saying that your engineers were under-resourced.
Engineering competence is largely composed of the skills to figure out what is causing problems e.g. slowness. If you can't figure out what is causing the slowness, your engineers aren't good enough to figure out what is causing the slowness, qed.
That's different than data science.
In other disciplines it is way more fuzzy. If you are in the conclusion business and there isn’t a clear path to test your conclusion in the short term you can bullshit away!
I see this pointing to any of the following:
a) DS teams overpromising the accuracy of their approaches b) marketing driving the narrative and DS getting pulled along c) incompetence from the DS team
I would imagine predictive statistics use more out-of-sample metrics like precision and recall.
That’s the problem: these metrics often come from overfitted or in-sample data, and are completely unrealistic when it comes to expected generalization performance.
I’m at the point where I never trust performance metrics anymore. Or rather, the worse they are, the more I trust them!
My reading of the OP's description is that the vendors were offering interpolative predictions, but did not use a test/train split of data. This is in contrast to extrapolative predictions which I would call out-of-sample.
Thus due to not using a test/train split, they achieved extremely good accuracy because they were testing on the same data they trained on. Even though this is "in-sample", you can't use the same data for testing and training.
Correct outcome? You totally predicted it correctly.
There is literally no way you can screw up something in statistics and not being able to make up a story to defend your approach.
Not being mean to you, just showing how typically the goal posts are moved.
To give you an example from physics, if you find just one experiment that goes against your model, you immediately invalidate the model. You don’t just make grand claims that the model in general works.
Usually isn’t it looking for an experiment that proves it and is repeatable?
If I discover a new element in one experiment, the results are published.
After publication, many labs will try to repeat and its not taken away if one can’t do it. Only if all can’t and it casts doubt on whether I did it in the first place.
Example of a test that invalidated our old theory of gravity and validated Einsteins claims:
https://en.wikipedia.org/wiki/Eddington_experiment
This is how science is done. But apparently not data science.
Maybe very far in the future there will be models of human biology that are as robust as classical physics, but right now there is such a large amount that is not understood, it's simply not feasible. A drug could work for one person and not another for reasons beyond the realistic scope of the original development hypothesis. It requires a probabilistic view to make any sort of statement about the efficacy then.
I suppose you could argue these models are just wrong and thus trivially disproven, but I don't think that's a productive framing. I doubt any biologist or doctor would claim they have anywhere near a complete model of how their specialty works. That doesn't mean a particular model isn't useful or isn't the best we currently have to work with.
Plus maybe the third best model will actually turn out to explain a separate puzzle piece in an eventual better model. Mechanistic models in biology aren't always well done in practice, but it's certainly not binary either.
Pierre Duhem would like to have a word with you:
https://plato.stanford.edu/entries/scientific-underdetermina...
> Holist underdetermination ensures, Duhem argues, that there cannot be any such thing as a “crucial experiment”: a single experiment whose outcome is predicted differently by two competing theories and which therefore serves to definitively confirm one and refute the other.
It’s like when people got upset about Trump winning when 538 only gave him a 30% chance. That one event tells us nothing. But if all predictions 538 says have a 30% chance of occurring happen 30% of the time, then they are spot on. That’s not apparent with a single event though.
The problem is that most managers, companies, and people (including yourself, apparently) are statistically illiterate enough to not understand this, and jump head first into data science initiatives expecting immediate results, which is usually doomed to fail, at which point they blame others and not their poorly formed expectations.
There’s plenty of bad data science out there, but most failed data science initiatives are doomed before anyone every builds a model or analyzes any data.
Who do you think you are?
There are three kinds of lies: Lies, damned lies, and data
I am being glib, I of course do not think all data is inconsequential, rather it is more often used from a place of ignorance or a place of ill intent it is rendered, on the whole, useless.
Unfortunately, as many posters here are pointing out, there's plenty of ways to do a correct-looking analysis of the data to get evidence to support your agenda.
Maybe your agenda is right and maybe it's not, but I'd love to hear a story of someone standing up and saying "your consultant submitted a report with glaring flaws, they should not be paid and you should reconsider X." It's more likely the little company just goes out of business or the big company buries the failure.
A typical VP will have an MBA and maybe took statistics in high school.
I've heard it a lot in situations where somebody is demanding a level of rigor that they themselves do not live up to. This is usually soon after they have framed the conversation around a solution that they want to pursue that also lacks any supporting data. That is to say, being data driven is on net good but it can also just be a thinly veiled appeal to status quo bias (which is itself not a terrible heuristic) or "highest-paid-person-in-the-room" bias.
I am not talking about in the instance of claim verification. I have seen a number of instances where a leader just wants to see data. Not any specific data, just all of the data. There is a belief that data can solve problems if only they had enough of it.
Not a data scientist, but it seems like a lot of people in business refuse to accept the fact that reality is generally boring, best practices are often "best" for a reason, and meaningful progress is hard. Of course it is possible to be too conservative, but 95% of ideas to improve a business or product are ego-stroking bullshit. Everyone wants the V10 engine to go down the highway at 65 mph, while towing a trailer, and there's only budget for an oil change every 15000 miles; don't look at the transmission fluid, just don't look.
One of the things I don't like about statements like this said in a Data Science context, is that they are true outside of Data Science as well. Executives make big decisions, managers make smaller decisions, nobody can evaluate how good/bad they really were for months or years. Engineers build something amazing, or build a house of cards, nobody cares as long as the money people are happy, even if the business use case turns out to be wrong in the long run.
>With a short-term focus they also won't really care, because they can still put these results in marketing materials and impress most outsiders as well.
Forget Data Science, you see this in KPIs as well. Say a crappy metric has to be moved by Q2 next year and people will destroy the company to move it.
I feel like Data Science is just one of those areas where you are exposed to a wider range of people and get to feel the full crapola of the insanity of working in a corporation. For lots of roles (e.g. Engineering) you get to hide in a hole behind layers of people and not see some of this insanity.
It seems like there's a general apathy/nihilism that's growing in society, whereas by contrast my entire education from childhood up I was held to strict standards and reliably punished when I failed to meet them, and this was in US public schools (albeit a highly ranked school district) and a public university. That or I was just raised in a bubble, and the historical examples I referenced growing up and reference to this day are just a case of survivorship bias, and all the bullshit that was alongside them back in the day has simply been forgotten. I'm not sure, but it is disappointing how little people at large seem to give a shit. Maybe it's a side-effect of the obesity epidemic and people just have less energy or something
This gives plenty of space for opportunists and tricksters to hide.
You don’t ever have to fear being beheaded by the people whose life savings you stole and you don’t have to face consequences if you have a good lawyer.
To do well in todays world learn all the rules and where the loop holes lie. Violating the spirit of the law is fine as long as you can lawyer around the letter of it.
The older I get, the more I realize how fragile a lot of human systems really are, but I suspect it has always been this way and it won't change significantly any time in my lifetime.
Your comment itself sound somewhat nihilistic, so I hope you're doing well mentally!
I agree that human systems have always been fragile, but have long been papered-over by things like "decency", "tradition" and "doing the right thing" and in extreme cases, mobs with pitch-forks.
I disagree that it won't change in our lifetime(s) - the extreme polarization and tribal politics will get worse and people will let systems break - or intentionally break systems just so that their team will gain a short-term win. I have no idea what new horror it will take to remind people to be decent to each other again, but looking back at how divisive COVID-19 was, I'm not hopeful.
Students of history and the arts can get an earlier exposure to this worldview. I think we engineering types can get too focused on technology and imagine everything is innovation and progress. You have to work uphill against your default interests to expose yourself to a longer view and consider that fundamentally modern people with modern minds lived for (many) thousands of years doing almost all the same cognitive things as us, just with different physical props.
Our lungs are constantly in flux as we breathe. But at the same time, we're just breathing and that doesn't really change until our end. I'd say human social systems are much like that.
For what you specifically experienced, my opinion, the bigger the organization the more inevitable this seems to become. To make things worse the size of the organization isn't limited to just a company or non-profit but to the size of all groups involved, i.e. a small charity or non-profit that's part of a huge government program is similar to a small engineering team in a huge tech company. They could do huge things or be completely worthless and so long as they pass along positive messages up the chain and the org or company as a whole is doing well then yay no consequences.
We're (hopefully) at the beginning of a cycle where companies realize they are causing apathy amongst the majority of the employed and hopefully experiment (and succeed) in providing meaningful pay raises to the lower echelons which will come at the short term costs of profits but are justified for long term productivity. Or we'll just keep divolving into a dystopia
In real life there's basically one absolute goal, and that's survival. And that's largely assured in developed western countries these days, unless you do something really stupid. Everything else is socially constructed, and pretty arbitrary. There are some decisions that are fairly consequential for what your life will look like (where & whether to go to college, what field to go in, what metro area to move to, which employers to work for, who to marry, whether & when & with whom to have kids), but you will still have a life regardless, it just might be a slightly smaller house or a spouse that you click with worse or less disposable income for travel.
That's also instructive for what decisions actually do matter. Don't do drugs. Wear your seatbelt. Don't get pregnant unless you mean to. Don't play with loaded guns. If you're staying away from major causes of death you're generally doing pretty well.
Or just get unlucky: no need to do anything stupid. One can easily die of cancer at 30 and leave a toddler behind.
Chances are it won't happen to you and your close ones. Perhaps try being grateful rather than dismissive?
[Edit: perhaps we have a misunderstanding as to the word "easily". I'm not saying it's likely, I'm saying it can and does happen without any warning signs and no amount of planning/preparation can save you.]
That cancer is what's known as a Hereditary Diffuse Gastric Cancer gene (HDGC). It just so happens that the E-cadherin control that suppresses those cancer cells is not processed properly. The diffuse part is what makes it particularly tricky. It's on the surface of the stomach epithelial cells and progresses from there. The only solution is a total gastrectomy (prophylactic if you do it early). No carcinogen necessary. It's found in populations all over the world and pathogenic lines don't even have to be related. The mutation can occur independently in the germline and is passed on. As long as you reproduce before it kills you nature really doesn't care.
Fun side fact. It also predisposes carriers to 70% chance of breast cancer. As a result many of those diagnosed are women who then find out they need to also have their stomachs removed.
But after growing up and having kids of my own as well as watching others' kids grow up with varying degrees of parental involvement, I have a whole new appreciation for adult caregivers who get involved and help shape healthy behaviors and habits in kids.
> your caregivers want you to mind your behavior, because then they don't have to, even if you would've been perfectly fine playing with mud or swearing in school or watching TV all day.
You've got it backwards. The easy way of caregiving is to just not care. Let kids watch TV all day, swear in inappropriate social situations, and whatever else they feel like doing. You don't have to get involved if you just don't care what they're doing.
But anyone who has worked with kids in an education setting can tell you that this doesn't actually produce good outcomes for the kids. There are occasional exception stories where students with minimal parental involvement lean heavily into becoming successful in life, but the more common outcome is that hands-off or absentee parenting styles lead to poor outcomes for the children, including social and personal issues. It's not just about getting good grades just because. It's about learning how to operate and function within a civilized society, as well as how to balance your own emotions, impulses, desires, and other behaviors they need to learn as they grow up.
Broadly the values I've tried to instil in my children break down as
- Take responsibilities seriously
- Apply your best efforts
- Be considerate
- Cultivate empathy
- Value yourself
- Don't be a dick
- When you screw up, admit it, and make amends
These, when applied, lead to the behaviours that make parenting easy, and I'm hoping will make them good members of society in general.
The behaviour that most of the school system seeks to embed in children is primarily "obey and do as you're told, don't question", which is far from good.
Raising children to care is good and takes lots of effort.
Raising children to leave parents alone usually means the children end up not caring or worse.
As I get older, I'm actually noticing more and more consequences catching up with people, albeit slowly. The people I knew who drank heavily through their 20s and 30s are in much worse shape than basically anyone who made an effort to stay healthy. People with poor diets and low physical activity are visibly worse off than others who paid attention to their inputs. I knew several people who got into recreational drugs in their 20s thinking they were safe because they educated themselves before hand, yet who ended up losing jobs, relationships, wealth, and a few who even lost their lives.
I've also noticed more peoples' career reputations catching up with them. It's not uncommon to interview someone only to later discover that they left a very negative reputation at a previous company where I happen to know someone.
I was very jealous of one of my peers who job-hopped his way up the salary ladder, joining companies and then immediately focusing on nothing other than interviewing at his next salary increase. He rotated through several of the big companies here until his reputation for demanding high salaries and then delivering nothing at all finally locked him out of any company with well-networked people who knew about him. He literally had to leave the state and go somewhere new to escape his past network and get new jobs after 10 years of this.
Consequences do catch up to people most times, but it's not immediately obvious. If you expect immediate justice or for people like SBF to go straight to jail the moment the headlines break, you're only seeing the beginning of the story.
Agreed. Few people defy gravity; in the end most hit the ground.
The phrase “slowly, then suddenly” comes to mind.
But Long term consequences always do.
Also, one could argue another interpretation of what you are advising is never take a risk, because it will have consequences. Well, in real life, it doesn't always. You can get away with a lot, and people do.
That's not at all what I was saying. I was referring to predictable negative consequences of unhealthy behaviors.
There are many risks that don't involve gambling away your health, your reputation, your credibility, etc.
With the economy contracting and inflation skyrocketing, consequences should be back in fashion relatively soon. We're already seeing it in mass layoffs and other areas of business.
> He literally had to leave the state and go somewhere new to escape his past network and get new jobs after 10 years of this.
That's not even that bad of a consequence. It sounds like his strategy was worth it tbh.
Personally I hate this kind of behaviour, but from a maximization POV (Especially in regards to career) it seems like the best move. There is likely some risk of ruin, but the upside appears to be much greater.
I largely agree with you: the big names attached to the resume, the pay, and the effort spent on interviewing skills likely offset the negatives of the reputation (though I also intuitively don't like it because the strategy is rather self-centred).
However, the consequence is rather significant if he has roots. It's harder to pack up and move if one has a romantic partner who is settled into a job at a particular place, and you could also possibly be leaving family and friends. Sometimes one has to move, but typically one has the option to come back, which wouldn't be practical for the person in question. It's still plausibly worth it for the person if he didn't have roots and collected a lot of compensation, but especially when one is older (the commenter mentioned 10 years of workin experience), moves can be tougher.
Separately, to put a positive spin on this, it often takes time for positive habits to pay off. When picking up a positive habit (e.g. exercise and especially learning a new technical skill such as a language), oftentimes much of the reward doesn't come until far later. This is important to keep in mind, especially if one has self-doubts or even a lack of encouragement for trying to adopt a new positive habit in one's life.
There is garbage and tent encampments thoughout much of my city, and I am told that nothing can be done about it. I've been invited to engrave something on my catalytic converter. I wonder what good that would do.
in 2018, the capitol was open to the public, no one broke in. 78 were arrested in the Capitol on Oct 5, 2018 and charged with Crowding, Obstructing, or Incommoding [1].
in 2022, the Capitol was closed to the public and people broke in. 12 people were arrested on Jan 6, 2021 and charged with Unlawful Entry or Assaulting a Police Officer [2].
Assaulting a Police Office is a felony; Crowding, Obstructing, or Incommoding is a misdemeanor. Seems there was a differential in severity of breaking the law as well.
[1] https://www.uscp.gov/media-center/press-releases/us-capitol-...
[2] https://www.uscp.gov/media-center/press-releases/us-capitol-...
I’m not sure how once can fail to see a difference between that group and the one protesting Brett Kavenaugh’s confirmation.
It seems that individuals often bemoan such a lack of consequences, but for some reason they are still quite prevalent in our “systems.”
I wonder how to harness the good intentions of individuals…
I wouldn't say fewer uniformly, but certainly very noisy. Some have their lives destroyed for minor or non-existent misdeeds, others get away with egregious crimes.
The Jan 6 riots are possible because again, the Capitol Police weren't ready to lay their careers and lives on the line "cracking skulls" to defend an old building. Most of them probably were taking in the spectacle and thinking about how exciting it will be to recount with their friends/family later.
its the "the higher the pay the easier the job" paradox.
I think you are also seeing the effect of the oligopolization of the world stemming from the bad rework of the antitrust laws relaxing antitrust enforcement significantly from the 1970's through now.Any sort of market power is really bad for this kind of behavior because almost noone wants to rock the boat if they don't have to and when you have an oligopoly/monopoly you can abuse you often can hide this stuff in slightly lower but still excessive profits.
I am not sure if apathy/nihilism is growing in the larger society. I think that things have always been like this because people have always struggled to find meaning in life. After taking an intro psychology class, I was exposed to the idea that society wants an individual to police him/herself. The "super-ego" that makes one feel guilty for breaking rules and want to aim for perfection.
This is purely anecdata, but I have found that this is more pronounced in a data science context. Managers and executives are (in my experience) more willing to admit they don't understand engineering work product and seek input from technical advisors, and executives and managers deal with decision making on a daily basis and understand that it can be nuanced. But since almost everyone reads financial reports or has to make a chart in Excel every now and then, they know enough to read someone else's analysis but not enough to recognize their knowledge gaps (particularly wrt advanced statistics).
When it comes to disastrous long term decisions, there's plenty of time to get input from multiple stakeholders. I always remember the armies of companies who went chasing after Hadoop because Big Data was going to transform something or the other. All the stakeholders were on board, from the CEO and CTO to IT and Engineering management. How much money and time got flushed down the toilet trying to implement and extract value from data with Hadoop. They only people who paid the consequences were the employees at Hadoop companies who thought their stock options would be worth something.
Relying on your data science or marketing department to tell you how good your data science or marketing department is doing, with their own metrics and their own evaluation methods that you don't understand, can only really lead to one outcome.
I guess data science is inferior to research in this way. People care about research methods, rigor, etc… Maybe data scientists should adopt stricter standards, like actual scientists.
This is to be expected from an information theory point of view. It's why "fake news" will always be a thing.
"Managers will say they want to make data-driven decisions, but they really want decision-driven data. If you strayed from this role– e.g. by warning people not to pursue stupid ideas– your reward was their disdain, then they’d do it anyway, then it wouldn’t work (what a shocker). The only way to win is to become a stooge."
In science, a good scientific result can be bad for business. There is often little appreciation for the "science" in data science.
It feels like even Google falls prey to this at times: they keep redoing the same A/B test until it comes up in favor of the change (or the designer whose pet project it is runs out of political capital, presumably).
Does management look at slides or AB test dashboards?
When OP talked about "the main bottleneck to my work" in terms of areas he would need to learn more about -- I was expecting him to talk about facility with statistical methods and using them appropriately!
I'm not sure what to take from the fact that he never did! I would like to ask him what he thinks about that!
And for the same reason that people tend to want pseudoscience instead of science in any other domain, too. Science is slow, tentative, and messy, and usually responds to questions with even more questions rather than with answers.
Pseudoscience tends to be much more concerned with exuding confidence and providing clean-cut answers. It's what happens when a desire for science meets a need for instant gratification. Along the way, things like blinding and controls and watching for bias and validating assumptions tend to get dropped when they're inconvenient or difficult to explain. And they're always inconvenient and difficult to explain.
Technically, I think investors & owners would want the company to use real data science to improve products & maximize profits.
Everybody in the middle just wants to use data to lie to get promoted faster - because you don't get promoted for actually doing a good job - you get promoted for convincing people you did a good job, and lying is a VERY useful / effective tool.
This is based on the assumption that companies are focused on long term profits and stability, and I’m not sure why anyone believes that to be the case anymore. The vast majority of companies are run based on next quarter’s stock price or growth metrics.
I worked on a newly formed data science team coming out of grad school that was tasked with taking some predictive initiatives that the company had relied on external consultants to produce, and implementing them in-house. The external team’s results always looked exactly like what the business wanted to hear, but they rarely played out in practice. This was in part because the underlying data quality was terrible, and the company wasn’t executing in a way that allowed anyone to actually answer the questions being asked. The consultants would just torture the data until they could come up with a report that would ensure the company would come back the following year. So we spent a lot of time trying pouring cold water into the business groups who saw data science as a magic wand that would conjure up more money at no cost. But we never were able to convince them to invest in anything that would take longer than a year. Anything that would require a change in their marketing or strategy executions that wouldn’t immediately deliver increased results was just a non-starter. But actual data science requires that kind of investment for long-term layoffs. So the data science team became figure-heads, never given the buy-in to actually make impact on business, but kept around so teams and leaders could tout being “data-driven” and throw “AI” and “machine-learning” into PR and marketing materials.
You aren’t wrong about middle management is looking to get promoted faster. But every single individual from the employee looking for a promotion to the executive suite to the investors are addicted to incentive windows no longer than 6-12 months.
LARGE INTERNET DATA COMPANIES. They want the real data science.
For them, data science actually allows them to perform a core business function (target their customers) in a profitable way (one way, asynchronous relationship. Note the complete lack of any "talking to a human being" in your relationship with big tech).
For everyone who isn't a large internet data company with an asynchronous relationship with their customers... what's the point?
Usually, they have only a handful of technical projects that benefit from data science.
In my experience, my multi-billion dollar organization got by with a shockingly small number of "real" data scientists.
Management is not your teacher at school, it is not there to check up your results make sense.
Management mostly assumes you’re competent at your job.
Our CTO did the "Quick, this is the future! I'll be fired if I don't hop on this trend" panic thing and picked up a handful of recent grads and gave them an obscene budget by our company's standard.
The main problem they were expected to solve - forecasting future sales - was functionally equivalent to "predict the next 20 years of ~25% of the world economy". Somehow these 4 guys with a handful of GPUs were expected to out-predict the entirety of the financial sector.
The amazing part was they knew it was crap. All of their stakeholders knew it was crap. Everyone else who heard about it knew it was crap. But our CTO kept paying them a fortune and giving them more hardware every year with almost no expectation of results or performance. It was a common joke (behind the scenes) that if they actually got it right, we'd shut down our original business and becomes the world's largest bank overnight.
At least it finally gave the physics modelers access to some decent GPUs which led to some breakthrough products, as they finally were able to sneak onto some modern hardware.
Some companies just don't have the data, or heck even the need, for data scientist yet try and hire them anyway.
Give smart people a fundamentally ill-posed problem and they won't get anywhere anyway.
Also, w.r.t hiring in cases like these, I think often the experienced candidates can smell that this won't be a good gig so don't apply, while the less experienced (or desperate) ones apply. This means the workers get stuck with an intractable problem, and the company gets stuck with workers who are too inexperienced to know better.
I think that vast majority of human organizational structures, individuals to large corporations and countries have no clue what they are doing. The most successful apply science and just barely keep their head below the surface by avoiding utter failure. Most people would describe the Olympics as a competition to find out who is the best amateur in a given sport. No, the Olympics is a competition to see who can make the least number of mistakes.
If you are going to make bold dumb moves, you need a whole lot of margin.
The business value comes from the stats guy.
but otherwise, yes, I see the problem.
There’s no point in fixing it. You can just pretend like you did. But if the stat work is quality, then it’s worth the effort to optimize.
Maybe a little harsh...
I dont know anything about Data Science but as a bystander with a mathematical background thats what I assumed was going on so its kindof interesting to see it spelt out like that. Like you've put words to a preconception that I didnt even know I had.
At the time I considered going down that path, but decided I did not have anywhere near the statistics & math knowledge to get very far. So I stuck with the path I had been on. Over time I saw a lot of acquaintances jumping into the data science game. I couldn't figure out how they were learning this stuff so fast. At some point I realized that most of them knew less than I did when I decided I didn't know enough to even begin that journey.
Of course, I was comparing myself against the giants of the field and not the long tail of foot soldiers. But it made for a great example to me of how with just about everything there's a small handful of people who are the primary movers, and then everybody else.
[1]: https://logicmag.io/intelligence/interview-with-an-anonymous...
That's a good way of putting it. I remember in my first calculus-based probability+statistics class in college, I felt incredibly challenged by the theory. I wondered why there are so many probability distributions out there, why the standard stats formulas look like they do, what "kernel density estimation" even is, etc.
On the other hand, my data science course did include some theory, but a big part of it was also learning how to type the right commands in R to perform the "featured analysis of the week" on a sample data set. Something about these lab exercises felt off because it felt more like training rather than education. The professor expressed something along the lines that if we wanted to go far with this in the future, he would expect us to design the algorithms behind the function calls. I think the analogy he used was "baking a cake from scratch rather than buying a ready made one at the store."
An important thing people miss is that shallow statistical knowledge can cause subtle failures, but shallow software engineering knowledge can cause subtle failures too.
A junior frontend developer will write buggy code, notice that the UI is glitched, and fix the bug. A junior data analyst will write buggy code, fix any bugs which cause the results to be obviously way off, but bugs which cause subtler problems will go unfixed.
Writing correct code without the benefit of knowing when there is a bug is challenging enough for senior developers. I don't trust newbie devs to do it at all.
Context here is I used to work in email marketing and at one point I was reading some SQL that one of the data scientists wrote and observed that it was triple-counting our conversions from marketing email. Triple-counting conversions means the numbers were way off, but not so far off as to be utterly absurd. If I hadn't happened to do a careful read of that code, we would've just kept believing that our email marketing was 3x as effective as it actually was.
So, it's impossible to know how much of a problem this is. But there is every reason to believe it is a significant problem, and lots of code written by data scientists is plagued by bugs which undermine the analysis. (When's the last time you wrote a program which ran correctly on the first try?) Any serious data science effort would enforce stern practices around code review, assertions, TDD, etc. to make the analysis as correct as possible -- but my impression is it is much more common for data analysis to be low-quality throwaway code.
This is true everywhere. As a professor, every semester I’m baffled by students who aren’t curious. But I’ve come to terms that there is a difference between those who will graduate and go on to be readers of hacker news and write this kind of article, and those who won’t.
I have more important things to do. The hacker mentality, imo, is about identifying what’s useful for you to explore to accomplish whatever you need. Often that’s a lot of glue between things that other people built. Other times it’s tweaking the internals to do something a bit different.
That said, sometimes I do like to read the code in libraries I use but often this is more for enjoyment with occasionally learning something interesting.
As an example, it’s more about understanding the statistics and linear algebra around estimating uncertainty in GLM regression estimates, than about reading the code for how the statsmodels library implements that.
I’d further argue that the nature of a hacker / power user is to break things apart once you want to get deep enough. If I need to know where in the cluster my instance of some software got lost into, I should be able to investigate all the tools I have available to somehow find it. Not just give up and say some garbage collector will get it for me.
With regards to your API statement, I'm just as guilty regarding reading the code, but I do run some manual tests to ensure that my script calling the database actually does what I think it should. Is that good enough? Who knows :)
I am a naturally curious individual but time limitations prevent further exploration in most circumstances. Additionally there is a relevancy factor weighed on top of it. If something looks curious I have to pre-determine if I think the time spent pursuing that rabbit hole has any value to it. Granted you never know the outcome - it is alway a gamble.
Also hard subjects at uni - there is only so much deep thinking you can do per day
I know very few university students with significant work commitments.
In the US, the stereotypical college student is not also holding down any kind of job. Maybe 5-7 hours of "work study" (light work running the reference desk at the library or working in the dining hall).
Frankly, I doubt the majority could do learn a lot and also work a significant number of job hours.
At a community college, it would be very different - most students also holding down jobs, I would guess. At a flagship state university, I would be very suprised.
Evidence in [1]... about 30% of full time students are working 20+ hours/week. Also apparently I was wrong about the low hours being typical; less than 10% are working but < 10 hours/week.
No kids, no sports, no community involvement, no side hustles, no expectations.
Both before and since I've had more free capacity to pursue learning for it's own sake.
Some semesters I was doing like 70-80 hours a week on average, split between managing clubs, homework, attending class, working part time jobs, studying. One week I remember being busy from 7am to 2am for 6 days straight. a few semesters I had a lot of free time, like second semester of senior year, and first semester of freshman year, but mainly it was the gaps - after midterms, during breaks, where I had obscene amounts of free time.
I learned a lot from my CS classes, but I actually felt like most of the value from the degree came from overhearing random chitchat between professors or other students and the reading more about those ideas and experimenting with them in my free time.
Since then workload has been intense of course but never comparable. I've had much more time to be able to explore personal interests since college.
Once I started full-time work it was like a revelation - finally I don't have to work on evenings and weekends! I actually get free time to myself! I can have hobbies!
I was. Didn’t have anywhere near the free time and the lack of stress I do post-college. Helps that I also make a good chunk of change rather than living off a relatively small stipend in one of the most expensive cities in the world.
I did my undergrad at a state school with a middling engineering program, where I had ample free time to explore topics in depth, pursue extracurriculars that taught me far more than my classes, and have a thriving social life.
Contrast that experience to what I saw as a teaching assistant at Georgia Tech: undergrads who are so full of classwork that they're punting on the least-valuable graded assignments, never mind extracurriculars. The level of rigor in courses is much higher, but it presses out freedom to explore independently.
Another datapoint: I competed against GT extracurricular teams during my undergrad years, and we beat them handily almost every time because their students couldn't justify high effort for work that wasn't graded. I once saw a GT team arrive a day late to a competition, work on a robot for three hours at the adjacent table, realize their robot did not work, and drive home without competing.
Contact hours at most universities are around 2-4 hours per week per 15-credit module. To gain a degree, you have to take 120 credits a year, typically two terms of 4 x 15 credit modules, or 8-16 hours of contact per week maximum with the entire summer off.
You therefore have at least 24 hours a week to study on your own to bring your working week up to 40 hours. Maybe you're working, fair enough. But if you don't have time to study subjects in depth then you need to reduce your working hours. If you can't, then by definition you are not a full-time student.
This is not a personal attack on you. Perhaps you were genuinely studious and spent all your time poring over the coursework. It is a commentary on the whole academic sector where we repeatedly see students do nothing for most of the time and spend the last 2 weeks cramming and putting in substandard assessments, then blame the course material/their lecturers/their anxiety etc. for their poor results. And of course the leadership teams lap it up and tell us to make our courses easier.
The difference probably belies in the rigor of the program. It sounds like you are working in a non-engineering based program. In our engineering programs we had 40 hours of class time + lab time per week.
I had a concurrent arts degree at the same time which is was, in comparison, incredibly light workload - though concurrently it took time away.
The only time that I will say was much lighter was in the final year of undergrad - the course load finally lightened up.
N.B. this whole conversation clearly excludes summer.
40 hours isn't normal even for engineering programs. Every engineering program I've looked at has higher course hour and credits required. Obviously I haven't looked at every single engineering program at every engineering school, so there probably exists some counter example showing it's no different than arts or science..
Where I studied, we had one semester with 40.5 hours of lecture, lab, and tutorials. One other semester was around 38 or 39 hours. The rest were in the mid-twenties for lecture, lab, tutorial. My program wasn't the typical engineering program, but all of the other engineering schools where I went (Western Canada) did require more credits and more class hours than science and arts and business programs. There may have been some exceptions with honors programs (meaning they have to take 132 credits vs 120 credits and write a thesis) in arts and science that put them closer to engineering programs, but these have limited enrollment.
This is anecdata of course, but my experience with a top-3 US undergrad aerospace engineering program in the late-90s, early 2000s was around 15-16 hours of class time per week, sometimes increasing to 18-19 or so with labs. Work outside of class was 3x this or maybe 4x around midterms or finals.
For example here is Purdue's handbook on credit guidelines:
https://www.purdue.edu/registrar/forms/Semester_Credit_Hours...
For context, I double majored in two adjacent subjects, physics and math. I went to a state school that has a very strong physics program. I also worked in physics lab for the last ~2 years, and graduated a semester early. While I did OK academically, I had no desire to run the gauntlet again in grad school, and left to work in tech.
I have never, ever, been as a busy as I was in college, nor do I ever want to be. I think that's a good thing! I have much more time to explore things that don't pan out, to do things I know are not "productive" (i.e, play video games), and am generally happier.
Apart from quality of life improvements, I think there are additional financial and intellectual benefits to not being overly burdened -- the time to explore topics that were not immediately adjacent to my field of study results in extremely useful skill development and better cross-pollination of ideas.
Did you go to a "good" school?
I went to a mediocre one for undergrad and a top school for grad. The one glaring difference I saw between the two: The top school's undergrad program gave students way, way too much busy work. All that work didn't give any insights, and was merely used to artificially distinguish students for grades. Their grad program was nothing like this.
Really glad I went to a mediocre school. Still learned everything, but had plenty of time to explore.
In my case, when I was studying, I had all the time in the world. During that time, I did try to learn many things but I didn't go deep or weren't consistent. I mostly wasted times in goofing around. Looking back, the amount of time I wasted during college years has become a biggest pain of my current life when I don't have any skill or time to learn those skills.
The first issue is a lot of acquiring broad-spectrum knowledge involves risking quite an amount of money. A good DSLR cam can easily rack up a few thousand euros, a fully spec'd Mac Pro or larger drones cross the five digits without blinking. Messing around with gas and electricity can kill you, messing with water pipes can cause immense water damage. It takes a lot of ... let's say recklessness to even think about dealing with this if you're not a professional, and you have to have the resources in the first place.
But the real issue is time. Students, at least here in Europe, don't have the luxury of taking six or seven years for their basic diploma - "thanks" to the Bologna reforms, you're fucked if you can't make it in the designed timeframe as you won't be eligible for most kinds of financial aid. That means you simply cannot afford "wasting" a week to get that deep level of knowledge, you simply are happy enough if it runs well enough to get a passing grade. And once you've entered the workforce, it becomes even harder to have actual hobbies. It's one thing if you live alone, no one will bat an eye if you pull in an all-nighter on a weekend with just yourself, a crate of beer and a laptop and that's assuming you're not completely drained from your average 40 hours work week, 10 hours of getting to the workplace, and another 10 hours on domestic chores. When you live together with another person, the game completely changes: they also want time and attention from you - bonus points if your s/o has roughly the same interests that you have (which is why I suspect so many people meet their s/o at work). And with children... forget about hobbies of any kind if you don't have enough resources for either yourself or your s/o to be a stay-at-home parent.
This is why I so strongly advocate for a four-day and six-hour work week, a proper minimum wage and government-subsidized affordable housing for everyone. Just imagine what useful things people could run as side projects if they actually had the time to pull them off, not to mention the obvious physical and mental health benefits of not having to struggle with survival every single day. Add to that the elimination of "bullshit jobs" and an end of wasting the best minds of the world on financial bullshit (i.e. HFT, "quant investment funds") or advertising... or getting rid of racism and other discrimination. We as humanity could make so much more progress if we were not so hell-bent on exploiting each other.
Is this true in Germany?
Rather, you have friends with money. My university circle was not "rich kids", but those who either had to fend for themselves with government assistance or with joint work-study programs ("Duales Studium").
In the end, it's clear that the Bologna reforms were not just about creating a common European academic standard - which is a good thing - but also about turning universities into factories with no room left anymore for the eccentrics who wanted to spend more time on their education than the system wants, those who need more time for specific subjects or those who need a job to survive and can't keep up the same workload as rich students can.
[1] https://www.xn--bafg-7qa.de/bafoeg/de/das-bafoeg-alle-infos-...
[2] https://www.tuhh.de/tuhh/studium/im-studium/rund-um-den-stud...
[3] https://www.deutschlandfunk.de/zwangsexmatrikulation-fuer-ho...
[4] https://www.gesetze-bayern.de/Content/Document/BayHSchG-49
Because there's a lot of things out there which are also interesting, and you don't have time to do all of them, so you choose. And different people choose differently.
It was a revelation to me when I realized that, no, it’s not that “most people” lack intellectual curiosity. Their interests are just different than mine.
I did student representation while I was at college, so I had quite a bit of contact with teaching staff around discussing the learning process. There were a lot of complaints from their side that students weren't engaging with the course and were rote learning answers for exams.
My perspective was that most of the courses were badly taught (students were given little guidance and struggled to learn the basics) AND badly examined (you had to guess at what the professor wanted in order to score well - it wasn't actually assessing learning accurately). The courses where you found truly curious students were the ones that taught the basics in a way that other professors would consider hand holding (which meant they could get passed that onto more advanced material), and gave clear advice on what was expected in and how to approach the exam (so that students didn't have to worry about that and could focus on learning and their interests).
You'll always get some students who just aren't interested (perhaps they picked the wrong course, or simply aren't that academic), but you'll also find that the same students respond dramatically differently to different environments.
Story time.
There was once a junior data scientist at Shopify that had learned Python and SQL and was tasked to figuring out how to fix their "broken app store recommendation engine" but since they didn't know Ruby, they asked for my help in figuring out what was going on.
Well somewhere in the soup of math was a fuzz factor at the very top. Think of it like
factor = 0.something # Not 100% sure what the decimal portion was.
some_complicated_math_that_maxed_out_at_one_pt_zero() + rand(factor)
Now the thing about ruby is that rand is basically broken for floats. Negative or floating point values for max are allowed, but may give surprising results.
https://ruby-doc.org/core-2.4.0/Kernel.html#method-i-randSo basically what they thought they were doing was introducing a bit of randomness that would hinder others from reverse engineering their algorithm. What they actually did was make the recommendation algorithm fifty percent total noise. Yes, it's true. On every load half the recommended app scores were noise.
They fixed the bug and I'm sure a ton of balance sheets for businesses around the world are markedly different know because of it, but I never heard of it again.
This is one of the core problems with data science.
The lack of feedback.
Having a subset of totally random recommendations wouldn't be a totally terrible idea---especially if you know which they were! It could help push the system out of local minima and it's the obvious benchmark to beat.
But over time, the higher up I climbed the more I realized the job had marginal business impact. Usually a big company would hire a bunch of PhDs with fancy degrees and stick them in some "Advance Analysis" department and leverage them as internal consultants, which just meant creating some models, writing a powerpoint deck, get a pat on the back from the execs - not a single model would ever see the day of light. I got all the way up to Director this way, before calling it quits this January, at the end I had basically nothing to do except work on "corporate AI strategy", which meant writing presentations and white papers for upper management.
It was comparatively easy job, one could coast their entire life in some of these corporations - especially in government sanctioned oligopolies like banking.
Do you have any observations why? I'm a pretty lowly business analyst, but my observation is if you don't own the decision making (usually by having profit and loss responsibility), you can't have much impact. Possibly it's the companies and industries I've worked at, but at the end of the day if the results don't meet expectations, it's the business owner that gets fired and not the people providing the recommendations.
All the "intelligence" takes place in the humans that design experiments to collect unambiguous data. "data" absent a profoundly intelligent (, expensive, fraught, ...) experimental design is basically useless.
Basically, like some other people have already said, companies are inherently political - they do not want data-driven decisions they want their decisions to be data-validated. If their view of reality aligns with the data that is all the better, but if it doesn't, their alignment takes priority. Moving up as a DS then involves delivering "evidence" that fits whatever narrative your boss and senior management want. Sometimes that evidence will be rock solid, other times there is no evidence. That's why I suspect in the beginning they loved hiring STEM PhDs from "elite" universities. If your degree is from Harvard Astronomy Dept, people will borrow your credentials to further their agenda - because you got a golden halo.
TLDR: science is not gospel, it's just a method of thinking to deduce natural laws, if you keep digging you can find your initial assumptions proven wrong, sometimes completely wrong, - in business and politics if you dig too hard, you start finding things that nobody wants to hear.
Regarding your point about owning profit and loss that is very true as well. In my second job I was in a center of excellence team and it was extremely hard to get any traction because we didn't own any sources of revenue so we were a cost center like HR or Accounting. Teams that owned LOBs want to hire their own analytics rather then "outsource" to a COE team as a way to retain control and expand their own power base.
Would I ever do it again? Who knows, maybe, I still believe it's possible to do good scientific work outside of academia (not to say good science always gets done in academia either). I am living off investments and savings right now and working on hobby projects that may or may not pan out. People always take less than ideal jobs for want of reality.
I think there is real value in scientific analysis in business but it's closer to operations research where you solve complex optimization problems that are directly pertinent to the core business (like traffic routing or container packing) than in busting out the latest DNN techniques.
Ooofff. This is too true. How often is the case that data is collected to test hypotheses vs confirming priors?
Find me some evidence of WMDs in Iraq! Yessss Sir!
I've seen so many analysis tasks where data scientists without questioning went away for a few weeks to crunch data and come back with some random graphs and statistics that are completely useless as decision support.
The preceding sentence is a hilariously cynical zinger:
“Those who have seen my Twitter posts know that I believe the role of the data scientist in a scenario of insane management is not to provide real, honest consultation, but to launder these insane ideas as having some sort of basis in objective reality even if they don’t.”
Of course, in many situations the business totally lacks what it needs to correctly do the "data-driven" stuff they want to, and it'd take a good deal of up-front effort by competent people to get it, amounting to entire new projects or deep modification of existing projects.
So, given the choice between: going without that stuff and acknowledging that a lot of what they're doing is guesswork and gut decision making, or simply arbitrary; putting a smaller but still-large amount of work into finding out what they can glean from what's available; spending the time and money to collect what they need, the right way, to do the data-driven decision making they claim to want to do; and insisting they're doing things "data driven" but having all their data hopelessly ruined by e.g. selection bias and comically-bad experimental construction that can't possibly be yielding reliable results, so they can cheap out and get no actual "data-driven" benefits aside from falsely claiming that's what they're doing—they tend to go with that last option, nearly every time!
I hated working with bad code and dealing with arrogant phds who don't value good code. I've seen so many terrible Jupyter Notebooks just copied and pasted into VS Code and the data scientist just washed their hands of it calling it "production ready." Here's a conversation I've had multiple times:
Me: have you ever considered not making every variable global scope
Them: that's just software engineering. We do machine learning
Me: if it's just software engineering, then why can't you do it?
Meanwhile, automated data science tools are getting halfway decent. If you know what algorithm to pick and you don't need to run millions of records through the model every minute, your standard business analyst could probably get a solid model going--at least as well as most data scientists for all the reasons the article mentions.
And I like that I know I can do data engineering. With data science you can never really know if you can hit your target metrics given the data you have. So data scientists end up encouraged to fudge their results or make sloppy decisions. With data engineering I can say "yes this is doable or no that's not" and people believe me.
My prediction: there's value in the massive volume of data but most of it can be had through standard dashboards, some summary statistics, a graph network, or maybe a linear/logistic regression. Most data science is BS and companies aren't getting the return they need to pay for these guys. (And good God, you almost certainly don't need a neural network.) Meanwhile, data engineering will get integrated into software development, and machine learning—by virtue of its proliferation through academia—will just become another tool for software developers. Data scientists won't get laid off enmass but they will go the way of the webmaster: either pick up new skills and evolve or move on til they end up with new titles
So many data scientists are full of themselves thinking they are magicians and software developers are blacksmiths who are beneath them.
Incrementally at my company the SWE's have automated so much of the data scientists workflow that they end up just as you describe, using the tooling and being relegated to becoming analysts.
After 3 years coming back to this field, I see the writing on the wall: In the 90's most models were created by software developers, in the 2030's most models will be created by software developers.
I say so because I've had time to read some of the reports that DS teams produce to drive decisions in my BIGCORP and it makes very little sense most of the times.
And we suffer from it because we have direct contact with clients, but nobody cares about my department opinion, they will rather believe in some model where I can see insane dispersion in datapoints when they plot em in reports, conclussions by people who clearly has zero understanding of our business.
I'm forced to make decisions on how to treat certain customers, by entering data into some software and being given an output I can't challenge, which produces lots of insane and unfair situations.
Also, IDK how they clean and treat their data, but if they're relying on our ERP's data, good luck. Our CRM if full of BS because most employees rush to put whatever it allows to continue as they need to keep up with KPIs, so they aren't trying to make nice comments and check everything is ok.
Data is only helpful when it is directly and clearly tied to the problem.
* Good: "Our customers are complaining of random drop-outs. We've noticed X% of requests to Y service take longer than Z time. We believe that's the problem".
* Bad: "Companies who are most successful on our platform upload X things in their first Z days. We must find a way for everyone to upload X things in Z days".
However, it seems this person's biggest gripe is with good old crap management; the bane of business for hundreds of years.
This line stood out:
> Companies all over were consistently pursuing things that could be reasoned about a priori as being insane ideas– ideas any decently smart person should know wouldn’t work before they’re tried.
That pretty much summarizes why I have been told that today's companies want only young people. Us "olds," are "negative naysayers," who say things like "You know that the laws of physics forbid this, right?" or "I tried that, a couple of years ago. It didn't work out, and here's why...".
Apparently, young people are able to do the impossible, because they haven't been told it's impossible, and mixing "olds" with them, spoils the soup, by telling them it's impossible (or maybe a lot more difficult that they imagine).
As an ML engineer you might need to do some data engineering work as part of your job but not the other way around.
This very much depend on the company. From experience DE is used as a catch-all title.
I haven't heard about this before and now I'm curious- can you elaborate on the differences?
Obviously it's not a perfectly clean separation but it's a trend, and people sometimes end up really talking past each other. You can see on r/datascience which is very US-heavy how people often recommend to beginners not to bother with advanced ML, stick to SQL, basic Python and analytics, and in the UK data science job market that's outright bad advice (it's fine advice for the UK analytics market which is a separate thing).
1. The person who developed the notebook is responsible for productionizing it. (No, it's not all crappy notebooks and some data scientists can indeed write high quality code).
2. You have someone like an ML engineer whose job it is to do this.
What you're describing seems like the least likely option; at least on the teams I've worked on "I can write tensorflow" would get you nowhere if that's not already a part of your job description.
Anyway, I'm going to go back to my 5K+ lines of code for an upcoming conference submission - almost all of which involve data cleaning and aggregation - and think about how I could be making a 2x more than I am now.
Thanks Hacker News.
Data Engineers are the people who take raw data (e.g. what lands in S3) and put that into data systems that can be used by other systems (e.g. Dashboards) and people (e.g. Analyst, Data Scientists, BI people). Data Engineers clean data, but they are really looking at cleaning out systemic issues (e.g. some data that is missing in one field is in another field, and that needs to be consolidated) and not the scrutinized row-by-row cleaning that Data Scientists end up doing. Data Engineers also do the data steps (e.g. creating a performant stored query) required to support things like business KPIs and reporting.
ML Engineering has a lot more variety based on the company and org, but generally it's about building an automated pipeline that includes ML. In smaller orgs you do everything - build a data pipeline, train a model, deploy that model, score new data, etc. In larger orgs, ML Engineers take a model built by somebody else and make it run at scale while meeting certain SLAs (e.g. making recommendations on a social media website).
I love a lot of it, but there's still plenty of bullshit to deal with. Just in the technical side, dealing with Python is a perpetual gong show, and most of my team's work seems to revolve around configuration of secrets and K8s.
I'm fortunate to be the guy that nerds out about performant code, so when something inevitably turns out to be a perf bottleneck, I can turn back into a regular old software engineer who trades in big data. Which I think is a better title/charge than data engineer, anyway.
I've talked with plenty of ML engineers, and they seem to immensely enjoy what they do. It seems that the periphery of data engineering is great; the core of it, not so much.
There are a lot of naked emperors walking around with lots of folks standing as close as they can to shield them from the cool winter wind.
I could not agree more with the overall sentiment that data science is overblown in terms of reproducible results because the people doing it just don't have an actual process or good leadership focus (which is not just a startup problem...).
So much so that I stayed stauncihily on the "data engineering" track because it was much more concrete in terms of technology, performance drivers, and business outcomes than the folk who sold pipedreams of magical AI models that would provide amazing analytics overnight.
Turns out that if you can't get at, scrub and actually _use_ the data, figuring out trends or training models doesn't happen, so I focused on making at least that 50% of the project happen and leave nice, tidy infrastructure, workflows and schemas for the data science folk to go through.
I also had the good fortune to work with some very organized, knowledgeable ML folk who actually understood how things worked, but some partners and customers had... incredibly disorganized "data scientists" that would leave stuff scattered all over the place (including private copies of datasets on their laptops when we had nice, secure remote sandboxes for them that even did data masking to avoid leaking sensitive data).
Personally, I blame a lot of this on lack of certifications or professional training that emphasises _process_. Otherwise it's exactly the same problem we've had for the past 20 years in BI departments: People doing their own Excel sheets because "SQL is hard" and nobody can do ETL properly.
(Full disclosure: I am an MS FTE, spent something like 10 years doing analytics almost full time, and have presented on how to do Data Science at scale a few times: https://carmo.io/talks)
He lost me here. Something I've always loved about being an engineer (and now in product) is that something small we do/tweak can have big impact.
If you tuned a parameter and that actually had tangible impact on the business, that's like the best case scenario and should be celebrated (vs doing some cool rocket science stuff that ends up unused and doesn't matter)
If you want to really follow the same compensation structure, we would then give engineers a really low base salary and make 80% of their compensation performance dependent.
Be careful what you wish for :)
Besides - this would drive some strange incentive structures. If you incentivise people based on cloud savings for instance, it will really only be the teams with unnecessarily large cloud spend in the first place that ‘get’ that bonus. If you incentivise on sales, engineers doing great work on back office tools don’t get any cake. Etc.
Maybe you are on $150k per year today, but in three months time you are back to $45k per year because you didn’t make some minimum sales threshold. Might be fine for some people, but depending on your mortgage…
I very much doubt my "going the extra mile" will really affect anyone at all in any major way. It may make some made up numbers go up -- or down -- but realistically it will have no major effect on anyone at all, except myself (and negatively).
Whatever effect it elicits in another will be short-lived, and forgotten next quarter -- least of all recompensed sufficiently for the sacrifices made.
2. Fact: You need to deal with bigger problems in ML like data set class imbalance, calibrating responses to the right scalar range(figuring out what that range even is in terms of domain). This is not taught in schools, just like writing Software is not taught in schools. One needs to be in the field to learn these skills and an ML Engineer can pick these up as much as a DS.
Various older information intensive fields (medicine, insurance, finance etc) knew the benefits and pitfals long ago. These examples show also the survival strategy for the generic "data scientist": specialization. The role of the human in the loop is to blow some context and relevance into an otherwise dead body of data. You can only do that if you really know your domain.
I was a data scientist who moved to engineering.
In a lot of orgs data science is there to help decision making, but human nature often makes it such that the decisions are already made.
I can relate to this so much. I've worked in multiple projects with great people who shifted away from solving the problem at hand, to instead construct some sort of generic problem-solving platform. In one project this actually happened twice: after refactoring the beef out of our SpecificProblemSolvingService into GenericProblemSolvingService, the generic problem-solving platform was then rewritten with one extra level of abstraction, so it could run any models designed to solve any task. As far as I know, neither service was ever used to solve any other problem except the SpecificProblem that we were solving the first place.
But it's fun writing platforms, I guess?
Personally I went for ML Engineering. My company at some point hired people as data scientists (some of my more senior colleagues still have the title, despite doing the same work I do), but started hiring people as ML engineers, i.e. people who can do half-decent SW engineering and also do ML. Just a filtering thing I guess.
I have a suspicion the term will start to fall out of fashion as things become more specialised.
I've run a "data science consultancy" in some form or fashion for three years now.
When people say "data science" they mean one of three things:
(1) MLE
(2) Data Management
(3) Data Analysis or Business Intelligence (applications of the same skillsets).
(1) has a lot of ongoing innovation, be it in MLOps, autoML, mapping frontier ML to business cases, etc. Innovation is expensive if the investment strategy is unprincipled. (2) is a critical and essential part of making data a usable asset. Management is expensive if it exists solely as a control process and gatekeeps access and use. (3) is core and will never get away from the adhocs and the standard flows, but the inferences are often dubious or not logically justifiable and requires depth of statistical knowledge (rare) to do well -- and courage to call out BS.
Very few people have the depth to do all three. What I have found is that many businesses hope for capacity in all three, plus some basic SWE, in the hope that they can decrease labor expenses. Not an irrational hope, to be frank, but ultimate the iron law of business holds: you can have it good, fast, or cheap -- pick two and be happy with one.
My core observation (and one I see validated based on client interest and experience) is that this is not new and has happened before -- it is the hype cycle in action. The digitization process (including moving to digital and then moving to Web) had a similar cycle. When you treat "data science" like its a silver bullet it will generally fail to do anything but suck budget. When you embed it with your technology teams and treat it as an iterative add, as useful as devops, etc., you have a better chance for value add.
I've found three core customer sets that helped us define a sustainable business:
(1) government agencies (which tend to put most expenditure under labor categories, so they hire a lot of long-term consultants and contractors)
(2) mid-to-small sized non-technology firms that want better data science strategy or want to build data-driven features into applications/products (especially in novel ways)
(3) smaller technology companies that don't have the MLE and data management system capabilities.
My career has been in heavily regulated industries, so our customers often have an appreciation for the management and governance portion after experiencing negative data science outcomes from maverick types.
The serious places don’t want you… so you end up at the place that can’t tell the difference, and the self fulfilling prophecy begins.
For me therefore, data science is the epitome of Graber's 'bullshit job' -- if the position didn't exist, the company would go on just the same.
There is a null hypothesis here, that the average person in role x is just average at that role. It is an extraordinary person who has high level skills across multiple domains like maths/science and coding, maybe so extraordinary that they wouldn't be working with you...
I'll admit that the article rings true, but I think there is an implied intentionality that I don't agree with. We are all just plodding along, doing our best with limited information and skills.
Never got the feeling the author blamed the data scientists (he was one himself), but rather management, and not their bad intentions but their incompetence.
> Shitty code & shitty data science
In my opinion the bar should be higher for code quality but also general engineering know-how in data science. You'd be surprised how many are uncomfortable with git, using the command line, interacting with APIs, managing environments, etc. Being able to only work within a jupyter notebook is not good enough, at all. Otherwise, you end up with people who's entire job it is to productionize and deploy the code which is a waste of time and effort.
> Poor mentorship
There is either a lack of quality leadership and mentorship or an inability for upper management to see the value in hiring for it. What ends up happening is you have a team who doesn't know how to grow, scale, or work together. They instead focus on building models and learning statistics when they should be focusing on building systems and process for helping the business scale analytics and building models when appropriate.
I enjoyed data science but found it to also not matter in the implementation that everyone thinks it should be. Data science isn't building nothing but ML models. In most companies, in my opinion, it is actually about scaling analytics. Being able to reach further up into data engineering, get raw data, explore it, shape it, give it back to DE to automate, and then automate the delivery of data to upper management and guide them through using it. If the team thinks their job is to just build models, everyone is going to have a miserable, miserable time.
I agree that many companies hire data scientists with only a vague idea about how to utilize them, but the same is true of software people in general. "Software is eating the world" and so is the practice of extracting value from data.
The margin on software is high - often more than 95% - so there's a lot of room to screw up and "figure it out" as a business. I think that's why there's a low bar for software and data management compared to, say, an automotive manufacturing line manager.
But that's where the opportunity is, if you're a budding data scientist:
- The business might not know how to effectively use/manage/train/mentor you.
- Upper management might have 20+ years of line of business experience, but will need your help to understand how your team can impact the business.
- You're going to need to seek out ways to impact the bottom line of the business.
All of the above is a recipe for leaping forward in your career. Since data science is a relatively new field, the demand for senior leadership FAR outstrips the available supply.
If you can learn how to effectively manage yourself, your team mates, and your function within the business - you have a ton of negotiating leverage and can name your price.
Source: I'm a data person who "retired" in their early 30s. Now I do all the research and hard science I want. ;)
This is so very, very true. Most of the "bad" data science orgs I've spent time with, are bad because leadership is either bluffers or a data engineer/BI type person. It's generally hard for these types to run effective DS orgs as the skills needed are very, very different.
I know some people cringe (mostly infra) when they think of data scientists having direct access to databases and infrastructure but honestly you should have a level of understanding and responsibility to get there.
The data scientists that do data engineering are usually much more valuable to the company and definitely earn more.
It turns out that data work is limited by all the same things every other part of the business is limited by: the need to make quick decisions, institutional imperative, the beliefs of decision makers, the ability to communicate well/influence, and so on.
Having better access or skill with data doesn't give you a pass on these things, despite the suggestions otherwise from laments such as this.
Also agree about the simple tools but it's really hard from a career perspective. If I deploy XGBoost in production and put it on my resume, I'm making double my salary next year. If I can find a simple ruleset or linear regression that performs 90%+ as well as the XGBoost and put it in production then nobody cares even though it feels like distilling the complex down to the simple is really where the value is.
Also, modern tooling makes a lot of these models more than explainable enough for a lot of cases… 10% is a lot
I feel some good field training in statistics(Look up Andrew Gelman) a couple of good courses on Linear, Bayesian Regression is all you need, rest is just engineering skill.
The dichotomy between ML Engg and Datascience is as stupid as was between Systems Engg and Application Engg before Devops came along.
But of course, too many view DS as some abstract skill where domain knowledge is not needed, and where the methodology will solve all problems / provide insight.
I used to think an MLE was a solid engineer who also had a strong quantitative and numerical computing background. The kind of engineer that always has a copy of Numerical Recipes handy, and if needed, could reimplement core components of statsmodels and sklearn in javascript.
I think after this current contraction in tech is over we'll see that most of the remaining "data scientists/MLEs" will be the type of engineer I imagine an MLE to be.
Application of Computational Stats/ML Models are not all that hard to aquire but essential. I think we need a fundamental rethink of how applied stats/ML is taught to engineers to make them effective. Here are a few things I can think of:
1. Getting a solid understanding of actually coming up with a simple enough model to do the job 2. Do Power Analysis to figure out how many samples we need. Creating datasets with Hard Negatives and overcoming sampling bias. 3. Using things like Multiple Regression to do EDA. i.e. using models as a tool vs the end goal to understand a problem space.
Could most CS folks actually implement Linux or Chromium from scratch?
That said, while Linux and Chromium are massive projects each with years of development with thousands of engineers behind them, so of course it would be ridiculous to expect a single engineer to build such a thing. I also wouldn't expect an MLE to build SKLearn entirely as is from scratch on their own.
However, I do certainly hope most CS folks could implement an OS or Web browser from scratch.
I am way, way less optimistic than you then.
I doubt even 10% of CS grads, let alone people who have been out of school for a few years, could tell you what a page table is.
The same is true of automation incidentally, there's lots of big companies doing a lot of easily automatable work, but the guy who manages all those people doing easily automatable-work is hardly going to be scrap his own area by calling in some SWEs.
On a completely unrelated note, the author's data eng sounds like nothing I've seen. Hell veto power over code? I'm not even sure all of them can code. They're just glorified sys-admins who now can provision some cloud infra. Somehow data engs are probably even more incompetent on average than data scientists, and the reason you move upstream is because it's an easier job and you don't need to spend your weekends grinding through maths-problems or learning new languages while probably still having higher value-add.
I've been a data scientist for quite awhile now at many different places, but every time I start interviewing again I always make sure to include a few pure software engineer roles in the positions I'm interviewing for. Even for some pretty elite teams, I'm still able to get to the final rounds but so far have always realized I still personally prefer the DS roles I'm looking at.
Any data scientist who wants to keep working on quantitative problems in the future should aim to be a solid software engineer.
so I jumped ship and became a software engineer. better pay and more interesting problems.
nowadays I do -- well, fudging slightly but you could describe it as "industrial automation control". writing libraries to provide convenient abstractions for controlling industrial equipment, writing robust scripts to drive that equipment, run physical tests on the $widgets we make, aggregate the experimental data, store it, etc. in the interviews they liked how (in my DS job) I had taken existing inefficient excel based workflows that had human-in-the-loop, and automated them, made unit tests, wrote docs, considered failure modes that nobody had considered before, things like that. and I just read about a fuckton of different stuff. for example in the interviews they wanted to know if I had worked with concurrency, I said I hadn't because it just didn't come up in the work I did. but I knew a little about it because I read voraciously, then I was able to answer all the theoretical questions they posed about locks and threads and async and so on. obviously that didn't mean I really knew about concurrency (that's a kind of deep metis that can only be acquired by practical experience and I'm still only scratching the surface of it), but it demonstrated that I had curiosity to learn about the field outside of the immediate things I worked on day to day.
during that job hunt I also had a strong offer from a company that wrote software for the visual effects industry and they wanted someone to improve their automated testing and continuous deployment frameworks. I didn't know much about CI but I knew about testing (pytest and hypothesis and things like that). they liked me talking about that kind of thing.
I guess the lesson is, if you are right now a data scientist and you want to be a software engineer, you can just decide to be that right now. be proactive and find a software problem to solve, and solve it. you don't have to ask permission to do this .. what are they going to do, tell you to stop being useful? note what you did, then figure out how to do the next thing better based on what you learned. your pay stub will say you're a data scientist, but you should just think of it as clandestine self-directed on-the-job training for your next job, so you can talk about it in the interviews. does that make sense?
This seems to be a problem with the industry as a whole. I'm speaking as a SWE, but I've observed similar things with PMs. I don't think it's impossible or even very hard to appreciate the right things, it just requires a bit of thought and the correct value system. Both of those seem to be a bit too far of a reach though.
The biggest, hardest bridge to cross was an appreciation of the importance of metrics. For some reason, getting a business person to grok something as simple as precision/recall/F-beta is a near-impossible task. You can do multiple presentations on it (after having honed those presentations over years with multiple manager audiences), and it never sticks. It's always "what's the accuracy?" It's impossible to do good work when your bosses insist on measuring and therefore optimizing for the wrong thing (which in my experience consulting for multiple Fortune 500 businesses, they always do).
Even worse, many organizations have such broken politics/cultures that the managers can't even tell you the big picture of what the project is trying to accomplish. Once you finally piece it together from the people who know their roles in-depth, it becomes clear that what they're trying to do is totally infeasible. At least that was my experience more than half the time.
The biggest point I’d emphasize: “there is a general industry-wide need for people who are good at both data science and coding to oversee firms’ data science practices in a technical capacity.”
Even more, most of the data tech leaders I hear strongly suggest this is not possible without C-suite representation of a data/engineer expert (I’m talking about non-tech companies)
Finally, it’s actually not uncommon to see people shift from data science to data engineering due to similar motivations. It’s actually the technical leadership that is sometimes surprised at the shift. You hear about this in podcasts from DS/DE people.
50k+ lines of R, 10k+ lines of Julia, 5k+ in Python, C, and who knows what else. Most of it for what is, essentially, data engineering work.
Where do researchers with social science degrees fall on this scale? Less money, less clout. The projects are certainly interesting though (which is why I do what I do).
Everyone in IT who likes to do some math and statistics at his workplace, even if it simply a linear regression or some histograms should go for data science. Also instead of nonsense discussions about what is agile and what not, I enjoy talking with my colleagues about the newest papers in ML, even if nobody understands the details.
Yes, there are a few very meaningful dashboards that are high value to the business, and then there is analysis meant to justify a project.
After core dashboards have been built, a lot of data analysis is a political weapon and the data science people are designers of those weapons.
Why bother to understand how either software development has solved a problem, or how maths+stats has solved a problem when you could just ignore operational practices and “train a neural network to do it”?
A lot of good engineers stand by the rule "use boring technology". Data scientists should adopt the rule "use boring math".
- demand forecasting for supply chain processes - credit scoring for loan approval - portfolio risk analysis in financial settings - some misc optimizations work on operations - maybe even six sigma can be listed here
Those have always had a data-science-like feel to it. The problem I see is when companies try to implement a data science team is:
1. push out subject expert knowledge and requirements just for the "freedom"; 2. have stake holders to be out of touch with the solutions; 3. too litle focus on putting stuff "in production", tracking, beeing able to experiment, in whatever sense those have for the company; 4. too much focus on numbers that can come out of the ds's computer, and treating operation related numbers as an after thought; 5. no basic knowledge of simple/common/classic solutions for their problem at hand.
So yeah, making business impact is way harder than is sounds, and too out of the skill set for the 23yo STEM graduate to actually make impact. And too buzzwordy and impressive for the typical decision maker. I mean, I've heard countless times things like: "if an AI knows how to tell a cat from a dog by using one of those neural nets logarithms, surely it can know how to partition my marketing budget, optimize my coupon giving logic and determine the strategy we should have in order to achieve a very obscurely constructed OKR". (yes, I have worked for people that used the word logarithm to mean algorithm).
In my case, I was in organisations that wanted data science, but had no capability or interest in supporting the role, so a lot of my time and effort has been having to put down the data science tools, and learn devops, software development and data engineering so that I can get back to the point where I do my data science work.
I’ve also become frustrated with my data science peers lack of knowledge about the surrounding fields- I get that being a top tier software dev isn’t the primary responsibility of a DS, but it would certainly make their life, and the life of everyone around them a lot easier if they did make an effort. There’s a sense of “learned helplessness” in parts of data science (and parts of data engineering too) in which “if some third-party tool can’t do it for us, we just can’t do it” and imagination is limited to the features the latest framework de jour offers.
obviously there are some examples like computer vision that require ML.
Made me LOL. First time I did that while reading a tech post/blog in years. Also neatly describes half the HN audience fawning over the latest AI thing.
1) To recognize bad management, there needs to be awareness of what good management is. What is the common source of that awareness? How do one know good manager from bad one?
2) If good management is so crucial for tech company success, why there is no worldwide trend to help engineering managers become better? There's an awful lot of courses, bootcamps, learning videos, tutorials, git repos and so on for those who wants to be software engineers, but all I see for managers is self-help style books and articles, centered around typical situation "oh, shit, they appointed you to a managerial position, how do you cope with that?"
If a 23 year old manages to get a data science job at a startup, and then actually delivers the results that the start-up expected of them, the learning experience there is infinitely more valuable than going to Google and using a bunch of tools that don't exist in the real world, on unrealistic timelines because you're not on the ads team and don't need to make money.
You can go learn "best practices" later, but working in a startup is an exercise in pragmatism. You deliver results, or you die.
After a few abortive attempts to learn statistics and linear algebra in isolation, I decided to sign up to that famous Andrew Ng online course.
I was bored out of my mind. I just did not find it at all interesting.
Not sure what my point is - maybe that a really easy way to learn if a seemingly lucrative and interesting thing is for you or not is to go ahead and learn the very basics of it directly, and see if you are at all motivated. I was not. And after that it was out of my mind.
Data Science, Political Science, Social Science, Scientology
Compare to Physics, Biology, Math.
Not everyone is doing regression and classification all day.
The above were studied in the Computer Science curriculum at my university. I work with a statistician, like someone with a degree in mathematical statistics. They know nothing about any of those.
Yeah, some of them are doing unsupervised learning (recommender systems) too!
I dunno, I personally think we'd all have been better off if we'd called it statistics as at least then people would realise that the field wasn't created yesterday.
However, the data that data scientists want to use is often messy and comes from varied sources. Hence, data engineers do supporting infra work like cleaning/loading data from different databases, etc.
"Data engineering" means building systems that can manipulate data (e.g. storing, retrieving, and delivering it). There are usually fairly well-defined functional requirements about what the system is supposed to do, plus goals about performance and reliability that might be slightly more nebulous.
"Data science" means building systems that can draw conclusions from data. The functional requirement is usually some form of "accuracy", as measured somehow against some kind of human evaluation of the same conclusion.
Concretely: a data engineer might be asked to build a system that can ingest every tweet posted to Twitter, and return the 10 most widely-used hashtags in the last hour. A data scientist might be asked to build a system that looks at a tweet and figures out what language it's written in, or whether it's spam, or whether an attached image is pornographic.
> The median data scientist is horrible at coding and engineering in general. The few who are remotely decent at coding are often not good at engineering in the sense that they tend to over-engineer solutions, have a sense of self-grandeur, and want to waste time building their own platform stuff (folks, do not do this).
> It was obvious that there is a general industry-wide need for people who are good at both data science and coding to oversee firms’ data science practices in a technical capacity.
The job of overseeing a crowd of stubborn self-important over-engineerers sounds pretty thankless.
> 23 year-old data scientists should probably not work in start-ups, frankly; they should be working at companies that have actual capacity to on-board and delegate work to data folks fresh out of college. So many careers are being ruined before they’ve even started because data science kids went straight from undergrad to being the third data science hire at a series C company where the first two hires either provide no mentorship, or provide shitty mentorship because they too started their careers in the same way.
Startups are a low-paid job with a lottery ticket for a little dash of excitement. You get what you pay for.
> ...I live in constant anxiety that someone will pop quiz me with questions like “what is the formula for an F-statistic,” and that by failing to get it right I will vanish in a puff of smoke. So my brain tells me that I must always refresh myself on the basics.
Focusing on the basics is better than pretending to understand fancy things, but even this level of "continuous professional training" or whatever you want to call it is, to me, a bit off the mark. We can look up formulas whenever we want these days. We need more meaningful ways to test our understanding of things.
Ageism is disgusting and I cannot believe such blatant discriminatory language is seen as OK for a link posted to hackernews. How would you all say if he wrote that 40+ year old programmers should xx?
To me the proper context of the story is "cheap money have been flooding the market for _decades_".
Anecdotally, this is something I've thought about many times, that an awful lot of books cover like 80% of the important stuff in the first say 4-5 chapters, and then are just full of filler.
I find this extremely frustrating as I'm at the same time fearing to miss out on some important insight in the latter chapters, while reading the full length of most books is simply super hard to do in any sensible quantity.
Otherwise, great post. Going through a similar transition, for some of the same reasons :)
This is something that bothers me about the industry in general, but thankfully seems rare inside tech giants.
My perception is that VC companies in pursuit of the next unicorn idea believe it can only come from a young mind, and so pile money and opportunities onto young people who aren't ready for that responsibility. Either they crash and burn, or they are carried over the finish line by multiple investment rounds then swan off, thinking they did it all themselves, to go through the experience again as an angel investor.
there are different flavors of DS, there are people who are doing diffusion models, doing top stuff that may or may not yield anything, but they are doing it because they know math well and enough code to put new maths stuff into new products. deep knowledge. so called ML engineers. maybe they are even good at coding at the lowest level, but in their point of view, why? these people are at huge companies that have vast resources downstream, they make people like OP work their work actually..
there is T shaped folks (unicorns, everybody wants one even if politically not ready [most arent]), where they know some concepts of many topics, perhaps so called full stack DS, which i consider myself to be... and i wouldn't be able to read thru most scientific papers, but I can put stuff together from start to finish including deploying it as an API that's scalable to top performance because of cloud. i do go back to basics often and I think its only natural! i think its like being a pilot, why not check the basics that actually, if forgotten, will take everything down lol... and you will use that the most as well!
i think also many people who are too much into one thing, math, code, whatever it is, start to call non basic things that are basic to them, well --- basic... BUT THERE IS NOTHING BASIC about multi linear reg and how to set it up all proper and how humanity spent thousands of years getting to this point..
there is also DS thats like data analyst on steroids, knowing middle basic and middle tier algos and stats well and can deliver mad value with a bit of business knowledge. hell, they could even use excel for their stuff, but proper understanding of the question at hand will most likely allow you to downgrade to lower, simpler tools. and simple is awesome! people often misinterpret complicated for advanced, not the case whatsoever.
once you know the land you accept your weak points and strong points and points you need to know enough to put stuff together. at the end if you know how to make sure stuff works and it works, hey, it works. and the only thing at that point between messing around and science, is "writing it down"... ;) push that code up , make it reproducible end to end.
Still, I got hired as Data Scientist. The money was good, but I didn't believe in the field and was kinda ashamed of the title (though it did seem to impress people - I think rural Europe is a few year behinds).
After 3 years I left for SWE SRE and my only regret is not doing that earlier. Decent grasp of SQL and a bit of handwavy knowledge about stats comes in handier than expected.
No time for that, buddy. CXOs need results, ASAP! (This was basically the attitude of my managers at the last place I worked).
>The work was often very low value-add to the business (often compensating for incompetence up the management chain).
The chain is what is important. This is not a data science specific issue.
Swap in or rearrange any team names to the original list and the last position or two will find this article true.
Has been my experience as ML engineer too. Decision making being intuition- and not data-driven was one of the largest shocks to me when I went from academia into industry.
How upper management and the board determine the course of the company was based more on emotion than anything else.
I just wanted to add that my personal transition from DS to DE has allowed me to work with a wider variety of data. DS is mostly tabular; data out in the wild comes in all shapes and sizes, and learning different techniques has been interesting.
If your DS employer isn't making real use of the capabilities of a skilled data-scientist and that makes you sad, consider looking for a company that will.
Right. What can you learn to over come crappy infrastructure?
This made me laugh :) I'm in academia and this post does remind me of certain students...
Yeah, dude, good luck with that. That is emphatically not how layoffs happen. :(
This is funny. Great article with lots of truth.
Shouldn't the author conclude that "Goodbye, shitty companies"? How were the two problems unique to data science?
I sometimes feel like that as a software dev, though, when management pushes some changes or "fixes" that aren't useful and won't fix anything.
As someone who has just moved from Data Science to Software Engineering I feel very much the same way, liberated.
I worked in DS for 5 years and had varying degrees of success in working at companies that understood the proper use and application of Data Science. What killed my passion for it was a few things:
1. Data Science is a dubious field - Data Science can certainly be applied correctly but I and others have used the underlying statistical methods gung-ho at times. Part of this comes down to something that W.D said. That Data Scientists are generally early on in their career. We have been captivated by the shiny new field and want to use it as quickly as possible without fully understanding it. Throughout my career I've been met with varying degrees of scepticism about my profession by people because Data Science offers more than it can give.
2. Data Science professional development is poorly understood/completely neglected - If you look for resources to grow in your Data Science skills online you are invariably drowned out by the sheer volume of crappy "Intro to Data Science" courses online. As far as I can find there is very little advanced Data Science professional development resources out there. Compounding the problem is that Data Science teams are invariably managed by people who aren't native to the field. This has the effect of the manager letting Data Scientists self direct their learning which will hit the problem mentioned previously.
3. Support for MLOps is non-existent - I think this problem will change over the next couple of years but Data Science has had to go through cycles of being integrated into a business. The first "wave" of Data Science was met with the realisation by companies that they couldn't get Data Scientists to magic money out of the poorly maintained data they kept. This has caused a huge increase in Data Engineers (not just Data Science has spawned this), now we have Data Scientists who have access to nice data (thanks Data Engineers!), they can build some interesting models but how do they get it deployed? This second "wave" is seeing the rise of MLOps tools, engineers, etc but Data Scientists currently don't have the know-how to get their own models in to production. This inability is incredibly demoralizing from my experience.
4. Educating fellow Data Scientists is too difficult - Unfortunately the perception that is given to people coming in to Data Science is that you can just do model engineering and call it a day. Bootcamps, courses, tutorials are all geared towards getting people good at building models, not about considering how those models fit into the bigger picture. There is little to no knowledge about good programming practices, source control (a lot of Data Scientists I worked with only knew git as a swear word) or deployment strategies. You could argue that a Data Scientist's should only be concerned with building models, I would agree but the reality is that companies will hire a team of Data Scientists but will likely not provide complementing teams to get models in to production. When trying to upskill others on my team it's been an uphill battle. Either people don't care as they just want to build models or they have come from an adjacent field with no software engineering experience.
Apologies for the stream of consciousness but it feels good to get it off my chest. My move to Software Engineering started in my last role where I was a Lead for a Data Science team. Thankfully my boss (head of Data) understood the need for developing a whole data system from good Data Engineering all the way through to MLOps for deployments. I was very fortunate to be able to move to being the Lead MLOps Engineer and develop our capability to deploy models with CI/CD mechanisms using AWS. That really gave me the taste for building systems rather than models. I really do think Data Science has a place and can provide great value but it's still a long way off. If we can make it so that Data Science teams can deploy to production quickly and safely we can really start to reap the rewards.
For Software Engineers looking at getting in to Data Science I would suggest looking at MLOps first. You get to combine existing experience with tackling new problems (how do we keep models live and continuously learning? how do we ensure the tracking of experiments?) and will have a tremendous impact.
This line is some amazing writing! Also extremely applicable to other areas of tech.
Mainstream data scientists don't do data science, they are either ML engineers or data analysts who use ML Python libraries and promoted to data scientists with bigger paycheck. With machine learning and AI being the trend, the job is sexy and it is a chance for the company to market itself as it uses AI and cutting-edge tech. I worked for a data solutions company and while I work together with data scientists to propose a design to our clients, it was the AI part narrated by the data scientists that makes clients ready to throw money. Even if the requirements seem impossible or at least very difficult to achieve. These projects often fail because the data scientists couldn't reach the accuracy goal written in the contract and the project ends up in the trash. These data scientists eventually leave the company or get fired, only to find another job within a short time with even bigger salaries.
Now for the data engineering part; I wish the OP all the best with his career, but he is still in the honeymoon period and didn't witness the misery of being a data engineer.
- The job can get very repetitive very quickly, unless he'll work on infrastructure, his tasks will be mainly focused on maintaining existing ETL pipelines or building ones, and both are labor tasks. You'll end up a data plumber who makes sure data goes from A to B and then C and that's it. You'll seldom find something new or revolutionary to work on and you'll keep using the same tools as long as you're in the same company. Even when moving to a new job, you'll pretty much be hired because of your knowledge of the same data warehouse or ETL tool you used in the previous job.
- Data engineering is an underappreciated job. If things go right, nobody pats you in the shoulder. When shit gets loose, you'll be the one to clean up the mess. What makes it worse is, you can't leave this mess long because data is a snowball effect; if you leave it unprocessed, you'll end up with more data clogging your pipelines and what can be fixed within an hour can quickly take long nights and even days to resolve, and do you know what this means? Managers won't have their fancy dashboards updated and they'll start panicking.
- Data engineering is not rewarding. Again, you're doing your job. No one cares.
- Data engineering has nothing to show for. Yes, data is being crunched and processed and baked. It's the data analyst who builds the fancy dashboards for managers, and the data scientists who create fancy graphs for managers, and the ML engineers who create fancy products for managers. You're just a plumber who, instead of fixing toilets and sinks, fixes data pipelines.
- Just recently, data engineering salaries are rising thanks to low supply and higher demand thanks to better awareness from CTOs and heads of data about the importance of the role, but until a few years ago, they were paid less than a software engineer.