However...
The coronavirus is a case where it's very easy to unintentionally misinterpret the data, and casual analysis will likely add more noise/false information which is not what is needed now. (and any analysis will become out-of-date very quickly).
Arrogant and condescending don’t mean “wrong”. Djikstra (the king of accusations of arrogance), once pointed out that his native language has no word for “egghead”: https://www.cs.utexas.edu/users/EWD/transcriptions/EWD12xx/E..., principally because his native culture has no history of dismissing the educated. OP is using (I interpret) “bootcamp data scientists” as a shorthand for people who know far less than they think they do. That seems like a fair characterization, honestly: that’s entirely the philosophy behind bootcamps in the first place.
The idea behind them is that a four-year education in computer science is unnecessary for routine software development jobs, and one can be trained up to useful programming ability in a much shorter time (and less cost) than a full-blown university education takes. While it’s difficult to say how true this is, OP (I interpret) and I have both observed that bootcamp types seem to presume that a few-month intensive bootcamp experience actually does cover everything that a four-year degree does, and are often surprised (as well as defensive) when they come across something that their bootcamp didn’t prepare them for.
A close analogue would be a paralegal thinking they know the law better than lawyers or nurses thinking they know medicine better than doctors - not that they don’t know a lot, but there are a few things that, while not as commonly useful, do come up in actual practice that are important when they do.
I can run circles around some programmers with a CS degree.
CS degrees have always been an odd qualification, mainly because so much of it is completely unrelated to professional programming. You can get a high mark and still be unable to write a non-trivial program on your own.
When I originally started doing my Chemistry degree, we did huge amounts of lab work compared to the odd class or two CS students did.
You don't come out of uni a good progammer if you study CS, while you will come out of uni a good chemist, biologist or engineer.
And that really sums up the problem with this line of argument.
Do you think that this is untrue of boot camps?
And you're not going to be able to pass it because you can pontificate at length about compiler design or explain what a linked list is, but flunked the practicals because none of your code even compiled.
One guy I worked with a long time ago that we hired would give us all these huge lectures about the right DB design and when to use structs or classes. Had an opinion on the right way to do everything. But he didn't seem to be clearing many bugs, and then our company fired him after 3 months when they realized in his main project to get our dropbox-esque system talking with Word he'd written a whole 10 lines of code.
Could talk the talk, but couldn't walk the walk.
They represent an infinitesimal minority of the population. Most people, even people with CS degrees lack that discipline. At least the CS degree forces them to learn some fundamentals so that they can potentially understand what they're looking at and learn on the job, and demonstrates at least a moderate interest in the subject as a vocation.
In my experience most people who graduate from bootcamps (in general, some are better than others) are all about getting a job as soon as possible. They're expecting a 6 figure salary within a year, and are willing/able to put in 70+ hour weeks for a few months to get there. Not knocking that perspective, but they'll lack a lot of background outside of their very specific niche and will likely be less adaptable, particularly on any highly technical subjects where some level of theory is required.
IMO the wave of "Software Engineering" degrees that are trickling out into the universities are probably the ideal. Full 4 year accredited degree that focuses on practical software development and less on the theory and mathematics. Let the CS-degrees focus on research/academics in the fashion of schools of Arts and Science and leave the Engineering to the Engineering schools.
I had on average about one lab class a semester as a Biochemist, and on average one programming class a semester as a CS. You couldn't pass the CS class without being able to write code that worked and got more complicated by the end of the semester. By no means did this mean I (or anyone else) was an awesome software engineer by virtue of doing the degree, but I'd put almost anyone with a decent grade from my undergrad program into an entry level position. Same deal with Chemists or Biochemists. I've done a bunch of hiring and can say often bootcamp grads are less able to come up with a sophisticated answer to a problem they haven't studied before.
I say this also as a grad of Insight, a DS bootcamp of sorts.
(also I say this fully appreciating that you were explaining the parent poster's view and you may not necessarily support it)
I say this as a graduate of Insight which is a kind of DS bootcamp.
What makes you think epidemiologists are so great? The one who has guided the UK's response has been severely criticised by fellow academics in the past for bogus low quality modelling. The Imperial College paper has basic errors and flawed assumptions that are obvious even to untrained laymen.
I think you're wildly over-estimating how much statistical and logical training most academics get. This is one of the basic underlying problems that the replication crisis has exposed: an unending flood of academic papers, especially those doing modelling, that fall apart when examined by people with statistical and mathematical training.
You're right that a 4 week bootcamp won't make someone a flawless handling of data. But it might still be a more rigorous form of training than epidemiologists get.
I think the most effective approach, which adds to your argument, is ensemble modelling from numerous, independent researchers from many fields. This is akin to how many people guessing at the number of gumballs in a container at the fair will average to the correct amount, while any single guess will not. There is a Freakonomics episode, Superpredictors, which demonstrates the use of this statistical approach for anticipating the outcome of voting in foreign politics.
Some people thrive in the bootcamp scene, while others enjoy the structure of a university. Both can be great tools for self-learning (I am thinking specifically of the resources a large university can provide, like media and electronics labs for creative projects that involve expensive cameras, 3D printers, etc), and both can be misused. University based education does not have to be ancient.
People that claim to be data scientists after attending a coding boot camp must be clowns. I’ve never come across such a person. All the data scientists I know are smart self taught people, not people who went to a boot camp.
I've met smart data scientists and engineers who've gone to a bootcamp. People who are good and bad at their jobs come from a huge variety of backgrounds, and it's not helpful to look down on folks whose background is different than your own.
As I am sure you know, data science is a specialized field, one which draws upon a combination of statistics and computer science practices. You may be able to get some of the computer science practices down in a boot camp, but there's no way you're getting a boot camp to fill in the Bachelor's in statistics you need to be a good data scientist. Years of hard work and sincere dedication, however, I am sure can get you there.
In academia, science is the pursuit of postulating a falsifiable theory, then gathering facts to verify that theory, and then going through a process of intensive peer review that either confirms, dispels or amends those findings.
The formality of the academic process is a necessity. The value of scientific findings entirely depends on the trustworthiness of the research. That is, how were the results obtained, which line of thinking was followed, did the research exclude crucial biases, etc.?
The importance here is that academic research is used as a pillar to produce products and services we all use in daily life. For instance, if you need a hip replacement, you want to be sure that prosthetic was designed based on rigorous scientific findings and studies that can vouch for safety and comfort.
The difference then with "data scientist" is that they often they don't apply the same rigorous research practices. It's easy to pick a dataset and bang off visualizations; it's a different story to actually come up with relevant questions, assert the quality of the data at hand and publish your findings towards a community of domain experts who are actually able to review your findings.
One needs in-depth domain knowledge to do that. Good data scientists will understand this limitation. They often work in a specific domain in a supporting capacity: bringing technical skills and capabilities to domain experts that don't have those skills.
Then there are those who purport to practice data science while grokking datasets, creating visualisations and cobbling a blogpost together at the end of the day. That's when you need to be really wary of what they publish, even if the bigger picture contains truthiness.
Hence why I have extremely mixed feelings about what https://medium.com/@tomaspueyo is doing.
To be sure, the core points of what he's telling are in line with what domain experts are telling us. But the extreme number juggling is quite mind bending. Moreover, the man is not a domain expert. He's an entrepreneur who happens to know how to write viral blogposts such as "how to deliver your funny speech" and "How to become the best in the world at something". What he does is anything but scientific. And so, even though he's making a heartfelt plea heard by many, one should be careful to not take the precise details in his pieces at face value.
At the moment, we all are victims of our own confirmation bias. Each day yields another data point, and given our desperate state, we want to see trends that confirm improvement, a probability that one will survive this, low mortality and so on. The reality is that we only have so few datapoints and it's still far too soon to make conclusive assertions about how this will pan out for the world at large and you in particular.
For all intents and purposes, there are authors who write similar massive viral pieces that may - and likely will - end up killing thousands of people inadvertently.
As I said, the core message is absolutely right, but his method - the way he packs his message with solid looking graphs and number juggling - is questionable. Sure, it gets the job done; whereas many others fail to push the very same message. But it's still a questionable tactic of convincing people.
Does it matter? Perhaps not. At the end of the day, he saved lives. Morality is a luxury presently. And yet, dismissing critical considerations outright equally opens the door to unintended consequences if we turn it into a habit each time someone is purported to have saved lives.
I'm not saying random data science blogs aren't often wrong. But you've probably been burned just as often by publish science and simply haven't realized it. And at least the data science blogs aren't behind expensive paywalls, aren't couched in meaningless vernacular, and present the code/data for reproducing their results, none of which can be said for a lot of science these days.
Open science and open access are important movements as to the relationship between publishers and academic research.
However, that debate doesn't negate my point as to citizen or data science.
For all the discussion about paywalls and the use of vernacular, the same timeless critical considerations need applying: who wrote the article? what is their background? are they a domain expert? where they did get their sources? how did they come to conclusions? is their method sound? are they asking relevant questions? etc. etc.
And I don't always see that happening online. On the contrary. The fluidity and the speed at which information flows online seems to be a justification to give in and accept what's being said at face value. The past few years should have made it clear that such indulgence can lead to dire consequences.
With powerful tools comes responsibility. Sure, it's great to see people apply free and open source tools to come to a new understanding of observations. But that's only half of the story: you still have to apply critical thinking to those conclusions. That's not something throwing more code or technical skills can achieve.
Having a proper, critical debate takes time and experience.
This seems to be an improvement over academia imo. The person writing the article should be immaterial. The same paper written by an unheard of researcher should be treated no differently than the same paper written by a well-known, tenured researcher at a prestigious institution. Unfortunately in academia, who you are and who you know is often as big a deal as what you do and what you know.
> What is their background? are they a domain expert?
Again, the methodology should stand alone. Being a "domain expert" or having a particular background is only relevant if the readers of the article are incapable of judging the results on their own merits, and have to instead rely appeals to authority. But appeals to authority introduce their own problems, including a moat that protects the status quo and the ingroup from potentially important new ideas and newcomers. And its hardly scientific.
> where they did get their sources? how did they come to conclusions? is their method sound? are they asking relevant questions? etc. etc.
These need not be restricted to academic work. If a blog fails to properly cite sources, you can move on. Plenty of academic papers are built on flimsy premises, poor methodology, asking incorrect questions, etc. etc. There is nothing differentiating the ability to assess the quality of something written in a blog vs something written in an academic journal here.
> And I don't always see that happening online. On the contrary. The fluidity and the speed at which information flows online seems to be a justification to give in and accept what's being said at face value. The past few years should have made it clear that such indulgence can lead to dire consequences.
And yet, as noted, the more conservative approach of academia has resulted in a huge amount of unreproducible "science". And now, in a time of crisis, the stringent model is being set aside in favor of openly available preprints shared online, without peer review, via social media. How can the traditional model be considered useful when, in times without urgency, the validity of its results are highly questionable, and in times of urgency, its thrown aside for the sake of actual progress? It seems to me the modern academic method is largely a facade, something closer to an ornate religious ritual that has been divorced of its actual intentions.
> But that's only half of the story: you still have to apply critical thinking to those conclusions. That's not something throwing more code or technical skills can achieve.
This assumes that those coding or utilizing technical skills are doing so without critical thinking. Perhaps this is true some of the time, but I don't see the situation being any different in academia.
> Having a proper, critical debate takes time and experience.
What constitutes a "proper, critical debate" is subjective. And from my position, it would appear that the advent of the internet and freedom of information is making the current academic model unsustainable and obsolete. And so academia has an entrenched interest in defining "proper, critical debate" in a way that protects their livelihood.
I'm not saying there isn't plenty of poorly written science and analysis being done on blogs out there. I'm simply saying I don't think traditional academia and science is much different.
Self motivation is much harder than you think, and a little nudge from a bootcamp may be what most people need before they become 'self motivated'.
There is a time and place for naive, exploratory, or just plain wrong-headed exposition. Making plausible looking but wrong "scientific" prognostications in the middle of a global emergency isn't one of them.
Some people can profit from this. Hence, the incentive.
I don't think the critique was of self-learning or bootcamps. It was a critique of a certain sort of cringe-worthy self-promotion.
And not just cringe-worthy. Spreading information based upon a cursory analysis of some CSV files without any training or experience in epidemiology, public health, public communications, etc. seems... irresponsible.
Also, off-topic, but most of my peers at university taught themselves how to program years before starting college -- typically in middle/early high school. And I mean actually taught themselves, from books and zine tutorials, not 'attended a formal course of instruction that was offered by a venture-backed firm instead of a formal course of instruction at a traditional university'. I guess times have changed, but the characterization of university students in CS as 'not self-taught programmers' is definitely the opposite of my experience. You came into the CS degree knowing how to program and learned how to do CS. And the characterization of "anything not university" as "self-taught" -- even formal courses of instruction that cost five figures -- is even more strange.
But SEOd medium posts have different reach potential than a facebook comment/phone call with friends/family. That's kind of the whole point of them.
Now, if a total non-expert had come out of nowhere with an analysis that contradicted the public health establishment, definitively showed that the danger was massively overblown, and was right, that would be one thing. But this medium post was incorrect in its conclusion, quite confident in its conclusion, and advised courses of action that would put many people in danger (such as reopening schools).
I'm also not sure that distinguishes the case I mentioned. It might also cause problems to assume some particular anecdote is true and representative.
Put more simply: a portfolio is a demonstration of your work. Anything you put out in public with your name on it is part of your portfolio. Obviously, don't put bad work in your public portfolio.
There's a reason you see very few serious data scientists publishing hot takes on their medium blogs -- it's a serious threat to your professional integrity and brand if you get it wrong. The only people I know & respect who are writing publicly about this topic have SME collaborators. (Of course, we're all playing with the data and talking about it in private with friends/coworkers over coffee.)
Self promotion specifically by spreading your amateur take on a health crisis seems a bit... immoral? Or at least the sort of thing that makes me question the candidate's judgement.
If you want to write publicly and in a formal way on this topic, get a subject matter expert to serve as a co-author. Anything less is more risky than it's worth, even from a purely selfish perspective.
Never mind that most of the people learning before uni did so from opportunity - they often had an engineer in the family or lived in a well-to-do area where such skills were apparent. (This magnified even further the older you are).
FWIW, I studied CS at a great cs uni (cal), and most of my peers were not self-taught. They still ended up at the same companies in the same positions as those who did ️.
Do you feel that your time there was wasted, and that you would have been better off attending a more focused bootcamp?