The Big Data Brain Drain: Why Science is in Trouble
jakevdp.github.io
jakevdp.github.io
But you'd only have to figure it out once and then learn to trust numpy, instead of rolling your own version every time.
So looping in a high-level language rather than using vectorized functions.
MATLAB syntax is ugly but the underlying principles are pretty cool. Well-written code scales automatically on newer hardware, or at least it has the potential to. That's not true in languages where higher-order vectors are built from discrete scalars.
The most vile aspect of Matlab is the faith every researcher has that producing something in Matlab is enough when the reality is code coming from Matlab will never escape, will never be as useful nakin-style pseudo for the creation of any larger system.
He's been doing it this way for years because that's what he was taught. That's the level of software engineering acumen you'll get in academia. But it "works". I've offered to help him modify the code so it will accept command line arguments, and we're going to sit down and do that so he can run several instances in parallel and utilize all of those fancypants cores on the computer I loaned him, but... he didn't know you could do that. No one told him! How would he know where to start looking that up? How reasonable is it to expect him to grok all that, when he's deep in math-land?
So it was blatant to me, software developer of four years, that something was pretty wrong, but for him: he's about to finish his PhD. He's been published a couple of times. They're not running horribly inept software development, they're running mathematics the best way they know how.
There are opportunities to build standalone tools which blow away their predecessors by multiple orders of magnitude, though; after getting enough researchers to use one such tool, you might attract sustained curiosity from a few people wondering "how the hell did s/he do that?!" and organically grow a small library with a real user base. That's one of my own long term goals, anyway.
I started in physics and there someone could make a great career corroborating for or disproving conceptual contributions. This is not a track in CS and is practically career suicide.
From experience most CS research can not be trusted to be correct, and enabling people to build a career on replicating or corroborating studies would in my opinion be of great value. Even the research that is correct is often not fully implemented so you not only have to implement their approach, you also have to discover how to realize it. That work is not publishable in CS, and it is a non-trivial amount of extremely risky work.
Psychology is probably one of the worst sciences for the attitude described in the article. Being in the most "mathy" corner of the field doesn't really help.
The example problem domain ("automated language translation") is actually a stellar counter-example to the claim. Has anyone actually tried to use Google Translate for anything sophisticated? It's still truly horrible, by human standards. The field needs more research and deeper conceptual understanding, not less.
There may be some problems that can be solved by throwing software/hardware/data at them, but I don't think this is a good paradigm for the big unsolved problems in general.
Have you compared Google Translate with the previous attempts to do automated translation based on conceptual understanding?
There is a reason why Google Translate would be claimed to be a success.
Being the best of a bad lot isn't enough.
Perhaps it is only a human reflex to believe that some contemplation is needed to solve problems that have resisted mounds of data being thrown on them. But being not-coincidentally human, I happen to find it plausible.
teams using the technique have won some competitions in kaggle , while doing little feature engineering(which is usually the part which you put domain knowledge) .
And it had shown better results than current systems for big problems like voice recognition and image recognition(the famous google cat experiment).
"I believe that science still has only two legs—theory and experimentation. The "four legs" viewpoint seems to imply the scientific method has changed in a fundamental way. I contend it is not the scientific method that has changed, but rather how it is being carried out."
Giving a few hundred nerds in an ivory tower access to all that hardware and data just makes the black box blacker.
A big part of the problem is that current researchers did have to manage their programming, IT, etc. during undergrad, grad, and post-doc periods. They did a lot of hacking with C, Perl, and Maple/MATLAB/etc.
As with many managers who were promoted because they were stellar engineers, those skills fade but they continue to think that they are qualified to judge the difficulty of the work. The fact that they have "done it"[1] before leads to a logically incorrect assumption that it's not that difficult.
After all, they learned those things on their own while thinking about the tough stuff.
[1] Except for, of course, in a predictable, repeatable, safe, auditable, etc. way with fifty other users asking for high priority changes.
I know many academics who spent a lot of time throughout their school career writing code to get their research done.
They looked at this as a necessary evil, and as such learned the bare amount minimum to get by. They are smart people and were able to make the code help them solve their research problem. But they are busy thinking about their research so usually their algorithms are fairly simple + straightforward (lots of nested loops and n^2 sort of things).
The main problem, in my experience, is that many of the research problems are actually fairly simple (algorithmically) and most research departments have access to fairly powerful computing facilities. Coupled, this means you can brute force a lot of solutions-- there is no real push for understanding of the algorithmic complexity.
As well, most academics are on a much MUCH longer timeline than your average business or startup. Did your algorithm break 1 month into processing your simulation? Fine, fix it and run it for another month. Or just take what you had and publish it anyways.
Just as academics look down upon the technical side of things, we are just as much to blame for idolizing the academics. Science (even in math and engineering) is a lot more 'sloppy' then we like to imagine. There are oodles of papers out there that are just downright incorrect-- and not on purpose!
(* My creds: I have participated in academia as both a student, researcher and software developer )
There's a huge resistance to using source control, so lots of time gets spent searching through deep folder structures and finding just which 'file1_v3 (4).doc' is the right one. Data gets lost due to simple mistakes.
They spent 20 minutes coding the Runge-Kutta algorithm every time they need to run a numerical simulation without realizing or caring that (a) they could spend 25 minutes to create a function that they could reuse or (b) that the function already exists.
In short, the issue isn't computing time, it's researcher time. But the idea of spending some time now to save time later is so foreign because of the focus on getting the publication out as quick as possible.
In industry, what I've seen is that often engineers are scrambling to please managers or customers, with work divided among multiple people, so the code is usually poorly written and undocumented.
In academics, publications are of primary importance, so everything is documented. The longer timescale means there's more time to refine code that's designed for a single, focused problem. The limited scope of the programs used means code quality isn't an issue most of the time.
Also, in theoretical computer science at least, the focus is entirely on rigorous proofs and finding optimal algorithms. While in industry, it's more "get practical things done quickly so we can sell it".
That's a pretty short sighted example of industry - I'm sure those examples are out there, but I don't think they're common (or the companies long lived).
Most places I've worked know that they'll have to maintain that code well into the future.
Not really so in academia (publish and forget) - which is why it's rare to see even basic measures taken for modularity and abstraction, e.g. the creation of types to represent entities in the problem domain. I think I've seen that done in Matlab, once.
The situation is simpler.
Academia has become abjectly miserable and abusive in its practices. It no longer offers good but low-paid jobs for smart non-conformists, it just offers it's special brand of misery based on some long-past promise of this.
Given this, only the mediocre stay (and compounding it, anyone who stays has no reason to be better than mediocre). And that is a huge, huge loss to the whole project of the development of human knowledge, something that has a long history in Western society.
Yes, there are shitty bits in academia, like the intense competition. But we should focus on where the problems are and how to fix them, not just write the whole enterprise off as shit.
* The incredibly workaholic
* The incredibly specialized into a well-funded field
* The incredibly skilled
* The incredibly good at the scientific method
* The incredibly savvy at the academic game
* The lucky bastards
* The even more incredible scientific polymaths
I'd say that any combination of two of the above factors will suffice to keep someone on the academic track for a while longer. Note how rare those factors actually are.
Hey, parent here. Love your "incredible" assumptiveness about my experiences. "Respect" to you too, baby, dude.
Actually I've never been an academic myself, not even slightly. OK, I have had friends and relatives that who've been failed as well as very successful academics so I have some idea of the culture but be that as it may.
Your post seems likely a way to showcase the versatility of the adjective "incredible", usefulness with new age cretinisms and your "implied success" there, well, except I don't see any.
See http://mg.co.za/article/2013-11-01-universities-head-for-ext...
Here's an example. The Globus Online grid ftp service web page intended for users adopts an overtly apologetic tone [1]. Users of this service are promised freedom from "low-value IT considerations and processes"--considerations and processes that the Globus Online team has humbly sought to undertake on their behalf. I have to laugh at the claim that there is "No need to involve your IT admin—all you need is Globus Online." The message is that information technology is of low academic value--unless you happen to have been one of the authors of publications that came out of the Globus Online project. If not, your career is sidelined.
Software development, system administration, network administration and desktop support have become somewhat specialized in the past 30 years, but in the minds of some principal investigators and academic administrators, these very different activities are conflated. An expert in numerical methods, computational fluid dynamics and dynamic downscaling methods for climate assessment models is a seasoned web developer with a portfolio, fluent in jQuery, underscore, backbone, responsive websites with bootstrap, CSS3, HTML5, PostgreSQL, PostGIS, the Google maps API, Cartodb visualizations, as well as an Android developer conversant with the SPen library for the Galaxy Note 10.1. It's as much effort to stay current technically as it is to keep up in the scientific literature.
There are faint signs of improvement. On January 14th, the NSF revised the biosketch format by changing the Publications section to Products [2]. "This change makes clear that products may include, but are not limited to, publications, data sets, software, patents, and copyrights." The previous biosketch format was awkward for software developers, inventors and producers of data sets.
Recently, a number of prominent computer scientists, and scientific software developers affiliated with the Climate Code Foundation [3], published a Science Code Manifesto [4]. The manifesto includes the recommendation that "software contributions must be included in systems of scientific assessment, credit, and recognition." Software developers in the digital humanities may wish to add their names to the list of signatories.
Whether these developments reflect a broader understanding that software developers ought to enjoy greater recognition and opportunity for advancement in academia that they do currently remains to be seen. Greater career advancement opportunity for software developers, inventors and data set producers working in academia might do something to address the Ph.D. overproduction problem.
But these developments were too little and too late for me. I left.
[1] https://www.globusonline.org/forusers/
http://wattsupwiththat.com/2013/07/27/another-uncertainty-fo...
Flipping a compiler flag gives completely different results! How can anyone trust this code, or any research made off of it regardless of their personal feelings about climate change?
this should have landed on the front page...
Anyways, I wonder if there is any explanation of this phenomena.It's the basic principle in real-world science that for stuff, people get credited who didn't invent/discover it. That's the reason why your submission wasn't noticed; you simply would have had to wait for the Nth time around.
Also consider it game theoretically -- that you're competing vs lots of articles... perhaps other articles at the time were more interesting/compelling.
As a staff programmer at a prestigious institution, I was making about $50k. I left for a large company and a few short years on I'm at around $200k.
Really?? Certainly, academia has its inefficiencies, but if there is an area of academia that has an undersupply of PhD's, please correct me!
That's where they undersupply is. I've seen some very low quality code from PhDs, even CS PhDs. One thing that is missing from a lot of academic curricula is software engineering techniques. Hell, I've had to fight in order to get people to use git.
I'm more familiar with medical/biology research, where you have people that understand the domain, but not necessarily how to properly code. And if they, as a PI, need to manage a programmer (tech) / computational researcher, it can be difficult.
And all that would merely constitute black box ignorance except that often if you try to help them make their work more reproducible or point out that even 32-bit floating point roundoff error can be initially indistinguishable from a bug that makes the results useless, a lot of them become agitated and tune out the possibility of any of this being relevant to their work.
I left academic science for industry over a decade ago because I caught a big shot who then proceeded to threaten me instead of fix what I found.
This means the software you write should be open sourced and the papers and articles you write should be freely available.
I'll use the basic idea of this article as more fodder for that argument (it's even more important now because...Big Data ZOMG) in the future :)
I don't fully agree with the thought that writing software and papers are exclusive. With a little creativity you can get a paper or two out of most academic software development projects (granted they won't usually be the A-level "this advances my discipline" kind of papers but one can't only write those anyways imo).
The part that is always unclear to me is: there is lot of evidence that research/university budgets are growing at a high rate. Where does all that money go to. It doesn't seem to go to the people who actually do science.
That seems to be the fundamental problem here.
I mean, I come from an excellent pedigree, too, top 10 undergraduate, top 10 grad school, and I currently work at a place where my boss is a nobel laureate. (He makes at least 250k, according to the place's 990 forms)
However, my going price gets diluted by the idiots running around with rubber-stamped PhDs. Maybe there needs to be a 'brain drain', if we can drain those people selectively somehow.
(The bottom line is that the money goes to administration costs as mentioned, leading institutions to expand to stay relevant, thus building more buildings and hiring more PIs with full expectation for them to be funded via grants. Meanwhile, the average researcher's earnings are mostly unchanged.)
[1]http://www.amazon.com/Economics-Shapes-Science-Paula-Stephan...
I have found myself thinking many times that a position in the industry, where I can use my teaching, data processing and analysis skills to further some business goal, seems like a much more preferable option than sitting around writing research papers and applying for grants all day. Not to mention that academia pays less and has worse overtime conditions than any industry job I could concievably get.
This article really nails the key issues for why I am feeling this reluctance towards an acedemic career.
Presumably you have some for one. Why are you still uncertain?
Trying to Cure Cancer: $25k/yr, bump to $40k after 6 years
Engineering Medical Devices, Airplanes, etc: $60k/yr
Trying to Build the Next Twitter: $100k/yr, $150k/yr after 6 years
Helping Rich People Game the System to Get Richer: $150k/yr, $300k/yr after a few years if successful
The desire to stay in academia comes from having different priorities than the market. Many/most people do. The people who don't usually happen to specialize in a field that the market is currently smiling upon. It's great if your dream can be gently tweaked to be compatible with market considerations, but please understand that for most people this is not the case. There are many Big Problems lurking on the horizon where intermediate progress can't be monetized. Academia lets you work on them. The market doesn't.
This sounds reasonable to me. Most people never have the experience of tripling or quadrupling their salary in a single career pivot, but that's what happens when you decide to bail on your sci/tech graduate research program and start working in the software industry. It takes a pretty crazy level of passion for your research domain to accept such insane opportunity costs. When I realized that by staying in science/academia long term I might never be able to afford to buy a house in a decent neighborhood or have kids it made it really easy for me to quit.
Would warmly recommend it. Visualization is a very large field, so you'll have to carve out some niche. I am in information visualization, which is the branch most generally applicable if you want to do data mining-related things. But there are many other variants: Visualization of scientific/simulation data, e.g. flow rendering, combustion processes, climate simulation, as well as medical fields: CT/MRI volume rendering, real-time 3D ultrasound and quite a few others. Central themes at the moment are using GPUs to implement more advanced 3D volume rendering techniques, or even using GPUs to draw data which is not 3D but where there are performance issues when using CPU alone. For instance, drawing dynamic (25FPS, interactive) scatterplots of large (>1 million records) datasets.
I guess the definition of a "large" dataset varies by context, in visualization you hit this limit earlier than in statistics and non-visual data mining if you use "discrete" methods where every item is drawn on screen.
Would it be possible for you to treat the vast amounts written on the internet regarding career choices as your "large dataset" and use your knowledge and tools to explore and analyze this dataset, and visualize it to us?
1.) Scientific research is hard. It is very frustrating to spend all day working on something you can't be sure is going to even work. On the other hand, programming is easy, it's mostly monkey work. Sure there are places where you have to be clever and think things out carefully, but at the end of the day, when you write code it just feels like you have so much more to show than when you do research.
2.) In research, you need to get grant money to do anything, because research is expensive. So much time is then spent writing grant proposals, and even when you get money you still can't do everything you want because it's just too expensive and/or time consuming. I'm not talking about LHC money here, but just the standard money for a professor running a lab in a university.
In programming, you can work on (I'd estimate) 95+% of problems with nothing more than your computer and some old fashioned hard work. There are so many good, free to use open source libraries out there, making it pretty easy to jump into whatever field is interesting to you. Best of all, when you finish a project, there is no need to spend weeks writing a paper about it or any of the nonsense associated with that, you can just publish your code.
If we make the leap and assume that this insight can be at least partially extended to fields beyond natural language processing, what we can expect is a situation in which domain knowledge is increasingly trumped by "mere" data-mining skills.
This is a great point. For many years, domain knowlege was merely such experience ex ante. In otherwords, the barrier to domain knowlege was access to data. In lieu of this, perhaps theory to estimate it. As data becomes larger, more free, and more amenable to analysis by a larger group of talent the "domain knowledge" itself as a barrier to entry (prestige, effectiveness) seemingly declines. Whle this is in theory good (more access, more analysis), there is probably a corralary that we should expect turf-wars and restriction to access, as those previously in positions of privledge fight to retain their status as "keepers of the keys".
Besides this, it's important to recognize the difference between identifying correlations and patterns in data vs understanding the mechanisms behind the phenomena. Krugman makes a strong point: "The problem is that there is no alternative to models. We all think in simplified models, all the time. The sophisticated thing to do is not to pretend to stop, but to be self-conscious -- to be aware that your models are maps rather than reality." [1]
Data-mining will help generate and support hypotheses, but this is complementary to model building.
If you think about the huge increase in computer power available today, it seems a field ripe for disruption. However, the guys that run the fed, and guys like Krugman are basically from the "slide rule era" of Econ.
I'm sympathetic to the power of modeling and techniques, but I do think its a case of the more you know, the more you understand what you don't know. And I think for most people the deeper you go into the field of Econ, the more this becomes apparent. Its getting better, though, and I'm sure in another generation (once the current tenures expire) things will look a good deal different (hopefully better).
Not sure if your views are similar, though.
[1] You see no little to no respect for dynamic systems that are chaotic or non linear; bounded rationality and its antecedent effcts; the role of institution underpinnings to markets, etc. Just to name a couple that are glaringly relevant and empirically important, but not subject to "hand math".
it is clear from the "slide-ruler-y math" jibe that you have no idea about the technical ability required to do research, especially in theoretical microeconomics and econometrics. you'll need more than mastery of "Mathematics for Engineering" and MATLAB to disrupt the economics profession.
Um, I don't mean to be offensive, but I have never heard that.
The drain towards industry is to solve industry's problems: make a profit. And yes, these can be extremely interesting problems, and hard ones too.
But what would our current world look like if the scientists from the last 500 years had bent their minds to solve the problems of merchants? (and don't get me wrong, some of the problems merchants had evolved great solutions for mankind)
I'm not convinced Big Science is in trouble though. Those who have the motivation and talent to stay in the academy will continue to do so. Yes, outstanding people will be lost but Science can progress from their contributions to commercial efforts. Geoff Hinton ends up at Google but we'll keep hearing from him the best use of his talents - the improvements to products we all use every day at a massive scale.
Society has reached a point where the academic route means practically begging for a job that won't even pay for a house, while the startup lottery offers, at least, a chance. (And finance, better yet, offers a high likelihood of being well-off.)