2,251 karma · joined September 11, 2011
It is often thought of as a prediction about what level of contribution is expected, but that's basically a farce. In every organization I've seen that makes distinctions between junior and senior developers, the only thing different between them is that senior developers are paid more. They don't do more, manage more, design with more foresight, or anything like that.
The exception is that developers who are brand new to a job function as junior engineers until they get calibrated and learn some parts of the codebase in which they can be effective. But this is just as true for a brand new senior dev with 10 years experience in your company's primary domain application as it is for a kid straight out of a bachelor's program whose only prior work experience has been theoretical REUs or something.
Basically, it's just a status tool.
In my experience, the problem is that interviewers have no idea how to correctly value a candidate's performance. Maybe the candidates are closer to being well-calibrated, but their self-assessments don't match up with the interviewers' because the interviewers don't know how to gauge what they are looking for?
Making the assumption that an interviewer knows how to measure the response of a candidate, even in cases of extremely quantitative questions with well-defined answers, is highly suspect to me. I think virtually no one knows how to do that effectively.
I experience the same problem with shortage-at-price-X in the field you describe. I'm a machine learning engineer with experience in MCMC methods, but I also have a lot of low-level Python and Cython experience, some intermediate experience with database internals, and lots of experience writing well-crafted code for production systems.
There are basically zero companies willing to pay what I'm seeking (which is a salary based on my previous job and a few offers I got around the time I took that job). In fact, in some of the more expensive cities, the real wage offered is far lower than other markets.
I've seen reputable, multi-billion dollar companies offering in the $140k range for this type of role in New York. That's wildly below anything reasonable for this sort of thing in New York. I've seen companies in Minneapolis offering $130k for the same kind of job -- and even that is still too low for Minneapolis! The same has been true in San Francisco as well.
Because these companies value you more for simply looking good on paper and looking good as a piece of office ornamentation when investors stroll through, and they view you as an arbitrary work receptacle closer to a software janitor than a statistical specialist, their whole mindset is about how to drive wage down.
Frankly, given the stresses of the job and the risk of burnout, I think it's actually a terrible time to be in the machine learning / computational stats employment field, despite all of the interesting new work and advances being made. The intellectual side is good, but the quality of jobs is through the floor.
I imagine they would say that your statement about crashing vs. e.g. launching the missiles is a false dilemma. You don't crash and you don't incorrectly launch the missiles.
I'm not a C++ developer so I can't say it with certainty. I more agree with what you're saying. I'm just relaying that my experience has been that out of many different language communities, C++ actually seems adamantly the opposite of what you're describing.
The software is junk software. There's no other secret thing going on -- no misdirection or duplicitous motives. A certain class of high-paying customers responds more to the Goldman brand name -- or at least believes it buys them cache with regulators or investors. For that class of customers, vetting the reliability and quality of the tech stack is at best an afterthought. Since that pile of money exists as a thing for Goldman to target, they do target it.
I advocate that more people should prioritize vetting the technology. If so, they would see it is not of sufficient quality to justify its use, let alone paying to perpetuate it elsewhere. But I'm not naive -- the political approach will always matter more to a wide range of people than will a more objective assessment.
This whole topic is not all that newsworthy. The team within Goldman that had architected and developed this years ago had spun out into a consulting group that essentially reimplemented the same thing in Bank of America (Quartz), JPMorgan (Athena) and many others, now including Morgan Stanley, and even trickling down to smaller banks like PNC.
I consider it one of the biggest ripoffs in modern finance that those organizations have paid untold fortunes to adopt the Goldman-like approach, sometimes even with new or additional proprietary languages brought in on the project. It also adds systemic risk for society because it further correlates these internal banking systems between the largest banks. If something goes systematically wrong with it in one place, there's a comparatively high risk the same sort of thing can or will go wrong in another too.
If we were bearing that risk for a good reason it might be OK. But really we're only bearing it because of the superficial branding of Goldman, and the pressure on banks to hand wave and appear to be doing something in the aftermath of the 2008 crisis. And so they go for what looks politically defensible (e.g. "well, this is what Goldman did and they survived the crash" -- despite it being widely researched and reported that Goldman's position in the crash truly had nothing at all to do with superior risk management systems and was a mixture of political favors and luck) instead of anything sensible from a system design point of view.
> Staying in the depression doesn't help you to take action at all.
It definitely can. Precisely when you lack the power to change your own extrinsic circumstances, and it is those circumstances causing the harm that leads to depression, then staying in the depression keeps sending out that distress signal for help. Taking you out of the depression without also addressing the circumstances makes it seem like you're managing it, when really you're still being harmed.
Sometimes depression is a correct, reasonable, and even necessary reaction to un-live-with-able life circumstances. It's a way of your body generating warning signals that others might see and then offer you help, not unlike shooting up a flare if you're stuck on a life raft.
If the circumstances truly are un-live-with-able, then medication that simply makes you superficially feel like the circumstance is live-with-able, when really it's just continuing to be destructive to your life in every single same way apart from your now-masked-with-medication feelings, then medication can be counter-productive.
Many people with depression don't actually have un-live-with-able circumstances. Instead, they have live-with-able circumstances and either they have medical issues preventing them from doing the actions necessary to manage that living, or else they have what I'll call psychological issues preventing them from it, where 'psychological' here is meant to mean the subset of mental health issues that are not well-treated by medication, but may be treated with counseling or changes to other life habits.
The problem I've always faced when seeking counseling or mental health help is that every single mental health professional I've ever interacted with, every single time across many years and highly varied geographical situations, has always, always, always dismissed completely the possibility that a person can actually have un-live-with-able circumstances within which the depressive reaction makes reasonable sense. Instead, they begin from a point of view fundamentally rooted in the belief that that cannot ever happen to a human.
I mean, if they were pressed to think about like a child soldier forcibly addicted to cocaine or something, maybe they'd agree people really can have circumstances such that depression is the correct reaction. But in general, with first world people, they just assume it's impossible and have a pervasive Occam's Razor sort of filter, before ever meeting you or even talking to you the first time, that, nope, you're wrong about how you view your own problems, that there's no way you could be thoughtful enough to have done meaningful introspection before seeing them, and that your circumstances are never a justifiable reason for feeling depressed.
And from this attitude, the next step is almost always to suggest medication right away. And any time I've tried to say something like, "well, I'll consider medication, but I'm not just going to jump right into it. I want to speak more about my circumstances and explain why I feel like it's a real life Catch-22 that truly, utterly is depressing, rather than the depression being a part of me as I respond to it," then it's like talking to a brick wall. The counselor / therapist / psychiatrist doesn't want to hear about about. They already know you need drugs within one 1-hour visit, and now that you're saying you won't just rush right out and get them, that means further that you are a problem because you won't just go get the drugs you need.
And what's crazy is that the underlying rationalizations from the different mental health professionals have been all over the place. One person thinks it's because I have issues about my childhood and my father. Another thinks it's because I grew up in a relatively more religious community. Someone else thinks I have PTSD from an abusive relationship in my late 20s. Yet another thinks that it's related to overwork and job stress.
All of these are important issues and some mixture of all of them is affecting me. But every counselor I see has their own colored opinion about which magic answer it is, all those magic answers are different from each other, and yet, every one of them thinks that drugs, drugs, drugs is what will magically solve the problem.
It really gives me no faith in the mental health treatment infrastructure, and causes me to be even more guarded about when or if I will consider trying anti-depressants.
From some of these statistics, it makes me feel like anti-depressants are way over-prescribed, and that counselors are just extremely lazy. They don't want to hear the whiny, tearful narratives of their hurting patients' lives -- just like friends and family also don't want to hear it. And so shoveling out some drugs is an easy way to focus on something different than the thing they find unpleasant (i.e. actually listening).
If you are low-income, you can get the fees waived, and the majority of her school books are still free. But not all.
Of course, developers don't have to contribute to an open source project at all. And if they do, it doesn't have to be in a "volunteering" sense in which their work actively offers benefit to others. They are free to choose instead to more selfishly only prioritize contributing in ways they personally want or like. But if they do, then they no longer can turn around and act like others should be gracious for their "generous" contribution to the project.
I guess I would say that developers can ignore the useful stream of +1s in order to more selfishly work on aspects that bring them personal satisfaction, instead of deriving satisfaction from the benefit offered to wide ranges of users. It's not that it's reasonable or unreasonable. It's just literally a particular option they could choose.
If they choose that, in cases where no serious rationale is given to explain why there is some more beneficial project agenda in the short term, then such petty disregard for a massive pile of evidence about what your project users needs are is a pretty strong indication of a bad library/project, with alarming dysfunction.
So in this sense, the +1 stream is a really useful barometer of the project.
Either the core developers say, "whoops, my bad, I had been intending to volunteer time which necessarily implies that my efforts should be focused on whatever the users benefit from the most -- let me switch to work on this hugely +1'd issue I'd been ignoring" ... or they say, "I hear you everybody, but since I'm more familiar with the project internals, let me offer a technical rationale for why we need to delay addressing this in order to work on other stuff first" ... or they say, "Screw you all for bothering me. I'm 'volunteering' my time on this, so I'm going to do what I want. I don't like this big +1'd issue, so I'm going to ignore it and just do whatever."
Either way it gives you a ton of information about the health and reliability of the project.
When the project maintainers fail to prioritize the things that huge streams of users are +1-ing, I think that's a strong signal that the project maintainers just want to poke around their preferred aspects of the project, rather than address actual needs.
In that sense, they aren't actually "volunteering" time -- they are receiving the compensation of their satisfaction of poking around only the parts of the code they like. They aren't doing it "for" anyone but themselves, and it's this idea that all open source contributions are "volunteered" and are given infinity free passes from criticisms about prioritization that causes a lot of the trouble.
In order for the time to be "volunteered" it has to be directed at things that others are benefiting from. If few are benefiting and it's just a personal quirk, curiosity, or interest of some random developer -- who wants to tune out the +1 noise of people asking for things they actually need -- then that developer deserves zero praise or kudos for "volunteering" time or anything. There's no sense in which it's generous to fail to support aspects of a project that people need in order to instead support aspects of a project you personally happen to like.
I think a lot of the +1 streams are about this kind of tension. Developers saying, "leave me alone ... I'm already doing this 'for free' -- what more do you want." And then cranky users saying in reply, "But you're not doing anything 'for free' -- you're picking the parts that bring you personal satisfaction and then trying to argue that those are higher priority than the things project users rely on and need."
In general, the +1 sort of issues that I've seen mostly arise because (a) it's an implementation challenge that is not at all suited to a new contributor and needs significant attention from devs who are already very familiar; (b) it's a dev task that the existing devs don't want to do, for various reasons; and (c) it's a dev task that a huge and vocal contingent of users feels is really critical and that it's almost a dealbreaker in terms of their usage of the open source tool in the first place.
I feel like a lot of the replies on this thread that focus on pushing it back to the person who asked or the people who +1 are missing the point. The far more common scenario is when it's entirely implausible that any of those people could do it without intense and significant handholding from devs already familiar enough to do it more quickly themselves.
In your comment there again is this mention that the dev time is volunteered, but this is a red herring. The time isn't useful just because it's volunteered. It has to be both volunteered and directed at effort that addresses something people need. If it's volunteered, but not directed at something people need, I think it's perfectly reasonable that they use some mechanism to indicate they feel dev resources aren't being allocated in a satisfactory way.
Generally when you begin working on problems where this sort of thing is relevant, you have to make choices to get something going. You want to avoid premature optimization and you need something on the ground. In short, you have to assume O(1) is O(1) ... and it basically always is, except when there's evidence that it's not.
The value of the article is to point out cases when these abstractions break down, and the value of performance testing. But carrying it to an extreme such as, "Never make any assumptions about how any data structures work until you've performance tested every single thing in your application" is of course wildly unproductive.
Making it annoying for them, e.g. by a constant reminder that many people support some action to be taken, is the whole point. That way, projects can't simply define away critical changes that users have good reason for wanting.
If it's easy for project maintainers to make a dictatorial or unilateral decision to ignore something a large body of users wants, and they can do that without paying any sort of annoyance penalty, it's super bad for the project. The mechanism of user needs no longer steers development priorities.
I see this happen often on projects where someone wants publicity or credit for their work on something open source, and so prioritizes demo-ware aspects of the project, or showy new features, over critical long-term problems, refactoring workflows, or basic utilities that are sorely needed.
Typically there are arguments of the sort, "I'm giving my development time for free, so leave me alone to work only on the aspects that I want" -- and these broadly form the basis of wanting to disallow +1-like pinging, annoying reminder behaviors.
The trouble is no one cares what your motivations are for choosing to contribute to the project. No one who uses the open source project has any reason whatsoever to care that you found some cost/benefit tradeoff to be favorable, for personal reasons, and to motivate you to contribute.
The project ecosystem generally just wants implementers who will prioritize things as-needed by large sections of the user base, and who will not complain if that means they don't get to use their "donated" time to work only on aspects they personally want.
So it creates a natural tension. Getting rid of +1-like pedancy would be bad, IMO, because it puts all of the prioritization power into the hands of the people who are choosing what to do by their mere wants rather than project needs. I'd like there to be a mechanism that penalizes want-pursuit a little more.
What this all says to me is that for an extreme claim (e.g. using an array and doing a linear search will be faster than a hashmap lookup) it should require equally extreme evidence to substantiate it (e.g. the results of performance tests that stand up to heavy scrutiny).
I feel there is a trap in which someone might look at this sort of thing and say, a-ha I don't actually need to care about big O reasoning or standard understanding of data structures at all! That would be a tragic misreading of this sort of example.
One of the criticisms applied to software engineers -- the one about bad abstractions like "DriverController" and "ControllerManager" etc. -- is a huge pet peeve of mine because it's basically a manifestation of Conway's Law [0]. It indicates that the communication channels of the organization are problematically ill-suited for the type of system that is needed. The organization won't be able to design it right because it is constrained by its own internal communication hierarchy, and so everyone is thinking in terms of "Handlers" and "Managers" and pieces of code literally end up becoming reflections of the specific humans and committees to which certain deliverables are due for judgement. This is not a problem regarding best practices at all -- it's a sociological problem with the way companies manage developers.
Domain specific programmers aren't immune to this either. You'll get things like "ModelFactory" and "FactoryManager" and "EquationObject" or "OptimizerHandler" or whatever. It's precisely the same problem, except that the manager sitting above the domain-specific programmers is some diehard quadratic programming PhD from the 70s who made a name by solving some crazy finite element physics problem using solely FORTRAN or pure C, and so that defines the communication hierarchy that the domain scientists are embedded in, and hence defines the possible design space their minds can gravitate towards.
There is definitely a risk on the software development side of over-engineering -- I think this is what the essay is getting at with the cheeky comments about too much abstraction or too much tricky run-time dispatching or dynamic behavior. But this is part of the learning path for crafting good code. You go through a period when everything you do balloons in scope because you are a sweaty hot mess of stereotyped design ideas, and then slowly you learn how only one or two things are needed at a time, how it's just as much about what to leave out as what to put in. The domain programmers who are given free reign to be terrible and are never made to wear the programming equivalent of orthopedic shoes to fix their bad patterns will never go through that phase and never get any better.
For me, this is what's so scary about Uber. It's not "disruption" as many people seem to claim. It's engineered regulatory capture.
The goal isn't to disrupt a market, but rather to wipe it out and have the financial and political backing to secure some sort of regulatory environment in which lower-priced options cannot emerge after the VC-subsidy phase ends and the monopolistic price increases begin.
It's double sad that as the sort of flagship start-up of the era, Uber leads the way in deplorable executive behavior, shady business practices, and questionable labor policies ... and despite it, they've managed to win the PR war that has every naive tech youngster singing about how they are "disruptive" and singing how all criticisms against them are invalid because of precious, precious "disruption."
Given all this, what's the best way for someone to pick up and start contributing to Cython?
Sadly, I think it's just age-old politics. You're hired for the political effects it has on your boss -- including looking like a l33t h4x0r in your violently-collaborative open plan office doing Agile. Kills productivity and morale, but managers & executives aren't compensated for actually producing anything, so it doesn't matter. And HR'll always spin some other story about turnover because the one thing they have to avoid at all costs is actually providing a healthy workplace.
[0] < http://suitdummy.blogspot.com/2015/05/why-hire-underemployme... >
The only way that such a thing could be commoditized into a front-end interface is if you could devise an interface that allowed for all the different kinds of questions that will be asked, the ways they will change, the new features that will be wanted -- because most of the work is taking some backend pipeline that is already optimized for being able to answer questions of type X, and then figuring out how to generalize it without losing any performance in order to also answer questions of type Y.
It's almost always highly specialized to the specific company and line of business involved, so consulting companies can pop up to take away some of the in-house work, but in general there is no conceivable "as-a-service" thing that could.
Thus, you're left with needing to manage your own backend, probably for quite a long time to come.
Finance, ML for search interfaces, small-data statistics consulting (like political statistics, ecology, and other fields), education analytics, and many other fields offer work that falls into these categories.
Basically, anywhere that there is a business or domain science researcher that needs ad hoc computer programs whose lives as programs will generally only serve to answer scientific questions for that researcher. The ad hoc nature of how the questions change most often mean that no service that pretends to put a front-end API over top of the science questions can adequately capture the variety of things that are needed, especially once the further need to heavily optimize them is added in.
When hiring for a web position, hire a web developer or a person whose engineering skill leads you to believe they will solve the web development problems. That's my whole point. Web developer != "full stack". Hire someone else for the database side, and have them work together. The two specialists, say, are worth much more than someone you venerate as "full stack" and make do both things.
I think specialization of labor applied to the software stack is actually so much more critical now than it ever has been, and that in truth most places that seek to structure themselves around the idea of "full stack" only do it under some misguided belief that it's somehow cheaper or more efficient to employ people who can supposedly "do it all." It's further dysfunction when you see postings asking for very inexperienced recent grads who are also somehow gurus in 5 or 6 full-stack domains.
"Full stack" is generally an outgrowth of bad management ideas, and sends up red flags about companies and teams that believe it's super important, or that very competent engineers with experience wouldn't be able to pick up the skills they need quickly because they have "never gone full stack on web" ...
I see a few of these kinds of things that are somewhat tied together:
"Doing things at scale" -- but doesn't actually define the scope of their specific problem in the job listing or the interview. They seem to believe there is some singular platonic thing that is "at scale" for all problems and all situations.
"Full stack" -- but list conflicting skill sets for a position, or (worse) basically admit that they have no idea what you'll be doing for them. If you're trying to hire a machine learning Ph.D. with 6 years of experience in Javascript, then something up the hiring pipeline at your company is messed up. The drive to do a ML Ph.D. (generally speaking) is not really compatible with the drive to acquire 6 years of Javascript experience.
"Fast-paced" / "constantly-changing" environment -- The job of business developers and managers is to present a stable double-sided interface. One side faces customers and the stream of business problems that Nature creates. The other side faces the employees who then implement solutions to those things. Yes, you can't control the problems that nature throws at you. But you can control the way those problems are ingested, broken down, analyzed, and presented to the workers who will solve them. When a business punts on this and basically says anytime Nature throws us something tricky, management will just jerk you around under the infinite excuse of "fast-paced environment" then you should have serious, serious concerns about whether those business managers are actually going to be successful, or whether they will show respect to the intrinsic human need for adequate work/life balance and professional respect for your position within the company.
The paper "Let's Put Garbage Can Regressions and Garbage Can Probits Where They Belong" by Achen [0] is a great discussion about some particular properties of this, and the tacit assumptions used to ignore it.
In that paper, it's demonstrated that with just a tiny bit of coding error in the covariates, you can end up with a fitted regression coefficient that is statistically significant and has the wrong sign -- even when there is no noise whatsoever in the target variable (i.e. you can set up a toy example in which the target variable is synthetically generated as a true linear function of two covariates with positive coefficients, then perform a slight non-linear distortion on one of the covariates, regress the synthetic target variable on the clean covariate and the distorted covariate, and get wildly incorrect coefficients that appear to be statistically significant).
People seem to think these toy example are some kind of alien phenomenon that could never happen with real-world data, but the paper is very explicit in the construction of the example data set. It's not harebrained or contrived, like Anscombe's Quartet or anything -- it's very much a plausible data set.
I think it's not hyperbolic at all to say that results like this more or less conclusively show that naive linear regression cannot be trusted. If you're careful with model validation, using randomized hold out data, lots of diagnostic plotting and sanity checking, then regression is a fine tool. But if you do something shocking like take two different univariate models with the same target, fit their regression coefficients, and then select the model with a more favorable t-stat as "the winner" then you are committing an egregious statistical fallacy that often, in real world situations, is giving you not just an inaccurate answer, but an answer pointing totally in the opposite direction of the truth.
What's frightening to me is that across many industries, even in places like high finance -- where "real money is on the line" -- it is extremely common to see huge business intelligence systems predicated entirely on this type of fallacious statistical approach with regression. Sadly, it's often because the regression approach was historically more tractable and the fallacies weren't as well known. And so as certain people gained more senior positions and sought to retain political control of the business tools that they oversaw, they grasped for convenient fictions like "interpretability" to justify their political choice to shun modern techniques.