Professors’ unprofessional programs have created a new profession
economist.com
economist.com
I could keep complaining, but I don't want to give people the impression that I don't like what I do. It beats the hell out of writing CRUD apps.
Another way of looking at it: People doing STEM PhDs are extremely busy. They have to master the content in their field, keep up with new research, conduct their own research, write papers, go to conferences, give talks, teach, keep up with whatever department service responsibilities they have, etc.
They're smart people so they're going to be able to slap together some code that does what they need it to do. But to expect them to learn git and how to use it well, to learn about managing dependencies, about unit tests, about build systems, about best practices for documenting code, about programming abstractions or OOP whatever else... that's a lot to ask.
(And I say this as a STEM PhD who did learn all that stuff, more or less.)
If I had to do this kind of things again, I might do them properly, but I'm not 100% sure. Back then, I had more pressing things to do than writing excellent code, since I was writing code to drive the research (as in, numerical investigation), not as proper research (the end results were theoretical results)
I think the real problem in academia is the same as the problem in some organizations that also don't follow best practices: the people empowered to make the decisions about whether to invest people's time into improving testing, building, technical debt, or documentation do not actually believe that an increase in productivity, increase in access or decrease in bugginess will be achieved.
The time to learn all of this is almost like a field in itself say Software engineering.
I just need to get shit done.
I had a project to do bayesian hierarchical modeling. I wrote it from scratch in R without using a MCMC framework.
Yeah it's ugly, yeah it's slow. But I got initial results to satisfy my mentor and within the time limit of the summer internship. This is on top of learning two semesters worth of Bayesian statistic in 3-4 weeks.
Time isn't expendable and I can't devote my life to every single thing while neglecting other aspect (love, health, etc..).
Yeah your theory is nice only if you don't have a deadline, a life outside of work, etc...
Not all academic disciplines are as hectic as yours. When I was in a top grad school, most of my colleagues were not under the stress you describe. And yet their code was still crap.
Incentives are a bigger reason than workload.
Serious questions: Where was this amazing place? When were you there? Was it a STEMish PhD? Really.
I've never heard of a grad school in STEM these days that is anywhere near like that. I'd love to know because the way STEM PhDs are going in the US these days, I can't recommend them in general, only specific PIs and labs.
Yes. Engineering. I don't like to reveal too much about my past, so I won't say the name. Suffice it to say that it's ranked in the top 10 for pretty much all engineering + many of the sciences. EE, CE and CS are all in the top 5.
It was a decade ago, though.
Workload all depended on your advisor and funding landscape. Some were very demanding, and they had miserable lives (mostly pointless, too - most of academia seems to be converging towards a race to the bottom). But not all professors were that way. Of course, standards still have to be met, and laid back professors' students usually took longer to get a PhD because they were more relaxed.
To give you an idea, two of my fellow group mates are now faculty members. And the one whose work was most like what you describe, and who did the best work IMO, never was able to secure a tenure track position.
The lifestyle you describe is really not worth the PhD you get in the end. The ROI is just not there - financially or otherwise.
I also feel like it's a race to the bottom now. It used to be 'publish or perish', now it's 'funding or famine'. The paper numbers, impact, and truthfulness have nothing to do with anything. I see so many PIs that essentially end up as one income households, their spouse's, not their own. The 'good ones' just get pushed out or leave, an evaporative effect, leaving only the liars. The academy is in real trouble these days. It isn't on life support yet, but it's not able to climb the stairs anymore.
I once had a quite respectable lecturer working for me, and his code was awful. But that was ok - we knew that. He was hired for his conceptual skills and domain knowledge, and we paired him with a junior developer that could take his raw output and turn them into cleaner code.
It's some engineer's job to do it properly in an actual production setting.
No. One's point is to communicate one's ideas. Code is the only specification of a system detailed enough to reproduce the desired results. You wouldn't present an unfinished paper written in the equivalent of crayon so that particular attitude is the worst kind of lazy bullshit for people actually trying to make use of your hard work.
> It's some engineer's job to do it properly in an actual production setting.
Fuck this attitude so damn hard. As an academic, one _is_ the _first_ engineer who has a duty to lead one's colleagues - to mark as clearly as possible the path for those who come after. This "fuck you, got mine" attitude is unacceptable and unsustainable.
The idea is everyone writing everything the "right way" from the get-go would be premature optimization.
The right way in this case doesn't mean AdapterFactoryBeanFactoryService-land. It means taking care that your code is clear, concise and correct for others skilled in the art. The fact that academic code exists at all, as an example, means that it _is_ in production. There's no excuse for shit code that isn't written for other human beings first and computers second.
We're all busy. How long do you think it takes for time spent learning good software development practices to pay for itself? Surely not longer than the length of a PhD.
Sure, but if the core of your job, the key skill you bring to the table, is writing software, you are (hopefully) going to find the time to learn best practices for writing software.
If that's not the core of your job, if your key skill is something else and writing software is just an occasional means to an end... maybe the incentives aren't there.
> How long do you think it takes for time spent learning good software development practices to pay for itself?
If you almost never write a program longer than 1000 lines and rarely use a program six months after you've written it, it may never pay for itself.
Don't get me wrong, I personally have an interest in good software development practices. I just understand why not everyone can or will take the time to learn them.
When is your first and last meeting of the work week? 6am Monday to 7pm Friday? Have you taken more than 3 weekends off in the last 12 months? Do you measure your daily coffee consumption in gallons? And yes, I'm serious.
PhD-land is crazy, perhaps this is why sooooo many drop out.
I wish the state of affairs was prettier, but I doubt it will change any time soon. And frankly, I don't think it has to. Software engineers will always find something to complain about others' code (coding style discussions come to mind). If there's value in commercializing scientific prototypes, experts should rush in and do it. There's no need for us to invest in perfect code just in case someone may want to reuse it. Having said that, NSF is starting to fund initiatives to improve the state of affairs. For my field, it's EarthCube: https://www.earthcube.org/ but this seems to be fumbling along; all of their goals would need to be adapted by scientists outside of that community (standards, etc). Without the right incentives (funding, publications), it's just not going to happen.
Can you identify (say) five things that academics can do with an existing code base that will make your job easier (and therefore cheaper to the downstream users)? Perhaps pitch the top five pick at the abilities of a computer science student looking for a project?
Can you identify (say) three key practices to adopt when starting a new project given the likely skill levels of postgrad students in your domain?
You could get an ebook out of those I imagine.
1) test suite - this shouldn't require explanation
2) documentation - both top level docs with examples and well-commented code. keep track of units
3) one-step build - I should be able to download the project, type 'make' (or something equivalent), and start using it
You have a different view of how academia works than most academics.
Most PIs these days do NOT collaborate and are actively hostile towards it. It's a shame, but it's true.
Also, the niche-ing of academia is real, in many fields, there may be only 3 other people on the planet that understand what you are actually doing, and many of them may not speak English and you may not know they are out there until 2 years from now (publishing takes a long time).
Most code is written by grad-students that barely know what a for loop is, let alone how to use git, and they only write it to do it once for a specific paper. Look into MatLab, that is the most used language in bio, by far. It's basically psuedo-code that compiles, and it's still a mystery to most grad-students.
Well, yes and no. Speaking as someone who has done this as well: if it's UI stuff, or gluing processes together, it's fine. But sometimes they want you to fix up a custom algorithm they implemented. Which can be one of those cutting-edge scientific things that actually require a PhD-level of knowledge and insight into the topic. So you can only get half-way with untangling the spaghetti mess without sincerely not knowing for sure if you broke something or fixed a bug, because you don't know what the correct output should be.
2. Have significant test coverage
3. Have at least a rudimentary form of Continuous Integration, even if it's just running a script on every check in.
Good developers get paid enough, that it indicates that maybe our job isn't quite that easy to learn.
Our basic programming abstractions and software engineering practices are things that we as a field developed over time, based on experience; it's not like they're something you just automatically know without having to sink time into studying if you're above some IQ level. They also depend to varying degrees on how big your program is, which parts are expected to change more or less rapidly, what kind of program it is, etc.
Unit tests assume you know what the code is supposed to do before you write it, that the code will change significantly more often than the required functionality, and that the tests are easier to read and check than the code being tested.
In practice, at least for simulation code, functional, integration and regression tests are useful when employed judiciously. Most importantly, you verify your results using published benchmarks in the scientific literature or analytic results where possible. Obsessively covering every trivial bit of code with a unit test of its own has always struct me as rather a weird fad.
The advantage unit tests have over regression tests is that the time between the moment developer implemented a change and finding out they broke something is as small as possible. When tens of people work on the same codebase this saves enormous amount of effort.
Fethisizing over some rule for it's sake is silly of course. My org would waste a lot of money and time without unit tests.
Another feature unit tests provide: Think of the unit tests as a living documentation. Breaking something leaves a breadcrumb trail to the invariant which was sullied.
I am maintaining several end user critical projects related to data transform and transport, mostly by myself. I have no idea how I could develop them as efficiently without unit tests to verify my changes don't brake some corner case as new features are added.
The cadence for changes can be fairly low - I might return to some component after doing something else for six months, and I probably need to add some feature without breaking something in the process.
Sure, we have integration testing and smoke testing, but the further a bug travels the release pipeline from the developer the more expensive it is to fix. This cost cascade is quite easy to visualize - the work of whomever caught the bug is stopped, and depending where they are located in the product ecosystem their stall can cause lots of work for other people. If they are a tester they need to file a report. If they are a customer, their work is interrupted, they contact local sales, who then contact global helpdesk, who then identify and log the bug.
Much simpler and easier if there is a unit test to catch the bug in the first place.
Now, there can be domain specific flavors to this. My domain is computational geometry and transport of 3D modeling data between domains. But in my domain anyone not securing their code with unit tests is wasting their employers money and setting their end users at risk by increasing the likelihood of bugs.
There is lot of cargo cult nonsense in software engineering. Unit testing is not one of them. It saves time and effort by catching a range of bugs not caught by e.g. compiler for statically typed language and it secures the program logic for future changes.
Assuming this is serious, FWIW I'm a STEM PhD and I find that even in the smallest scripts, even just web-scraping stuff, if I don't write unit tests then I get it wrong and later discover I need to do it all again. Or worse, I figure out later, after I've built lots of code over this, that there is an error and the subsequent code all needs to be redone.
In short, I find unit tests to be a big time and effort saver. And I get the right answer, for a change. :-)
I've found that unit tests are less important when you have good CI and your codebase isn't shit. On high quality code it takes so little time to fix bugs and redeploy that unit testing is a hard sell.
Do not write unit tests just because the methodology requires it. Write unit tests when the unit needs to be tested. Typical examples:
* Smoke tests (does this code work at all?)
* Tests for known edge conditions
* Tests for known bugs (to catch regressions)
(For predictive models) evaluate the output on a test set? Depending on what you're trying to do, the contents of the black box may not matter so very much if the results on an agreed test set are good (...and you're confident it's independent from any training data...)
Manual eyeballing and sanity-checking of output? (in my experience important however many layers of automated testing you're using, and undervalued by people who focus on software as an engineering process more than an art-form).
...and, circumstantially, probably a bunch of others. I don't think this is a field where making lots of rules is especially helpful (except, perhaps, the "always eyeball" one...)
Even then, logic mistakes can and do still happen.
Funny, that sounds exactly like my IT Dept.'s approach to coding.
I work in Operations and was exposed to their approach when helping them translate some business rules into SQL for the warehouse management system. During the course of interaction the IT Director lost the code we had worked on together twice, keeps the requirements docs in his email etc. etc.
Me: your engineering practices are terrible
He: We don't have time to do it properly
Me: but somehow we have time to do it three times
I told my manager but no-one really understands or cares enough. The Warehousing Director calling out the IT Director about his approach to IT - where would that all end ?
Then there's the group of folks who tend to read those papers and turn it into products. Yes, those are the ones I expect to have repos, unit tests, build systems and all that.
I used to work as a scientific/medical translator and one of the things I loved about my work then was that with each new project, I had to learn enough about a new topic/subfield to become a bit of an pseudo-expert in it. I greatly enjoyed that research and would love to do some work in software development that would require the same type of constant learning with each new project. Diving into some convoluted, 20-year-old Fortran driving highly-specialized scientific software actually sounds kind of exciting to me (yeah, I'm not normal either :) ).
Is there much demand for this kind of work, I wonder? I'm guessing the best way to get your foot in the door is to be in academia, and those days are long passed for me.
As an undergrad, I had a good friend who was doing his PhD in Atmospheric Physics. It turns out, most of that field works in Fortran '88. This is not a very useful language, seeing as it uses GOTO statements to function as a loop. Fortunately, it does have comments. My friend managed to sweet-talk an older PI into giving him the old code for use in nuclear blast atmospherics (Exp: say you nuked all of France, what happens to Greenland's ice). At about 3 am before a project was due in the morning, he was pulling through the spaghetti that was the code, tired, jittery, and over-caffeinated. In this mess of logic diagrams he had to draw out by hand, he finally got to somewhere he thought was going to really cement all the code together for him. He follows a GOTO statement, and there was only a set of another GOTO statements. This went on for about 30 (my recollection of his words) GOTO statements, all 'nested'. Eventually, he gets to one that only has a comment line: 'HAHA MADE YOU LOOK'.
His laptop was defenestrated and he had to buy a new one with me about a week later.
You can staff projects off of grants, so it removes the need to sell the product. And the companies are probably going to be more academic, with a high percentage having graduate degrees. But companies will be under pressure to win grants. And because they have less of a need to sell a product, your work might not be widely used after the project ends. Also, the DoD is a major source of grants, so it would be helpful if you were comfortable with working on military research projects.
But it might be something to look into if you want to get into software development in an academic environment.
Academia is being pushed more and more toward industry collaboration, so if you can get the exposure and connections it shouldn't be too hard to find the opportunities from the inside.
You really have to enjoy this kind of work and some of the frustrations that come with it. I did, but ultimately had to leave when funding for my group ran out. If life circumstances allowed me to work for less I might have stayed with them.
This is far from an original thought:
In the good old days physicists repeated each other's experiments, just to be sure. Today they stick to FORTRAN, so that they can share each other's programs, bugs included.
-- E. W. Dijkstra, 1975 apparentlyIf you have original dataset and .Rmd file, you can then bisect both analyses to find the bug.
Grants are time-limited, and at some point usually the money for developing the software runs out. The PhD students working on the software move on to something else, and you have yet another abandoned piece of scientific software.
There are of course exceptions, but in general it's much harder to get money for maintaining and improving software over a longer timeframe than for building something new.
This sounds negative but I don't mean it that way; it's just how it is, no judgement meant. But it's not something you really hear anyone teaching new PhD students as an option, and even if they would, it's highly uncertain and success depends on many factors out of your control. So I wouldn't call it a career path, or even viable career advice.
Dear God how do you manage to do this?!
PS: I couldn't read the article, presumably because of some anti-adblocker mechanism.
> I couldn't read the article, presumably because of some anti-adblocker mechanism.
That is so stupid of news sites to do. They only thing that makes me do is leave your site immediately.
It's almost as if they don't want anyone to read the content without also viewing the ads...
I keep Chromium in 'full fat' mode for sites with rich content and for coping with those public wifi connections where you have a javascript based terms and conditions page.
(I think it's "firefox -profileManager")
Make code review part of peer review.
I think the next generation of scientists at least understand that things like unit tests are useful, but passing peer review is the only incentive to do anything, and reviewers don't read code.
I review a paper in an hour or so. If somebody comes to me with some fancy new symbolic execution framework it is going to take me days to review the code. Unless you are planning on paying me for a months work to review 20 papers, this is impossible.
I also cleaned up my own mess from 1996, when I recently re-released a General Relativity package for Maple, GRTensorIII (https://github.com/grtensor/grtensor) that was part of my PhD work. The 1996 code is not that tragic, but is far from my best work - although I did earn my PhD.
http://yosefk.com/blog/why-bad-scientific-code-beats-code-fo...
Computer science is something that seems underrecognized in the field I work. People clearly acknowledge its importance, but then turn around and basically ignore it when talking to potential grad students or mentoring undergrads in prep for grad school. We don't offer any kind of course like "programming for X" even though most of the students need it.
My experience closely parallels the "why bad scientific code beats code following best practices." I've had comp sci students come in and what happens is they clearly understood python and java, but had difficulty understanding the problems with inheritance, and wrapping their head around other more functional languages we were using. They also were unfamiliar with the statistical/content areas, so had difficulty implementing things. I had thought it would be great to have comp sci students involved (and still do) but it didn't solve my problems like I thought--so instead of having students who understood the concepts but not the programming, now I had students who understood the programming but not the concepts.
When you're dealing with really intense math and statistics, it's difficult to separate out the programming from the math. It's not like web development where you have an "insert text here" kind of approach that works often; the algorithms and the problems are really wrapped up in one another. This might all be changing with data science DL and AI and that kind of stuff infiltrating comp sci's assumed territory, but I'm not really seeing it much so far.
It seems the prototypical situation in software design is some software that's team-developed for mass consumption. In science, you have the reverse often, which is software designed by small units that might be a one-off thing. These constraints put different kinds of pressures on the process, such as intense pressure on getting something to work correctly at all costs, including elegant design.
Also, the unit testing thing is kind of confusing to me. Every time discussion about a new language comes up in the context of numerical/scientific computing, one of the big questions is "does it have a REPL"? It seems one of the big reasons for doing this is basically unit testing. It might not be unit testing in the formal sense that you might have at some software design companies, but anything someone complicated involves feeding each tiny separable part of the code something with known expected output, sometimes in strange, boundary-testing ways, so that seems pretty similar to me. There's also a plethora of test-case datasets out there for this very purpose.
To me the bigger problem is homogenization in software in science, that is, a domain being dominated by a single piece of software. I think it leads to unrecognized errors due to lack of replication across implementations, and problems typical of monopolies (even when something is open source). There's a kind of development benefit:cost supply:demand problem that leads to dominance of single works of code that is really unhealthy for science (replicating with standardized methods is good too, but to me that's a slightly different issue).
Unit test vs REPLs is an interesting one. Agree that there are similarities there (although I'd argue that your tooling needs to be pretty damn good for unit tests to offer the responsiveness that a good REPL can). For me, I think part of the difference is that a REPL session is personal and nobody sees the blind alleys, while unit tests are an enduring part of the product and something others will see, use, and potentially critique. So while they can address the same kinds of questions, I'm not too shocked that people feel differently about them.
In each case, the people being given responsibility for the programs where in way over their heads. They were smart and educated and generally viewed programming as a skill somewhat akin to typing, something anyone could learn to do adequately, if not quickly, in a few days of practice.
The ecosystem simulation made unnecessary oversimplifications (assuming that an exponential relationship could be modeled as a linear one because the engineers didn't know how to handle the integration of an exponential(!!!) and how that should be handled in a program).
The flood control modeling program, used by the Federal government, was some of the worst code I had ever seen. It was written in a fashion where variables were treated a bit like a small finite set of registers. Ten or twelve global variables in the program were reused over and over for different purposes, sometimes to return computed answers, sometimes as temporaries inside of some function, and sometimes as iteration counters. It was a complete mess.
A graduate student that had never learned C or C++ was being given a previous grad student's C++ simulation code to use as the basis for his dissertation work. That code was pretty useless and poorly documented. Software engineering principles played no part in the work these students were doing, they were from a different discipline.
In the case of the real time machine control, programs were written in thousands of lines of assembly code and would control perhaps a dozen asynchronous activities though a combination of interrupt handlers, polling loops, and time outs. They were frustrated that the machines would simply hang every day or two. What a mess.
It was definitely one of those programs that does this, than that, the sometimes the other thing, and it went on forever. I had trouble understanding it myself.
I didn't end up rewriting it because, well, startups. The whole thing went kaput, and nah, didn't have much to do with the code.
I would agree with you that these programs are more understandable to someone with expertise, and that can be why it's so hard to refactor them. It's difficult to become this expert in a branch of science and also take the time to learn good software design and programming practices.
But overall, I'd say that these programs could become vastly more manageable and easier to understand through better programming practices.
> "Those who do put effort into producing good code risk being seen by their colleagues as time-wasters."
Producing good code takes time and effort. It would seem a complete culture shift is required before any significant changes will happen.
I wrote this course as a basis:
http://learngitthehardway.tk/learngitthehardway.pdf
Finding the right pitch point for someone to learn in the right way is really hard. People come to things like git and build systems from all sorts of angles.
But maybe a (2015) in the title would be a good idea.
Both appear to be hyper focused on solving the problem that little else matters.
Thanks! Do you do this kind of work as a contractor, presumably?
Do you also typically write documentation for the projects you're working on (inside from inline comments)?
How did you get started in this line of work?
No I don't usually write docs - I get the authors or maintainers of the models I work on to do that. I do guide them, provide templates and examples, and ask for clarification when they're missing parts. I have standardized methodologies ready for that, which I developed myself mostly (this is one of my USP's, as long as I manage to convince people of the value, which is quite hard and which I often fail at). I don't think it's good practice or very efficient for programmers to reverse-engineer the whole thing because you have to become a domain expert to do so. I also think that this is why it's not for everybody - too many people let themselves get sucked in too deep, making it very time intensive. I understand the temptation, it's much more intellectually satisfying to go deep yourself. But I think you need to be as much project manager as programmer, so that you can get the actual domain experts to figure out the complicated (domain knowledge) parts, and limit yourself to factor out/replace the plumbing and introduce good software engineering practices. Those usually don't last after you (I) leave though, so it's also important not to get too worked up about that.
I started out at a research group that found itself accidentally too heavy on software people, the group got into projects doing software stuff because of that, developed a reputation for being 'the software guys' and failed as a research group because of that (it's a lot more complicated than that, this is the Cliff's Notes obviously). Through many coincidences that can't be replicated on purpose, I'm now hyper-specialized in doing the thing the OP describes in a tiny, narrow field.
The 'trick' (well it's not a trick really) to get work is to be very well connected and work hard to remain that way (being well connected is not something you find yourself in, it's the result of many years of thankless, feedback-less grinding), make your work visible to the outside (i.e. marketing, although obviously the 'buy Adwords' type of marketing is 100% useless here) and to know the science funding processes very, very well to understand incentives of all parties involved. This last part is vastly underestimated; not just for what I do, but also for researchers themselves. For example, the reason I'm usually in is when the project asks for something with demonstrable real-world application (this is a very common requirement the last decades, even for highly theoretical fields). So knowing how to put a veneer on theoretical work is a very non-obvious but highly valuable skill. ('veneer' is not 'hiding things' or 'faking', which will work maybe once or twice - I'm talking about (essentially) science communication more than 'writing papers' science).
Furthermore, being realistic is also important. I'm never going to be rich doing this, nor will I ever employ large amounts of people (or any people at all apart from the occasional 1 or 2 day freelance subcontractor). It's also something for the long haul - 10 years to become established. Other downsides are the sometimes infuriating academic politics, the eternal 'I'm a mathematician/physics/CS PhD so I'm God's gift to mankind and everything that cannot be distilled down to a theorem is 100% useless' characters you run into (they're not that common tbh but they still annoy me endlessly), and not having clear goals or even goals at all. It's like being a PhD student except worse, and with no end in sight or even a thesis to work towards. The upsides are long-term contracts, lots of freedom, intellectually more stimulating work than writing marketing websites or CRUD apps, and being respected as an expert at something even if it's a tiny sliver you're considered to be an expert at.
It's something I'd like to do, but the networking/politics might keep me out of it. (My family are in the sciences, so I'm familiar with how academia works.)
It's a shame it doesn't pay more. On the other hand, getting exposed to a lot of different things is more valuable long term than specializing in corporate niche.
I left a pretty high-paying job to work on open source (for $0) so I can relate.
To be honest, it's likely more survivor bias than skill or successful execution of a preconceived plan. This is important to recognize as I'm asked a few times a year how to end up in the place I'm in, but I don't think there is a solid path to do so, nor am I in a position to give advice beyond the standard 'think hard, work hard, keep your eyes out for opportunities, pay yourself first'.
Just saying - in case anyone is reading this in the hope of steering their career one way or the other, don't take the answers I gave here as anything more than an anecdote :)
Follow-up: say I have a lot of free time to patch it up, write better documentation, unit tests, etc. What's the first, most value-added thing to do?
That and running only on real-world, way too big datasets. Turnaround time for running them is too long, only major problems show up, the effects of options x and y are only small so errors in them are largely unnoticeable unless you look for them - which you can't/don't on large datasets.
As to nr2, depends on what your goal is. Do you want to make this into a more polished product someone who doesn't know how to code can use, do you want to open it up so that others can try variations of your algorithms, do you want to make it into a tool that becomes standard in your field and will act as a citation-generator or at least -attractor? Something else still? Defining your goal with the software to set a framework to make future decisions in is actually the most value-added thing of all, I'd say.
For 1 and 3 you need a self-contained download, i.e. no 'download this and this compiler' etc. - just a zip with an exe. I haven't tried the tool, I can't tell if there's any functionality missing to do this (like a 'first use' tutorial, or even just some training slides/pdf and -exercises). Furthermore for 3 you need to have a command line version. It depends a bit though on how ubiquitous Matlab is in your field. Probably ubiquitous enough that it won't be a deal breaker for anyone sufficiently motivated, but some people have a natural dislike for it from their undergrad days, or others just are philosophically opposed to using for-pay tools (even if they don't have to pay for it themselves). I'm not, just saying that you might get more citations if it adheres more to a certain 'ideological purity benchmark'; you have to decide for yourself what that is in your field. This is not generally something most people are concerned with, but the type of person who would put their own stuff on Github is likely to be more sensitive to this. Again, it's a social thing - just bringing it up.
For 2, I'd start with writing some documentation (not too much, just a few paragraphs) that explains the structure of your program, and where someone who'd want to add a new method should start; or better yet, explain the plugin framework you've build if you have one.
Also for 2, you need some unit tests if any new functionality can't be implemented completely separate like through a plugin dll. Not so much for your potential contributors, but more for yourself so that you know everything keeps working when you get someone else's contributions that touch a lot of code.
In general, you're more likely to be taken seriously as a thought leader in your field if you have other users and/or some social validation that others are using it, like when you're taught workshops at conferences or elsewhere. So some introductory material, exercises that teach you all aspects of using it, maybe a link to a video or mention of you giving a workshop somewhere. In the life cycle of your software (from what I can tell), that sort of effort will yield much higher ROI than most coding will (especially the 'refactor to detangle spaghetti' coding which is only needed when you actually have a reason for it, not just for the sake of it).
All of this of course assuming you're still in academia and/or want to continue there.
There's a much bigger validation dataset that I used later on that I just haven't uploaded yet. In the latest iteration I have manually traced examples that are used for training and validation, which makes me a lot more confident in the results.
Adding an instructional video / workshop was one of the top things on my to-do list. I agree that that would help any potential users in a huge way. I did video chat with some people who used it, and they ultimately published a paper using it, so that was pretty cool.
In the end, though, this was just a semi-polished byproduct of the actual experimental work I was doing, so like others have pointed out, I didn't have to make the code nice to pass peer review or graduate.
I mostly write software, re-implementing simulation models to be more robust, or faster, or working with other software, or some variation on that. Part of that is also data-wrangling, presenting my/our work, and teaching others how to do software engineering rather than programming. That last is more an ambition than something that is actually (usually) successful, and I do it mostly because I like it and because the occasional person you can get to see the light is so rewarding.