A Call to Build Models Like We Build Open-Source Software
colinraffel.com
colinraffel.com
https://salsa.debian.org/deeplearning-team/ml-policy https://lists.debian.org/msgid-search/8bc4a2fdb2a0619d1d8214...
From your first link: the (experimental, work-in-progress) policy document is a worthwhile read:
https://salsa.debian.org/deeplearning-team/ml-policy/-/blob/...
People like to code their own work, instead of contributing to others. Incremental work does not count. This is more of a flaw in scientific output counting I believe. Also, our technology stack is completely "zero-server" that makes everything on Android much harder and complex. Disclaimer: I'm the thesis adviser of this student
Lots of truth here. Actually, maybe the OP could change the title to make it clear you're talking about machine learning models? When I first saw the title I thought it would be about the deeply problematic state of scientific coding and "scientific" modelling in general. The ML world isn't so bad regardless of how it may seem, because it's got deep roots in computer science and industry. The moment you start looking at the code for things like epidemiological models, climate models etc it becomes clear that the incentive structures in science are completely broken.
I've not only heard from others but also seen it with my own eyes that people who call themselves "scientists" will happily write models in C that e.g. use pointer values in an equation instead of the dereferenced values, leading to incorrect outputs, and when this is pointed they claim there are no problems, or that they checked the results and the bugs make no difference (i.e. they lie). Or they'll have race conditions that cause unstable outputs and pretend it's because of a (pre seeded) PRNG they used, apparently in the hope that other scientists don't understand pseudo-randomness properly. Or they'll mix up their variables because everything is a single letter and there are thousands of lines of code, or they'll do out-of-bounds reads in a sorting algorithm because they don't know about standard libraries etc.
Nobody ever seems to retract papers because of bugs like these, and when the results of their model don't match reality they'll make arguments like "the model was validated against other models, which is a reasonable way to prove validity" or "we don't make predictions we make scenarios".
If researchers were more willing to collaborate on shared codebases they'd start to learn programming better, be able to recruit wider and more diverse teams, they'd share infrastructure and best practices and generally things might stand a chance of improving. But, as you say, there are major flaws in how the output of scientists are evaluated (by governments/non-profits), and this has a nasty habit of converting well meaning scientists into what are effectively pseudo-scientists.
Tensorflow and many more are already open source, so it looks like the problem is not softwareside... Same thing like we have OpenGL but we don't have the 3d models of games...
So you need some "Open-Source-Data" which we don't have, but still need otherwise the tenser-flow model will be useless...
I have an AMD GPU which means I need to use TensorFlow-ROCm, but I've tried a few times and never succeeded in making it work.
A new layer could be initialized to act as a 1:1 passthrough, then trained independently of all the existing layers, in situ. Once any gains were optimized, you could then train the whole model to realize any improvements.
hmm, maybe I'm a pessimistic misanthrope.
They were talking about making models in a more collaborative fashion, which is a non-obvious problem. Models are heavy, single-task oriented and expensive to train from scratch.
That's the great thing about open source. It's for everyone without discrimination.
edit: i'm probably being a bit hyperbolic about the usecases for ML. But the hyperscaler thing still holds: most of the interesting usecases i know of could be done more lowtech with more expert knowledge, with more parcimonious usage of stochastic optimization.