Composability in Julia: Implementing Deep Equilibrium Models via Neural ODEs
julialang.org
julialang.org
I have never once been able to follow one of these blog posts. Seems like universally these posts have terrible curse of knowledge [1]. To be fair, I know only a light-to-moderate amount about machine learning and very little about differential equations, but still— I'm a long-term Julia fan, professional data scientist, mathy PhD, would hope that's at least table stakes. Maybe I just need to...try harder? I wonder how many people would be excited but are totally lost by these posts.
To that point, if anybody has a recommendation for a gentle introduction to these topics (preferably Julia), I'd be most appreciative.
ps: I need to read those adjoints articles.
I don't understand this part. Are you saying you're a "professional data scientist" with a "mathy PhD" who has a "light-to-moderate amount" of machine learning knowledge? How did you get the job?
I would expect anyone with a mathy PhD to understand ODEs and PDEs, and neural ODEs are commonly understood (by those who read the papers, where I think it is an appropriate assumption that people who want to understand this stuff would do) to be effectively infinite-depth neural networks where every layer represents the same function.
Processes initiated by data scientists during the execution of their role will tend to fail silently. What is meant here is, throwing an inappropriate model at otherwise good data produces unreliable (catastrophic in certain situations) results, but produces results nonetheless. Without the proper discernment of the reliability of the results, we have an unequivocal failure to execute the role. This is the oft-unmentioned companion to, but decidedly more insidious than, the "garbage in, garbage out" (i.e., right model, wrong data) aphorism.
It is up to the person performing this operation to deduce whether or not the conclusions are trustworthy. I don't see how someone can be confident of this without either relying on a pre-defined workflow verified by someone else qualified to assess the consequences, or to have those qualifications themselves.
What follows is a contrived example, but illustrative of the problem:
Consider e.g. user privacy: it is by now well-known that e.g. embedding vectors (or even merely the relationships between them) can leak a lot of information about the person or object it represents. It is not enough to understand how the forward pass of such a model commences, but also what is stored in those representations, which, having gone through a master's with quite a few people who now call themselves data scientists, I am not confident is commonly understood.
Sure, if they spent a lot of time wading through the literature they'd probably understand fine, but the point they were making was that the post was quite unapproachable without having delved into the specific literature on neural differential equations.
I think this is a reasonably valid complaint, and does not warrant you implying that they don't deserve their job.
The way any article is written reflects the audience it is suited for. If this article was intended for people unfamiliar with neural ODEs they would have put more effort into writing it in a suitable way.
You also seem to be quite unfamiliar with breadth of the topic of mathematics. It is quite possible to take a mathematics phd without touching differential equations, except as an undergrad.
BTW, implying that people don't deserve their job is just shitty behaviour, and way out of line.
Indeed! My job involves approximately zero machine learning, at least for a narrow or stereotypical definition of machine learning. I work on optimization, domain-specific models inspired by queueing theory, various phsyically-motivated structural models, writing production code for data-heavy products, testing hypotheses in data, brainstorming how existing data can be used to solve new customer problems, writing documentation, communicating with customers, etc.
If you think this isn't data science but know of better nomenclature: please share! I've struggled to write a great job posting for this kind of work, and surely improved nomenclature would help :)
By "light to moderate amount" do you mean that you haven't catalogued the menagerie of statistical models that are commonly taught in machine learning / deep learning courses? Because that to me is secondary to having the underlying principles (optimization theory, measure theory, etc.) down pat. Recognizing the fundamentals in the theory which describes how machine learning proceeds is invaluable for comprehension.
If you wanted a flashier name for your role I might suggest "Solutions Architect" or even "Machine Learning Engineer" based on what roles you want to aim for. Because honestly you're doing a lot of what is already entailed by those titles. "Data Scientist" also fits for sure, being such a broad title nowadays.
I interpreted what you said more glibly(?) than it seems you intended, and apparently I expressed my surprise snarkier than I intended.
And also the 18.337 lecture notes (probably also has videos available): https://github.com/mitmath/18337
I'm asking for sources outside of Julia because I find the coupling of algorithm types to tools kind of strange and the whole SciML trend is kind of opaque to me. (Are people applying ML as a solution to newer problems? Are they using new approaches to solve ML problems? How legit is the whole thing? I just don't know.)
I suspect neural ODE work was done in Julia earlier because it was easier given some language features and libraries. But there does seem to be some work on neural ODEs in Python/Pytorch.
(There are good possible explanations - it could be very new, have only niche applications, Julia is somehow uniquely suited for it etc. I don’t know)
There's over 100 dependent packages: https://juliahub.com/ui/Packages/OrdinaryDiffEq/DlSvy/5.64.1...
I don't think there are any gatekeepers limiting it's use. Articles like the one highlighted here help to get the word out to more potential users.
In terms of the developer team, throughout the SciML organization repositories we have had around 30 people who have had over 100 commits, which is similar in number to NumPy and SciPy. Julia naturally has a much lower barrier to entry in terms of committing to such packages (since the packages are all in Julia rather than C/Fortran), so the percentage of users who become developers is much higher which is probably why you see a lot more developer activity in contrast to "pure" users. With things like the Python community you have a group of people who write blog posts and teach the tools in courses without ever hacking on the packages or its ecosystem. In Julia, that background is sufficient knowledge to also be developing the package, so everyone writing about Julia seems to also be associated with developing Julia packages somehow. I tend to think that's a positive, but it does make the community look insular as everyone you see writing about Julia is also a developer of packages.
Lastly, since we have been focusing on people with big systems and numerically hard problems, we have had the benefit of being able to overlook some "simple user" issues so far. We are starting to do a big push to clean up things like compile times (https://github.com/SciML/DifferentialEquations.jl/issues/786), improve the documentation, throw better errors, support older versions longer, etc. One way to think about SciML is that it's somewhat the Linux to the monolith Python packages's Windows. We give modular tools in a bunch of different packages that work together, get high performance, and become "more than the sum of the parts", but sometimes people are fine with the simple app made for one purpose. With DEQs, there's a Python package specifically for DEQs (https://github.com/locuslab/deq). Does it have all of the Newton-Krylov choices for the different classes of Jacobians and all of that? No, but it gets something simple and puts an easily Google-able face to it. So while all it takes in Julia with SciML is to stick a nonlinear solver in the right spot in the right way and know how the adjoint codegen will bring it all together, the majority want Visual Studio instead of Awk+Sed or Vim. We understand that, and so the DiffEqFlux.jl package is essentially just a repository of tutorials and prebuilt architectures that people tend to want (https://diffeqflux.sciml.ai/dev/) but we need to continue improving that "simplified experience". The age of Linux is more about making desktop managers that act sufficiently like Windows and less about trying to get everyone building ArchLinux from source. Right now we are currently too much like ArchLinux and need to build more of the Ubuntu-like pieces. We thus have similarly loyal hardcore followers but need to focus a bit on making that installation process easier and the error messages shorter to attract a larger crowd.
The basic idea is inspired by the “adjoint method” for ODE solving (so you don’t have to hold in memory all the intermediate layer outputs — which is otherwise necessary to compute the backpropagated gradient signal).
The recent one that I found funny was the Second-Order Neural ODE (https://arxiv.org/abs/2109.14158). The second order adjoint is rather old, with a canonical implementation in Sundials (https://github.com/LLNL/sundials/raw/master/doc/cvodes/cvs_g...) based on a very good analysis of how to do second order adjoints fast (https://epubs.siam.org/doi/abs/10.1137/030601582?journalCode...). But what about putting a neural network in there? In Julia it's just composing forward-over-reverse to get the optimal adjoint that matches Sundials, so there's been a tutorial on it in DiffEqFlux since 2019 (https://diffeqflux.sciml.ai/dev/examples/second_order_adjoin...) with both Newton and Newton-Krylov methods. The rest of the optimizations from the paper follow using BacksolveAdjoint and relying on dead code elimination (DCE), which is a compiler pass that wouldn't exist in Python so I guess they have to care? I assumed that was too trivial to publish given that history of prior work, but somehow that paper got a NeurIPS Spotlight with "a novel computational framework for computing higher-order derivatives of deep continuous-time models". At this point I just kind of shrug though, I think it's a symptom of conference culture not giving people enough time to review the literature thoroughly so random things seem to seep through. Take it as one reason among many to treat any ML conference paper similar to an unreviewed preprint.
All of that together, that's why the Julia universe of tools has mostly been focusing on improvements and performance vs Sundials, PETSc, etc. since those are the real challengers. You can see that we spend our time benchmarking against Sundials all of the time in our stiff ODE benchmarks and recently started outperforming it with QNDF pulling 2x-5x wins in various ways (https://benchmarks.sciml.ai/html/Bio/BCR.html, see https://sciml.ai/news/2021/05/24/QNDF/). The stiff neural ODE paper describes an 3 ways to achieve an improvement on the Sundials adjoint in terms of complexity (https://aip.scitation.org/doi/10.1063/5.0060697), and our adjoint benchmarking paper shows how putting all of these various pieces together leads to about 2-3 orders of magnitude improvement over the naive CVODES adjoint (https://arxiv.org/abs/1812.01892) (against the method, the follow up of course will be against the direct wrapping).
What's funny though is that people then get upset when we show benchmarks against some of the Python tools, like >100x improvements in solver speeds on physical and biological problems (https://gist.github.com/ChrisRackauckas/cc6ac746e2dfd285c28e...) and adjoints (https://gist.github.com/ChrisRackauckas/4a4d526c15cc4170ce37...). I don't understand why anyone would be surprised though: we've spent years "competing" against the C and Fortran codes and only recently started pulling ahead due to the combined effort of a whole community, while the Python tools were just a few people with simple methods who never benchmarked against the previous tools. If they benchmarked enough they would see that we're not an outlier claiming to be 100x faster than everyone else, instead we're in and slightly ahead of the pack but they are the outlier that is 100x behind the whole group. Personally, I would require every paper to at least have a benchmark against Sundials (and/or its methods) as a baseline which is the standard we tend to hold.
I'm also slightly salty because a recent paper (https://arxiv.org/pdf/2011.03902.pdf) claimed "Differentiable event handling generalizes many numerical methods that often have specialized methods for gradient computation,...", citing one of our results. We instead went through the trouble to find a semi-complete list of prior work, which would make it abundantly clear that what they claim as new is in fact traceable to a paper by Rozenvasser in ~1965 (we translated it from Russian to make sure), with a long history of more recent work. In fact I think DifferentialEquations.jl just supports this out of the box and has non-broken event-handling (their implementation tests against a product of jump conditions).
I am aware of those benchmarks :). I will probably adopt your library for this very reason, although I have to say that I really don't like Julia as a language. It is pretty clear that besides Sundials, PETSC and other C++ based libraries, you are rapidly becoming the only game in town. As much as I am tempted I really shouldn't be handwriting integration routines :). I don't think you need a random internet stranger to tell you that, but it is abundantly clear that DifferentialEquations.jl provides far more long-term value than these types of machine learning papers.
Has anyone successfully applied DEQs to larger-scale cognitive tasks or benchmarks, as opposed to MNIST, which is a tiny trivial task by today's standards?
Think ImageNet-1000, COCO, LVIS, WMT language translation, ..., SuperGLUE. There's a long list of datasets and benchmarks that regular boring fixed-depth NNs tackle with remarkable ease these days.
Has anyone anywhere applied DEQs to any of those datasets / benchmarks?
This really does not jive with my understanding. Each layer of, e.g. VGG 16 [1] does not implement the same function. Each layer has its own weights. There are certain architectures that tie weights across layers, but not all of them.
(I mean this as a question, but don't see an obvious place for a question mark...)