If you draw the distinction between Julia the language and Julia the ecosystem, then that's not a damning verdict for the language, even if we accept the premise that you can't trust the computations you mentioned.
The way that Julia composes is unprecedented. Typical examples include computing with measurement uncertainty, or big floats through functions that were not explicitly designed to support those types. This is great, but the correctness of this composition cannot be assumed.
One good proxy for how this composition occurs is in many type combinations. As Here is a concrete, made-up example. There is a function you've been using without any problem that takes an argument of type `DenseArray{Float64, 2}`. (Meaning, the function gets called with an argument of that type. The function definition is not constrained to work on that type.)
If you've never focused on the correctness (e.g. reviewing the code, unit tests) of this function to see what it does on `SparseArray{BigFloat, 2}`, I would not take it for granted that it works.
This is perhaps too big a burden on a user. Some might say they rather not have the compositionality unless it's guaranteed to be correct. But, if you're actually asking me to imagine using Julia in a situation where I really want to be sure about the correctness of a Julia program, I would at the very attempt to characterize all the expected types to go into any function. If at runtime a function is called with a new type, I would log that, and possibly error out.
Note this is related to, but not equivalent to what's necessary to do "static" compilation in Julia.
Language and ecosystem are so intertwined that I almost never do this. Maybe when there are multiple well developed ecosystems, but when there’s only one, then it and the language are inseparable to me (for any language, not just Julia).
But the original critique also mentioned several correctness issues in the core language itself, so the issue does seem to go deeper.
Also depends what you count as correct, see undefined behavior.
I feel that should be the job of the authors to think about what correctness means for their application, and under what circumstances they can give a correct answer. Ideally, also warn the user when the inputs don't fall within that domain or give them options to handle such cases.
Let's talk about an example: say you have a neat algorithm for fast multiplication of sparse matrices, approximately correct up to a defined Frobenius error. Say we have some reasonable treatment of dimension mismatch and NaN, inf,... with option flags already. Now if I wanted to know whether your algorithm has the same numerical guarantee for a dense matrix, I'd have to read your paper, run some tests, ask around, and I might never be quite sure.
It seems much more sane to put the onus on the authors to be explicit about the behaviour in that case. Your and others' comments seem to put the onus mostly on the user for thinking about the behaviour of others' code. That does not seem like a good practice, especially in the context of complex math and nested composition.
plus_one(x::Number) = x + 1
function plus_one(x)
@warn "untested type usage of plus_one"
x + 1
end
That's cumbersome, but a macro could make it less annoying.I think this is it. Or rather, I think to some people, myself included, compositionality should imply correctness. If you /can/ compose things, but the result is wrong, can you really say that you've got composition?
How useful it is really to be able to plug in your own types in other peoples libraries if you have to trace through the execution of that library, figure out which libraries they transitively use, to ensure that all of your instantiations are sound? How do you even test this properly?
It's a really hard problem, and from what I can tell, Julia gives you no tools to deal with this.
[0] is probably relevant here, although I'm not sure I share the positive outlook.
- easy unrestricted composability
- lack of well-defined interfaces
- lack of an effective way to use the strong typing system to validate correctness
is not a very good idea.
Some formalized notion of "interface" is certainly important, but it seems no clear formulation has emerged.
I think it's fairly consistent with the dynamic nature of Julia and fairly inconsistent with the static nature of Julia. I don't know what a good solution to the interface problem is.
I think one can get pretty far with writing tests to check interfaces. If you are a library that expects user-defined types, you can expose an interface to these tests so that a user can check if they've implemented everything.
This is a very generic approach, and aside from the key limitation of only giving results at runtime, is much more powerful.
If you don't want composition, then there's no issue and Julia can be as weak/strong as other languages.
If you do want composition, then I see two ways (but I'm sure there are others): you do the "typical" thing with glue code or you use the more "automatic" way that Julia provides. Which one is better? If this is too subjective: Which one is more correct?
Yes, Julia can propagate errors in unexpected ways, but how would you implement this in another language? You'd probably have to spend X hours writing glue code and Z hours writing tests to make sure your glue code is working. This also raises issues with maintainability when one of the two packages you're connecting /composing changes.
Julia offers a reduction on the X hours for writing glue code (sometimes with X = 0) and maybe a similar time Z writing test code. The maintainability, I'd argue, becomes easier.
The cost is the unknown unkowns that might creep up when doing this composition. My (extremely) subjective sense is that this doesn't happen that often to me (I don't usually pass approximation intervals to sparse matrix to auto differentiation to neural networks), which means the benefits outweight the cost in this regard. YMMV
Edited a few things for clarity
Two points
1) the glue code you HAVE to write to make the composition in other languages and the mantianability issues you get also may introduce wrong results 2) MAY doesn't mean it will and we trade-off speed / convenience and risk in many other areas. Surgeons (or the health system in general) do trade off % of success for cost / speed, just because they don't have the resources to do everything / spend 100 hours on everyone
Bugs are everywhere - including Python and R libraries. You would not design a safety critical system in Python using libraries and codepaths that are not heavily tested. Even then, you still have to carefully design testsuites that give you confidence that what you are computing is what is expected.
See these talks on how a collision avoidance system was designed with Julia. https://www.youtube.com/watch?v=rj-WhTL_VXE https://www.youtube.com/watch?v=19zm1Fn0S9M
At the same time, correct implementation of statistics is hard and testing it to be correct is harder.
One of the working groups is dedicated toward submissions: R Submissions Working Group [1]. They recently announced: "The R Consortium is happy to announce that on Nov 22nd, 2021, the R Submissions Working Group successfully submitted an R-based test submission package through the FDA eCTD gateway! The submission package has been received by the FDA staff who were able to reproduce the numerical results." [2]
[0] https://www.r-consortium.org/all-projects/isc-working-groups
[1] https://rconsortium.github.io/submissions-wg/
[2] https://www.r-consortium.org/blog/2021/12/08/successful-r-ba...
Additional note, we recently did a overhaul of SciMLSensitivity (https://sensitivity.sciml.ai/dev/, as part of the new https://docs.sciml.ai/dev/) and setup a system which amounts to 15 hours of direct unit tests doing a combinatoric check of arguments with 4 hours of downstream testing (https://github.com/SciML/SciMLSensitivity.jl/actions/runs/25...). What that identified is that any remaining issues that can arise are due to the implicit parameters mechanism in Zygote (Zygote.params). To counteract this upstream issue, we (a) try to default to never default to Zygote VJPs whenever we can avoid it (hence defaulting to Enzyme and ReverseDiff first as previously mentioned), and (b) put in a mechanism for early error throwing if Zygote hits any not implemented derivative case with an explicit error message (https://github.com/SciML/SciMLSensitivity.jl/blob/v7.0.1/src...). We have alerted the devs of the machine learning libraries, and from this there has been a lot of movement. In particular, a globals-free machine learning library, Lux.jl, was created with fully explicit parameters https://lux.csail.mit.edu/dev/, and thus by design it cannot have this issue. In addition, the Flux.jl library itself is looking to do a redesign that eliminates implicit parameters (https://github.com/FluxML/Flux.jl/issues/1986). Which design will be the one in the end, that's uncertain right now, but it's clear that no matter what the future designs of the deep learning libraries will fully cut out that part of Zygote.jl. And additionally, the other AD libraries (Enzyme and Diffractor for example) do not have this "feature", so it's an issue that can only arise from a specific (not recommended) way of using Zygote (which now throws explicit error messages early and often if used anywhere near SciML because I don't tolerate it). tl;dr: don't use Zygote.params, the ML AD devs know this its usage is being removed ASAP.
So from this, SciML should be rather safe and if not, please share some details and I'd be happy to dig in.