Every programming language doesn't need to become _the_ language to do something. They are experiments in how to best express what you want to compute. Even if Julia never takes off, they explore multiple directions other languages might want to implement - multiple dispatch for polymorphism, nested parallelism, macros so you can create DSLs from regular Julia code, and so much more. Asserting that Julia is only successful if everyone is using it is just super reductive.
I think it's just a reaction to constantly seeing Julia being oversold and overmarketed in the geekosphere.
There _are_ legitimate complaints about the language (or more exactly, its runtime). Julia 1.9 is a significant milestone on addressing those.
Maybe Julia will not suceed on overtaking anything, but that is not because of a lack of merit.
I'm old enough to remember when every Python came with the mandatory "oh no, significant whitespace" post. Now it's just the most mainstream language.
“Well, Ive got nothing substantive to say, so I better start whining about surface level syntax stuff I don't like”
There was recently a tutorial on Julia for biologists in Nature Methods:
https://www.nature.com/articles/s41592-023-01832-z
That said, I imagine things might get interesting again if/when Modular open sources their Mojo/"Python" compiler.
To me, it looks like, to the contrary, that Python is so full of counter-intuitive bits... like doing a[begin:end+1] to take a slice.
A few years ago, I worked in a neuroimaging research lab and taught some basic Python programming to quite a few research assistants with psych degrees. At the time, a scientific computing or intro to programming course wasn't part of the psych curriculum at the university that the lab was. 0-indexed arrays were often a big sticking point. It just didn't mesh with their intuition.
For example, if `len(some_list)` returned 2, they would try to get the last value with `some_list[2]` and then get confused. Or for another example, if they used `for i in range(n)`, they'd be confused when `i` didn't start at 1 and end at n. They eventually caught on and learned further syntax such as `some_list[-1]` for getting the last value in a list. But, it was a bit of a frustrating experience for both them and myself because I only had time to teach them the basics and some google-fu as well as give them some example code. While 0-indexing was a source of a lot of their bugs in the beginning, I learned that the best method for them to get the hang of it was to encourage them to do a lot of print debugging, which was more than adequate for any scripts they'd be writing. They just really needed more feedback from what they wrote and what it actually did to get a better grasp of the idiosyncrasies of programming vs. their intuition.
Side note: The frustration of teaching non-techies Python was nothing compared to teaching them LaTeX. I was luckily able to eventually convince the PI to adopt Overleaf and organizing the text into separate files and using `\include` or `\input` to reduce the headache of collaborative writing as well as move all the internal docs to markdown. The latter became a huge help when we eventually moved to Jupyter notebooks for a lot of our python-related scripts as the team members already became adept enough with markdown. It was also easier to teach everyone how to use pandoc for making pretty PDFs of the internal docs or converting their markdown drafts into LaTeX.
It is all a matter of what one is used to. Quoting myself:
> I guess counter-intuitiveness is in the eye of the beholder.
1-indexed arrays are not common in programming languages. There is nothing wrong with them, but most languages are 0-indexed, so for most programmers a 1-indexed language is something slightly unexpected.
For non-programmers it might be different.
> Given the premices on which the Python language was designed (giving access to programming to non-programmers)
I had college lectures using 0 as start index, others using 1, both were fine, no problem.
However taking a slice with [start:end+1] is unequivocally counterintuitive.
I see your point, but personally find it a matter of preference, not of objective truth.
`range(end+1)` is consistent with that slice notation that you disapprove of.
At the same time, on 0-indexed vs 1-indexed you seem to have no preference (I prefer 0 for historical reasons).
How about `len(items)` vs `items.length`? I have my own preferences there, but I would not say any or the other is intrinsically better.
Anyway, I already wrote way too much about such a small syntax thing :)
They could have added, in parallel to range(start,end+1), some interval() function that would adopt the more sensible syntax interval(start,end), and then design the slice syntax on this new interval() one.
And to begin with, why on Earth is range() using the syntax range(start,end+1)? I honestly can't understand why it was so important for the language designers to calque it on the (barely) most common idiom in C "for (int i=begin; i<end+1; ++i)" which can by the way very well be replaced by "for (int i=begin; i<=end; ++i)".
While they were at designing a sugar syntax for slicing (and the same applies for range), why didn't they internalize the "i<end+1" quirk by doing the conversion in the implementation code???
> How about `len(items)` vs `items.length`?
This, poses no problem to me. Just a different language, nothing hurting my sense of logic like [start:end+1] does. Same for 0 vs 1, just convention, fine by me.
> measure theory
First of all, even for a specialist, that extremely hard to be delusioned into thinking we're doing measure theory, when we're manipulating indices of arrays/matrices/tensors.
But anyways, assuming you think it's measure theory. I'm trained in the field and that's impossible to expect half-open intervals in this particular place!
To the contrary: the support of base elements of major functional spaces used in measure theory textbooks are always compacts.
I didn't mean to advocate at all... Just explained that I never peronally felt the weirdness of it.
The point about integers is taken but Python style also gives you len(range(n, n+k)) == k and len(range(n, k)) == n - k which is very convenient at times.
> To the contrary: the support of base elements of major functional spaces used in measure theory textbooks are always compact.
You can take either type of intervals to define Borel sigma algebras.
Half-open intervals, however, additionaly form a semiring while finite unions of half-open intervals additionaly form a ring - the starting point of Caratheodory's extension theorem [1] which is kind of essential at least in some expositions.
[1]: https://en.m.wikipedia.org/wiki/Carath%C3%A9odory%27s_extens...
In PyTorch you have a graph that is created on runtime by connecting the operations together in a transparent manner.
Jax may feel a bit magic, but all that's done is sending / splitting tracers and recording the operations and compiling; by limiting the language, you have controlled branching with the proper semantics.
---
The main reason I have had such with Julia is simply because of how early / soon I ended up needing to use them whereas with the other languages, you can get away without getting into the very messed up things.
You've jumped the shark here mate because autodiff in PyTorch is implemented using compile-time generated code in libtorch - it's not only the very definition of opaque but also pretty close in spirit to a macro.
What is opaque is jax's tracer objects in the sense that the system exists, is there, the user is aware that it is there, and you can't peek inside it without actively going out of your way.
With PyTorch you can implement your own functions with their own forward and backward passes; whether the underlying operations use libtorch is irrelevant, that is simply a backend.
The computation graph __is__ created at runtime, whether the scaffolding for this is done in libtorch or directly in python is not relevant.
lots of your other claims are just complete misunderstandings e.g. whether the graph is created at runtime - it's not and you can read adam's original paper to find that out if you don't believe me - and the idea that libtorch is just a backend - it's not - it is, for all intents and purposes, pytorch.
I work for AWS; these are the definitions we use, more or less.
https://stackoverflow.com/questions/17384020/what-do-transpa...
> whether the graph is created at runtime
I was specifically referring to the computation graph of the model that is used by autograd.
Taken from [1,2]:
> Instead, PyTorch uses the operator overloading approach, which builds up a representation of the computed function every time it is executed.
> Internally, a Variable is simply a wrapper around a Tensor that also holds a reference to a graph of Function objects. This graph is an immutable, purely functional representation of the derivative of computed function; Variables are simply mutable pointers to this graph.
> A Function can be thought of as a closure that has all context necessary to compute vector-Jacobian products. They accept the gradients of the outputs, and return the gradients of the inputs (formally, the left product including the term for their respective operation.) A graph of Functions is a single argument closure that takes in a left product and multiplies it by the derivatives of all operations it contains. The left products passed around are themselves Variables, making the evaluation of the graph differentiable.
But what do I know?
> libtorch is just a backend - it's not - it is, for all intents and purposes, pytorch.
Sure, libtorch is PyTorch; that doesn't mean one can't take the semantics of PyTorch and implement them in just numpy backed python. Doesn't mean it will be fast.
[1] PyTorch: An Imperative Style, High-Performance Deep Learning Library
[2] Automatic differentiation in PyTorch
you're just being asinine - we're literally talking about binary code that's never seen by anyone that doesn't compile from source and goes digging around in the build dir - how could you possibly call that code "transparent" in any sense of the word? are blob drivers also transparent according to these "AWS" definitions?
>I was specifically referring to the computation graph of the model that is used by autograd.
it's literally right there in bolded text on the first page of the original paper (the 2017 neurips paper):
>Immediate, eager execution. An eager framework runs tensor computations as it encounters them; it avoids ever materializing a “forward graph”, recording only what is necessary to differentiate the computation
autograd has absolutely nothing to do with the graph - autograd is literally 10s of thousands of lines of generated, templatized, code that connects edges one op at a time. you can argue with me all you want or you can just go to repo tip and see for yourself https://github.com/pytorch/pytorch/blob/main/aten/src/ATen/t...
Either I am an absolute fucking moron or this reads that PyTorch keeps track of a graph.
Please explain.
>PyTorch (and Chainer) *eschew* this tape; *instead*, every intermediate result records only the *subset* of the computation graph that was relevant to their computation.
>Either I am an absolute fucking moron
you said it not me
If my comments implied that, then I am sorry for miscommunicating, but by calling it "dynamic" I meant that there is no compiling of a tape or tracking it as done with TF and the likes.
> I nternally, a Variable is simply a wrapper around a Tensor that also holds a reference to a graph of Function objects. This graph is an immutable, purely functional representation of the derivative of computed function; Variables are simply mutable pointers to this graph (they are mutated when an in-place operation occurs; see Section 3.1). A Function can be thought of as a closure that has all context necessary to compute vector-Jacobian products. They accept the gradients of the outputs, and return the gradients of the inputs (formally, the left product including the term for their respective operation.) A graph of Functions is a single argument closure that takes in a left product and multiplies it by the derivatives of all operations it contains. The left products passed around are themselves Variables, making the evaluation of the graph differentiable.
Please explain how it doesn't hold a graph when it literally says it is holding one, dynamically created and kept alive via ref counting.
I am losing my mind here.
Such what? Fun? ;)