Julia's Arraypocalypse Release
github.com
github.com
Within a few weeks, Julia 0.4 will be shipping, which is not focused on changes to arrays.
1. Array slices will create views rather than copies. This is just a matter of performance – although possible, it is currently way too annoying to write copy-free array code. This change will make it the default.
2. Slicing will behave more like APL: the rank of a slice of an array will be the sum of the ranks of the slices. This is a regular and powerful scheme, but sometimes at odds with how slicing works in Matlab.
These changes will make the language easier to write high-performance, general code, while, with some care not making linear algebra harder.
Except for breaking any existing code which mutates slices (unless slices are COW?), and implicit sharing can have disastrous effects on memory usage.
A small long-lived slices keeping a original arrays alive implicitly. Oracle recently undid the same optimisation in Java's strings because of that exact issue (which obviously broke code the other way around, code relying on cheap substrings would now take O(n) instead of O(1)). Erlang developers are also warned early and often about that behaviour in binaries.
Still something that should be in bold text in Julia 0.5's readmes, though; slicing is cheap now, but it comes with this very important caveat!
Iterating over the columns of an in-memory matrix, for example, is the use case that this change is meant to speed up.
Like advanced indexing (http://docs.scipy.org/doc/numpy/reference/arrays.indexing.ht...) in the Python world?
I am saying this change is not quite "just a matter of performance", that the documentation and changelogs should be very very clear about the behavioural changes and the possible issues with the new behaviour, and that there should be an easy way to "unshare" slices and force a copy.
2. there are gradations in breaking changes, the breaking changes I outline are insidious because they are silent, they won't cause the program to fail loudly, but they will cause it to behave very differently than it used to and possibly to corrupt data.
As it stands if you translate Fortran into Julia without a deep understanding of what is happening below the surface, your code will be much, much slower.
I've tried out Julia a few years back, but didn't find it ready for serious use (yet). Am I right to assume that this means the language is still in serious flux?
For most users, this should not change much. However, it gives us the ability to do some really amazing things with our array infrastructure going forward.
What kind of things? Super curious.
It is still a bit wanting.
But, a lot more remains to be done, and contributions would be really welcome in these areas that take Julia's statistical computing to the next level.
The amazing thing about Julia I've come to enjoy is how clean the source code, how easy it is to find functionality, and it's all written IN Julia. That means I can modify libraries that are normally C or C++ for perf reasons. It's changed the prospect of needing to add functionality from one of hesitation to feeling that it's actually enjoyable.
The key is to get it all done sooner so that we can keep getting that much closer to 1.0.
In short, there is an inherent disconnect between arrays-as-containers and matrices-as-linear-algebra-operators. There are many ways that you can define array operations such that the universe is closed and self-consistent, but there always seems to be some awkwardness. Right now, Julia uses some Matlab-ish semantics (sometimes ignoring trailing singleton dimensions), but there is an interaction with Julia's types that makes the status quo a little strange.
So there are "just" two major breaking changes proposed here: changing our array semantics universe to be more APL-like (instead of Matlabish), and making indexing return a view instead of a copy.
We had no idea where Julia would go. In some sense, it has been a personal journey of maturity in my own thinking, and if I can take the liberty to say so, the same is true for the community as well.
I'm not from a Matlab/Octave background, so one-based indexing feels really foreign to me.
Btw, Matlab also uses 1-based indexing, no? I thought that Julia chose 1-based indexing to get brain backwards compatibility with Matlab.
I think zero based is "prettier" because of half open intervals, but it really doesn't matter.
Wouldn't it be possible to create array classes in Julia that all indexed zero-based? I would welcome a 'Julia0' fork :-)
Coming from a regular programming background, as opposed to a technical computing background, can certainly make this look alien. However, I think it starts feeling natural quickly. There are many other things in Julia that can take a lot longer to get used to.
Edit: For a history of 0-based indexing see http://exple.tive.org/blarg/2013/10/22/citation-needed/ My understanding is that with modern compilers, there's no need for 0-based indexing; you can use 0-base, 1-base, or 69-base indexing as you prefer. 0-base is applicable to people creating chips or writing hardware drivers. Most scientific programmers are not so closely tied to hardware; 1-based indexing is one of the things that helps make Julia more accessible to high-school programmers or PhD students who don't want to learn binary arithmetic.
It's never been needed, Pascal had arbitrary array bounds (including negative) 50 years ago.
> 1-based indexing is one of the things that helps make Julia more accessible to high-school programmers or PhD students who don't want to learn binary arithmetic.
What does binary arithmetic have to do with anything? And have there been any study pertaining to the accessibility of 0-indexed versus 1-indexed arrays to high-scool programmers which would warrant such an assertion?
One based indexing counts, or labels boxes. Counting implies no "holes", but abandons any implication about how they are laid out in storage. Labelling allows treating numerical arrays as sparse hash tables, with as many "holes" as needed. It's a high level view, suited to mathy ways of thinking about data.
If I want to build a hash table, a power of two size for the underlying array is a nice performance optimization since you can use a bitmask instead of a modulo operation, and it makes a big performance difference, what does one-based give you other than some nebulous "it just _seems_ simpler to think about" kind of nonsense? I love how we expect people to grok the type system but we're afraid zero-based arrays will scare mathematicians off.
See: https://en.wikipedia.org/wiki/Modular_arithmetic#Integers_mo...
In my experience, Julia is uninteresting to most statisticians because of the lack of mature libraries and because R is already good enough for most of the stuff they want to do. I don't think that relatively small language changes like the one discussed in the original issue are even on their radar, and I doubt that these changes are driving away people that would otherwise be maintaining packages.
Hopefully Julia can work out all the language features and mature into a good technical language. It would also help if there is a guide about the canonical programming style. I prefer static typing but thought using type annotation I could come close to static typing while still retain the freedom of quick-and-dirty code. I believe for the current Julia I was wrong.
Can you provide some concrete examples of the functionality you're looking for? What exactly does "quick-and-dirty" code look like to you?