JDF.jl – Julia DataFrames serialization format
github.com
github.com
> Development Plans
> I fully intend to develop JDF.jl into a language neutral format by version v0.4. However, I have other OSS commitments including R's {disk.frame} and hence new features might be slow to come onboard. But I am fully committed to making JDF files created using JDF.jl v0.2 or higher loadable in all future JDF.jl versions.
It's beyond v0.4 and at v0.5.1 right now so it seems delayed, but the repo describes intentions to be a language-neutral format after the dust settles. I think we're just taking a snapshot too early in the timeframe to really know what this project is truly all about. And given that the author is an OSS dev for multiple languages, this intention makes a lot of sense.
The NY Fed have their main model of the economy in Julia (ported from Matlab)
Have you seen the type system used to generically dispatch matrices to GPUs or cpus.
Or the dispatching to give you auto-differentiation?
Maybe it's not your definition of "new ideas", but they are really useful and original.
Being fair, I have seen dozens of not hundreds of products and mass deployed projects written in similar languages. Even Haskell made its way to a lot of desktops in the Linux world, and Haskell is super niche research territory imo.
I guess I'll be curious once I install some software or an OS and see it brings in Julia in as a dependency or something. Otherwise I worry it's Matlab 2.0 with less of a mindshare...
Python is also a very bad language for many of the things it is used for (anything to do with numerics, e.g. machine learning) and having a competing format from a language which is, at least in that regard, far superior seems like a good thing.
It's a big stretch to say that a tiny Julia library which does a thing we can already do well in Julia and Python and just about any other language has any impact on the language ecosystems and how they compare to each other.
Sure, but it should.
I typically use Python for data-related tasks.
Julia is just okay. It lost its steam due to long beta, slow time to first plot after 1.0 release, and generally not being a strategically adopted language by ML innovation groups. A great example of what is versus what was hoped is John Myles White's work[0], where the headline is "worked on large community Python projects" but is one of the originators of Julia's fundamental data packages.
I've had lots of intelligent people who use code as a tool instead of as a craft -- for data folks, their craft is working with data and not the tooling around it. For folks like this, Python is simply a better option in today's environment.
By what metric? Me, being involved with the infrastructure of research projects, I see how much time and effort is wasted just by working around the sheer stupidity of Python and its tooling daily. Because researchers generally don't see the infrastructure work as important, they also tend not to associate the effort it takes to get the infrastructure decently functional with their own work. So, they tend to think that it doesn't matter what language environment they are using (and thus prefer the familiar one).
Inevitably, this spills into researcher's work anyways. One of the typical problems (related to Python tooling) I see is this: the project worked for a while with a set of dependencies with which it was initially created, and one day everything explodes because some dependency screw something up. Non-infra people usually have hard time figuring out what broke and how to fix it, so they choose the path of least resistance: either not build / test project automatically anymore (only on individual developer's machines) or try to "freeze" the dependencies, preventing the update from breaking the old (broken) stuff from being exposed, or "fixing" by mindlessly copying some StackOverflow "solution" that does some asinine thing, but allows for the project to keep "working".
In the end of the day, Python-based projects tend to have very short shelf life. Often by the time they near completion they are already so behind the "latest and greatest" that others, who might have wanted to adopt them, don't want to do that because they'd have to make special downgrades in their infrastructure to even try it. Larger Python projects are almost guaranteed not to work with other projects because of overlapping and conflicting dependencies. Making Python-based projects available to non-programmers is another painful experience that usually ends in a failure.
So, you might achieve higher development velocity on a particular stretch in your development journey, but you definitely didn't account for the whole thing, definitely not with an eye for sustainability / longevity.
EDIT: That said, I just looked up the TIOBE index and it appears to be growing. My subjective take is likely in error here!
I think Julia's adoption is still not at all comparable to Python or R, but neither of those languages were comparably popular at this point in their own lifetimes. I think people underestimate how slow change is here.
Julia absolutely could fail to compete in the long run, but it is still growing and it is improving.
I appreciate the clarification and correction. Apologies for muddying the water.
- It appears that there's substantial community effort invested in smoothing the pathway for scientific modeling and related ML. So that's one area where I expect anyone picking up Julia will see quick+substantial wins.
- Recent Julia versions (1.9, 1.10 on the way) have made tremendous progress in TTFX: https://julialang.org/blog/2023/04/julia-1.9-highlights/
- I happen to work on problems (with more of a computation/simulation flavor, rather than "data related") where Python is extremely painful and Julia happens to be the best fit by far (and rapidly getting even smoother!).
> I've had lots of intelligent people who use code as a tool instead of as a craft
Agreed -- so there's no point being fanatical about it. Folks can keep an eye out and periodically re-evaluate whether the ecosystem has reached maturity for their needs. For those who want to "use libraries", it will take a bit longer than those who want to "write libraries".
Julia aims to solve the two-language problem. If you live in Python and never have to touch another language like C or Fortran, then the value proposition of Julia is going to feel much more tenuous. But there are a set of people (library developers) who experience the ecosystem in the diametrically opposite way: https://twitter.com/dillonniederhut/status/16791406806799728...
> and generally not being a strategically adopted language by ML innovation groups.
I see the tremendous engineering effort being sunk to massage different (mutually incompatible) subsets of Python into some shape amenable for ML -- supporting new algorithms on top, and more efficient program execution at the bottom (translating from software to hardware).
I've heard the claim that it's sometimes better to delay a solution and let the users feel the pain before they open their minds to the solution. That is what I'm reminded of when I think of ML and Julia. If ML users don't see the value of Julia yet, that's okay -- they might once they dig themselves in deeper.
If they manage to solve all ML problems with Python (with C and what not), that's fine too! I think the world is a bigger place where there are also other interesting things going on, and Julia is helping a lot of people do things they couldn't accomplish otherwise :-)
--
There's no fundamental reasons for languages to "age", unless they happen to be tied to some unsuitable assumptions in how they model the world -- aspects that cannot be rearchitected without fundamentally changing the language. Barring that problem, languages only get more mature.
The problem with Python is not that it is three decades old, but that its revealed priorities (model of the world and software) is out of sync with many of today's needs. Even if we forget advanced things like ML, there's still a bunch of really basic stuff that make Python painful: https://twitter.com/dmimno/status/1679474354579488771
I think the design of Julia is much more robust in those aspects, but we also need to see how multiple dispatch plays out for very large codebases+teams. Meanwhile, I expect Julia will keep improving rapidly in the use cases it supports.
My point earlier was that people who solely use Python aren't DK exhibits in action simply because they know one language. Further, they may do so as a consequence of tech stack used in their roles or in their organizations. Given Python's adoption in the data domain, many people in the space use code simply as a tool to get things delivered.
Python is useful for people who have no formal training, but they are leveraging off the work of others. I hope you can understand how experts do not attach much weight to people still trying to figure out how computers actually work.
> experts do not attach much weight to people still trying to figure out how computers actually work
is a much weaker form of this point:
> It is impossible to have an intelligent discussion with someone who only knows Python, FSVO know. Dunning-Kruger in action.
The first point is true for anyone focused on a single language, outside of perhaps Assembly. The second is simply wrong.
NumPy does have an official documented format which is straightforward to read and write from other languages if you write some code yourself, but still doesn't have very wide support outside of Python.
I think there’s even an Excel plug-in for parquet these days.
While CSV comes with plenty of known problems, its one advantage, which apparently remains a significant one, is that it doesn't require any plugins or special software support at all, just a text editor. It seems that's a big enough feature that no attempt to fully replace it has really worked out. Amazing when you think about it, the power of plain text encoding.
Of course I agree that it's often not the best choice but somehow it remains incredibly useful and can even be hard to make arguments against it on a team project because it just kind or works well enough in a lot of cases and is simply a lowest common denominator when people can't agree on what format to use.
That is not a thing. The closest would be the NumPy binary format, which is "nothing to write home about". It's not good in any particular way. Some sort of a stop-gap solution for when you really want to save the contents of a NumPy array very much, and you don't want to deal with writing some commonly supported format that's (hopefully) optimized for some particular task.
That felt bit surprising.. is parquet slow to read in general or is the julia implementation just slow?
Acronyms are lossy compression. I’d bet there’s a JDF that even predates the one you reference in one of your descendant comments.
I’d be interested in knowing (though admittedly not enough to look) if there’s even a three letter acronym using the English alphabet that’s not already in use.