Embracing Swift for Deep Learning
fast.ai
fast.ai
They talked about using Swift for differentiable programming, is there something unique in Swift that enables that? From what research I've done it seems the techniques of Deep Learning are much more adaptable than the standard sklearn fit/predict template.
I was a bit surprised since fast.ai seemed to be big on PyTorch, it seems a paradigm switch to TensorFlow, but that's what research is all about.
In addition to the TensorFlow / Pytorch dichtomy there seems to be Julia's Zygote, still in early stages. Deep Learning is like The Thing, you look away from it for a few minutes and it's sprouted legs out of its head and walked off into a different room: https://en.wikipedia.org/wiki/Differentiable_programming
"Some day this Singularity's goin' end..."
See links below. Did you look at or consider Julia?
--
[1]:https://github.com/tensorflow/swift/blob/master/docs/WhySwif...
Later on from your link: "[We] picked Swift over Julia because Swift has a much larger community, is syntactically closer to Python, and because we were more familiar with its internal implementation details - which allowed us to implement a prototype much faster."
Honestly, I think the best argument (and the only one I buy) for Swift is the latter part: “…we were more familiar with its internal implementation details - which allowed us to implement a prototype much faster.” – just like I said about a year ago [1]. If you are sitting on a team deeply familiar and passionate about a language – Swift – what kind of managerial fool would not let them take a stab at it? Especially with Lattner’s excellent track record.
[1]: https://news.ycombinator.com/item?id=16939525
The ideas behind Zygote dates to somewhere around spring 2017, but I think it took about a year to hammer out compiler internals and find time to hack, so you are still right that nothing was public when Google settled on Swift – I think there been at least one Mountain View visit though over XLA.jl, but do not quote me on that one.
The race is still on and I am looking forward to seeing what all the camps bring to this budding field. I have worked with an excellent student on SPMD auto batching for his thesis project and we now have some things to show [2]. This is still a great time to be a machine learning practitioner and endlessly exciting if you care about the intersection between Machine Learning and programming languages.
[2]: https://github.com/FluxML/Hydra.jl
My only request would be for Jeremy to explain “Swift for TensorFlow is the first serious effort I’ve seen to incorporate differentiable programming deep in to the heart of a widely used language that is designed from the ground up for performance.” to me. Is it the “serious” and/or “widely used” subset where the Julia camp is disjoint? =)
There must be points when a technology takes off and becomes mainstream, what predates those points? That to me, this is the interesting question. In 2015 Torch (Lua) dominated the mindshare, why did TensorFlow succeed then? I think Lua itself caused it, lack of a coherent object model, etc. – sure as heck it was not the speed as I joked around by writing `import TensorFlow as TensorSlow` in my scripts for at least a year past the initial release. There was resistance against writing Python bindings for Torch, but in the end it happened and PyTorch was born; at this point TensorFlow dominated the mindshare. Why have PyTorch now become the favoured framework among all my colleagues and students then? Despite them being solidly in the TensorFlow camp prior to this. I think the answer is eager execution and the move in TensorFlow 2.0 to mimic exactly this speaks in my favour. So what would the Julia moment be then? If I knew, I would tell you, but I and several others in the Julia community are at least hard at work cracking this nut.
Just like I said a year ago, I am biased in favour of the bazaar and big tent that Julia represents. Swift will not have physicists, mathematicians, ESA employees, HPC people, etc. present and I would miss them as they bring wonderful libraries and viewpoints to the Julia bazaar. For example, I think TensorFlow with its initial graph model was designed that way precisely because it fit the mindset of non-practitioners with a compiler background – or perhaps it was the desire to “write once, deploy anywhere”? PyTorch could then take a big piece of the pie because they saw eagerness to be essential due to their academic/practitioner background. Only time will tell who is right here and I think me favouring Julia is not a safe bet, but certainly a reasonable one given the options available.
E.g. both Swift and Julia (and Rust, and Clang ...) have a common underlying backend (LLVM) in common, whereas in Common Lisp, it's ... Common Lisp all the way down, and it could potentially be a much better world, but it would just require duplicating too much stuff to be a viable candidate?
And for Julia in particular, there isn't really a need for bindings over C/C++ (the whole purpose of the language is not requiring another language for performance). Many of the most popular libraries for numeric computing are 100% written Julia, including Flux.jl [1] for Machine Learning and DifferentialEquations.jl [2]. Plus even real-time applications can be done with careful programming, since Julia allows writing programs with either no or minimal allocations, for example [3].
[1] https://github.com/FluxML/Flux.jl
[2] https://github.com/JuliaDiffEq/DifferentialEquations.jl
[3] https://juliacomputing.com/case-studies/mit-robotics.html
Herb Sutter has a quite good CppCon talk about it.
Many engineers prefer it due to cargo cult and anti-tracing GC bias, despite years of CS papers proving the contrary since the early 80's.
Just as one would take any marketing/promotional material for why a company's products are better than competitors with a grain of salt, and would want such material to be clearly marked as marketing material, I think bringing to light the author's relationship to the Swift helps, and does not hinder, further discussion.
If we choose to elect Swift as an ML language, we're handing keys to a certain platform vendor. I don't want to be tied down like that.
Keep using Swift for app development--it's great at that. But keep it far away from research and development.
Swift is open source
Not to mention the fact that mac doesn’t have any technology on the server, so if any code is supposed to reach production one day, it’ll have to at least run on linux.
EDIT: there is a new drop for macOS from April 18 which I am now downloading.
I really like modeling with TensorFlow but I have a like-sometimes/hate relationship with Python. If Swift support gets stable and well supported then that would float my boat.
That said, I remain optimistic given Jeremy's involvement — he/Rachael/fastai's been nothing but a source of good for ML in general.
To be clear, coming from mostly python/c++, I was looking for something like `import pathlib` or `#include <filesystem>`. Lot of modeling work I do involves boring things like moving files around, plotting them etc.
Again, its likely that I did not look around enough.
By the way, Jeremy Howard spends several minutes on one of his videos explaining the style and why despite seeming strange it is the right choice for fast.ai. I just can't remember which video, but I would guess it is in dl 2 2019 first video.
When decisions made by an algorithm can result in real monetary costs, one must always err on the side of caution. Only an irresponsible person would choose Python or other unsafe, dynamically typed language when implementing a crucial piece of business logic.
Sometimes python is appropriate for production code, weak type system or no. Sometimes it isn't. As a project gets larger, more complex, and more interconnected with other things, static type systems become more useful.
End of story, surely? Didn't we all know that already?
You can write perfectly fine production level code in Python. It like lego block. Start with simple and then add on.
Adding compile type checks comes with its own demerits. I guess here Swift is trying to offer more tool chain on compiler level for model building rather than just being type safe.
> recent versions of Python have added type annotations that optionally allow the programmer to specify the types used in a program. However, Python’s type system is not capable of expressing many types and type relationships, does not do any automated typing, and can not reliably check all types at compile time. Therefore, using types in Python requires a lot of extra code, but falls far short of the level of type safety that other languages can provide.
Currently, programmers are forced to pepper the code with lots of ignore annotation comments and redundant union types, just because the type checkers aren’t smart enough.
In actual languages people us e(e.g. that have more than 2% of the job market), type errors caught by the compiler don't catch "all computation" errors.
Not even in Haskell...
So genuinely wish to know what best practices you follow that makes developing large projects in dynamic languages make as inexpensive as typed languages to backup your comments.
Have you written 50K LOC C/C++? How do you manage finding errors like buffer overflows, use after free, and so on? And did the program have the same functionality as 50K Python, or had 1/10th the features?
(Also: you can run a linter and find typos, and you can have tests and know about bad types. And that's assuming those are actual errors people have and matter in the first place...)
I've written those huge codebases in java, where our weekly github commits pushed like 5-10k lines of code per user. It would have been worse if it weren't for Lombok, and totally untenable if it weren't for IDE-assisted refactoring. I've since worked on much larger and more impactful problems with smaller clojure teams and seen work consistently get done in 1/10th of the LOCs.
I used to be a fan of static types -- I still am in limited contexts, but I've come to realize that statically-typed languages (especially object-oriented ones) usually lead to projects that _require_ static typing to even be maintainable. In contrast, a philosophy common to functional languages (which don't always have to be dynamic) is:
> "It is better to have 100 functions operate on one data structure than 10 functions on 10 data structures." —Alan Perlis
I think discussions focused around type systems often miss the context of what languages' type systems we're actually talking about, and that context is way more important than making sweeping judgments about types.
I can count _on one hand_ the number of times I've defined a custom type in clojure in the 3 years I've been writing it professionally. Yes, sometimes I've called a hash map function on some kind of a list, but that cost has felt really small (and caught early) compared to the cognitive cost of keeping track of class hierarchies, interfaces, access rules, and annotations. And for what win? I still get NPEs, ClassCastExceptions, and RuntimeExceptions in Java, but now I have a whole class of errors and processes introduced around playing well with its inheritance model.
So, great, it's stopping me from directly instantiating an abstract ApiClient class, and quickly catching that I tried to get an HttpApiClient from an HttpsApiClientBuilder.build() call -- but that pattern would not have existed to present an issue were it a different language