Swift: Google’s Bet on Differentiable Programming
tryolabs.com
tryolabs.com
Who knows, perhaps this will become the greatest thing since Lisp... What do I know.
The greatest possible barrier to adoption.
And speaking of Lisp - wasn't symbolic differentiation a fairly common thing in Lisp? (basically as a neat example of what you can do once your code is easy to manipulate as data).
Also, many optimizers that are popular in ML only need gradients (in which case the Jacobian is just the gradient vector). Second order methods are important in applications with ill-conditioning (such as bundle adjustment or large-scale GPR), but they have lots of exploitable structure/sparsity. The situation is not nearly as dire as you suggest.
Somehow, I ended up with a calculus equation for determining the right number of bits per entry and rounds to do to winnow the lists, for any given pair of machines where machine A found n entries and machine B found m. But I couldn’t solve it. Then I discovered that even though I did poorly at calculus, I still remembered more than anyone else on the team, and then couldn’t find help from any other engineer in the building either.
Eventually I located a QA person who used to TA calculus. She informed me that my equation probably could not be solved by hand. I gave it another day or so and then gave up. If I couldn’t do it by hand I wasn’t going to be able to write a heuristic for it anyway.
For years, this would be the longest period in my programming career where I didn’t touch a computer. I just sat with pen and paper pounding away at it and getting nowhere. And that’s also the last time I knowingly touched calculus at work.
(although you might argue some of my data vis discussions amount to determining whether we show the either the sum or rate of change of a trend line to explain it better. The S curve that shows up so often in project progress charts is just the integral of a normal distribution, after all)
Only if you treat it as a tree, not a DAG.
edit: Sorry no, it's still linear, even for a tree.
Smashing your ML models into git or other text-based VCS probably isn't the best way to do it
https://dvc.org/doc/understanding-dvc/related-technologies#g...
Mind you - DVC seems to be a platform on top of git rather than a replacement for git. So I'd argue that it's not really a new revision control system.
In essence there are cases outside the well developed uses (CNN, LSTM etc.) such as Neural ODEs where you need to mix different tools (ODE solvers and neural networks) and the ability to do Differentiable Programming is helpful otherwise it is harder the get gradients.
The way I can see it being useful is that it helps speed up development work so we can explore more architectures, again Neural ODEs being a great example.
Erik Meijer gave two great talks on the concept.
Differential programming is about building software that is differentiable end-to-end, so that optimal solutions can be calculated with gradient descent.
Probabilistic programming (which is a bit more vague) is about specifying probabilistic models in an elegant and consistent way (which than then be used for training and inference.)
So, you can build some kinds of probabilistic programs with differential programming languages, but not vice versa.
If we already have useful phrases like "embedded programming", "numerical programming", "systems programming", or "CRUD programming", etc, I'm not seeing the pretentiousness of "differentiable programming". If you program embedded chips, we often call it "embedded programming"; likewise, if you write programs where differentials are a 1st-class syntax concept, I'm not seeing the harm in calling it "differentiable programming" -- because that basically describes the specialization.
>the quotes are ridiculous ("Deep Learning est mort. Vive Differentiable Programming.";
Fyi, it's a joking type of rhetorical technique called a "snowclone". Previous comment about a similar phrase: https://news.ycombinator.com/item?id=11455219
The whole point is you’re supposed to have the same “something” on both sides (X is dead, long live X), to indicate it’s not a totally new thing but a significant shift in how it’s done
The most well known one:
> The King is dead. Long live the King!
If you change one side, you’re removing the tongue-in-cheek nature of it, and it does sound pretty pretentious.
> Le Roi (Louis ##) est mort. Vive le Roi (Louis ## + 1) !
Using the sentence with Deep Learning and Differential Learning just suggests that Differential Learning is the heir/successor/evolution of Deep Learning. It does not imply that they are the same thing.
As a French person used to the saying, Le Cun probably meant that.
... where did I imply it means they're the same thing?
From my comment:
> indicate it’s not a totally new thing but a significant shift in how it’s done
You could say, an evolution?
-
A snowclone is a statement in a certain form. The relevant form that English speakers use (I'm not a French person, and this is an English article) is "X is dead, long live X", where both are X.
That's where the "joking" the above comment is referring to comes from, it sounds "nonsensical" if you take it literally.
If you change one X to Y, suddenly there's no tongue-in-cheek aspect, you're just saying "that thing sucks, this is the new hotness".
I suspect the author just missed that nuance or got caught up in their excitement, but the whole point of a snowclone is it has a formula, and by customizing the variable parts of that formula, you add a new subtle meaning or tint to the statement.
The saying works because, as you say, it suggests a successor, but the successor has to use the same title, because what people want is a new king, so that nothing changes and they can live as they did before the king was dead, not a revolution with a civil war tainted in blood.
If you do volontarily change the title, it's because you think the new one will be better, which is pretentious.
Differential programming would be less flashy but may be confusing.
I wouldn’t actually be interested in this topic much except for the top level comment complaining about them wanting a new version control system for this and now I’m a bit curious what they’re on about this time, so will probably get sucked in.
What do you think will have more impact on the economy, crypto or cyber?
If you are talking to someone about cryptocurrency, referring to it as crypto later in the conversation in context is perfectly valid and doesn't lessen the meaning.
I do however agree with you when outside of it's context that these shortened names are horrible and effectively buzzwords.
Really? I seem to shoot myself in the foot a lot with jupyter notebook. I can't count the number of times my snippet was not working and it was because I was reusing some variable name from some other cell that no longer exists. The amount of bugs I get in a notebook is ridiculous. Of course, I'm probably using it wrong
Jupyter notebooks are painful to read, the allow you to do silly stuff all too easily, they’re nightmarish to debug, don’t play well with git, and almost every. Single. One. I’ve ever seen my teammates write eschewed almost every software engineering principle possible.
You’re not using them wrong, they shepard you to working very “fast and loose” and that’s a knife edge that you have to hope gets you to your destination before everything falls apart at the seams.
That’s my issue with Jupyter notebooks, between them and Python they implicitly encourage you to take all kinds of shortcuts and hacks.
Yes, it’s on my teammates for writing poor code, but it’s on those tools for encouraging that behaviour. It’s like the C vs Rust debate right: yes people should write secure code, and free their memory properly and not write code that has data races in it, but in the majority of cases, they don’t.
Quality engineering mostly comes from people, not languages. It is about your own personal values, and then the values of the team you are on. If there were a magic bullet programming language that guided everyone away from poor code and it did not have tradeoffs like a hugely steep learning curve (hi Haskell) then you would see businesses quickly moving in that direction. Such a mythical language would offer a clear competitive advantage to any company who adopted it.
What you are looking at really is not good vs. bad, but tradeoffs. A language that allows you to take shortcuts and use hacks sounds like it could get you to your destination quicker sometimes. That's really valuable if your goal is to run many throw-away experiments before you land on a solution that is worth spending time on improving the code.
a) ugly, unlabelled plots
b) including tons of uninteresting code that labels the axes (etc)
c) putting the plotting code into a separate module.
There are some extensions that do help with this, but extensions also kinda defeat the whole purpose of a notebook.
I agree, and this is why I built -
- ReviewNB - Code review tool for Jupyter notebooks(think rich diffs and commenting on notebook cells)
- GitPlus - A JupyterLab extension to push commits & create GitHub pull requests from JupyterLab.
There are a ton of tools trying to fill this void, and they usually provide things like the comparison of different metrics between models versions, which git doesn't provide.
There is definitely some need for an EDSL of some sort, but I think a general method is pretty useless. Being able to arbitrarily come up with automatic jacobians for a function isn't really language specific, and usually much better results are obtained using manually calculated jacobians. By starting from scratch you lose all the language theory poured into all the pre-existing languages.
I'm sure there'll be a nice haskell version that works in a much simpler manner. Here's a good start: https://github.com/hasktorch/hasktorch/blob/master/examples/...
I think it's pretty trivial to generalize and extend it beyond multilinear functions.
Swift was railroaded in Google by Chris Lattner, who has since left Google and S4TF is on death watch. No one is really using it and it hasn't delivered anything useful in 2.5 years
This is not what Google has found, actually. Teams who wanted to use this for research found that a static language is not flexible enough when they want to generate graphs at runtime. This is apparently pretty common these days, and obviously Python allows it. Especially with JAX that traces code for autodiff
Or has at least found that existing solutions for statically typed "differentiable" programming are ineffective, and I'd agree.
But having some way to check types/properties of tensors that you are doing operations to would really help to make sure you don't get your one hidden dimension accidentally switched with the other or something. Some of these problems are silent and need something other than dynamic runtime checking to find them, even if it's just a bolt-on type checker to python.
There are a lot of issues with our current approach of just using memory and indexed dimensions. [0]
Sure, the 90's were a rough period for them, but I think a series of failed OS strategies and technical debt are more responsible for that than just what language they used.
You could argue that their ambitions re Swift scaling from scripting all the way to writing an entire OS might never grow substantially outside Apple, but there's also the teaching aspect to think about.
"Objective C without C" removes a whole class of problems people have in just getting code to run, and I'll bet it shapes their mind in how they think about what to be concerned about in their code v what's just noise.
I would add that verbatim text is sometimes hard to read because it doesn't wrap long lines and small screens require the reader to scroll horizontally so try not to use it for large/wide blocks of text.
Also bullet lists are usually written as separate, ordinary paragraphs for each item with an asterisk or dash as the paragraphs first character.
Google is still taking themselves very seriously while everyone else is starting to get bored.
The problem with being 25 is that you have about 8 years ahead of you before you figure out how full of shit everyone is in their twenties, and maybe another 8 before you figure out that everyone is full of shit and stop worrying quite so much about it.
Make that ±80…
You can use CasADI to do automatic differentiation in C++, Python and Matlab today:
Tight integration with the language may be beneficial in making it simpler to write, but its not like you can't do this already in other languages. Baking it into the language might be useful to make it more popular. Nobody should be doing the chain rule by hand in the 21st century.
Boy there is some truth right here. It wasn't until long after I graduated undergrad that I realized just how much bullshit is out there. Even in the science world! When I started actually reading the methodology of studies with impressive sounding conclusions, I realized that easily 30-60% were just garbage. The specific journal really, really matters. I'd say 90% of science journalism targeting laymen is just absolute bullshit.
I started actually chasing down wikipedia citations and OMG they are bad!! Half are broken links, a large fraction don't support the conclusions they're being used for, and a massive fraction are really dubious sources.
I realized that so many people I respected are so full of shit.
I realized that so many of MY OWN OPINIONS were bullshit. And they STILL are. I hold so few opinions that are genuinely well-reasoned and substantiated. They are so shallow.
Yet, this is just how the world works. Human intuition is a hell of a drug. A lot of the people I respect tend to be right, but for all the wrong reasons. It's SOOOO rare to find people that can REALLY back up their thinking on something.
Those wolves were the ones buying N95 Masks in January; Buying Flonase (OTC Glucocorticoid) 'just in case'.
If the alternative non-BS opinion/belief is un-popular (i.e. coronavirus is serious), it's easier and safer to just check out; tend your own garden.
http://www.artandpopularculture.com/Suus_cuique_crepitus_ben...
For those who actually spent time with Swift, and realize its value and potential, consider yourself lucky that large portions of the industry have an ill informed aversion to it. That creates opportunity that can be taken advantage of in the next 5 years. Developers who invest in Swift early can become market leaders, or run circles around teams struggling with slowness of python, over-complexity of c++.
Top three comments paraphrased:
> 1) "huh? Foundation?"
but you have no qualms with `if __name__ == "__main__"` ?
> 2) "The name is pretentious"
is that an example of the well-substantiated and deep technical analysis HN is famous for?
> 3) Swift is "a bit verbose and heavy handed" so Google FAILED by not making yet another language.
This is not really about swift. Swift seems to have been chosen because the creator was there when they picked the language, even though he left.
You use of the word "seems" is very apt here.
Have you considered that Google might have hired Lattner precisely because he is the founder of LLVM and Swift, and they hoped to leverage his organizational skills to jump start next generation tooling? We know google is heavily invested in llvm and C++, but dissatisfied with the direction C++ is heading [0]. They also are designing custom hardware like TPUs that isn't supported well by any current language. To me it seems like they are thinking a generation or two ahead with their tooling while the outside observers can't imagine anything beyond 80s era language design.
[0] https://www.infoworld.com/article/3535795/c-plus-plus-propos...
If the answer to all these questions is "No", why should I care about this "new generation tooling"?
EDIT: and I'm not really attached to Pytorch either. In the last 8 years I switched from cuda-convnet to Caffe, to Theano, to Tensorflow, to Pytorch, and now I'm curious about Jax. I have also written cuda kernels, and vectorized multithreaded neural network code in plain C (Cilk+ and AVX intrinsics) when it made sense to do so.
Those that are interested in machine learning tooling or library development may see an opportunity to join early, especially when people have such irrational unfounded bias against a language, as evidenced by the hot takes in this thread. My personal opinion, that I don’t want to force on anyone, is that Swift as a technology is under-estimated outside of Apple and Google.
There aren't major benefits to using Swift4TensorFlow yet. But (most likely) there will be within the next year or two. You'll be able to do low level research (e.g. deformable convolutions) in a high level language (Swift), rather than needing to write CUDA, or waiting for PyTorch to write it for you.
[0] https://course.fast.ai/videos/?lesson=13 [1] https://course.fast.ai/videos/?lesson=14
Not sure I understand - will Swift automatically generate efficient GPU kernels for these low level ops, or will it be making calls to CuDNN, etc?
Even then I'm not sure what granularity MLIR will allow.
On the other hand you can do it in Julia today. There is a high-level kernel compiler and array abstractions but you could also write lower level code in pure Julia as well. Check out the Julia GPU GitHub org
As for Julia, I like it. Other than the fact that it counts from 1 (that is just wrong!). However, I'm not sure it's got what it'd take to become a Python killer. I feel like it needs a big push to become successful in a long run. For example, if Nvidia and/or AMD decide to adopt it as the official language for GPU programming. Something crazy like that.
Personally, I'm interested in GPU accelerated Numpy with autodiff built in. Because I find pure Numpy incredibly sexy. So basically something like ChainerX or Jax. Chainer is dead, so that leaves Jax as the main Pytorch challenger.
This would really give me some incentive to learn the language.
But it also gives reason it shows signs of promise.
So you should get involved if you are interested in contributing to and experimenting with a promising new technology, but not if you're just trying to accomplish your current task most efficiently.
Given the ML and Modula-3 influences in Swift, and the Xerox PARC work on Mesa/Cedar, it looks quite 80s era language design to me.
You have to use something like CFAbsoluteTimeGetCurrent while even in something not very modern like C# you would use DateTime.Now()
I think you have bought into the coolaide pretty hard here. Everything you are saying is a hopeful assumption of the future.
#include <numeric>
#include <iostream>
#include <chrono>
#include <vector>
int main(int argc, char** argv) {
for (int i = 0; i < 15; i++) {
std::vector<int> result;
auto start = std::chrono::system_clock::now();
for (int j = 0; j < 3000; j++) {
result.push_back(i);
}
auto sum = std::accumulate(result.begin(), result.end(), 0);
auto end = std::chrono::system_clock::now();
std::cout << (end - start).count() << " " << sum << std::endl;
}
}
I think you underestimate what a multiple decades of coding 40hrs a week gives you in terms of development speed.But I'd like to point out that while Google has some of the top C++ experts working for them, is heavily involved in C++ standardization and compiler writing process, in 2016 they claimed to have 2 billion lines of C++ running their infrastructure...
.. and yet they don't suffer from familiarity bias or the sunken cost fallacy I hear in your comment.
Instead Google C++ developers are sounding an alarm over the future direction the language and its crippling complexity:
http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2020/p213...
Just like with Go, their monorepo and internal tooling deturps the understanding how everyone else actually uses C++.
It's a problem.
C++ and javascript are languages of professional software engineers, of which there are many many more languages with various pros and cons.
Python has been the defacto standard in scientific/data/academic programming for decades. The only other language you could say rivals it would be MATLAB which is even more simplistic.
My point is that simplicity and clarity matters to people who don't care that much about programming and are completely unfocused on it, they are just using it do get data for unrelated research.
'if __name__ == "__main__"' is not in the example code nor is it a required part of a python program so not really sure what your point is here.
In my experience (Genomics) this is simply not true. Python has caught on over the last 5 or so years, but prior to that Perl was the defacto language for genetic analysis. Its still quite heavily used. Perl is not a paragon of simplicity and clarity.
(a) While I'm being honest that my observations are based on the fields I have experience, there is no such justification that "It is true broadly for computation in academia" in your comment.
(b) Interpreting "niche" as "small" (especially given your "true broadly" claim): Computational genetics is huge in terms of funding dollars and number of researchers.
I feel like trying out various languages/frameworks would affect compsci labs a lot less than other fields, since the students probably have some foundational knowledge of languages and have already learned a few before getting there. Might be easier for them to pick up new ones.
Doesnt really change my argument though, R is also a slow but simple language that is popular among academics but not professional software engineers. My whole point is that Swift is never going to be popular with academics because the syntax isn't simple enough.
The person you are replying to didn't call Python simplistic (and it certainly is not simple IMHO), they called it slow.
Yet swift is open source, and Apple and the community can fork it if they so choose. This is great news for me personally as an iOS developer and an ML noob who doesn't want to write Python. I can't comment on Julia because I have no experience with it, but I applaud the efforts to build the Swift ecosystem to challenge Python.
I think a lot of the criticisms so far are that it's early days for Swift in ML, and that's one point the author is emphasizing.
Does anyone force anything on Google? This seems to express little confidence in the competence of Google and their people. Perhaps Google chose Swift and brought Lattner in for his obvious expertise.
While other people will do the sensible thing and learn Rust. Because it runs circles around Swift, it offers many paradigms and can be used in almost any industry, operating system and product, not just developing apps for Apple's ecosystem.
Swift will take over the world when Apple will take over the world which is safe to assume it will never happen.
I am not saying at all that is bad to learn Swift and use Swift, but have correct expectations about it.
Some ideas in the language or parts of its core concepts are really good. First class optionals and sum types, keyword arguments, etc., I liked all of those.
Unfortunately, by and large, Swift is lipstick put on a pig. I have never used any other language that regularly gave me type errors that were WRONG. Nowhere else have I seen the error message "expression is too complex to be type-checked", especially when all you do is concatenating a bunch of strings. No other mainstream language has such shoddy Linux support (it has an official binary that works on Ubuntu... but not even a .deb package; parts of Foundation are unimplemented on Linux, others behave differently than on macOS; the installation breaks system paths, making it effectively impossible to install Python afterwards[1]). Not to mention, Swift claims to be memory-safe but this all flies out of the window once you're in a multithreaded environment (for example, lazy variables are not threadsafe).
In addition, I regularly visited the Swift forums. The community is totally deluded and instead of facing the real problems of the language (and the tooling), it continues to bikeshed minor syntactic "improvements" (if they even are improvements) just so the codes reads "more beautifully", for whatever that is supposed to mean.
But the worst thing is how the community, including your post, thinks Swift (and Apple) is this godsend, the language to end all language wars, and that everyone will eventually adopt it en masse. Even if Swift were a good language, that would be ridiculous. There was even a thread on that forum called "crowdfunding world domination". It has since become awfully quiet...
I don’t see people running over here to write numerical libraries like you see in Julia, that’s largely because of the crowd around Swift. The language is also a bit verbose and heavy handed for what data scientists would prefer. Latner was too close to Swift to understand this. The blame really falls on google project management.
Even though the article talks about the "why not Julia", which is the highest comment at time of typing ... Choosing a cross-compatible language would have kept more people interested in the long-run. Why should I as a Windows user want to learn a language just to use it with Tensorflow; when I'm not sure if such language support and other tooling will generally come to Windows?
I know this is a hot take but... I doubt Google has the capability to be frank. They created Dart and Go (a.k.a generics ain't necessary). They created Tensorflow 1 which is totally different from Tensorflow 2.
Swift may not be the best but Swift is starting to become such a large part of Apple so it will have backing no matter the internal politics.
The language is not where the battle will be, it will be the tooling.
People are searching for a better language in this space and it's something that often needs a corporate backing. Google is aware of this problem and hired Chris Latner to fix it, its just a bit of unfortunate oversight, I guess we'll keep using Python for now.
Hopefully Arraymancer will help increase its reach, I wish the implementors all the best.
Which is unfortunate, because it would probably be the best language if it were controlled by a non-profit foundation like Python. As it stands it's basically unusable.
And C# and Java have a benefit of JIT VM by default, meaning you only build once for all platforms, unless you need AOT for whatever rare reason (which they also have).
Not saying it wouldn't work, it definitely would, but I think I'd rather switch profession than deal with Maven and Eclipse in 2020.
Swift culture is more about having non-mutable structs that in turn is extended via extensions and heavy use of copy on write when mutable structs are needed. It's a small difference but it's there.
You have a weird notion of mutable by default in either Java or .NET. The former is notorious for builder patter because of that exact reason. Does Swift have special syntax for copy + update like F#: { someStruct with X = 10 }?
Never had problems with Maven. How is Swift different?
People have not been using Eclipse much for a while. There is IntelliJ IDEA for Java and Resharper for C#.
In Swift the copy on write happens as an implementation detail: https://stackoverflow.com/questions/43486408/does-swift-copy...
I don't really know why but the coding patterns (what I call culture) that are popular for each language are very, very different even when they can support the same feature-set.
What do you mean by "types are too shallow"?
Yes, jit can be slow to boot, but I think this is an area they're going to be focusing on.
"the tooling is poor" Not sure I agree here. I think it's great that I can easily see various stages from LLVM IR to x86 asm of a function if I want to.
So -- as I don't have experience with languages with much richer type systems like Rust or Haskell -- it's hard to imagine what's missing, or conceive of tools other than a hammer. Mind elaborating (or pointing me to a post or article explaining the point)?
Boot times still aren't ideal, but I find it takes about .1 seconds to launch a julia repl now. First time to plot is still a bit painful due to JIT overhead, but that's coming down very aggressively (there will be a big improvement in 1.5, the next release with differential compilation levels for different modules), and we now have PackageCompiler.jl for bundling packages into your Sysimage so they don't need to be recompiled every time to you reboot julia.
I also think the tooling is quite strong, we have ana amazingly powerful type system and I would classify discovering multiple dispatch as a religious experience.
I also don't think I could ever go back to using a langue that doesn't have multiple dispatch, and I don't think any language out there has a comparable out-of-the-box REPL experience.
If I'm a person who wants to do some data science or whatever and I have very little software background I want there to be libraries that do basically everything I ever want to do and those libraries need to be very easy to support. I want to be able to Google every single error message I ever see, no matter how trivial or language specific, and find a reasonable explanation. I also want the environment to work more or less out of the box (admittedly, python has botched this one since so many machines now have a python2 and a python3 install).
I think a big advantage of Julia is that it has a unusually high ratio of domain experts to newbies, and those domain experts are very helpful caring people. It's quite easy to get tailored, detailed personalized help from someone.
This advantage will probably fade as the community grows, but at least for now, it's fantastic.
Perhaps a more interesting question is whether the ML community needs a better language than C++ for _implementing_ ML packages. TensorFlow, PyTorch, CNTK, ONNX, all this stuff is implemented in C++ with Python bindings and wrappers. If there was a better language for implementing the learning routines, could it help narrow the divide between the software engineers who build the tools, and the data scientists who use them?
So I think that's the good news--because of the more independent nature of the work, you generally can win data scientists over to a new language one at a time, you don't necessarily need to win over an entire organization at once.
Getting a company or a large open-source project to switch from C++ to Swift or Rust or whatever, seems much harder.
That said I love Python as a language, but if it doesn't fix its issues, on the (very) long run its inevitable the data science community will move to a better solution. Python 4 should focus 100% of JIT compilation.
I really try to take advantage of the database to avoid ever having to munge very large CSVs in pandas. So like 80-90% of my work is done in query languages in a database, the remaining 10-20% is in Python (or sometimes R) once my data is cooked down to a small enough size to easily fit in local RAM. If the data is still too big, I will just sample it.
Numba gives you JIT compilation annotations for parallel vector operations--it's a little bit like OpenMP for Python, in a way.
* Ubiquitous.
* Not owned by a single corporation.
* Fairly performant runtime characteristics, with multiple implementations.
* Optional typing for quick explorations.
* Quite pleasant to use in its 2020 incarnation.
* The community has a proven process, tooling and track record of incrementally improving a language and its ecosystem.
Please don't shoot, I'm interested in constructive criticism.
And if Python, Julia and R don't cut it, then there's no reason to think another scripting language would. Instead you'd be looking at a statically typed and compile language with excellent support for parallelism.
On Differentiable Programming, you might be aware of a tensorflow counterpart in javascript:
I tested this to a certain extent and its not a toy. Its well thought out product from a very talented team, and has the ease of coding that we love about javascript. It can run on browsers!
This being said, we should note the strengths of a statically compiled language with the ease of installation and deployment like with Go, Rust, Nim, etc. in enterprise scale numerical computing.
Going from Python to JavaScript is a step backwards, not forward.
Frankly, I'm surprised that Go has made it this far - I mean, it's a great language, I get it, but Google is fickle when it comes to things like this.
Seeing how Chrome team pushes for PWAs, Android team woke up and is now delivering JetPack Composer, and we still don't know what is going to happen with Fuchsia, the question remains.
Both are relatively young languages with rapidly growing adoption. Dart was on a downward trend but has seen a rejuvenation in the last few years thanks to flutter.
However it still remains to be seen how long they will fund the team.
Chrome team cares about PWAs and Android team about JetPack Composer, and eventually having it compatible with Kotlin/Native (for iOS).
So it is still a big question mark why bother with Flutter, specially when it still lacks several usable production ready plugins for native features.
Julia has done a great job at marketing itself as the language built for modern numerical computing by numerical computing people. They have effectively recruited a lot of the scientific community to build libraries for it. I think the language is flawed is some deep ways, but there is a lot to learn in how they positioned themselves.
I like Kotlin, but the garbage collector isn't really meant for numerical computing I'd guess, and I doubt having to think about LLVM and JVM and JS at the same time is going to work out well for it when it needs such a heavy, heavy focus on performance.
Kotlin is a language by a small company that only grew out of sheer love by the world. Even Google had to throw in the towel and officially support Kotlin. Flutter is a reflection of "creating a new language" - Dart.
Kotlin doesnt have the heavy handed lock-in of Swift. Wouldnt Swift fundamentally by handicapped by Apple's stewardship ? What if Lattner wants to add new constructs to the language ?
1 - Many InteliJ plugins are now written in Kotlin (Android Studio)
2 - They really screwed up with Android Java dialect and people were looking for alternatives
3 - They need an exit story for when the hammer of justice finaly falls down on how they screwed up Sun
"Java / C# / Scala (and other OOP languages with pervasive dynamic dispatch): These languages share most of the static analysis problems as Python: their primary abstraction features (classes and interfaces) are built on highly dynamic constructs, which means that static analysis of Tensor operations depends on "best effort" techniques like alias analysis and class hierarchy analysis. Further, because they are pervasively reference-based, it is difficult to reliably disambiguate pointer aliases.
As with Python, it is possible that our approaches could work for this class of languages, but such a system would either force model developers to use very low-abstraction APIs (e.g. all code must be in a final class) or the system would rely on heuristic-based static analysis techniques that work in some cases but not others."
[1] https://github.com/tensorflow/swift/blob/master/docs/WhySwif...
They can't argue for Swift over Julia due to community size, given that Julia is far more portable, and more familiar to users in the scientific domain. 'Similarity of syntax to Python' is another very subjective 'advantage' of Swift: Later in the same document they mention "mainstream syntax" - that is, Swift having a different syntax from python - as an advantage.
I wonder whether they just decided on the language in advance, which is totally fine, but we could do without the unconvincing self-justification.
> and picked Swift over Julia because Swift has a much larger community, is syntactically closer to Python, and because we were more familiar with its internal implementation details - which allowed us to implement a prototype much faster.
I think it's very debatable to claim Swift is more similar to Python syntactically, as Julia looks more like a dynamic language to the user. Also, Julia is closer to languages like Matlab and R, which many mathematical and scientific programmers are coming from.
Swift has a much larger community, but it's not clear how big the overlap is between iOS app developers and Machine Learning developers. It probably would make deploying models on iOS devices easier, however.
If a Windows dev can target other platforms using Linux subsystem or containers, the only downside becomes an inability to target Windows desktops and servers, which are not hugely important targets outside of enterprise IT.
But there is of course a sense in which Swift is a general purpose language as opposed to something like SQL if that's what you mean.
Unfortunately, right now Swift is not (yet) a pragmatic choice for anything other than iOS/macOS apps.
I disagree that this is common. Most people are on web + browser + Office these days.
This is the actual port which has the CI and Installer for Swift on Windows: [0]
Python:
import time
for it in range(15):
start = time.time()
Swift: import Foundation
for it in 0..<15 {
let start = CFAbsoluteTimeGetCurrent()
This is why people like Python:- import time: clearly we are importing a 'time' library and then we clearly see where we use it two lines later
- range(15): clearly this is referring to a range of numbers up to 15
- start = time.time(): doesnt need any explanation
This is why academics and non-software engineers will never use Swift:
- import Foundation: huh? Foundation?
- for it in 0..<15 {: okay, not bad, I'm guessing '..<' creates a range of numbers?
- let start = CFAbsoluteTimeGetCurrent(): okay i guess we need to prepend variables with 'let'? TimeGetCurrent makes sense but wtf is CFAbsolute? Also where does this function even come from? (probably Foundation? but how to know that without a specially-configured IDE?)
EDIT: Yes everyone, I understand the difference between exclusive and inclusive ranges. The point is that some people (maybe most data programmers?) don't care. The index variable you assign it to will index into an array of length 15 the way you would expect. Also in this example the actual value of 'it' doesn't even matter, the only purpose of range(15) is to do something 15 times.
Here’s a paper by Travis Oliphant describing SciPy that has >2500 citations. https://scholar.google.com/scholar?hl=en&as_sdt=0%2C33&q=pyt...
In many fields of science Python is already the dominant language, in others (like neuroscience), the writing is on the wall for Matlab. Approximately all the momentum, new packages, and new student training in systems neuroscience that I’ve seen in the last 5 years is in Python.
At least in my university, most people really do use Python + C and Julia for many, many cases and MATLAB and such are used mostly in mechanical and civil engineering, some aero-astro (though a ton of people still use Python and C for embedded controllers), and Geophysics/Geophysical engineering (but, thanks to ML, people are switching over to Python as well).
I think even these fields are slowly switching to open versions of computing languages, I will say :)
The issue I see is with the undergraduate curriculum in many Universities. This is where I see the legacy use of MATLAB is really hurting the future generation of students. Many still don't know proper programming fundamentals because MATLAB really isn't set up to be a good starting point for programming in general. To me, MATLAB is a great tool IF you know how to program already.
Yes, and the government still needs COBOL programmers.
Going forward, I believe Python has far more momentum than either MATLAB or Mathematica. I think far more MATLAB and Mathematica users will learn Python than the other way around in the future, and far more new scientific programmers will learn Python than either of those.
However, when I left academia (engineering) 8 years ago, its use was already declining in graduate level research, and right before I left most professors had already switched their undergrad instructional materials to using Python and Scilab. I observed this happening at many other institutions as well. Anecdotally, this trend started maybe 10 years ago in North America, and is happening at various rates around the world.
I'm in industry now and MATLAB usage has declined precipitously due to exorbitant licensing costs and just a poor fit for productionization in a modern software stack. Most have switched to Python or some other language. My perception is that MATLAB has become something of a niche language outside of academia -- it's akin to what SPSS/Minitab are in statistics.
The University I work at still teaches MATLAB to new engineering students still.
Same for any other field that uses ML extensively.
Using `..<` and `...` is pretty simple to figure out from context. The former produces an exclusively-bounded interval on the right, while the latter is an inclusive interval. This is functionality that more languages could stand to adopt, in my opinion.
I agree that the names themselves are not very transparent. However, they become less opaque as you learn the Swift ecosystem. Admittedly, this makes them not as immediately user-friendly as Python's simple names, but it's not as though they're some gigantic obstacle that's impossible to overcome.
Personally, I like Swift a lot (even though I never use it). It has a syntax that has improved on languages like Java and Python, it's generally fast, it's statically typed, and it has a large community. The fact that implicit nullable types are discouraged directly by the syntax is phenomenal, and the way Swift embraces a lot of functional programming capabilities is also great. If it weren't so tied to Apple hardware, I would likely recommend it as a first language for a lot of people. (I know that it runs on non-Apple hardware, but my understanding is that support has been somewhat limited in that regard, though it's getting better.)
IMO that's essentially the problem. Most people* don't want to have to learn the ecosystem of a language because it's not their focus.
The other issue is that when you start googling for information about the Swift ecosystem, you're not going to find anything relevant to academic, mathematical, or data-science programming. All the information you will find will be very specific to enterprise-grade iOS and macOS development, which will be a huge turn-off to most people in this community.
EDIT: *academics
Yet.
The question is whether Google and other Swift enthusiasts can change that over time.
Javascript, Swift, and VBA have let.
C, C++, Java, C#, PHP, Python, Go don't have it.
I'm also willing to bet that if you haven't studied math in English let is a non-obvious keyword.
10 LET A$ = "Hello world"
20 PRINT A$
30 GOTO 10I find this notation slightly clearer than the python version. It took me some time to remember whether range(15) includes 15 or not.
Once I knew these two facts, it didn't add much confusion.
1. Indexing starts from 0
2. Thus, range can be thought "from up to one before x", x here would be 15.
And I learned this pretty early and did not get confused later on.
If you're anything like me, you'll end up spending quite a bit of time looking up the documentation to the range operator to remind yourself how this week's language works again.
A current and more readable way of expressing this would be
let start = Date().timeIntervalSinceReferenceDate
If you don't need exact seconds right away, you can simplify further to just: let start = Date()
which is easily as simple as the Python example.CACurrentMediaTime() / CFAbsoluteTimeGetCurrent() are first of all not deprecated (just check CFDate.h / CABase.h) but return a time interval since system boot so they are guaranteed to be increasing. It's just a fp64 representation of mach_absolute_time() without needing to worry about the time base vs seconds.
Date() / NSDate returns a wall clock time, which is less accurate and not guaranteed to increase uniformly (ie adjusting to time server, user changes time etc)
That is true for CACurrentMediaTime, but that time stops when the system sleeps (https://developer.apple.com/documentation/quartzcore/1395996... says it calls mach_absolute_time, and https://developer.apple.co/documentation/driverkit/3438076-m... says ” Returns current value of a clock that increments monotonically in tick units (starting at an arbitrary point), this clock does not increment while the system is asleep.”)
Also (https://developer.apple.com/documentation/corefoundation/154...):
Repeated calls to this function do not guarantee monotonically increasing results. The system time may decrease due to synchronization with external time references or due to an explicit user change of the clock.
Also CFAbsoluteTimeGetCurrent explicitly calls out that it isn't guaranteed to only increase. CACurrentMediaTime is monotonic though.
CFAbsoluteTimeGetCurrent also returns seconds since 2001 and is not monotonic, so there's really no reason to use it instead of Date().timeIntervalSinceReferenceDate. The most idiomatic equivalent to the Python time method is definitely some usage of Date(), as time in Python doesn't have monotonic guarantees either.
[1] https://developer.apple.com/documentation/corefoundation/154...
Because they deal with calendar-related stuff that is better accessed through Date.
https://developer.apple.com/documentation/corefoundation/154...
https://developer.apple.com/documentation/foundation/nsdate/...
So for this comparison, it is better to use Date().
But anyway that's not a Swift problem. Swift itself is pretty easy to get into.
Since we have custom attributes I will investigate soon if there's a nice framework around that can make it work a bit like the usual JSON libraries in for example C# work.
It's more important that the language can accurately and succinctly represent the mental model for the task at hand, and the whole point of this article is that Swift can offer a syntax that is _more_ aligned with the domain of ML while offering superior performance and unlocking fast development of the primitives.
I think functional programming advocates underrate simplicity of procedural languages. Programming is not math, algorithms are taught and described as a series of steps which translate directly to simple languages like Fortran or Python.
I think ML is great, but I’m skeptical if it is a big win for scientific computing.
Plenty of excellent programmers are not mathematicians. How would that work if programming were just math? That’s like saying physics is just math while ignoring all of the experimental parts that have nothing to do with math.
Most of the concepts in Python come from academics and mathematics, so it's an easy transition. I don't think math has a time concept in a straight forward way, so time is an edge case in Python.
Python's no saint when it comes to time stuff either. I had some code using time.strptime() to parse time strings. It worked fine. Then I needed to handle data from a source that included time zone information (e.g., "+0800") on the strings. I added '%z' to the format string in the correct place--and strptime() ignored it.
Turns out that if you want time zones to work, you need the strptime() in datetime, not the one in time.
BTW, there is both a time.time() and a datetime.time(), so even that line that needs no explanation might still cause some confusion.
I think the main reason is that people keep thinking that a simple API is possible, and that more complex stuff can initially be ignored, or moved to a separate corner of the library.
The problem isn’t simple, though. You have “what time is it?” vs “what time is it here?”, “how much time did this take?” cannot be computed from two answers to “what time is it?”, different calendars, different ideas about when there were leap years, border changes can mean that ‘here’ was in a different time zone a year ago, depending on where you are in a country, etc.
I guess we need the equivalent of ICU for dates and times. We have the time zone database, but that isn’t enough.
>>> list(range(15))
[0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14]
The important thing with syntax is to avoid the illusion of understanding. That's when the language user is confident that the syntax means one thing when it actually means something else. If the user is not sure what something means, they'll look it up in docs or maybe write a few toy examples to make sure it does what they think it does. Python's range() is ambiguous enough that I did this when I was learning the language. I was pretty sure it would create a range from 0 to 14, but I wanted to make sure it wasn't inclusive (0-15).Examples of the illusion of understanding abound. These aren't all true for everyone, and HN users have been writing software long enough to have internalized many of them, but every language has them:
- Single equal as assignment. Almost every newbie gets bitten it. They see "=" and are confident that it means compare (especially if it's in a conditional).
- x ^ y means xor, not "raise x to the power of y"
- "if (a < b < c)" does not do what newbies think it does.
- JavaScript's this.
Sometimes syntax can make sense on its own, but create the illusion of understanding when combined with another bit of syntax. eg: Python newbies will write things like "if a == b or c" thinking that it will be true if a is equal to either b or c.
The illusion of understanding is the cause of some of the most frustrating troubleshooting sessions. It's the thing that causes newbies to say, "Fuck this. I'm going to do something else with my life."
>The illusion of understanding is the cause of some of the most frustrating troubleshooting sessions. It's the thing that causes newbies to say, "Fuck this. I'm going to do something else with my life."
About 14 years ago (give or take up to 4 years) I read about a study that was done at a prestigious CS university, where some tests were given entering CS students at the beginning of the course to see who was ok with arbitrary logic and syntax and who was not, IIRC it was 40% of the class who would get hung up on "but why" and "it doesn't make sense" and would end up failing but the ones who were able to cope with the arbitrary nature of things would graduate and the others would end up dropping out, changing studies.
About every couple of months I wish I could find that damn paper / study again.
NOTE: my memories of this study might also have been faded by the years, so...
It seems like the conclusions of the study were overstated, but the general idea is correct: Those who apply rules consistently tend to do better at programming than those who don't. This is true even if their initial rules are incorrect, as they simply have to learn the new rules. They don't have to learn the skill of consistently applying rules.
1. http://www.eis.mdx.ac.uk/research/PhDArea/saeed/paper1.pdf
2. http://www.eis.mdx.ac.uk/staffpages/r_bornat/papers/camel_hu...
And embarrassingly this thing I've gone around believing for the last 14 years isn't so.
It's more like saying people don't understand all of how pointers work in C.
Is it? Does that range start at 0 or 1 or some other value? Does it include 15 or exclude it?
> - start = time.time(): doesnt need any explanation
Doesn't it? Is that UTC or local time? Or maybe it's CPU ticks? Or maybe the time since the start of the program?
You've basically just demonstrated the assumptions you're used to, not any kind of objective evaluation of the code's understandability.
disclaimer: I use swift, but also I have used python.
The details (e.g. that list indexes and ranges start at 0 by default and are half-open) are consistent and predictable after just a tiny bit of experience.
Maybe most people share your bias, and so it could qualify as a reasonable definition of "intuitive for most humans", but there's little robust evidence of that.
I claim that it is easier to achieve this level of understanding in Python than most other programming languages. (And not just me: this has been found to be true in a handful of academic studies of novice programmers, and is a belief widely shared by many programming teachers.)
Using words that the reader is already familiar with, sticking to a few simple patterns, and designing APIs which behave predictably and consistently makes communication much more fluent for non-experts.
There are deeper levels of understanding, e.g. “have carefully examined the implementation of every subroutine and every bit of syntax used in the code snippet, and have meditated for decades on the subtleties of their semantics”, but while helpful in reading, writing, and debugging code, these are not the standard we should use for judging how legible it is.
What does range mean? Is it English? Attacking every single possible element of a language is not compelling. The defacto standard, is 0 indexing. The exceptions index at 1.
> - start = time.time(): doesnt need any explanation
People often oversimplify the concept of time. However, on average, the cognitive load for Python is lower than most. Certainly less than Swift. In one case I would look up what time.time actually did and in the case of Swift, I would throw that code away and work on another language with less nonsensical functions, like PHP. /s
"Defacto standards" are meaningless. The semantics of any procedure call are completely opaque from just looking at an API let alone a code snippet, especially the API of a dynamically typed language, and doubly-so if that language supports monkey patching.
So the original post's dismissive argument claiming one sequence of syntactic sugar and one set of procedure calls is clearer than another is just nonsense, particularly for such a trivial example.
> However, on average, the cognitive load for Python is lower than most.
Maybe it is. That's an empirical question that can't be answered by any arugment I've seen in this HN thread.
No they arent. A mismatch between what is expected and doesnt happen, within a specific context contributes to cognitive load. "Intuitive" is a soft term with a basis in reality. The only language (Quorum) that has made an effort to do analyses was largely ignored. Usability in languages exist, with or without the numbers you wish for. Swift is less uable than many some languages and more than others.
This is like reading a novel that says, "And then Jill began to count." and then asking the same questions. A non-technical reader does not need to know these details. The smaller details are not required to grok the bigger picture.
>Doesn't it? Is that UTC or local time? Or maybe it's CPU ticks? Or maybe the time since the start of the program?
When is the last time someone asked you, "Know what time it is?" and you responded with, "Is that UTC or local time?" Same thing, these details do not and should not matter to a non-technical reader.
Keep in mind, the audience is for non software engineers from people who barely know how to code, to people who do not know how to code but still need to be able to read at least some code.
For everyone else, they will have used something simpler like Excel, C, R, Python.
So, is that evaluated when that statement is run, when the value is first read (lazy evaluation, as in Haskell), or every time ‘start’ gets read. For example, Scala has both
val start = time.time()
(evaluates it once, immediately), lazy val start = time.time()
(evaluates it once, at first use), and def start = time.time()
(creates a parameterless function that evaluates time.time() every time it is called)CF clearly is a prefix for a library. I'll take an educated guess it means Core Foundation? Pretty common pattern of naming things, with +/- to be certain. And once you've seen it, it is just there, and you know precisely what it means. So 10 minutes of your life to learn what CFxxx() means.
Let. I like lets. Some don't. Surely we can coexist?
x..y is also not unique to Swift. It has a nice mathematical look to it, is more concise.
Btw, is that 'range' in Python inclusive or exclusive? It isn't clear from the notation. Must I read the language spec to figure that out? .. /g
It does, and the prefix is only there because C's namespacing isn't great.
Furthermore, it's obvious what `time()` does in this context, but if I was writing this code I would _absolutely_ have to look up the correct time function to use for timing a function.
First of all if you are doing "for i in range(N)" then you are already doing it wrong, for ML and data analytics you should be using NumPy "np.arange()", Numpy arange doesnt even run in "python" it's implemented in C. So it may even be faster than swift '..<' . Let me know when you can use swift with spark.
time.time() (as well as datetime.datetime.now() and other stuff like that) always looked extremely ugly to me. I would feel better writing CFAbsoluteTimeGetCurrent() - it seems more tidy and making much more sense once you calm down and actually read that.
"range(15)" vs "0..<15" could go either way.
"let" vs "var" in Swift is indeed something that adds verbosity relative to Python, and adds some cognitive load with the benefit of better correctness checking from the compiler. Very much a static vs dynamic typing thing. That's where you'll see the real friction in Swift adoption for developers less invested in getting the best possible performance.
#[derivative]
fn cube(x: f32) -> f32 { x * x * x }
// expands to
// impl Derivative for cube {
// type Args = (f32,);
// type Return = f32;
// fn derivative(x: Self::Args) -> Self::Return {
// 3 * x * x
// }
// }
let cubeGrad = gradient(cube);
assert_eq!(cubeGrad(2), 12);
where `gradient` is just a normal Rust function: fn gradient<F: Derivative>(_: F) -> impl Fn(F::Args) -> F::Return {
|x: F::Args| Derivative::derivative(x)
}
My original attempt was accidentally more powerful, in that it allowed `cube` to be a stateful function like a closure.To implement this, the only thing one needs to do in the `#[derivative]` macro is:
* parse the function into an AST
* fold the AST using symbolic differentiation
* quasi quote the new AST
Swift can probably differentiate functions in other translation units because it keeps the bytecode of all functions around. A proc-macro based approach wouldn't be able to achieve that, at least, not in general, but a Rust plugin could, since plugins can drive compilation to, e.g., compile whatever function is being derived, and differentiate its bytecode instead of Rust, and if the function happens to be in a different translation unit, access their byte code.
I feel there is an impedance mismatch here.
For parsing Rust into an AST you just use syn, which also supports doing AST folds, and for semi-quoting you just use the quote crate.
For the symbolic differentiation part, you can use any of the hundred symbolic differentiation libraries out there to obtain '3 * x * x' from 'x * x * x'. Pretty much any machine learning framework out there supports AD and symbolic differentiation in some form, by using some libraries already. Swift people are standardizing standard practice, not discovering the americas.
If any of that looks like magic to you, then you might just be lacking the background.
You can have a code using a GPU array library and just differentiate it (which ends up being more flexible / composable).
- parses the Rust AST
- folds it into a CUDA C AST
- writes it to a temporary .cu file
- compiles it with nvcc
- links it into your Rust binary
and that allows you to launch your function as a CUDA kernel.
The Rust emu crate does this (more or less), but targets WebGPU instead of CUDA.
Note that to be useful you need backward differentiation (a function with one output but several outputs) and compatibility with tests and loops.
See the Zygote.jl library (Julia) for a working example in a different language (their whitepaper is a great starting point if you want to implement something like that).
That question can be generalized to "since <Turing Complete language> can do <anything-we-want-via-library/function/macro/codegen/customcode/extension/whatever>, can't we just do that instead of modifying the core language?"
Well yes, but the advantage of having 1st-class language syntax includes:
+ more helpful compiler error messages because it has intelligence of the semantics
+ more helpful IDE because the UI has intelligence about the intentions of the programmer
+ interoperability because programmers are all using the same syntax instead of Programmer Bob using his custom macro-dialect and Progammer Alice using her idiosyncratic macro-dialect.
Not really, since that usually requires using weird syntax or APIs, but the Rust proc-macro solution has the exact same syntax as the Swift one:
@differentiable
...cube function
let grad = gradient(at: cube);
vs #[derivative]
...cube function....
let grad = gradient(cube);You're focusing on syntax but I was talking about the higher-level semantics. When programmers write custom code (e.g. custom macros), as far as the compiler/IDE is concerned, it's just an opaque string that it checks for valid syntax. Those tools have no "intelligence" of what the higher concepts of derivatives/deltas etc.
E.g. the IntelliJ plugin for Rust will have no idea what a the higher semantics of a differentiable function is. It just sees a recursive "gradient()" which is just an arbitrary string of code to parse and it might as well be spelled "abcdefgzyxwvut123()".
I.e. the blog post is talking about an symbiotic ecosystem of tools (e.g. compilers, IDEs, etc) that treats differentiable programming as a 1st-class language feature. The ultra flexibility of Rust or Lisp macros to reproduce the exact syntax doesn't really solve the semantics understood by the ecosystem of tools.
edit reply to: >You seem to be at least suggesting that if one adds differentiable programming as a first class language feature, IDEs would automatically be intelligent about it,
No, that's an uncharitable interpretation. I never claimed that IDEs will automagically understand higher level semantics with no extra work required to enhance the IDE's parser.
I'm only saying that your macro proposal only recreates the syntax but not the semantics -- so you're really not fully solving what the blog author is talking about.
That's not true. IDEs would need to add support for that feature, and whether the feature is a language feature, or a custom Rust macro, for Rust at least, the amount of work required is the same.
Rust IDEs already understand many ecosystem macros (cfg_if, serde derives, etc.). Somebody would just need to put in the work for a new macro, but that would need to be the case anyways.
The same would apply to Rust static analysis tools like clippy. Somebody would need to write lints, and whether those target an ecosystem macro or a keyword doesn't really matter with respect to the amount of work that this requries.
So what are you saying? Because that's what I'm honestly understanding from your post.
That somehow this feature being a Rust macro means that it cannot be as good as if it was a first class feature. You are making this claim in general, as if first class features are always strictly better than user-defined features, yet this claim isn't true, and you haven't mention any particular aspect of automatic differentiation for which this is the case.
Swift doesn't have proc macros, so they have to make the feature first class. For languages with proc macros, the question that one should be asking is: what value does a first class language feature adds over just using a proc macro.
What do you think these semantics are? The only semantics I see is being able to compute the derivative of a function at some particular point (for some particular inputs). The proc macro solution has this semantics: it automatically computes the derivative of a function, and gives you a way to query its value at a particular point.
Point 3 is simply that programmers aren't sharing interfaces, which is indeed a shame, but that's independent of whether you choose to put the interface in the language or the library. Didn't TensorFlow fork the Swift compiler?
https://github.com/apple/swift/tree/tensorflow
Forking the compiler is hardly my idea of getting all programmers on the same page.
Maybe that's too niche of a subject for people to really care about, and maybe it's going to fail now that Lattner is gone (though there still seemed to be a bunch of activity last I checked), but the problem it is trying to solve is real.
Add some hyphens and commas and shed a bit of cruft, and it's fine IMO. (I hope you don't mind my speculative editing.)
I did realize that if I have to throw a million cores at something vs. 1000 perhaps it would make sense to spend the effort or buy the time from an expert in C++ so as to save on the compute cost. But then, what if those million cores are only needed for a a day or an hour? Then python or some other rapid prototyping language would make a bit more sense imo.
prospective politeness is not a basis for justification
And when you start talking about buying 10s of millions in additional hardware to support the project you're trying to launch you start thinking pretty hard about whether you can speed things up.
And for practitioners most of this is not crazy high level math, it's first year linear algebra and multivariable calculus at most, and generally just lego blocks, intuition and data work.
While C++ might not be the ideal language, choosing one that only works properly in Apple platforms, dismissing the complaints of lack of any solid roadmap in platforms that researchers actually use and hand waving any usability problems to just use notebooks in Google Compute cloud is not how a language gains adoption beyond its initial niche.
Switching from garbage collected and dynamically typed to manual memory management + static typing + borrow checker + trait-based generics is going to be a deep dive for these people. Maybe they'd end up better off in the end though, who knows. I personally think that dynamic typing and GCs are overrated and their benefits overstated but that's obviously not the consensus out there.
Either way, people will find a way to justify their preference. The people in question liked Swift, and there's nothing wrong with that.
(I can't comment on that because I don't know Rust nor am I a data scientist!)
[1] https://github.com/tensorflow/swift/blob/master/docs/WhySwif...
Note that the recent Tensorflow Summit had zero annoucements related to S4TF, including post event blog posts.
From the outside it looks like there is a small group trying to push S4TF, without regards for usability outside Google Cloud, while everyone else is doing JavaScript, Python and C++ as usual.
Regarding the tensorflow dev summit, it was supposed to be 2 days long initially, with the S4TF talk taking place on the second day. One week before the summit, the whole day 2 got scraped due to covid19 though. So the intention was there at least.
In a nutshell, differentiable programming is a programming paradigm in which your program itself can be differentiated. This allows you to set a certain objective you want to optimize, have your program automatically calculate the gradient of itself with regards to this objective, and then fine-tune itself in the direction of this gradient. This is exactly what you do when you train a neural network."
Isn't this just declarative programming as we know it from e.g. Prolog, SQL or other places where the programmer declares what their objective is, and it's left up to the interpreter, compiler or scheduler to figure out the best way to achieve that? And now that's being applied to ML (which probably makes sense, since it involves a lot of manual tweaking). Sounds like a great use case for a library, but hardly worthy of being called a new programming paradigm.
• I fail to see how differentiable programming (the idea of expressing the desired computation in terms of differentiable objective functions) is any less of a "paradigm" than logic programming (the idea of expressing the desired computation in terms of logical predicates).
• Depending on the expressiveness of your programming language, every paradigm can seem like it's "a great use case of a library, but hardly worthy of being called a new programming paradigm."
It's also not quite declarative programming either. In the same way that you build ordinary executable programs out of smaller executable parts. You build differentiable programs of smaller differentiable parts. You aren't declaring what outcome you want, you are simply restricting what you are building your program out of so that it has properties that allow you to interpret the program differently from how it will run "normally".
Declarative programming is more or less orthogonal to differentiable programming. You're right that declarative as a paradigm leaves the implementation details up to the "compiler," and so you can have arbitrary implementations created in response to one declared specification. And often what that means is that under the hood it's possible that the compiler can be tweaked to output better results. But the thing is that those compiler changes and optimizations, are just that: arbitrary. You can't "know" or "prove" anything about them, and that means it mostly requires human creativity and intelligence to make progress.
That is fine and all, but differentiable programming is asking/answering the question: Ok, but what if we COULD prove something here? What if the units of computation could be guaranteed to have certain mathematical properties that make them isomorphic to other formalisms?
And why do we care about that?
Well, in math what happens is that some people will start with their favorite formalism and prove a bunch of stuff about them, and figure out how to do interesting calculations or transformations on them. Like "Ah, well if you have a piece of data that conforms strictly to the following limitations, then from this alone we can calculate this very interesting answer to a very interesting question."
But a lot of the time those mathematical paradigms can't talk to each other -- in programming terms, their apis just aren't compatible. Like raster vs vector images. Both image storage/display paradigms are "about" the same thing, but a lot of the operations you can do on one don't even make sense to try on the other, and our ability to translate back and forth between them is a little wonky. Math formalisms are a bit like that, a lot of the time.
So it's very interesting in math world when someone proves that a formalism in one paradigm can be transformed perfectly into a formalism from another paradigm. All of a sudden all the operations available in either paradigm become available in both paradigms because you can always just take your data, transform it, do the operation, then transform the answer back.
(Side note: this is why some people are excited about Category Theory: it's like a mathematical rosetta stone. Ie. A lot of things can be translated into and out of category theory, and in turn all those things that were previously in separate magisteria are interchangable.)
Ok so, back to differentiable programming. If you suddenly have a way to conceive of your program / unit of computation as a differentiable function, then right off the bat you get access to all the tools ever created for calculus. The optimization thing where you find the gradient of the program and follow it toward some target is just one of the things. You also get a huge suite of tools that let you enter the world of provably correct programs, for example.
You also get access to all the tools of all math that can translated to and from calculus, which is... a lot of them. I wish I knew more about math so I could rattle off the 100 ways that would help, but I can't, so instead I'll just say that I think it would be a game-changer for creating optimized, robust systems that work way, way better than our current tech.
Though there might be potential for extending the frameworks to e.g. differential cryptanalysis - I'm not knowledgeable enough about it to say how much differential cryptanalysis can be done programmatically.
You'd like to be able to write your functions using the language's standard function syntax, but have access to both the function and its differentiated form. You can achieve it with macros (in an ad-hoc way) or as a custom language feature (what's being done here, again kinda ad-hoc), or you can use an algebra+interpreter style but at the cost of having to use a less natural syntax (Haskell do notation or similar). The thing that I'd say is closest to the ideal solution is something like stage polymorphism (which is genuinely an exciting new paradigm in my book: it squares the circle of macros versus strong typing), where you can write a function definition in natural function syntax and have access both to the function itself and the AST of the function in a much richer form than what a macro gets (which can be interpreted to produce a differentiated version of the function).
Sure, there are some analogies, but it is stretching the definition a bit.
It's a bit of an outlier, most ATP episodes are Apple focussed news with adjecent tech interests. Episode 371 is not that.
To be fair, he does answer the question a minute later.
doStuff({ foo: 3, bar: 7 })
def f(x, **kwargs):
return g(x, foo=1, **kwargs) function f(x, ...args) {
return g(x, ...args);
}
Or did you mean forwarding object fields? function f(x, params) {
return g(x, {foo: 1, ...params});
}I'm sick of this pointless and never-ending upgrade cycle. I started a new project this week, and it's in Clojure. I find it a much more productive language, and I'm confident I can run it on nearly any computer from the past 20 years.
(This is probably not my most productive comment ever, but I'm pretty frustrated right now with having chosen Swift in the past.)
I have stuck with C++ but I did buy a book on Swift, which is now hopelessly out of date. People can say bad things about C++ (plenty of material!) but at least it is mostly backwards compatible.
Can't you just reuse the SSD?
(Still, I can't complain too much - it is a free entire OS upgrade unlike OSX upgrades of yesteryear)
Some people like Swift on its own merits but I do not. It’s a terrible fit for the kinds of programming I want to do.
You do need to use the Xcode-bundled toolchain when archiving for the App Store, but not for development or off–App Store releases. Even for the App Store, you could get away with just one computer on your team having the latest macOS and Xcode, or with using a separate build server instead of archiving on your workstation.
Supplementals: http://taichi.graphics/wp-content/uploads/2019/10/taichi_lan...
Python already drives a significant market and captured talent in numerical differentiable programming, particularly because of the ease with which prototyping and tuning is done with it. Obviously when we want to scale we have the conversation on some other option.Personally its a bit frustrating that Go support for TensorFlow or one of the competitors is not quite satisfactory and this is surprising, I can't explain it.
Instead of inventing any new language why don't one of you -- Python, Go, C++ & Rust, Julia, or Swift -- complete the job with end-to-end differentiable programming. 'Complete' to me means a language level seamless GPU (or related distributed/parallel architecture) and language level deployment ease (the kind of thing done with Kubernetes) and integration with embedded hardware.
I believe in the language as a system programming language (and I’ve done systems programming in the past).
That said, like C#, it is strongly associated with Apple as a “private” language (despite being open-source).
That’s probably a big reason it’s slow to be adopted in a wider fashion, but I’m sure there’s other issues.
It would be nice to see it be more widely adopted, but I’m OK with it being confined to the Apple ecosystem, as that is my domain.
IBM may have counter-balanced this.
What do other think is the reason for the slow backend adoption of Swift?
It's not different enough. All the demos I've seen of it, felt like it was mostly an "also ran" language that modernised it to be a significant improvement over Obj-C.
But I've never seen any argument for it being a huge enough improvement over C#, Kotlin or others to be worth it to switch ecosystems.
More radical languages all tend to look a bit "read only" to me but I am aware that might just be unfamiliarity on my part.
C# and Kotlin both have nice features but they are more conservative Java-likes than Swift. Swift had enough modern features and syntax to feel like it would be reward the effort learning.
But I don't own an iOS device so I'm still waiting for a use case!
So while Swift is a welcomed improvement over Objective-C, hardly brings anything worthwile to dump either the Java or .NET ecosystems.
The only ones that are similar are iOS and TVOS, but there's a few differences in the way that you structure apps.
What does it have that Kotlin doesn't?
(As a Scala programmer, Swift and Kotlin look pretty much the same to me; as far as I can see they're both polished but conservative languages that don't fundamentally offer anything that wasn't in OCaml)
I don't think Kotlin has any form of the "ugly C" API Swift has.
Swift is a bit more secure in terms of "it compiles, it works" than Kotlin, but Kotlin is definitely a huge improvement over Java.
Disagree. Reified generics let you write functions that do different things based on the runtime type of an object, but such functions are confusing to a reader and best avoided.
> Swift has associated values with enums
Java/Kotlin style enums are my favourite from any language, since they're first-class objects; you can give them a value in a field if that's what you want, but you can also have them implement an interface etc.
> ARC is simpler to predict but can induce a performance penalty.
In the general case ARC has the same pause problems as traditional GC, since you can't tell when one reference is the last living reference to an arbitrarily large object graph.
> Swift is a bit more secure in terms of "it compiles, it works" than Kotlin
How so? If anything I'd have expected the opposite.
It's quite powerful, once you get your head around it.
I write about that here: https://littlegreenviper.com/miscellany/swiftwater/enums-wit...
And here: https://littlegreenviper.com/miscellany/swiftwater/writing-a...
But they don't make sense as enums. If you have to look "under the hood" to get a reasonable model of a feature, it's a bad design.
> It's quite powerful, once you get your head around it.
> I write about that here: https://littlegreenviper.com/miscellany/swiftwater/enums-wit....
Like I said, you can write the same kind of thing very easily in Kotlin by using a sealed class - your example would translate line-for-line. It's a useful feature. But it's very confusing to call it a kind of enum rather than as a kind of class; as your other post acknowledges, it means Swift essentially has two different things that are differently implemented but both called "enum".
In contrast, Swift has done an excellent job of making a real clean break from its Obj-C predecessor, despite sporting full Obj-C interop. It has for most practical purposes none of Obj-C’s quirks for pure Swift code, which is great.
I frequently find myself wishing that Kotlin shared Swift’s level of independence from its predecessor, or that I could just use Swift instead. While Kotlin isn’t a bad language, its ties to Java leave me less than enamored.
What necessary modern features do you feel Swift has and C# doesn't?
These guys are great: https://www.objc.io
I'm pretty sure their site is a Vapor server.
C# was successful because at the time it appeared people wanted an alternative to Java.
Now I do business web apps, but previously I did games. Which we published on all app stores + web. I'm also doing web sites. I did cross platform mobile apps with Xamarin. I can do even frontend web thanks to Blazor. I can use C# for embedded work - but that's not my thing. And you can use C# for machine learning, too.
With Swift I would be restricted. Which may be fine for some people who like to work only on certain kind of apps, but not for me.
Only Rust tempts me to learn and use as a new language.
I don't expect it to takeover the world, I just want it to have enough success to keep me earning money and be kept up to date. Which kind of happens. :)
I don't see how IBM will push for Swift in the future. They already pulled out (?) of the swift server working group, and didn't have the kind of pull to begin with.
IMO the most important reasons for the slow adaptation is lacking cross platform functionality (currently no upstream windows support + poor linux tooling) and questionable API design. Both make it difficult for new developers to give swift a try. Luckily, windows supportis supposed to come with the next version in a few months.
The API design is probably very subjective, but comparing the example in the blog post I don't understand how anyone could prefer:
import Foundation CFAbsoluteTimeGetCurrent()
over:
import time time.time()
The python way is much more intuitive IMO.
For web related work some other languages with their frameworks might yield better productivity: C#, PHP, Python, Java, Ruby.
In latest web framework benchmarks by Techempower, Java, C#, Kotlin beat it by a large margin.
It doesn't do very good against Java even in The Computer Language Benchmarks Game, where each wins in 5 tests.
You can't use Swift on Windows. You can target practically target most operating systems with C#. C# is not only open source, but an international standard. Ecma (ECMA-334) and ISO (ISO/IEC 23270:2018)
Although both started as proprietary languages, both are open source now, both are developed mainly by one big company, the difference is that Microsoft is interested seeing C# on many platforms, while Apple doesn't have that interest.
I think Kotlin, Go, Rust have greater potential to rise in popularity than Swift.
I would've been excited about Swift if it weren't for these things.
From my understanding, the author is someone from outside Google who doesn't know anything about Google bets more than what is publicly available, and the author is biased toward the Swift language.
JAX makes more sense, as it uses Python, but more importantly, Google works at a lower level of the stack with TPU, MLIR and XLA. It doesn't really matter what language is on top of that stack.
That said, for machine learning a static type system (as they are now) is definitely not as much of a boom as in most other areas. Most of it involves tensor operations, and embedding things like it's shape in the generics will often lead to an exponential explosion of possible monomorphizations that the compiler will be forced to create. Even JIT languages like Julia will have trouble even though it only needs to compile when it's used (for example StaticArrays will get to a point when the compiler will take so much time that it's not worth it anymore). And even then I feel like most of the issues that actually take time are deeper than stuff like shapes and not naming dimensions, like numerical instabilities, hyperparameter tuning, architecture search and other stuff that gets more benefit of a language allowing quicker exploration.
The difference is that whenever a function in called in Julia, based on the type of the arguments the compiler can infer every type of every variable and arguments of every subsequent function, immediately compiling them down to machine code (unlike Python there is no such thing as a Julia Virtual Machine or even a Julia Interpreter outside of the debugger). Whenever you enter a function in Julia it becomes a static program, with all the properties of a static language (for example, you can't redefine types, you can define functions but the program can't see them since they are not part of the compiled code, you can't import libraries, and all types were already checked before running). That's why Julia is fast, and the language was entirely designed for working this way (there are lots of things you can do in Python that you can't in Julia to make this work).
No they are not. At least not for my definition of the term "strongly typed". Of course you can use a different definition, but if, as you say, "there isn't really a language" that does not fulfill your definition you might want to reconsider its usefulness. The point of classifying languages is mood when every fan comes along and says "my favorite language also has that, if you tweak your understanding just a bit".
julia> foo(x) = x^2 - 2x + factorial(4x)
foo (generic function with 1 method)
julia> @code_typed foo(4)
CodeInfo(
1 ─ %1 = Base.mul_int(x, x)::Int64
│ %2 = Base.mul_int(2, x)::Int64
│ %3 = Base.sub_int(%1, %2)::Int64
│ %4 = Base.mul_int(4, x)::Int64
│ %5 = invoke Base.factorial_lookup(%4::Int64, Base._fact_table64::Array{Int64,1}, 20::Int64)::Int64
│ %6 = Base.add_int(%3, %5)::Int64
└── return %6
) => Int64
This is me querying Julia for it's typed intermediate representation of the function `foo` I defined if the input is an integer. As you may be able to see, this is a statically typed program. So while Julia is a dynamic language, the interior of function bodies can be static if type inference succeeds. However, if type inference fails, Julia is also perfectly happy to just let chunks be dynamic.If you're having trouble understanding that code_typed output and are familiar with C, maybe the output LLVM instructions are helpful:
julia> @code_llvm foo(4)
; @ REPL[8]:1 within `foo'
define i64 @julia_foo_17575(i64) {
top:
; ┌ @ int.jl:52 within `-'
%1 = add i64 %0, -2
%2 = mul i64 %1, %0
; └
; ┌ @ int.jl:54 within `*'
%3 = shl i64 %0, 2
; └
; ┌ @ combinatorics.jl:27 within `factorial'
%4 = call i64 @julia_factorial_lookup_17503(i64 %3, %jl_value_t addrspace(10)* addrspacecast (%jl_value_t* inttoptr (i64 140302197922128 to %jl_value_t*) to %jl_value_t addrspace(10)*), i64 20)
; └
; ┌ @ int.jl:53 within `+'
%5 = add i64 %4, %2
; └
ret i64 %5
}
Also, I encourage you to note that nowhere did I manually put a type annotation in my definition of foo. foo is a generic function and will compile new methods whenever it encounters a new type.Also what Julia does is not exactly type inference. It is rather a form of specialization. Type inference just gives you the type for an expression using just the syntactical representation. Nothing more.
As an example, if I was to invoke your function f with a floating point number, would Julia not happily specialize it for that too (assuming factorial was defined for floating points or taken out of the equation)? The thing here is, that a type checker needs to know about all the specializations a-priori.
Specialization will only occur if there are more than one "f" that handles your type (for example both "f" with a number and with float), and then the compiler will choose the most specialized (float). If you write your Julia program like you would with a static language (for example using over-restrictive types like just float and int) it will return the same error as the static language as soon as it can infer the types (when it leaves the global namespace and enters any function). Using over-restrictive types (and declaring return types for methods) is considered bad practice in Julia, but that's a cultural thing, not a limitation of the type checking.
It is not a type check if it happens at runtime.
But I understand your point, it's not really like a static language checking as it can't do without running the program, so it's better than fully interpreted (a unit test would catch errors even in paths it didn't take) but worse than static checking (as it can't compile 100% of the valid program at once).
Yes, but I could then define
bar(x::Int64) = x^2 - 1
and this would create a method which only operates on 64 bit integers.Now I want to be clear that I'm not saying Julia is a static language, I'm just saying that once specialization occurs, things are much more like a static language than one might expect, and indeed it's quite possible add mechanisms for instance to make a function error if it can't statically infer it's return type at compile time.
Here's a hacky proof of concept implementation, but one could get much more sophisticated: https://stackoverflow.com/questions/58071564/how-can-i-write...
If you're in a field where symbols have existing meanings, it's asinine to make your code clunky and harder to read by not using those existing meanings.
Python is my main language, and I struggled mightily to get my head around some aspects of Swift (e.g. randomly sticking question marks in different places until it was happy). As an aside, I also found the API incredibly verbose, the documentation poor, and Xcode to be a bad IDE.
In contrast, I enjoyed Julia - while the documentation isn't great either, after a single afternoon of learning (with plenty of Googling) I was able to port code over from Python and have it run perfectly.
Doesn't matter anyway, it's a google project, it will be abandoned before 2022
If people start working with Julia now, they'll be able to pick up that wave when it happens.
(I honestly can't tell if this is snark or a realistic and plausible prediction.)
As the article points out, Google did consider it. IIRC it came down to Julia and Swift in the end. And, given Chris Lattner was leading the effort, there was only really going to be one answer. There's clearly some merit to that: they were expecting to make changes in the compiler (again, if I remember, for e.g. optimising GPU code). If you're going to change the compiler, it's pretty compelling to opt for the language that one of your team designed. And it's not clear (to me) what the implications of commits into the Julia master tree would have been.
That doesn't generate a community though. It's yet to be seen whether that will happen. It would take the level of resource that very few firms can afford to dedicate. Google is one of them: though its patchy record on committing to long term endeavours means it's definitely not a slam dunk. And Lattner leaving further detracts from that confidence.
I'll be interested to see what they come up with. Still think it's a pity they didn't choose Julia. But it's not my project so I don't get to choose.
It might be compelling if that person sticks around. But given Chris Lattner already moved on, it doesn't seem all that compelling after all.
I still think it is not too late to admit that was a mistake, and change course to Julia. Especially Chris Lattner has moved on.
2 - S4TF is only usable in Google Cloud notebooks
Is this really true? Many data scientists don't come from a development background. Low level swift code (what's likely to be in TF) can be as obscure as C++. Will the 'industry lag' without users having to dig into the library and understand the code?
True, but I didn't encounter it in the wild yet. Usually Swift used in "Python mode" is fast enough.
It's not an inherent part of Swift the language, and efforts like those at Google and the open source development of Swift can develop more modern and suitable replacements for those libraries.
I like python, but man I really don't want to make large systems in it. Swift is a great language, and imo the biggest thing holding it back is that it's intertwined with a lot of Apple code. But it doesn't need to stay that way, and for that reason I applaud the efforts to move it beyond just an "app creation" language.
In retrospective you can really say Swift was a bad choice for the project because the time to market was much slower than it could be vs e.g choosing Julia. The other thing they didn't take into account was the actual market, that is, the Data Science ecosystem in Swift is non-existente, you have an excellent Deep Learning library standing alone without a numpy, a pandas, a scipy, a opencv, a pillow, ect, which makes doing real application with it nearly impossible.
That said, Swift as a language is amazing, doing parallel computation is so easy, not having a garbage collector makes it super efficient. Its the kind of thing we need, but the language right now is not in the right state.
But the python example doesn't make me trust the rest of the article. It is clearly a swift example, translated verbatim to python.
Idiomatic Python would be this:
import time
for it in range(15):
start = time.time()
total = sum((it,) * 3000)
end = time.time()
print(end - start, total)
Which is shorter, and way faster.Now of course, Python is slower than swift (although a numpy version would not be, but I get it's not relevant in general purpose machine learning). But misrepresenting a language is not a good way to make a point.
The objective of the demo was not to see which language could sum up a bunch of numbers the fastest. You could keep optimizing that until you are left with just `print(<the resulting number>)`. The objective was to have a simple example of looping over an array a bunch of times. The only reason I ended up summing the numbers in the array and printing them was so that LLVM wouldn't optimize it away and be unfair towards python. I actually wrote it first in Python tbh.
let clock = Clock.system
let now = clock.thisInstant()
there just isn't a nice core library interface yet.I find optional?s & protocol-oriented programming enable clear & concise mental models. I like swift's syntax better than Java/Kotlin, C++ & Rust imo. And it's damn fast.
This reminds me of concurrency and transactional memory in Clojure. You can have all of those things at a library level, but building it into the language... Well it kinda FORCES you to deal with them, for good and ill.
It's like trying to put a car into orbit or smth...
Because for mobile devices you don't want to ship LLVM. Plus it's easier to package binaries than scripts which usually require Docker to deal with dependencies in a sane way.
I'd argue that Julia is much closer to having a good static compilation story than Swift or Nim are to having vibrant scientific ecosystems.
> Plus it's easier to package binaries than scripts which usually require Docker to deal with dependencies in a sane way.
Julia does have a really good story for reproducible dependency management. We learned a lot of valuable lessons by looking at Python's dumpsterfire.
I've been doing software development in data science, large scale optimization and machine learning for over 15 years... I've needed automatic differentiation in my language .... exactly never. I mean, most of the languages I use regularly are capable of it, and it is a neat trick; it's just not that useful.
The best part of this article is Yann and Soumith twittering they need Lush back (and not because of automatic differentiation). I agree; it's still my all time favorite programming language, and I don't even fool around with Deep Learning. https://twitter.com/jeremyphoward/status/1097799892167122944
Hi, I'm working on making a tensor lib in Rust (think numpy + autodiff) to learn about these topics. There isn't much information online about how projects like numpy and autograd work under the hood.
Do you have any ideas/tips/resources about how it could be done?
http://www.win-vector.com/blog/2010/06/automatic-differentia...
I agree that cpython threads are not parallelism, but python still comes with built-in support for multiprocessing and I've been earning my bread using that within the past two years, so unless you have to use in-process parallelism for some weird reason, python and your OS scheduler of choice has you covered there.
Needing shared memory paralellism is a weird reason now? Pretty much any parallel algorithm that's not embarassingly parallel is going to perform better with threads able to share memory than with message passing between processes.
OTOH, those heroics do happen, and been OK so far. Accelerating differentiable programming is basically an extra transform layer on accelerating data parallel programming. Thankfully, our team writes zero raw OpenCL/CUDA nowadays and instead fairly dense dataframes code. Similar to async/await being added to doing a lot for web programming on Python, curious what it'll take for data parallel fragments (incl. differentiable.) If it wasn't for language resistance for UDF + overhead, and legacy libs around blocking, we'd be happy.
- Parallel matrix multiplication may be embarassingly parallel, you have a reduction step that is not trivial to parallelize across processes. Also you need to take care of register tiling, L1 cache tiling and L2 cache tiling. It is way easier to do this in OpenMP
- Parallel Monte Carlo Tree-Search: it's much easier and more efficient to spawn/collect trees with a proper spawn/sync librairie.
Is it because it's easy to write?
Honestly I am curious as an outsider to the Python world.
First: "All these usability problems aren’t just making it more difficult to write code, they are unnecessarily causing the industry to lag behind academia."
Industry lag behind academia? Seriously, have you any idea of the amount of data that industry crunches these days, while some in academia still think half a GB is "big data"? Or the amount of money there is in the kind of ML that industry does (which is the main reason why anyone cares about this field at all).
Also, your whole post is about innovation going on at ... google. Then apple. Not academia.
Secondly: "swift is fast", point taken. Then you go off on a drool about your favourite syntactic sugar. My own experience is that this is exactly the kind of thing you do not need in an enterprise-grade product.
For research this seem indeed like a bad fit.
(Side note, named functions in Swift are indeed also closures:
let x = 10
func f() {
print(x)
}
f() // Okay; prints 10 as you'd expect
)Way to bury the lead there, haha.
The following Common Lisp code:
(defun bench-array ()
(declare (optimize (speed 3) (debug 0) (safety 1)))
(let ((arr (make-array 3000)))
(declare (dynamic-extent arr)) ; stack allocate
(loop
:with sum fixnum = 0
:for i fixnum :from 0 :repeat 15 :do
(time
(setf sum
(loop :for k fixnum :from 0 :repeat 3000
:do
(setf (aref arr k) i)
:finally (return (loop :for l fixnum :across arr
:sum l fixnum)))))
(print sum))))
compiles in 0.01s and runs in 5us per inner loop on my system (SBCL).It has array bounds checking enabled (safety 1 declaration). If I remove it (safety 0), runtime improves to 2-3us per inner loop.
I'll take Common Lisp over Swift any day of the week :-]
This reminds me of the initial 'Dart' hype and promises.
Really? Doesn't that depend on the modules, the underlying code? The problem?
But even so, Python is a little slow when you start thinking about threading. You'd be better off using Rust or Go or something rather than the half-baked support found in scripting languages.
It's a long time since I've heard that phrase used disparagingly. Didn't we all decide to use the term "dynamic languages" just to avoid the judgemental overtones associated with "scripting".
I didn't mean to disparage. While I did intend to scrutinize Python's multiprocessing (my experience with it was less than fun), my use of the term "scripting language" was entirely subconscious.
But I do have an anecdote.
A decade ago I was using Python and Flask for building web apps. Now I use Rust instead, and many of my collages are choosing to use Go.
I still use Python for scripting and now also use it for ML. But I wouldn't use it for the web anymore.
I think the landscapes and use cases are shifting since there are new tools available. Python is doing new things (pytorch, tensorflow), but less of the things I used to use it for (Flask, Django, ...)
I think the "judgemental overtones" are in your head, not in the speaker of the word "script".
Just use something actually free, or create your own.
Differentiable programming in the end is just a way of making something that you can already do better (just like you could create neural network long before theano/tensorflow/torch but it was not as streamlined). With a differentiable programming approach you can get something as dynamic as pytorch, with the performance optimizations and deploy capabilities of a tensorflow graph and with an easy way to plug any new operation and it's gradient by writing in the same host language (so no need to learn or restrict yourself the tensorflow/pytorch defined methods/DSL).
You don't even need to change the compiler or define a new language for it. Julia's Zygote [1] is just a 100% Julia library you can import, for which you can add at any point any custom gradient even if the library creators never added them, and them run either on CPU or GPU (for which you can also fully extend using just pure Julia [2]). And of course, you can also use a higher level framework like Flux [3] which is also high level Julia code.
I think the heart of differentiable programming is just another step in the evolution, from early (lua) torch-like libraries that gave you the high level blocks to compose, to autodiff libraries that gave you easy access to low level math operators to build the blocks to the point where you can easily create your own operators to create the high level blocks.
[1] https://github.com/FluxML/Zygote.jl
There would be domains where this would be a major advantage and (probably) applications where it wouldn't be particularly helpful.
Best to think of it as another tool in the engineering toolbox.