Will Scientists Ever Move to Python 3?
jakevdp.github.com
jakevdp.github.com
As far as 2to3/3to2, that entire process in my opinion will be going away; once you've pegged Python 2.6 as your bottom version, it's straightforward (though not without effort) to produce a codebase that runs in Python 2 and 3 without any changes. I've already done this for my template language Mako, where version 0.7.4 has been enhanced to support Python 2.4 all the way through the latest 3.3 without any 2to3/3to2 step.
then read the source code (very short) to six: https://bitbucket.org/gutworth/six/src/64891e65200f9853aa64d...
I don't actually pull in six, I just include those functions which I need locally, usually with some tweaks.
Moving your current/new development to Python 3, to a scientist, sounds a lot like "So there's this new font. It sort of works like the old one. No it doesn't let you write any new equations, and it may or may not complain about being used near papers written in the older font. We know you don't particularly like fussing with typesetting, but how about you invest in this maybe-compatible thing instead of spending your time doing your science? We promise you, the professional typesetters have moved beyond that stodgy Yourfont2.7 you've been using to write your papers."
Letting major APIs return iterators or views instead of lists just introduces unnecessary complication. Most people doing data analysis don't even know what these structures are but they definitely know lists.
Scientists have to deal a lot with bytes or strings of bytes but never Unicode. Python3 treats Unicode as first class citizen as opposed to raw strings of bytes like in Python2.
Sometimes you have to convert a lot of clear text data formats and needing to use 'print(x, end=" ")' instead of a simple 'print x,' makes me cringe every time. Printing something is substantial why shouldn't be a statement.
Finally the loss of performance in Py3k is the straw to break the camel's back. I use numpy because it is usually faster than the stuff I wrote myself in C. I use python because I got the best performance without having to care much about programming.
I want to be able to reproduce results I or other people did 10 years ago. Maybe poeple in 100 years want to do that. A language that breaks backward compatibility for trivial consistency issues is definitely not suited for that. The easiest solution would be to stay with 2.x for ever or choose a more stable language.
However, one way to get people to upgrade is to have BACKWARD COMPATIBILITY. Yes, I know, Python developers consider backward compatibility some sort of noose, but it is the only thing that permits forward-going change without losing the existing user base. It would be easy, too: at the beginning of each file, have """Python 2.4""" (etc) and everything would "just work". There would be no need to wait for anybody to make libraries compatible with new versions.
The Python developers have done us all a huge disservice by ignoring full backward compatibility. And while I appreciate all of the time and effort they have donated to the community, I believe this one single issue, backward compatibility, matters most.
Python 2 Unicode support is not like that at all. Unicode strings themselves work pretty much like they do in Python 3.
The problem with Unicode & Python 2 is that there is also a non-Unicode string type that is used for legacy reasons in many places and mixing the two is error prone because of the implicit str<->unicode conversions that only work for ascii chars.
So yeah, 2018 may be a lower bound for the point at which the scientific community moves to python3. But that doesn't mean it won't happen.
Unfortunately, scientific programming is highly dependent on libraries. So, in my field (bioinformatics), I'm basically constrained to Python or R. Python is the less shitty of the two. You can see how thrilled I am about this situation.
The classic example is keeping multiple MPI implementations around; but I know scientists who keep multiple versions of several libraries in their home directories depending on which apps they want to use, and just select the right module groups.
That said, I've never tried this with anything as fundamental as libc...
Then, Google switched to having a redundant set of system libraries (including libc) installed, and periodic numbered releases of all of those libraries. I don't rightly recall if Google just had all of the distributed jobs, GWS, etc. run chroot or if some of them just had LD_LIBRARY_PATH not include the usual system directories. (I hope most things were run chroot.) The path to libc definitely contained a directory numbered for the release. That way, the OS of the servers could be updated independently of the libraries used to run GWS, the indexing system, etc. This helped reduce risk and testing effort when upgrading things.
Some third parties compiled libraries for us, against our setup. When one of the versions of the system libraries was about to be removed from all of the servers, I remember editing the RPATH in the ELF headers of a third party library to make everything keep linking properly. (Edit: ... and without needing to ask them to recompile everything for us.)
Anyway, it most certainly is possible (and often desirable) to have one set of libraries for the OS and all vendor-supplied binaries and a different set of libraries (including libc) for all of the code you write. Of course, Google has a lot more resources for this sort of thing, but it's not difficult to set up a chroot environment for your computations, or at least modify LD_LIBRARY_PATH.
1) http://python3porting.com/cextensions.html
2) http://docs.python.org/3.0/whatsnew/3.0.html#build-and-c-api...
In the scientific computing world C extensions are used everywhere and moving Python 2.x c extensions to Python 3 is a significant roadblock.
It's something to think about with your software choice: every line is a legacy line for someone else to understand and compile/interpret/run. Can they do it in the future? Do you care? If so, think hard and- I recommend- choose a language which has a formal standard with multiple implementations.
[1] Obviously innovation is useful. But reinventing wheels into the same design isn't. Obviously Haskell and other research-y languages will change. I'm talking about production languages for production environments or for other things persisting over the 5-20+ year spans such as science.
OTOH many Common Lisp people lament the fact that the standardization stopped with the 1994 spec, saying it led to the fragmentation of the implementations with mutually incompatible extensions.
Do you think it would have hurt Common Lisp if there had been one dominant implementation that everyone followed and reused libraries from, like with Python?
There is room for an update, I think. Probably something to do with thread memory semantics. But the language itself provides facility for a great deal of extension. E.g., the default threading API (bordeaux-threads) is be identical cross-implementation, although the exact Lisp system calls will vary.
You have to be careful to consider what exactly needs to be changed and what can simply be added by a macro & library function. Certain guarantees relating to memory will fall in that area.
Of course I do not think CL is perfect. For one thing, it could have used a respin post-CLOS to have a smoother type system and better interfaces.
> Do you think it would have hurt Common Lisp if there had been one dominant implementation that everyone followed and reused libraries from, like with Python?
Yes. There's only 1 python: python.org's python. Never mind Jython, IronPython, and PyPy. We all are tied to CPython. :-( It defines the situation, regardless of the docs. Really not ideal for the Python world.
Having a regularly evolving standard means that you have to port your code constantly to ensure you're up to date with the latest hotness... or even be able to interrelate with newer code. One of the great strengths of the POSIX standard & C is that C89 has stayed constant and available on *nix for the last two decades; C itself has stayed mostly constant for about three and a half. The cost of forcing regular change is enormous and may well destroy a community.
rachelbythebay, a blogger I follow, wrote a post I can't find right now, on the problems when you rely on this sort of situation. She does a lot of sysadmin work, usually (I guess) in C++, and was astounded when she learned that just upgrading your system might shatter your software in Perl or whathave you. I have that opinion (based on my experience) as well. Having to have had to contend with Python 2.4, 2.5, 2.6, and 2.7 all in the same company for the same codebase, I have this conclusion: do it right the first time, second if you can swing it. Common Lisp succeeds at that (n.b., before CL, lots of Lisp fragmentation existed). Python fails. C wins. Perl 5 is trying to have run-time machine switching based on specified version (may have details wrong there). Ruby fails. C++98 is "ok", C++11 may be a bear.
I do think that a language should be excruciatingly small and very effective, and libraries should be built around that language; this allows libraries to be evolved/replaced without having the core language altered. C got this right. Scheme from R1-R4 also went this road, but as it is an academically driven language, its been in the shadows by and large. I guess newer Schemes have bigger standards though.
In my opinion, Fortran (meaning Fortran 90/95/03/08) gives both the ease of native array syntax, like matlab, and leaner than NumPy, and the speed of compiled language (like C, somewhat faster), and the safety of a typed language.
For numerical computing with lots of 2-, 3-, or 4-dimensional arrays, I don't really see any other language to be on par with Fortran.
Don't get me wrong, I have nothing against Python 3, and many of the changes seem fairly sensible, but I shouldn't have to fight against it to do something that works perfectly well with minimal effort in 2.7. I guess that makes me a curmudgeon, but I don't want to dick around with Python from a theoretical standpoint, or take a great deal of time exploring its features. Since I usually only use it for quick-and-dirty scripts, this is my ideal use case. I honestly don't care how much better iterators and GC are handled with it.
I have nothing against Python 2 or 3. I'm just relaying my own experience. We used a lot of C++ and Java too. And the Java guys were fond of Groovy. I did mostly C++ and Python.
http://www.python.org/dev/peps/pep-3107/
(More specifically, Python 3 includes generic "function annotations" that look a lot like type declarations do in other languages, but can be used for other purposes as well. The idea is to offload typechecking to a library, so you could eg. include type inference, dimensional analysis, preconditions or DBC, etc. as desired.)
The argument is you need the major library authors to start primarily coding in Python3 and using 3to2 to backport to Python2. Until this happens, the quality of Python3 libraries will suffer and there will be disincentives to migrate.
(That's sarcasm, BTW.)
Since scientists are considered "good" or "bad" by their paper writing and fund raising productivity, and not the ease-of-use or quality of their data-analysis scripts ... well, you can guess what happens.
2. Their job is to solve scientific problems, not solve them in ways that make computing folk feel good.
3. Arguably the scientist who wastes time fretting about the technical details of software rather than their research or their grant, is the short sighted one.
Actually, it should be. Whole fields can and do suffer because of generally crappy software practices. (Medicine) It should be thought of as not pissing off your suppliers.
Fortran continues to be faster than C on certain types of problems and libraries like lapack continue to use it for that reason. Changing languages just to be "hip" seems to me to be an awful thing to do.
Indeed, seeing as you've just contributed an ignorant comment here yourself. After all, you couldn't possibly be referring to me, since I'm not a lay person, and actually have extensive professional experience in the scientific software community.
> Fortran continues to be faster than C on certain types of problems and libraries like lapack continue to use it for that reason. Changing languages just to be "hip" seems to me to be an awful thing to do.
You might be surprised to learn that it's possible to use FORTRAN in conjunction with C and Python, which removes the need to write all code in it. There are usually only a select few methods that need to be optimized. Those can be written in FORTRAN, and are actually often available in library form.
But none of this means you have to write entire software packages, including the high-level logic, in FORTRAN. That's something that I see shockingly often in the scientific software community.
1) Sage, which I use, is a huge project incorporating lots of other scientific software and using loads of Cython. It simply takes time to migrate all the dependencies and then Sage itself to Python 3, and it's not really a priority for anyone.
2) It's not really something that anybody seems to focus on a lot. Scientists mainly worry about science, minor differences in the programming language used are generally of little concern. Migrating to a different version is therefore not something most of them would spend time on, unless forced to by external circumstances.
And the main reason some scientist migrated to Python is probably that Matlab is non-free, and the language of Matlab sucks. Julia might attract both Python users and future Matlab refugees.
I can (kind of) understand your complaining about the print statement going away, despite its inelegance. However, if the semantics haven't changed, why change the name?
Don't exclude Perl! :-) In certain scientific subdomains, Perl is still widely used, also for historical reasons...