However, per the blog post, this version of Pyston is closed source. So CPython won't be interested, and many others won't be interested either.
However, per the blog post, this version of Pyston is closed source. So CPython won't be interested, and many others won't be interested either.
This has been discussed extensively: https://news.ycombinator.com/item?id=11125769
Pypy has done yeoman's work in improving performance while maintaining compatibility with an impressive amount of the ecosystem, and even still there are many important packages which aren't compatible with Pypy and for which Pypy-compatible analogs don't exist (or aren't supported/maintained). For example, the only Pypy-compatible Postgres drivers were unsupported last I checked.
Moreover, the Python community (or at least its leadership) seems to have very little energy around tackling these longstanding problems. Meanwhile, there are many other languages which are not only performant, but which are rapidly encroaching on Python's historically unique(ish) "easiness" (in the sense that Python is considered "easy", which is to say for people who don't have to manage build/deploy/packaging/etc or otherwise have performance issues). Further, many of these languages continue to improve at a remarkable pace, while Python is content to rest on the laurels of its scientific computing mindshare--and given the rather poor nature of the numeric computing package APIs and their somewhat low performance ceiling, I don't expect Python to be so dominant in this domain in another 5-10 years, especially as more companies need to figure out how to productionize scientific workloads.
https://commandcenter.blogspot.com/2012/06/less-is-exponenti...
If this version of generics finally makes it, then I care, otherwise only when dealing with Docker and Kubernetes eco-system.
Silence and distance seems to be a common echo of failed projects.
Don't forget Rob Pike created Sawzall at Google first (2005). [2]
[1] https://golang.org/doc/faq#history [2] https://research.google/pubs/pub61/
It doesn't really support or attack your claims. It's just a thing I wanted to share.
Google most certainly did not.
http://qinsb.blogspot.com/2011/03/unladen-swallow-retrospect...
And anytime you point this out, people will trot out a toy problem in Cython and try to prove you wrong, which has zero relevance to real world issues. Hell, Cython cannot even compile something as basic as Numpy FFTs, something that is absolutely critical in signal processing.
Google put some effort, not that much, into Unladen Swallow, and the people who worked on it admitted that it didn't achieve the goals they set. The Python community was well on their way to creating Python 3 before work started on Unladen Swallow (so it was not in any way "Python's response"). As other comments pointed out, Google did not hire Rob Pike to "do" Go. Go was initially a side project of a few people and had nothing whatsoever to do with Python. The Go team was surprised when Python users started migrating to it. Google never used Python for web development, with one major exception, Youtube, an acquisition. I also don't understand how Google has "moved on" from Python (this seems unlikely) and how what they use Python for internally has anything at all to do with the fate of Django and Python webdev.
I don't mean to pile on, but I kinda have to ask, what were you thinking when you wrote this comment? Do you believe all those statements, or were you taking huge liberties with facts and making guesses to tell a story that supports a statement you wanted to make? You seem like a valued member of the community and this kinda makes me distrust the accuracy of everything else you write. (And to be honest, makes me distrust more of what I read in general, which is probably good but sad.)
A couple references for history/dates:
https://en.wikipedia.org/wiki/CPython#Unladen_Swallow https://www.python.org/dev/peps/pep-3000/
WRT poor APIs, I'm talking about things like matplotlib or pandas or etc that take a whole slew of arguments and try to guess the caller's intent by inspecting the types of the arguments. The referent isn't "some other scientific computing API" (although I'm sure there are some sane scientific computing APIs), but rather "other APIs in general" since there's nothing inherent to any particular domain that demands this kind of 'magical' API.
WRT 'hardly anyone is CPU bound'--the context is numeric computing; what are people bound by if not CPU? I've seen several projects where web endpoints were timing out while grinding in Pandas, largely because there weren't good options for taking advantage of multiple processors. Based on prototypes I did, I'm confident that other languages could serve those requests in single-digit seconds if not sub-second.
> Nowadays, my rule of thumb for pandas is that you should have 5 to 10 times as much RAM as the size of your dataset
[1] https://wesmckinney.com/blog/apache-arrow-pandas-internals/
Web application using Pandas and "highly optimized web application" would seem to be nearly disjoint sets...
Also, if an endpoint is spending minutes to respond, then I would think actually profiling the application would be a good start. Maybe researching prior art in the problem domain would be good too. If nobody can be bothered to explore the several solutions to distributing pandas computations over multiple cores, like Dask, and get the NPV of just buying more or faster cores, then “Python sucks” isn’t your problem.
“My application is slow, the language sucks!” Doesn’t indicate a very serious investigation into the problem.
HTTP is pretty complex, but that's neither here nor there. The relevant bit is that there is no domain for which guessing caller intent based on reflection over argument types is appropriate.
> “My application is slow, the language sucks!” Doesn’t indicate a very serious investigation into the problem.
I was pretty explicit above and elsewhere in this thread about why Python's performance is miserable; I'm not sure why you would invoke such a poorly constructed straw man when everyone can look upthread and see my actual arguments.
But once you've done the cleaning/exploration, you should move any heavy computing to a high-performance library like numpy.
Plotting seems to tend towards magic because plots are basically art, with all the desire for aesthetic customization that applies, and it's a very common task so users also want brevity (magic). The result is a plot() function with a gazillion options hidden behind keyword arguments.
I agree that matplotlib has a sprawling interface, and this can be annoying, but I'm still not sure what "guess the caller's intent by inspecting the types of the arguments" means. Sure, the functions have multiple call signatures, but that's not exactly unusual in libraries or languages. I don't understand the context that brings guesswork into the picture. Skimming the manual—are you using the data keyword argument and hitting the `plot('n', 'o', data=obj)` ambiguity [0]? Or calling plot through `pyplot.plot` &c. (which rely on state) instead of `Axes.plot` &c.?
Asking because if there's an interface trap I'm unaware of I'd like to learn about it before walking into it blindly.
Pandas I sort of agree with; I personally find it harder to remember how to use pandas than dplyr, despite using pandas more often and spending more time reading the pandas documentation. I also find it inconvenient to represent missing values in Pandas (`None` and `NaN` are overloaded, and `None` forces the `object` dtype). But maybe the problem is on my end.
[0] https://matplotlib.org/3.3.2/api/_as_gen/matplotlib.pyplot.p...
R provides an interface to dataframes and matrices, whereas numpy is just for matrices (and their generalisations, arrrays). I think the appropriate comparison is between base R and Numpy + pandas. (FWIW, I agree with your major point, but then I learned R first so that may be biasing me).
I will say that the differences between python and R syntax are almost trivial; I've taught classes with python and R example scripts that are almost _exactly_ the same, and run in both languages
What are those languages? I may have a blindspot, but the languages that get enough buzz for me to notice are either not competing with Python in important dimensions (e.g. Rust) or have a narrower focus (e.g. Julia). Elixir maybe? JavaScript and its derivatives? But these lack the scientific programming ecosystem that's helped drive Python recently.
I have no doubt that languages exist that might fit the bill of being both as easy as and more performant than Python -- there are a lot of languages out there. But I'm unaware of any whose mindshare has been sufficiently growing that it threatens Python in the mid-term.
If I'm off base please let me know. I'd love to find a viable competitor to Python that's strictly better than it.
I strongly recommend Go as a better Python. Personally, I think it's easier to write than Python (although people who care very little about correctness will be bothered a bit by the type checker), and the tooling is many times better (single-binary deployments, great dependency management, etc are awesome). Also, the performance is about 100-1000 times better for serial execution, and Go's goroutines allow you to take advantage of multiple cores much more easily than with Python.
JavaScript and TypeScript are similarly easy-to-use, performant languages with a better-than-Python tooling story. I've also heard similar things about Elixir, Closure, and Kotlin.
I've heard Go described this way several times, but I've found it to be a significantly lower-level language than Python.
For example, it's much more verbose. In this recent blog post [0], the author converts some C++ code to Go - and it gets longer. 57 lines of C++ become 65 lines of Go. The same code in Python is about 20 lines.
In particular, the author's Go code requires five lines to do the equivalent of Python's `with open(path) as f` and four lines for the equivalent of `word_array = list(word_counts.items())`:
f, err := os.Open(path)
if err != nil {
return err
}
defer f.Close()
// [...]
wordArray := make([]WordCount, 0, len(wordCounts))
for word, count := range wordCounts {
wordArray = append(wordArray, WordCount{word: word, count: count})
}
[0] http://jmoiron.net/blog/cpp-deserves-its-bad-reputation/This is crazy. Many of those lines are closing brackets or whitespace. But moreover, optimizing for characters or LOC is absurd. Optimize for maintainability or readability, at which point Go is at least as good as Python (I would argue better). Optimize for tooling, especially package management and build tooling--Go is many times better than Python here. Optimize for performance--Go is literally hundreds or thousands of times better here. Optimize for breadth and quality of ecosystem. Optimize for deployment story (single small artifact vs hundreds of megabytes of dependencies). These are the things that matter, not lines of code.
The type-assertion escape hatch has always been largely sufficient for all but the most performance-critical projects (of which I have exactly 1, and it was a side project), and runtime panics due to failed assertions are quite easy to eliminate via wrapper types.
My advice: just go for it. You almost certainly won't miss generics.
(Edit: I've broken my own rule and given advice without first asking what kind of programming you do. My assertion holds for most run-of-the-mill stuff, e.g. writing REST interfaces, network servers, etc. If you're within 2-sigma of the industry, you won't miss generics.)
Dynamic languages do that implicitly.
Implement max in Go with interface{} without casts and reflection:
def max(a, b):
if a >= b:
a
else:
b1. The abstract idea of writing an algorithm that supports a variety of types. This allows for casting, reflection, dynamic typing, etc.
2. The specific idea of a type system that allows for parameterized types (aka "typesafe generics"). This definition excludes castng, reflection, and dynamic typing.
Typically "typesafe generics" is what people talk about when they discuss "generics", but since you chose to pick the "dynamically typed languages are generic" nit, I assumed you were talking about (1).
Some of Python's strengths that I work with regularly are dynamism, easy data exploration, visualizations, succinct & customizable syntax, extremely strong data science libraries, REPL/Jupyter, C bindings, easy to use packaging solution via PyPi (bet some people are going to disagree with that one), quick to prototype code. If you really need performance in Python, you can probably get it out of the box (if you can run deep learning with Python, performance isn't a limitation).
I love Go and use it for a bunch of projects, but I've only once wanted to move a project from Python to Go and that was a performance-centric CLI that was only written in Python originally because of how quickly it let us prototype in comparison to Go.
From my perspective Julia is sacrifising some amount of "duct tape UX" to gain speed, and that's the wrong direction.
Whenever we need more speed, we just pull the slow bits down into compiled languages, and scale them out to many cores with solutions like MPI.
R is another language that is mainly duct tape for stringing pieces of compiled code together. If it had a better UX for developers than Python, I think we would see it dominating much more today, without having any speed advantage.
This only works sometimes--for problems that allow you to do a relatively large amount of computation in the compiled language to justify the cost of marshalling Python data structures into native data structures. For matrices of scalar values, this works well. For many other problems (consider large graphs of arbitrarily-typed Python objects, or even a dataframe on which you need to invoke a Python callback on each element). If you rewrite a big enough piece of your Python codebase in the compiled language, then it will work, but now you're maintaining a significant C/C++/etc code base and the bindings and the build/packaging system that knows how to integrate the two on all of your target platforms. Python really doesn't have a good answer for these kinds of problems, and these are by far the more common case (though perhaps not more common in data science specifically).
Interestingly, R is probably a better UX for statisticians/data scientists than Python is (almost all the good parts of Numpy/Pandas were in R first), but it really suffers from not being well known by developers.
To be fair to R though, it's much, much easier to deploy than Python, which is a shocking indictment of the current Python packaging ecosystem.
What particular language features of Julia make that trade-off? (Not a rhetorical question, I'm not disagreeing with you, just curious; I'm familiar with Python, not really familiar with Julia.)
These aren't fundamental issues with the language and will be solved by tiered compilation (there's already an interpreter mode, just have to integrate that with normal use), separate compilation and incremental sysimage creation which works with the package manager.
More than performance (pure Python can be too slow, but NumPy is usually fast enough), this is what I miss when using Python.
There's currently no good solution for bundling Python code and the interpreter into a single binary. PyInstaller works, but the resulting binaries write lots of files to `/tmp` every time they run, which is a hack that leads to long startup times. PyOxidizer avoids this, but includes the entire standard library in every binary, making them too large by an order of magnitude.
It's still a little too wild-west for my tastes right now, but I want it to improve as I think it's got real potential.
Data science may be at the heart of Python's strengths, but it's simply not accurate to suggest Python (whoever that is!) is resting on its laurels, as evidenced by the very active, albeit sprawling, package ecosystem. And while the CPython devs seem to have found their stride in 3.x releases.
It's also inaccurate to suggest numeric computing in Python has a "low performance ceiling," considering you can get near-C performance via JIT or AOT using a package like Numba, and in most cases it's not even necessary because of Numpy and the many other highly optimized compute packages that can do most of the heavy lifting.
I think the main draw of Python is not just that the syntax and language features are approachable, but that the package ecosystem is so broad and active, you are likely to make lighter work of the same job done in another language. I think to displace Python you would have to displace the package ecosystem, which seems as big and broad as it's ever been.
As far as I'm concerned, CPython as an interpreter technology has not advanced in the past decade - and Python has grown thanks to the data science community efforts and excellent library ecosystem you mention, not due to PSF. PSF can only miss so many opportunities before something comprehensively better starts to eclipse it.
I agree performance is important, and I think we have reason to be optimistic, but with the understanding that that level of improvement, even if currently underway, will just take a lot of effort and time. Meanwhile, as someone fairly new to Python and working on several performance critical pieces, I've been pretty impressed with what you can do with the current compute packages (after taking months to work my way through most of them).
In a recent post linked on HN, Steve Yegge basically nailed it:
> How much new software was written in something other than Python, which might have been written in Python if Guido hadn’t burned everyone’s house down? It’s hard to say, but I can tell you, it hasn’t been good for Python. It’s a huge mess and everyone is miserable.
Note to future language maintainers: don't burn everyone's house down.
The main problem I have with Python is its maintainability, especially when you have multiple developers working in the same code base. I would never choose a dynamically typed language again to build something that is going to be more than several thousand lines of code, especially because of how it cripples your IDE which is essential for helping junior developers understand the existing codebase. Type hints help to an extent, but they're just a small bandaid over an oozing sore. Languages like Go are significantly better if you're going to build large codebases that need to survive for a long period of time.
For example, I do a bit of Swift in addition to Python, and sometimes I've heard people try to compare the two. But I vastly prefer Python to Swift when I can afford to.
I also will say that a lot of the complexity of making a Python JIT is all the weird edge cases and special interpreter functionality. To a VM engineer, CPython is a very bizarre codebase; all the complexity is tucked away in corners other than its C codebase.
Your complexity has to go somewhere.
Ironically this also makes CPython less interesting as a reference language - which keeps PyPy and co insignificant in adoption...
https://github.com/python/cpython/blob/master/Python/ceval.c...
CPython does use threads (on Linux, it spawns one pthread per Python thread). However it also has a lock (the infamous GIL) that prevents those threads from interpreting Python code simultaneously.
It's an interpreter implementation technique.
The plan looks very high level at this point, but it looks like Mark is an expert in interpreter VM and JIT technologies. All I can hope for is that he doesn't get blocked by the "keep cpython simple" obstructionism.
There's new boss(es) over the Python project, and things might be changing in the near future.