Why People Should Learn Python
iluxonchik.github.io
iluxonchik.github.io
Indicators for languages which are liked by the HN community:
- Positive stories get many upvotes and are instantly at the top of HN
- Most comments are positive
- Frequent coverage on HN about the language
Indicators for languages which are not liked by the HN community:
- Rants get many upvotes and are instantly at the top of HN
- Most comments are negative about the language
- Rare coverage on HN
Applying those indicators, following languages should have these like-levels (on a range from 1 to 5):
Python *****
Go ****
Node.js *
C#/.NET ***
Java ***
RoR **
Ruby ***
Elixir ***
Haskell ****
Closure ***
C ****
Rust *****
C++ ***
PHP *
Swift ****
This is just my feeling, happy to hear your perception. And this does not reflect my opinion about those languages in any way.Replacing Node.js with JavaScript doesn't help either as the frontend vs server-side vs npm ecosystem vs es2015 becomes a loaded question .)
A lot of HN users are for example startup owners/creators so it's normal that they will prefer easy and semi-automated language like Python.
Maybe something investors should think about? Ehh, who am I kidding...
Developers will handle that. And if they want jobs, startups must be made. As you might know, successful startups are made by less technical people mostly...
I think that's due to people having flash backs to terribly written code in both languages.
For example, PHP apps pre-2005 rarely if ever used frontend controllers, preferring instead to twine in configuration and helper functions into isolated pages that were directly served by Apache. This meant that there was no single obvious place to put things like caching.
Additionally, once upon a time, PHP didn't really have well developed ORMS, so everyone just wrote queries and hopefully parameterised queries---but not always. So if you had a hobby project that took off, then within 3 years it'd be a roiling ball of mud where it's difficult to do basic things like tracing which calls happen per page.
PHP came to prominence early in the web, and it was easy to use so everyone had a hobby project that took off, or had to work within something like that. Hence the hate.
To compare it to Javascript, I think the disdain for both languages is due to their initial nature---they were both written over an short period of time, and evolved in spurts rather than being carefully designed from day one or designed to overcome the shortcomings of another language.
The quality of the above really shows up in the inconsistency of the standard libraries, and the fact that certain semantics were inconsistent or danger zones in practice but couldn't be changed.
---
I do remember though when everyone hated C, and loved C++, then went back to loving C... so maybe programmers are just fickle and none of what I said makes sense. :)
And yes there is a lot of horrible PHP code out there, writting by bad developers. I guess newer or not as accessible languages don't have as many bad devs as PHP/JS.
But if you have good devs, I prefer a good PHP codebase to any other language that I have come across (I tried a lot of them).
In the end good PHP code ends up looking like good Java code I think. But you don't have to worry about everything being nullable.
But then horrible PHP code is something else... Luckily I don't really have to touch that very often.
I have never felt so uninspired by a language. My overwhelming feeling from the get-go was that it was a collection of libraries, which would be fine -- but why create a new language?
PHP has the rare distinction among languages of not containing any novel features or ideas. Not only that, but it borrows liberally from several other languages, resulting in an overabundance of syntax. Instead of a tool for attacking problems, it's a reheated leftover. There is not suggestion of a better or worse way to accomplish some goal.
Finally, the original premise of PHP, interspersed functional code and HTML, is high on my list of really bad ideas. Code written in this style is doubly-challenged in terms of reuse and maintainability. Again, there is again no semblance of better or worse with this approach. If you had Richard Stallman and Linus Torvalds produce the same website with PHP, there would be very little in common between their approaches.
Given the success of Facebook et al, I presume that the interspersed code/html thing has been superseded, and further that the egregious security issues in PHP have been fixed. But the picture in my mind's eye is something that lumbers along despite its thorough mediocrity. I will allow that making PHP adequately consistent and secure are impressive feats.
Sincere question: how am I wrong here?
====== edit ====== Drop extra 'the'
I'll counter that one of the main reasons PHP became such a used language (other than ease of deploy) is exactly the reasons you dislike it. There are multiple ways to do things - that coincide with the ways you do things in many other languages (which it steals from).
You dont have to learn anything radically new to use it, and that makes it very easy to do things quickly.
It doesn't. But if your language ideas are admittedly no better than the next guy's, what's the point?
For me PHP is like Java but with sane nullability. It's like python but it actually has interfaces so that I can do proper OOP.
Its the lowest barrier to entry web language - and honestly, I'm ok with people writing shit PHP. It gets good Ideas out there quickly - allowing the business to grow without a focus on the Development. This creates opportunities for sharp developers to come in and work on proven ideas.
That's one big important point.
I agree with this sentiment very strongly. I understand that there's a lot of hate, flamebait, etc etc about the language, with which from my own experience I agree with or disagree with in fairly equal parts. At this stage, I have written more large web projects in PHP than any other language (for varying reasons). I don't proselytize on why it is (or is not) a good language. But for me, personally, a huge chunk of whether I'm happy or not comes from whether I feel inspired while I'm working. Quite separate from the excitement of solving problems: while working with Rust or Python, I get excited, and am generally happy (even when fighting with Rust's compiler); while working with PHP, I tend to feel vaguely annoyed. Take from that what you will.
I'm of the opinion that the best tool for the job is the one you know how to use, and in the real world of deadlines and opportunity cost, if you get a great PHP dev to write your application in Python it's going to be far more expensive and far worse than if it had been written in PHP. But it's also going to be far less fun to both write and maintain. I think you have to make your language decision depending on your priorities.
> the egregious security issues in PHP have been fixed
This I strongly disagree with. And I don't think it's possible. It's not that the existence of the language includes vulnerabilities, it just makes it extremely difficult to write applications securely. I actually teach an upper-division university class on Network and Application Security, I am writing my PhD dissertation on the topic. I have worked on the internal security team of a very large company, and I have worked as a penetration tester / security auditor. I am very familiar with the vulnerabilities that plague web applications. And I still occasionally find the odd XSS vulnerability, timing concurrency attack, etc. in PHP code that I myself have written. And most PHP applications are not written by people with my background.
Exactly! I was never satisfied with my PHP code, as there does not seem to be a better and worse answer in many cases.
For the record I did not mean to assert that PHP has become secure, I meant that I presumed it had. Some years ago it was a laughingstock on Bugtraq.
Thanks for the feedback!
You can easily create a collection from data and process it. For example using the DB class to get data returns an array of Std objects.
For a legacy app I'm working w/ --it needs to be an array of arrays... I could for loop, and then add each to a new array.. or I can simply do:
$data = DB::table('something')->where('something')->get();
return collect($data)->map(function($x) { return (array) $x });
There's a talk at lara-con on using collections to replace just about all loops, and it's pretty insightful and makes code so much more readable and enjoyable.
PHP by itself can be ugly -- but the Laravel community has created something beautiful from what used to be akin to horse shit.
Static calls and active record are not good OOP. Neither are Std objects.
Have a look at doctrine if you need an ORM. Otherwise read up on DDD and SOLID.
Nobody sanes mixes HTML and PHP code anymore these days. Most people are using Twig for templating or at least something similar. Everything (decent) is object oriented these days.
So your views were correct 15 years ago, but have a look with fresh eyes and you will see something completely different.
But if you narrow that down to actual developers who have a job working on a product, not someone hacking a wordpress template, then it will look much different.
Not saying the average PHP developer is great at OOP, but at least things are moving into that direction. And the code quality has come very far since 2001, just look at at current open source libraries (see packagist.org) and compare them to old PHP code.
If you need anything more resilient server side then I'd recommend using a Golang or Elixir framework like Phoenix which has multi-threading.
Luckily I'm currently working with a PHP7 app, so all is good, but my next app will likely not be PHP at all.
This is my life right now... lol
It's not so bad, but the frustrating part is when I talk to people in interviews about all of the great work I'm doing modernizing (the surrounding) WordPress infrastructure and the tools I've built for it, their eyes just glaze over.
I frequently get told "you're a talented developer and you're wasting your time working at X". Well, at least at my job I get to remove technical debt constantly. I don't feel like that's the case at these places I'm interviewing.
Not a knock on php on my part, PHP is equally suited.
I would have personally probably chosen Flask (python).
Guzzle doesn't look that bad. It looks like requests from Python, but you have to create a client first.
> $client = new Client(['base_uri' => 'http://httpbin.org']);
> $response = $client->get('http://httpbin.org/get');
Or in Python
> response = requests.get('http://httpbin.org/get')
As a hobbyist Node developer (day job is not in Node), I find it hard to talk with other Node developers as there is some build-tool or even full framework that sits on top of Express that I "should" be using but I'm not. But maybe it's just me.
That shouldn't be seen as a mark against the technology. Nobody is forcing you to switch tools.
Except when they do because they're the front end team. :P
(Or just have a rule that as a team you try e.g. '1 new thing' per project).
Again this shouldn't be a reason to hate on the JS community.
I think the dislike comes from the idea that JavaScript-development can be or at least was very messy and that "the JS-people" now try to carry that over to the server-side.
npm also has many issues and node as a server-side VM wasn't very dev-ops friendly in the beginning, if I'm not mistaken.
Note: This doesn't mean I don't like JS or node, I use them daily, though I have never used a production node-server. I'd rather run my JS on the JVM.
This person could ROCK at C++, but who's going to look at his resume and give him an interview?
Sure, PHP has it's faults, but as your article points out, it's designed to give people who aren't full time programmers tools with which to build websites or scripts to perform tasks, and there is a large market for that sort of thing.
And just like that, the cyber-security profession was born.
Admittedly, the author of that post lost me pretty early on when he said that languages don't need to be predictable, it's on developers to learn everything about how and why a language implemented things they way they're implemented, and complaining just means you're lazy. Sorry, but when I can guess at syntax in Python or Ruby and be right(because the languages were designed to be predictable and consistent), I don't want to consult arcane documentation in PHP to figure out why things don't work the way I expect them to work.
The author is also pretty rude and flippant, often when he's not even right. He ignores things like the Fractal author explicitly saying that Wikipedia and Facebook have very smart developers, so he can say that the PHP "community of amateurs" comment was about those developers. Pretty disingenous.
He repeatedly complains that the author "doesn't understand" "loosely typed" languages, but never considers that maybe the author understands them but thinks they're bad design, hence the whole "fractal of bad design" in the name of the post?
But I rarely see any positive mentioning on HN about c#, .NET, C++ or Java.
I think that might be a slight exaggeration. At present, zero of the stories on the HN front page appear to be about Node.js.
If half of all HN front page items were about Node.js on average, the probability of getting (let's say, conservatively) 40 with none about Node.js would be (1/2)^40 or about 10^-12.
Of course people exaggerate and I am fallible. If 30% of all HN front page items are about Node.js on average, the probability of getting at most one such item out of 40 is about 10^-5. Which is, y'know, still much less likely than ten consecutive heads when flipping a coin.
Over time my excitement changed. I assume others may share my thoughts hence the less excitement:
* It is on the JVM and there are lots of languages on the JVM competing with Java.
* While the fusion OOP + FP fusion is interesting I have grown wary of kitchen sink languages. There are so many ways people write Scala. Scala from one project looks vastly different.
* The endless DSLs
* Java 8 is often good enough. Particularly for interop and community support.
* Every time I have to add Scala to a Java project I'm thinking here comes 9MB jar that some classpath scanning thing is going to take its sweet time on (Spring, Servlet 3.0, etc).
But I think most of the reason is Rust has filled that hard core typers niche with the added novelty of unique memory management and competing with other LLVM languages instead of JVM languages. Scala has become "old" like Java and OCaml :) .
* It also compiles to Scala.js (it's actually the independent best compile-to-JS language out there in terms of libraries, tooling and IDE support). Support for compiling directly to native is coming around nicely.
* Scala is quite far away from the kitchen sink, and many of the questionable parts have been deprecated and/or removed.
* I have never seen much DSLs while writing Scala. Sure, it's probably quite nice, but I have rarely seen endless DSLs in practice.
* The jar is around 4MB (2.12) and when writing idiomatic Scala (~ not Spring/Servlet) there shouldn't be any kind of reflection or classpath scanning happening.
As far as kitchen sink I'm still not sure what idiomatic Scala is especially with the likes of Scalaz. Maybe modern Scala is finally abandoning the OOP fusion.
Of course my biggest reason of not liking Scala is completely arbitrary and almost illogical (of course most reasons to (dis)favor things are arbitrary). If I write in an FP language I hate using curly braces. I prefer Haskell, OCaml, and even Lisp syntax for FP. For some reason curly braces makes me think hmm imperative shit. Not to mention Scala has lots of scoping and closure use so you see lots of curly braces.
And that brings me to another point. Scala code isn't really that much more compact (lines of code or even compressed size) to Java 8 (compared to Haskell or OCaml).
[1]: https://mvnrepository.com/artifact/org.scala-lang/scala-libr...
[1] https://mvnrepository.com/artifact/org.scala-lang/scala-libr...
I would disagree with the comment on verbosity compared to Java 8/Haskell/Ocaml. Most of the things I do regularly in Scala would require a magnitude of more code. Ocaml doesn't even have higher-kinded types or implicits, so I don't think it can win much on that front (except of saving a line where Scala's curly brace might be). With Haskell it really depends. There are certainly ways to write much more succinct code, but what I see in practice is that it's not that much different from Scala. Haskell has it's own unsolved issues though: broken typeclasses, lacking module system, poor tooling, and a weekly-changing, terrible dependency management.
HKTs are great for a generic library writer (aka standard library) but I think the over abstract nature ends up confusing your average developer. You basically have your type aficionados obfuscating the codebase for developers who can barely handle generics. It is sort of analogous to using massive reflection / meta-programming in scripting language (albeit far more safe).
Being extremely implicit and abstract comes with a cognitive load.
OCaml's Functors (ie super strict and explicit) on the other hand can get massively repetitive.
I vaguely remember HKTs being expensive compile time. Is that still the case?
The average programmer can enjoy completely consistent names and signatures across different libraries, without having to know that the consistency comes from one method being applicable to various types, not manual labor repetitiously adding the same method to different types which can drift apart at any time without anyone noticing.
HKTs are not that expensive. Slow compile times are usually caused by people doing Prolog at the typelevel in Scala.
I think they're a necessary evil. I need extension methods, ad-hoc polymorphism, and a way to use ad-hoc polymorphism while still writing nice DSLs (i.e. what spray calls the magnet pattern). I'm nervous about implicits, but I haven't seen a better solution.
> HKTs are great for a generic library writer (aka standard library) but I think the over abstract nature ends up confusing your average developer.
There are few things more frustrating than having to copy-paste because your language can't handle the abstraction. There are less cases where you need HKTs than where you need basic generics, sure, but sooner or later you do need them. And since everyone hates copy-paste, they end up using annotations or reflection or string-based templating to generate code, all of which end up more confusing to the average developer. At least with HKTs you still have ordinary values that follow the rules of the language, rather than wondering whether you can still factor out a common method because it's called from an annotated method and an unannotated method, and what annotations you need to put on it to make that work.
> I vaguely remember HKTs being expensive compile time. Is that still the case?
HKTs per se are not expensive. Type inference in the presence of HKTs can be expensive (though note that even in OCaml or pure ML there are pathological cases for type inference - http://cs.stackexchange.com/questions/6617/concise-example-o... ). Scala implicit resolution can be expensive especially when used with recursion; global-uniqueness constraints a la Haskell typeclasses would be one solution, but arguably introduce more trouble than they're worth. One way or another, Scala compilation ends up being pretty slow, which is unfortunate.
Poll: What's Your Favorite Programming Language? (2012) https://news.ycombinator.com/item?id=3746692
Poll: What's Your Most Disliked Programming Language? https://news.ycombinator.com/item?id=3748961
Lisp ***** ***** Python, Rust, Lisp
****' Go, Swift, C, PostgreSQL
**** Elm, Elixir, Typescript, Haskell, OCaml, JS-ES6, Kotlin, Crystal
***' Scala, Lua, Clojure, Erlang
*** C#/.NET, Java, C++, Ruby, JS-ES5, MySQL
** Node JS, RoR, Meteor, MongoDB
* PHP, VBasic, Cobol, CoffeescriptIf we're including MySQL and PostgreSQL, SQLite should be way up there based on the universally positive comments. MongoDB might be ranked too high at 2, though people have begun to acknowledge that was because they tried to use it for the wrong task.
* * * * * WASM
Dare to question the pragmatics of diverging web dev practices into a whole new realm of complexity and drown in downvotes - and a litany of people being downright personally insulting/condescending. This is sort of a corollary to JS being * (I think you ranked ES6 too high) and I've been inured to people hating/ranting against JS to a point - but even I was shocked by the personal rancor that oozes out if you try to have an objective discussion about whether the web really needs one more language.
Anecdotally I would have expected Javascript to be near the top of hours being used.
/ goes to a corner to cry...
Julia *****
(or maybe that's just my opinion :-)The thing that really stands out to me though is the poor state of python tooling and the library ecosystem. Having used Rubygems and bundler pip feels like taking an enormous step back. It's much less expressive. It doesn't handle the difference between production only dependencies, regular dependencies and development only dependencies in a good way. It's difficult to differentiate between locked dependencies vs desired dependencies(the Gemfile/Gemfile.lock distinction). PyPi and especially using private PyPi registers is more complex than it is in Ruby. There seems to be fewer nice libraries and they seem to be spread across the web whereas Ruby centralises around GitHub. I also find that Python libraries have lacking or hard to find documentation in many cases.
I wonder how much of this is due to the community spending a lot of time on the unproductive Python2/Python3 debacle instead of moving the language and the ecosystem forward. It seems like Ruby and Python would be in more similar positions had it not been for the problems around the Python 2/Python 3 upgrade. Ruby had the risk of a similar problems in the 1.8 to 1.9 upgrade, but it seems to have worked out better than the Python 2/Python 3 upgrade.
Could you clarify a bit more on this with the real world examples ?
Around documentation and usage I think it comes down to the fact that for Python libraries a google search usually takes you to the modules page on the PyPi registry, to a project website or a read the docs instance. I usually find that where ever I end up doesn't provide all the information I want, e.g what's the latest version, what are some typical use cases for this library. As a concrete example let's compare Flask and Sinatra.
For both Flask and Sinatra a google search gives me their respective websites as the top relevant result. In this case I find both websites don't really provide all the information I am after. When I navigate to their GitHub page however the difference is more stark. Sinatra uses GitHub flavoured markdown to render a nice guide with increasingly more complex use cases. Flask just has some non markdown text pointing me in a few different directions. I could probably find all the information I need to setup and use Sinatra from just the GitHub Readme whereas for Flask I'd have to go navigate their read the docs page where all the documentation and usage instructions are spread over multiple pages and sub pages. I guess it comes down to my preference for the more concise format of a long markdown README.md vs the experience with a read the docs like set of pages.
I recon this is a bit rambling, it's kinda hard to put into words what my feelings around this are and why I hold the opinions I do.
Also, I've been using Jupyter lately to brush up on data structures and algorithms. It's been pretty sweet to write code in a web notebook and render output including digraphs beneath each block of code. Is there anything like that for Ruby?
I don't share your experience on the documentation of libraries though. I've generally had no problems working with the major third-party libraries (requests, python-dateutils, Flask).
And I also am reluctant to feel strongly against how the 2v3 chasm has been drawn out, though it's easy to say that as someone who didn't have to make the jump. The main difference from Ruby 1.8->1.9, AFAICT, is that the Ruby community has a near-monolithic devotion to Rails. When Rails left 1.8 behind, not making the jump would basically be career-suicide for a Ruby dev. The downside is that it shows how dependent Ruby is on Rails' momentum. Whereas in Python, a scientist (among several other use-cases) can live on 2.x for a very long time no matter what all moves to 3.x.
products.select { |product| product.type == 'x' }.map(&:price).reduce(:+)
sum([p.price for p in products if p.type == "x"])
Excuse untested code from phone, but I like it better than the ruby version.products.lazy.select { |product| product.type == 'x' }.map(&:price).reduce(:+)
for a in x: pass
into a comprehension, you write: [pass for a in x]
Simple, right? So to convert for a in x: for b in y: pass
I obviously write: [pass for b in y for a in x]
Bzzt. Nope, you have to write it in middle-endian order to get the same behaviour.Python is pretty imperative, and the language design ethos is pretty against designing around functional programming. But if you try to absorb some of the python-specific stuff like list comprehensions, generators, and context managers, you can write some tooling to help you get a DSL a little closer to what you're used to.
The way my brain works (at this point in my career, 25 years of programming) is just against the Python grain. I think in terms of data flow graphs. Python wants me to think in terms of procedures and loops.
Personally I moved away from relying on language specific tools to do more than just "install library X" and manage my deployment and dependencies using docker. So far it works great for me.
> There seems to be fewer nice libraries
Depends on your domain. Ruby is very good at web development. Python is good at almost everything.
> Python libraries have lacking or hard to find documentation in many cases.
I didn't use Ruby, but JVM languages or Javascript don't come close to Python in the documentation coverage. Especially if you consider SO as a documentation, googling "how do I do X" in python usually gets you on the right track to the high quality solution. Javascript or Java SO post are more like a mine field in comparison.
Django~=1.10.0
And it keeps the env autoupdated with the latest patch revision but not the latest minor revision. pip-tools doesn't seem to do much related to that, no?
You put
gem 'rails', '~> 4.2.7'
in your Gemfile and when you run `bundler install` the exact version of rails you ended up resolving with the above is locked in your Gemfile.lock. Thus if you have a version in Gemfile.lock it is the source of truth and you have reproducible builds. If you want to update rails given the pattern in your gemfile you just issue `bundle update rails` and it will update to the latest conforming version and update your Gemfile.lock
The fact that pip doesn't support this in a good way is a major limitation imo. You can have either reproducible builds or "auto upgrading", but not both. pip-tools solves this, but requires you to use a different tool than pip and by extension you have to teach and explain it to anyone working on your project.
If I check out an old commit I can look up docker image version in the source control, download exactly the docker image that was running at that time and start shell in this image. Inside the image I can run "pip freeze" to check exact version of each package that was running at that time. I can then update the old Dockerfile with those specific versions by adding the "==x.y.z" e.g. "pip install oauth2client==3.0.0" in the pip install to reproduce the old build.
Additional benefit is that such scheme works for system packages, or for any other language libraries I can think of.
Ah, okay, this crystallized it for me, thank you. Other commenters were talking about one or the other as a limitation, which confused me because pip does do both (just not at the same time).
Yes, this is currently a limitation, and pip-tools seems to solve it. Hopefully pip will get a native solution soon.
Interestingly enough in Ruby it's the reverse, you use gemspec for gems which doesn't have the concept of locking. Instead you are supposed to specify semver conforming version patterns and then resolution of these happens when the gem is installed. Only end user projects use Gemfile.lock
For example, with dependencies, production dependencies and dev dependencies are defined through two requirements files: requirements.txt and requirements-dev.txt. requirements-test.txt for testing. No need for tooling, just pass those into pip.
Most python tooling is pretty limited in scope but works well. From my side, I've had more trouble with "turing complete tooling" that ruby has offered me. I have not had to use ruby a great deal though.
I think the difference in philosophy between the two communities is that python stuff will usually value ease of debugging/implementation over "ease of use". Most things in pyhton and its echosystem are straightforward. None of this "Library A adding methods to classes of Library B" stuff (not that it's impossible). Of course, you then end up with more boilerplate.
Py2/Py3 has definitely been a major pain.
gem ist superior to pip. Take a look at "pip-tools" if you haven't already, it eases some of the pain.
For web-dev Rails is more stable, more modern and has better maintained library ecosystem than Django. Django is okay-ish but can't live up to Rails. Django still lacks some stuff that i feel it should bring with it out of the box.... dev/test/production setups, asset compilation etc.
On the other hand, if you don't need a full blown web framework and are looking for something more minimalistic, Flask is awesome.
Heck, the tooling is better on the PHP side even...
ignorant of any prior ecosystem i just thought i was getting into the most current
i came for the prebaked arbitrary precision floating point arithmetic and now i use it for prototyping all of my projects, and all the other stuff that i would call general scripting
this 2to3 complaint made sense to me from an employed standpoint:i work as X and use 2.7's Y library; but you said general scripting
i wonder what libraries were you hoping to use but were unable, and for what general scripting purpose?
I recall several showstoppers in Python v3 including integration into Emacs, accessing a MySQL database, and basic visualization. The library system was also a mess. It didn't seem possible to develop on v3 and expect that the libraries you needed would be available. Or that anyone else would care since they were all using v2.7.
Don't get me wrong: I like Python and know it's truly good, but __why__??
Felt the same way about "self" argument. But then now I really like it and see it as a very nice idea.
There are certain areas of the language which may not behave exactly the way you'd be expecting, if you're new to it.
These usually have an unusual syntax, which will cause you to find out more.
They're basically the mini-roundabouts of python.
https://en.wikipedia.org/wiki/Roundabout#Mini-roundabouts
If you know what you're doing you cruise straight over, but if for any reason you're unsure, you're forced to find out more before continuing.
It's not designed to stop you, because "We're all consenting adults", but that doesn't mean it's going to encourage you to do something that might be harmful.
Driving in Belgium - the scripting equivalent of batch files
It doesn't stop there: Take something like __str__ and __repr__ combined with str(), repr(), vars(), dir(), print(), pprint.pprint()... I have no idea how anyone ever thought it was a good idea to have that many ways to output the content of a variable. Even after all the time I used python I never managed to find a consistent way to output variable values.
I generally like python as a scripting language. But stuff like this made me constantly lookup obscure details in the documentation which made it very hard for me to completely embrace the language. It also prevented me from diving into more arcane stuff like meta-programming...
Very true! Though since in practice, most people simply use a graphical debugger or pick one of the above as their champion (pprint.pprint is my personal favourite), then I doubt this situation will ever be changed.
If you want to affect how str(x) behaves, you override x's __str__() method. Operators behave this way too: for example, if you want to overload the == operator for a class, you override __eq__() in that class (for any Java developers reading: Python == works like Java equals(); the Python version of Java == is the is operator, which cannot be overloaded). Yes, that means x == y is just syntactic sugar for x.__eq__(y).
The dunders are there so you have the freedom to name your methods what you want without having to worry that you'll accidentally clobber a method that affects a global or an operator, as most people won't just up and decide to both begin and end a method's name with __.
The colon is so ironic for a language that takes pride in replacing braces with indentation. It's there because of readability, see https://docs.python.org/3/faq/design.html#why-are-colons-req... but so, how about an almost Erlang-like full stop?
if condition:
statement
statement.
That would make semantic indentation optional, IDEs and editors know how to reindent code, spare us with bugs introduced by careless copy and paste.Nevertheless, Python has been widely successful and it's here to stay for a long while. I hope designers will look at those bizarre features and think of better solutions for future languages. Remember that Python has its roots in the 80s. By the standards of the time it was very good. Compare it with the order of magnitudes worse design of PHP, from middle 90s, and there is a lot to forgive to Python. Maybe Guido Van Rossum would make different choices if he were to design his language now, under the influence of the languages of the last 20 years.
This is just brackets/braces in disguise, at which point you might as well follow convention and use {}.
This is obviously a very personal issue, but for me, the whitespace has less cognitive load. E.g.: in languages with brackets, people usually still indent the code for readability.
I've also found that when teaching people to program, consistency works well. I think this is one of the reasons Python is easy to pick up, just because Python code consistently looks like Python.
I'm asking because I'm using Python little and only for short scripts, so I didn't bother investigating much. Still I've been bitten a couple of times by bugs introduced by moving code around and not noticing that a line was not indented correctly. That in Python and in Haml (http://haml.info/) If somebody knows about those tools I will appreciate and it will make my life easier. BTW, I'm fixing somebody's else Python scripts right now :-)
Regardless of whom or what is doing the indentation, people still prefer reading code which is indented, even in a language which relies on braces or parenthesis for flow. The indenting is not for the compiler's sake, it has no problem understanding this: http://www.ioccc.org/2014/maffiodo1/prog.c
Also people get very picky and personal when it comes to indentation in ways that a computer can't always choose. I prefer to almost always indent and format my code myself, rather than letting the editor choose. Python is the only language I actually let my editor auto-format, because of PEP-8.
> Is there any automatical indentation function in some editor for Python?
Pretty much any editor will have a feature for this, or allow you to install a popular plugin for it in less than a minute.
if cond:
line
pasted line # where does it go?Correct me if I'm wrong, but there doesn't seem to be built-in support in Python for having attributes initialized by the the constructor.
For example in Perl 6, I can write
class Point {
has $.x;
has $.y;
}
... and I get a constructor Pair.new(x => 1, y => 42) for free, no need to write custom initializers. class Foo:
def __init__(self, *args, **kwargs):
self.__dict__.update(**kwargs)
And it will automatically assign any keyword arguments you use as attributes to the object. For example `foo = Foo(name='Bob', age=99)`If you still want to keep a strict list of allowed attributes, you can define them as parameters, and use a shortcut to assign all local variables to attributes.
class Foo:
def __init__(self, name, age):
self.__dict__.update(locals())
del self.self
So `foo = Foo(name='Bob', age=99)` will still work as will `foo = Foo('Bob', 99)`. But `foo = Foo('Bob', 99, True)` will throw an error, as will `foo = Foo(name='Bob', age=99, likes_cake=True)`. You can add kwargs back to the parameter list if you want to allow assigning any attribute.This isn't recommended though. So for all practical purposes, Python does require a bit of boilerplate in the constructor.
Edit: Realized a cleaner way to do the second example.
* Packaging is horrible
* Releasing python code is a non-standarised nightmare. Every solution has it's own flaws
* Big and complex projects in python are really hard to reason about
* Poor support for concurrency (fixed in py3)
Don't get me wrong, python is really cool as a proof-of-concept scripting language. But for mature and complex stuff, I'd recommend something with type checking and better support for release process.
* It would often fail randomly.
* node_modules has a fixed location.
* There is a ludicrous number of dependencies even for the simplest projects. ~1000 is not unusual. Reasoning about that dependency tree simply isn't possible.
* There was even a package to detect if a number was negative. It had a version 2 because there was a bug!
* In general even small projects took an age to build.
I'm sure this isn't all the craziness, it's just the stuff I ran into during my (fairly limited) exposure to npm packaging.
* type annotation + awesome tool for analysing code (dialyzer)
* immutability, list cannot become a dict out of nowhere
* great support for concurrency
* concept of application (OTP) gives out ability to run many loosely coupled apps on one VM. I believe "microservices" is the buzzword that matches the concept.
* distribution by design.
* great monitoring tools
* I could go on... really.
I think static typing alone doesn't make managing complexity easier.
I'd like to know what other features people think Python is lacking that make it much worse than C# or Java when handling large projects.
[0] https://glyph.twistedmatrix.com/2016/08/python-packaging.htm...
1. In-consistant syntax.
2. The language uses exceptions to control flow of logic.
3. Divided community, since Python 3+ included breaking changes to the standard.
4. Un-discoverable APIs. You better hope the documentation is bulletproof else the API could change in any which way during runtime.
5. Poor error messages. If an import goes wrong you are not told why, and so on.
6. Ultimately slow performance for anything marginally complex.
7. Poor tooling.
Admittedly it is mildly annoying when someone decides to use some random number of spaces for indentation, and you have to fix it to make it match the rest of your code (my preferred editor, Geany, has a feature for this). We should all just use tabs. But "very hard"?
This probably reads pedantically, but I do think there's a distinction. By analogy, if someone tried to post APL on a site that didn't support Unicode, you wouldn't say the issue is it's "hard to cut and paste APL".
In addition to the points posted in a parallel thread, selecting and copying leading whitespace reliably is more difficult than just getting the printable characters. Also: it's not something that anyone has any cause to do outside of copying Python text.
2. care to elaborate ?
3. it's not divided, it's evolving from 2 to 3. Java's community is under Oracle's grip, don't know if it's any better...
4. it's not because you can do it that it's worth to do it, so in practice API are just not changing at runtime
5. I can say the same of my industrial strength Java environment...
6. Not a problem since what you do doesn't require speed
7. Ever used PyCharm ? I wouldn't say it's poor tooling. The debugger is fine. What I miss is something like a good profiler to look at running code.
But I'd say :
8. Optional typing should've been there from day one, it'd made code editor, tooling better
9. The date/time handling is not good
10. SOA support is weak (and in the kind of setting I am, we use tons of XSD, WSDL, XML, security, signature you name it)
vprof is quite nice: https://github.com/nvdv/vprof
For example, rather than checking if a dictionary has an element before accessing it, you just try to access it and handle the KeyError if it doesn't. The code reads easier and the normal flow would be faster in this case.
When used in EAFP context only, I think it is fine to use exceptions as part of logic in Python. They are relatively less costly compared to C.
deque
defaultdict
Counter
After many years I often write DefaultDict myself.Edit: formatting
http://effbot.org/pyfaq/why-does-python-use-methods-for-some...
I guess is that.
I'm using Python since about half a year and for all I did this wasn't an issue. I just started with 3 and had the feeling everything around me was 3 as well. 3 was released in '08.
One of the main reasons I like python is that it is the most self-documenting language I've come across. Say what you like about significant whitespace, but in my last job I used to show salespeople snippets of code and they understood them - with no knowledge of the language.
[1]: https://www.python.org/dev/peps/pep-0257/(If neither happens, then I guess it isn't really a problem.)
There is no formal definition of the language. It is defined by the C implementation. So I'm not sure how meaningful that distinction is.
> If neither happens, then I guess it isn't really a problem.
Or people give up on Python and use another language.
The real-world distinction between the Python specification and its implementation started in the late 1990s, with JPython (now Jython).
Sklearn, tensorflow, pandas, Sqlalchemy, requests, Beautifulsoup, numpy, scipy, pulp.
If you have a language that had equivalent libraries that cover all these domains I'd love to hear it. Until then I can build really cool stuff in Python very easily thanks to the amazing hardwork and generosity of these library creators.
R[1].
Well, in this case it might be advisable to take a look at Perl. And never go back. Perl (especially Perl6) is hacker's paradise, where everything's true and nothing is forbidden.
As somebody who shares the same mindset, I truly enjoyed my years with Perl, and I occasionally return to this experience.
And support engineer's hell
Python is a strange case. The CPython code base seems to be quite well organized and well written, and about 10% of packages on GitHub or PyPI are extremely solid.
The problem is the staggering amount of packages (public or in-house) that have 100000+ LOC and use the worst features of OO, functional, and procedural programming at the same time.
The result is not spaghetti code, but ravioli code where there's zero control flow and "everything happens somewhere else". Add to that the "consenting adult" philosophy, and you have the modern equivalent of goto mixed with self-modifying code.
I'm sure Perl has all that, but it's far from the only language.
import sys
from collections import Counter
cnt = Counter([n.strip() for n in sys.stdin.readlines()])
for key, val in cnt.most_common():
print(val, key)
This isn't just playing code golf, it goes with point (2) - awesome libraries. One of the joys and curses with Python is finding the Pythonic ways to do things that you've hamfistedly written your own way for months. Counter(x for x in y) cnt = Counter(n.strip() for n in sys.stdin)
* no need for a list comprehension, a generator expression will do* file objects are iterable, and iterated line-wise
http://www.oreilly.com/programming/free/a-whirlwind-tour-of-...
I was looking for a dense & concise guide to the language, but one of the most popular books "Learning Python 5e" by Lutz was too long @ 1600 pages.
I found that learning Python is simple, mastering it is difficult, the magic which happens all the times makes it difficult to understand, which is the reason why static languages are awesome as less magic as possible.
Wait, what? How is its interpreter remotely relevant to its semantics?
> if you have a syntax error in C, the program will simply not compile, on the other hand, in interpreted languages, the problem might go unnoticed for quite some time
Hang on--Python allows you to run code that has syntax errors? I'm so confused.
#!/usr/bin/python
import time
time.sleep(5)
fdafadsfadsf
It'll take 5 seconds to crash with a "NameError: name 'fdafadsfadsf' is not defined". Likewise, if you do this: #!/usr/bin/python
import time
x = 1
time.sleep(5)
print "x is "+x
It'll take 5 seconds to crash with "TypeError: cannot concatenate 'str' and 'int' objects"In a language like Java, these would be syntax errors, and they'd be picked up at compile time as they're illegal in Java's syntax.
I assume they're legal in Python's syntax, because for all the syntax checker knows, time.sleep() might define fdafadsfadsf and convert x into a string. So they're runtime errors in Python.
Of course, Java still has runtime exceptions - null pointer exceptions, for example.
> Python 2 vs 3
It is mostly over. All important libraries are accessible on 3. It is clear that if you are starting from scratch you should go with 3. There are good tools to help with the migration (e.g. six). Some libraries start dropping python 2 support.
> Bad tooling
Having used Emacs with jedi for a while I switched to PyCharm to leverage optional typing. PyCharm beats my tooling experience in all languages I used expect Java:
- Auto imports work well.
- Code completion works very well, especially on optionally typed code.
- Even refactoring works very well, failing short only to my experience with Java, but beating C++ or even Scala.
- Debugging is actually a highlight - I can drop into full fledged REPL at break point. "Evaluate expression" in C++ or Java didn't come close.
Profiling is indeed not as advanced as what I was used to on the JVM, but IMO it is sufficient for 95% of use cases. For that 5% it is clear from the start that you shouldn't go with python. Language specific performance monitoring in production becomes less of an issue when you start running inside containers and it starts to make more sense to monitor containers rather than individual processes.
> Maintainability at scale
I am using Python 3.5 optional typing and I feel my project is much more maintainable as a result. Among others, refactoring in PyCharm works very well.
> Packaging and deployment
As mentioned in another comment, personally I moved away from relying on language specific tools to do more than just "install version X of library Y" and manage my deployment and dependencies using docker. It have it's own problems, but so far it works very well, and it's quickly improving.
> Performance
In domain of scientific computing by using correct libraries and things like Cython or Numba you can come close to C++ performance levels. In other domains it's becoming more common to be disk/network bound, and then it doesn't matter if you use Assembler or Python. Also JITs are coming to Python: https://lwn.net/Articles/691070/.
> Syntax/libraries/language features/etc.
It is the most subjective section, but my experience with python in this regard is highly superior to Scala/Java/C++/JS.
PS> Import-Csv .\names.txt -Header name | group name | select name, count
Name Count
---- -----
cat 3
dog 2
mouse 1
bird 1A 13-line python script is presented, followed by a smug "you can almost hear the UNIX folks complain about how verbose the PowerShell version is".
First off, the provided Unix commands _don't work_ (should have used `sort -rn -k 2`), while the provided powershell, as verbose as it is, does work.
Second, the entire python script along with the wrapping unix commands is encapsulated with a powershell 1-liner:
cat ./names.txt | group | sort -d CountThere is such a large breadth of fields that it has been successful in. Web apps, data analysis, machine learning, micro controllers, etc.
It has a long track record for being stable, reliable and secure. The popularity over the years has been steady and with the recent addition of data science field and machine learning, it's seeing another growth spike.
In my opinion, it's a highly versatile and very productive language that seems to be able to handle just about any job or industry you can throw at it. I personally switched to it full time and love it.
http://shop.oreilly.com/product/0636920032519.do
It's a bit pricey, but I managed to score it on a 50% sale, and I have to say I think it's a very good book for people already fluent in other programming languages.
There is just something so awesome about shoving your dependencies right in the script and not worrying about pip, easyinstall, system libraries etc [1].
Maybe Java 10 will fix the startup speed (or whenever jigsaw is done correctly) but I doubt it.
[1]: http://docs.groovy-lang.org/latest/html/documentation/grape....
Groovy needs two lines (@Grab and import) to get a dependency. Golang only requires the import. Perhaps Groovy needs to be rewritten in Go.
I'm not dissing Golang but it isn't even remotely the same. A major portion of the OP content was about using Python instead of shell scripts.
Golang is not scripting language. You need to compile for each platform and the code is opaque once it is compiled. Not to mention Golang might not have dependency issues at runtime but the overarching consensus is that certainly has issues with compile time dependencies management (auto checking stuff out from github is not effective).
> Perhaps Groovy needs to be rewritten in Go.
I normally don't say this... but that is just dumb fanboy comment.
EDIT: apparently you are working on a scripting language powered by Go called Gro. I'm not sure if you just mistyped and meant Gro instead of Go.... and reading your blog it seems you must have been severely burned by the Groovy community.
I didn't mistype. Go would make a good language for writing dynamic languages such as Apache Groovy in. My own Gro is just one example of such a dynamic language.
Wasn't the whole point of Groovy to have something Ruby-like that was more intimately tied into the JVM ecosystem? Seems to me that Go would be just about the worst imaginable [0] implementation language for that.
[0] Or at least, in the class of extremely poor choices that includes all languages that are not themselves especially targeted to the JVM platform.
I'm not sure how easy it would be to access ad-hoc libraries without recompile though since Go doesn't really have dynamic loading of libraries (or maybe it does now?).
Little tangental ... After reading his blog I'm a little concerned though that the dude (Gavin) has spent an enormous effort defending his stance on how poorly the Groovy organization is... I'm not saying his wrong but is almost obsessive and a waste of energy.
Right, my point was that a central part of the point of Groovy was being for the JVM, it wasn't a language in search of a runtime.
So, while there might be an argument that a language like Groovy for Go runtime might be useful (I'm somewhat unconvinced of that), it just seems not to be the case that Go is a good choice for a language such as Groovy (that is, it is not the case that Groovy is an example of a language for which it would make sense to use Go as the implementation language.)
In the current version of Gro, I recompile (just like "go run"), so there's a 1 second lag even for simple scripts. For the same reason, the very basic REPL provided has a 1-second process time for each command. There's talk on reddit.com/r/golang about building a VM for Go so maybe in future Gro can use those.
> Are you going to change your middle name now to "Gro"
"Gro" is short for the "Grover edition of the Groovy language".
> has spent an enormous effort defending his stance on how poorly the Groovy organization is... [...] almost obsessive and a waste of energy
I've averaged a blog entry a month for the last 5 years. It seems like enormous effort and obsessive to you because you saw them all at once.
"oh your script can just be rewritten in Go and while we are at it Groovy should be rewritten in Go".
But yeah I would be interested in platform agnostic scripting language with out GIL, cheap threads, easy dependency access. I'll keep Gro on my radar :)
Are you going to change your middle name now to "Gro" (I did find that amusing albeit I'm not sure I know your entire history with Groovy)?
Learning C later on helped me understand some basic concepts better and being able to speed up Python e.g. using Cython can come in quite handy.
It literally has everything you want in a language.
Granted, it's not on-the-metal fast, and it still has some gimmicks, like variables behaving like pointers or references (which can be a little awkward), but still, python is the language of the decade.
I think when I code, I want to own the machine, not just skate on the surface. Python doesn't even crank my engine the way BASIC did years ago.
I also think part of my problem is it doesn't solve any needs right now.
Python is a language that makes programming easy. That's a lot.
> While I'm on the subject of WordStar, I've heard a story that I haven't been able to verify. When Phillipe Kahn was asked why he chose to use the WordStar command set in the first version of Sidekick, he said that he asked a lot of people for their editor preferences. Almost everybody had a different first preference (back then it could have been Emacs, vi, Word Perfect, WordStar, Brief, Leading Edge Word Processor, or who knows what else). But almost everybody he asked knew WordStar and named it as their second preference. I don't know if it's true, but it smacks of truth. Certainly every microcomputer programmer I knew back in the late 80s was proficient with WordStar.
But since then, I've never really had a need to use it, and I don't even feel there is a good reason for it.
If you're writing something more than a quick shell script or a project that will last more than a few days, then there a dozen static languages that have better libraries, more powerful IDEs, etc. And misspelling a variable or missing a comma is not going to cause all sorts of problems.
On the other hand, if you do need a quick one-off script or some duct tape, things like Perl or AWK, sed, and friends seem like a better bet.
Python (and Ruby) seem bloated for quick scripts, and wholly lacking for larger projects.
If you know bash you probably have a pretty good idea how to manage Linux and roughly how things work on Linux (you could apply this to other Bash capable OS as well).
If you know Bash fairly well you probably also know how to automate many things and you probably already know regex which is fairly useful in GSD. If you know Bash you obviously know the command line well.
You should learn Bash (or some shell but Bash is so ubiquitous).
> Of course it can be done all with command-line tools, but the point was not to show you how you can count and sort word occurrences. I just wanted to show that it's possible to include Python scripts in your "command line workflow". I wanted to give a simple example that would give you a taste of the idea :)
I wonder how large projects such as OpenStack, Django, Edx, PyPy deal without static types?
If you want to learn how computer programs work, you don't have to use existing libraries - you can still implement the functionality yourself.
By your argument (I hate python...it has libraries) any modern & moderately successful language with a large library ecosystem is bad, and you prefer Go precisely because it lacks those libraries?
I'm glad I made that decision. The language was easy enough to pick up in my spare time. I hardly ever use pandas/numpy unless I'm doing some really serious statistical work; the core language and patterns in Python are now pleasant enough for me to use in day-to-day data wrangling and hacking. I'm now mixing in Ruby for a current project -- mostly because I like Rakefiles, and also, Ruby's static site generators like Middleman -- and it's been surprisingly painful to do things in Ruby. I used to hate how when doing a Google search for something in the Python stdlib would bring up 2.x before 3.x...but it almost seems even more problematic in Ruby, where 1.9.3 and 2.0.0 docs keep coming up, even though official maintenance of them has been dropped.
(2.0's maintenance ended in Feb. 2016 http://ruby-doc.org/core-2.0.0/Array.html)
If I didn't know better, I wouldn't even know what Ruby's official version was. For me, Google results show no other documentation for Ruby's Array, as the next results are from tutorialspoint and sitepoint.
And that's not even getting into the issue of how problematic the do-things-how-you-like freedom of Ruby is incredibly hard to discern as an outsider. To give one example I recently had: how do you iterate through a human-calendar-based time interval, such as a month (i.e. any interval that doesn't have a constant length in seconds)?
In Python, I remember how frustrating it was to figure out what the hell the difference was between datetime, time, calendar, the gmtime function, and then timedelta, and then the third-party libraries of pytz and python-dateutil. But once I understood them, which was largely a matter of realizing how naive I was about the nature of measuring time, doing something like iterating over calendar intervals is straightforward.
This is my first Google result for "python iterate month intervals" -- it's from 2008 and still works perfectly: http://stackoverflow.com/questions/153584/how-to-iterate-ove...
Now here's the result for Ruby; all of the answers, as far as I can tell, are not only not-very-Rubyish, but not reliable either: http://stackoverflow.com/questions/1724639/iterate-every-mon...
FWIW, I remember solving this problem by using the DateTimein ActiveSupport. Which is a credit to how much fun and joy it is to build things in Rails right out of the box. But that complexity catches up quick, and including ActiveSupport in day-to-day Ruby work can sometimes to unwanted monkeypatching.
That said, unless "recently" was years and years ago, Date#next_month is a thing:
require 'date'
date = Date.today
while (date < Date.today.next_year)
puts date
date = date.next_month
end
One thing that I think throws a lot of people is that the official Ruby docs are actually better than Google for figuring stuff like this out - it only takes a couple minutes to find #next_month by perusing the method listing at http://ruby-doc.org/stdlib-2.3.1/libdoc/date/rdoc/Date.html, but do a web search and you get a bunch of Railsy crap.Though I will grant you that even searches on ruby-doc.org keep bringing up results for Ruby 2.0.x and whatnot. That part is pretty bizarre.
if name in names:
names[name] += 1
else:
names[name] = 1
That can actually be written as just one neat, readable line: names[name] = names[name] + 1 if name in names else 1
It may not be as readable to some who are used to spelling it out as an if/else block, but I really prefer the one line way. It reads closer to regular English I think.In the original, at a glance I know exactly what is going on.
In your version, I have to read the whole sentence carefully to notice that it's even a conditional and not a normal assignment, and then I have to mentally unpack it to understand the logic that you're trying to implement. If/else should never be a one-liner.
Good code is boring code :)
from collections import defaultdict
def count_names():
names = defaultdict(int)
for name in sys.stdin.readlines():
name = name.strip()
names[name] += 1
...But today, I'm less excited about it. I've gotten spoiled by C#, that has static typing, and consequently, better tooling, as well as a good measure of the flexibility of Python, through dynamics, lambdas, LINQ.
It seems Python is finally getting over the 2.X -> 3.X schism, but that kind of threw a wrench into things for a number of years, as well.
Try running a nested loop on a non-trivial example, and you can end up spending minutes in what would take milliseconds in any other language. If you want to program in Python, you must get used to the functional paradigm. Not "should", "must".
Now, I'm all in for bringing more functional programming into daily life, but I find a bit dishonest that beginners are not told about this. My first serious program took about a week for a task, and shaving that time down involved a lot of pain changing simple nested loops into multiplications of Numpy matrices. There was a lot of caching too. I have no idea how someone without a CS degree would have dealt with this.
Moving from Scala to Python it was annoying at first.. but honestly the code is easier to read in most cases, so no complaints here.
I kinda agree. Certainly it guards against some of the code you see in overly nested javascript.
The downside is that the code is slightly removed from it's call site - but rarely by much.
Overall the increased complexity it would bring to the language has not been justified by anyone arguing for it's benefits. Giving the function a name isn't the worst thing in the world.
I think the argument against map and reduce was that there was better and more Pythonic ways to do both - not because they were 'too functional'.
Huh? Can you provide an example? There's nothing about writing nested loops in Python that is qualitatively different from other languages.
> If you want to program in Python, you must get used to the functional paradigm. Not "should", "must".
This is incorrect, much to my chagrin. Python's trend over the past few years has been not only not to encourage functional programming, but to explicitly discourage functional programming. For example, starting in Python 3, `reduce` has been banished from the language proper to the standard library.
Well he did mention numpy and obviously doing something that can be done in numpy with nested loops will be much much slower. However that is as much a case of numpy being really fast as python being slow.
For example, just tested elementwise multiplication of two 10kX10k matrices and with numpy and numpy arrays it took ~350 ms vs ~15 seconds with a nested for loops and python lists. Still minutes vs. milliseconds seems like extreme hyperbole.
edit: Figured I'd test how long the naive nested loop approach would take in a 'fast' language like Julia, and much to my surprise it took over 50 seconds.
edit2: cleaned up all my code to try to make it as similar as possible across both languages:
Nested for loops python 2.7: ~30 seconds
Nested for loops Julia 4.6; ~50 seconds
Nested for loops PyPy 5.3.1 ~0.9 seconds
Make of these numbers what you will
http://docs.julialang.org/en/release-0.4/manual/performance-...
s=10000;
a=ones(s,s);
b=ones(s,s);
c=zeros(s,s);
tic();
for i in 1:s
for j in 1:s
c[i,j]=a[i,j]*b[i,j]
end
end
toc();
Obviously in real code I'd simple write c=a.*b
and get basically the same performance as numpyJulia also stores its arrays column major, so put the outer loop over columns for another order of magnitude performance increase. (I'm not sure I love this feature, but presumably there's a reason...maybe for better performance on matrix operations?) My reply to your grandparent comment has a simple implementation that runs in 0.4 seconds on an aging machine.
function f(s)
a=ones(s,s);
b=ones(s,s);
c=zeros(s,s);
tic();
for j in 1:s
for i in 1:s
c[i,j]=a[i,j]*b[i,j]
end
end
toc();
return c
end
Note that you also might have to run the function twice, because the first time JIT-compilation kicks in.Julia 0.5 also now automatically devectorizes some code (eg. x = a .* b .+ c) so that you don't have to write explicit loops to get performance benefits.
Guess I need to take some time at some point to sit down and actually learn 'proper' Julia.
I wonder what the difference would look like then?
julia> include("elmult.jl")
eachindex
0.42587028
outer loop over rows
8.923323015
outer loop over cols
0.516270241
Code:
function elmult(m1, m2)
result = zeros(m1)
for i in eachindex(m1)
result[i] = m1[i] * m2[i]
end
return result
end
function elmult_rowmajor(m1, m2)
result = zeros(m1)
nr = size(m1, 1)
nc = size(m1, 2)
for i in 1:nr, j in 1:nc
result[i,j] = m1[i,j] * m2[i,j]
end
return result
end
function elmult_colmajor(m1, m2)
result = zeros(m1)
nr = size(m1, 1)
nc = size(m1, 2)
for j in 1:nc, i in 1:nr
result[i,j] = m1[i,j] * m2[i,j]
end
return result
end
elmult(rand(3,3), rand(3,3))
elmult_colmajor(rand(3,3), rand(3,3))
elmult_rowmajor(rand(3,3), rand(3,3))
m1, m2 = rand(10000,10000), rand(10000, 10000);
println("eachindex")
println(@elapsed elmult(m1, m2))
println("outer loop over rows")
println(@elapsed elmult_rowmajor(m1, m2))
println("outer loop over cols")
println(@elapsed elmult_colmajor(m1, m2))Here's an alternative example where I think it makes sense (checking if the size of all elements in a collection are equal):
reduce(lambda a, b: len(a) == len(b), ls) >>> reduce(lambda a, b: len(a) == len(b), ["a", "b", "c"])
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "<stdin>", line 1, in <lambda>
TypeError: object of type 'bool' has no len()
It's trying to do len(len("a") == len("b")) == len("c"), which becomes len(True) == len("c"), and len(True) doesn't work.Some variations which do work are:
len(set(map(len, ls))) == 1
len(set(len(x) for x in ls)) == 1
all(len(ls[0]) == len(x) for x in ls)
len(min(ls, key=len)) == len(max(ls, key=len))
Append "if ls else True" to handle the empty list.Heh, just proved my point. Reduce won't be missed.
class EqualLengths:
def __init__(self, size):
self.size = size
def __nonzero__(self):
return True
def __repr__(self):
return "True"
def equal_lengths(a, b):
if isinstance(a, EqualLengths):
if a.size == len(b):
return a
return False
if a is False:
return a
n = len(a)
if n == len(b):
return EqualLengths(n)
return False
>>> reduce(equal_lengths, ["a", "b", "cc", "d"])
False
>>> reduce(equal_lengths, ["a", "b", "c", "d"])
True
However, I could do something similar with sum(): class SumEqual:
def __init__(self):
self.size = None
self.is_equal = True
def __add__(self, other):
if self.is_equal:
if self.size is None:
self.size = len(other)
else:
self.is_equal = (self.size == len(other))
return self
def __nonzero__(self):
return self.is_equal
def __repr__(self):
return repr(self.is_equal)
>>> sum(["a", "b", "cc", "d"], SumEqual())
False
>>> sum(["a", "b", "c", "d"], SumEqual())
True
and unlike the first case, this sum() solution will work when there are fewer than two items in the list.Gladly. Check the nested loop in [1]. That's exactly what I fought against (although that's not me asking that question). No other language AFAIK suffers from that issue.
[1] http://stackoverflow.com/questions/8097408/why-python-is-so-...
2) With the exception of Luajit and JS, any dynamic interpreted language will take in the range of 5-20 secs to run equivalent code. I'm pretty sure Ruby and PHP are in the same range, no time to benchmark it now. Besides, if you use the Pythonic version of this loop (using list comprehensions), it will take more like 5 or 10 seconds and not 20. Run the same code on Pypy, and it takes under a second.
3) In any case, if performance is really an issue, the correct answer (as the parent discovered) is to use Numpy or Cython or C extensions for that part of the code. This is not uncommon knowledge.
4) I've been using Python in production for a decade and have never thought of this (~0.2 vs ~10 secs) as a problem because I never used pure Python for critical numeric code (or Ruby, or even JS). I did that in C or Cython, or Numpy (or OCaml, on one occasion). If I had to do something like that now, I'd probably use Rust and call it from Python (assuming I was writing the app in Python).
5) Any language has edge cases like this that can surprise you if you're not familiar with the ecosystem. The higher the level of abstraction, the more edge cases. Look at even the Julia thread directly below this one. It's a massive difference between the naive, Julia beginner version (no offense, I would have written it the same way, probably) and the experienced Julia programmer version. Even a language like Go has performance gotchas.
Python certainly has it's problems (like building large, robust programs, although that's being actively addressed with MyPy types), but I honestly don't agree that this is one of them.
Or did you mean something else?
[1] http://stackoverflow.com/questions/8097408/why-python-is-so-...