Navigating the Postmodern Python World
stephendiehl.com
stephendiehl.com
since i'm still in the early phase of the project, i know that python expressiveness is an edge, but i'm looking right now at what's going to be the "definitive" language i'm going to rebuild my product for the next 3 to 4 years.
Python badly needs optional typing. really. i'm pretty sure that would solve both the speed and tooling issues. right now, for me, it starts to become unsuitable as soon as you reach 5-10k lines of code and a team of 2.
Just because you can get around a feature doesn't invalidate that feature.
I'm not sure how I feel about static typing in Java, but it does do what OP talks about, because by default you're not playing devils advocate with reflection.
With Java + Spring + XML you can get one language for the price of two!
How useful is Java's type checking when one typo in an XML file breaks your program?
At least in Python I can follow the logic of the language and find where the mistake is. Spring's XML files don't have such logic and usually all you have is a confusing traceback to work with.
Static typing can guarantee the functionality of quite a lot of code, but it cannot do this for any arbitrary piece of code. As an example, no compiler in any language can know what whether a user will numbers or garbage, when you're expecting numbers. You need to check for an exception in Java even in this simple case.
Instead, you need to figure out what level of reliability you require and then work out a strategy to get there. Static-checking can be one element of this strategy, since you can use it to eliminate whole classes of bugs. But I've usually found that architectural solutions like canaries, automatic restarts, feature disablement, fallback codepaths, etc. are more effective at reducing errors per programmer-hour spent on them, and when you have those the reliability thresholds for the actual code you write can be much looser. (You probably still want to verify the hell out of the framework code that implements these architectural mechanisms, though.)
1_ I just coded something close to what i'm doing in java (using spring), and that made me feel to stay as far as i could from that bloated ecosystem (although the language in itself is ok)
2_ I looked at golang but was scared away when i saw that there wasn't generics or anything that would code some "meta level" functions. Plus, it still felt not broadly enough adopted at that time. Yet it IS the language i'm planning to rebuild my project for now (although i'll probably first rebuild just a small part of my api to see how the language behave with business process modeling and DB interactions, because i'm still very skeptical).
3_ I love C# , frankly the best language hands down, yet i don't like the idea of depending on the MS ecosystem (i've always felt like one day they'll come and make me pay big time). Plus windows seems like a falling platform to me.
4_ I just coded a moderately big website on django, and a big "sudoko-like" generator in pure python and loved the expressiveness and power of that language (as long as you're using the default data structures). Speed never was an issue, and since it was basically just a website, reading the django documentation was everything i needed. I had a look at ruby but frankly there wasn't a big enough difference to make me want to switch.
Here's the story...
i know some people (the european southern observatory here in santiago) that use zope interfaces to address your concerns. i find the whole idea somewhat odd, but they seem to like it and make decent software, from what i've seen.
http://stackoverflow.com/questions/2521189/purpose-of-zope-i...
I'm not sure i'm going to use it though. That's another library i'll be dependent on, for something the language itself should provide.
To me, the fact that some projects had to create those things is a symptom that the language needs to evolve.
Personally, I'd love to be doing multi person large project stuff in Haskell... I can dream :-)
Every time i have a look at Haskell snippets, i feel like i'm reading math. I'd like to read code like i read english...
I think Haskell has the same problem as Lisp. Since you can abstract lots and lots of things away, there are tons of codebase specific abstractions that developers must keep on their heads before they can start coding.
That said, the Haskell situation is better than Lisp's, because it has a big set of "default" abstractions that become part of the language, reducing the number of codebase specific ones.
see this for a quick video overview: http://www.youtube.com/watch?v=TuewiStNT0M
but basically you should do a "Create project from existing code", then wait for a while until the analysis engine parties on your code. if you run into any issues/questions, feel free to contact me at zatroms@microsoft.com (my login is spelled backwards). thanks!
And this isn't even a case of "If you do everything perfect like me, [Language] works great!" - We only have 45% code coverage, and our code architecture, in certain places, is really sub-optimal (Caused by us, not the language).
When I'm about to make a major change to a function (different inputs or different outputs) I'll always start by grepping first to get an idea of what I'm about to get myself into, write or adjust tests for the new 'signature', change the function, and code/debug/test until it works. My team is in the loop before my code is committed to a shared branch.
* We have other responsibilities besides the Python code.
Edit: If your team is stepping all over each other at such a low LOC, my hunch is that its related to a communication problem with the team, your code is too tightly coupled, or there's not enough architecture planning (Too much is bad, but so is too little).
I would expect to see more code reuse in Python than in any statically typed language.
It wins you more code reusability with less code written to accomplish aforementioned reusability.
I always run into silly inheritance chains, type coercion or adapter fns/methods to handle reusing code for unlike-types, rather than relying on a common base of functionality regardless of where it was derived from (function API, inheritance, mixin, etc).
Best way I've seen this done is in Clojure. Good mix of the best of both worlds in terms of static and dynamic typing, multimethods, protocols, essentially structurally typed arguments, etc.
Also I should clarify, 60KLS is not all 1 app: it's spread across 8 discrete apps, an API, and data layer (models).
- tree-structured sqlalchemy managed objects comparaison , generating diffs, then applying diffs to those trees, and persist everything. I'm using sqlalchemy declarative approach.
That diff applying is performed in a celery background task, reusing my flask configuration.
So, in the worst case, i have to deal at the same time with : - Business logic on a bunch of SqlAlchemy ORM object ( declarative approach) - a Flask request context - a celery task context - and sqlalchemy session
At that point, i'm changing the signature of a function that takes pieces of those three parts to perform some business logic. Now i'm telling that the IDE (pyCharm, the best one) and python "compile phase" doesn't give you a CLUE on what you're doing.
You're dealing with so much "magic" that it becomes unmanageable. You don't need that many lines of code to reach that point.
EDIT : you've got 6 people working on 60 LOC spread on 8 discrete apps. That's about the same amount of isolated group of LOC per person than me (a person having to deal with a group of 5-10K LOC)
How is your app structured? I'm guessing that Flask is the top level glue and everything else is scattered around the Flask app. That's the general approach (in most modern MVC frameworks) and I think it's also the root cause of complexity.
Celery and Flask (and sql alchemy too) should really be asides to the main codebase. The code should be layered and discrete libraries for handling different parts of the system. If you have 6k loc that all cross reference one another then you have problems in any language. Presumably there are a number of different components in there. Each should stand on its own with as simpler api as possible. As ever, too much coupling is going to make it impossible to reason about your code.
If you're about to change a signature for a function, it should already be fairly obvious as to where it is called from. If not, you need to ask yourself why. What is this function that's so fundamental to the system that it could be called by any module? Why is it buried in another module an being accessed from elsewhere?
My current app has about 4k loc in python and the same again in js (angular). It's broken into dozens of parts that I only connect where needed through a simple api.
At the core is a sort of image processing library (that itself contains lots of different components). On top of that is a system that works with the image processing. Above that another system that interacts with the data models, uses the system below and farms out processing to picloud (though could use celery). Finally, the Flask layer just provides a web interface to talk to the system that handles that business processing. I can tap into any of those layers to drive them. The point is that I can operate at a high level without needing to consider any details of deeper parts of the system.
These are the layers of abstraction that make a system understandable and stop it from being brittle.
Python needs a way to make those interfaces automatically verifiable.
It's even worse once you start to use big libraries. If you're using top level functions, then maybe your IDE can help you, but as soon as you're dealing with magical properties or parameters, that becomes a mess.
Take for example the "desc" magical function in SQLAlchemy, on things like "order_by" on relationships. That's extremely useful and clever, but i'd really like python to give me some "ok, you're not doing things wrong" message as I'm typing.
Even better, once i enter a relationship declaration, it should give me a list of all the parameter i can use, along with the things thoses properties accept for value. This way i wouldn't have to check the documentation every time i'm writing one. It could also let me discover new things as i'm typing ("hey what's that property doing ? that looks interesting..."). Autocompletion is another way of discovering APIs.
But for that, you need type declarations.
EDIT : as for my api, it's really nothing fancy. It's structured in three big parts "admin / common / public". They each have their "model / business / service" layers, and each have their modules. Only the service layer is impacted by flask. I have some "utils" modules for very low-level stuffs (json serialization, etc). Flask configuration are used a bit everywhere, because i want my api to have only one configuration file. Nothing special, really.
Changing a function signature is the kind of thing that tests can & should catch. You need them anyway to catch edge cases in the logic and document the code, and then once you've done that you usually get pretty good coverage for free.
To give you an example on how i do small refactoring in objective-c :
I change the function (or class) signature i know i need to change. I compile, then xcode shows me every single line of code i need to change.
Most of time, it shows me places i didn't remember also used that function (i'm talking "utils"-like functions).
Python has an unfortunate culture of "just an editor, please", so the main IDEs for it are light-years behind Java IDEs, but you could probably easily whip up a vimscript or .el that does this and highlights the line in the editor.
On a separate note: sqlalchemy is one of the best pieces of software I've ever used in any language!
Biggest benefit is probably the flexibility. SA allows every pattern from raw table access right up to complex object hierarchies mapping to joins or views. If you've already got an existing database structure it's invaluable to be able to hide the "implementation" in this way (I know that's not really the right term) and just present the API that better reflects the domain.
On the other hand it's also a really good tool for creating the tables yourself from the python table declarations.
I've used SA in both enterprise environments mapping really hairy old database schemas to greenfield web apps and I've never found anything that it can't do well. I'd go so far as to say If you're doing relational database stuff in python and you're not using SA - you're doing it wrong :-)
http://docs.cython.org/index.html
Cython can turn Python source into a binary `.so` that you import, thus removing the overhead of the bytecode compiler/interpreter.
You can also annotate any or all of your Python variables with C types. This way, Cython will generate raw high-performaning C using Python syntax. Other Python functions can call your optimized C versions, allowing you to optimize only the hotspots.
Further, if you need to integrate with other languages, Cython can import other libraries' C functions and call them. Conversely, it can generate a header file for your Cython code, thus allowing other languages like C to call your existing Python or Cython functions.
I have wanted to play around with implementing some type checking, but none of my stuff is Python 3 unfortunately.
Do you know how mature it is ?
"The only way that annotations take on meaning is when they are interpreted by third-party libraries"
and
"Following from point 2, this PEP makes no attempt to introduce any kind of standard semantics, even for the built-in types. This work will be left to third-party libraries"
means the author really isn't serious about it. If you don't have a standard semantic you've got nothing useful.
This is specially an issue in the enterprise world where very few people care about unit testing.
Love the contact page [1] of this website. What are some other effective filters?
After reading those weirdly-written snippets, I'm not comfortable with buying more code from the author.
edit: To be more on-topic, Dwolla had a pretty challenge for their hackathon last year: http://venturebeat.com/2012/06/29/dwolla-etsy-join-up-for-ny...
tufqifo.n.ejfim@hnbjm.dpn stephen.?.diehl@?????.???
since m=l in the name, assumed the same for domain, considering it's the most popular email service, gmail.com my first guess.
since n=m in the domain name, I assumed the same for the middle initial, which gave away the whole thing.
I imagine webmaster@stephendiehl.com would also get to him.
Yet, looking at his two scripts, the Haskell was so transparent as to not need running, whereas the Python was opaque (probably because I don't use numpy).
HOWEVER, In most of the compiled versions you set speed/space/safety/debug levels 1-3 and get different uses of the type hints.
In SBCL, if i set safety/debug at high priority, I will get warnings or errors if I make a type mistake.
Furthermore, there is an option you can turn on to make the type infrencer tell you what it is doing.
Here's some example doc:
http://common-lisp.net/project/cmucl/doc/cmu-user/compiler-h...
You'll notice that CMUCL's compiler is called 'Python' so technically, Python supports type inference. :)
You can't say that Python has both speed and great libraries. It has one or the other. Hopefully this will change at some point and I'll be able to reap the benefits of both.
If I have code that wants those libraries, PyPy is not an option. As an end user, I don't care who should fix it. I just know that I can't get my job done with PyPy.
And the authors of numpy, scipy, pandas, etc seem to agree that it would be impractical to rewrite these libraries in the way you are suggesting.
It wasn't clear to me what in particular the author thinks is Python's "niche", so I didn't understand the point of the article as an article, although the content was interesting.
I'd say the biggest challenge to Python in that niche is the emergence of type inference for statically typed languages. Not too long ago, the difference between a statically typed language and Python was
List<Foo> foos = new List<Foo>();
foos.add(new Foo(2));
foos.add(new Foo(4));
versus foos = [ Foo(2), Foo(4) ]
Now, a static language might look more like this: val foos = List(Foo(2), Foo(4))
As far as I know, there still isn't a statically typed language that matches Python's simplicity and low barrier to entry, but I say that as someone who knows little about Go. def foos = [ new Foo(2), new Foo(4) ]There's also this classic bug:
def printOne(Collection c) {
if (c.empty) {
print("Collection is empty")
} else {
print(c.iterator().next())
}
}
Can you spot the bug? This code works for all Collections... except Maps. If c is a Map, Groovy translates c.empty to c.get("empty"). Constantly having to be on my toes to avoid stuff like that is a pain.> let maybe_use_haskell = [ Just 2, Just 4 ]
also, in you could do your original example in a single line,
List<T> ts = new List<T>(List.fromArray({T(2),T(4)}));
even though it's still messier than something with type inference.
foos = [2, 4]
even the fact that the 2 and 4 refer to foos can be inferred, as long as Foo behaves like a number and you make it part of the Num typeclass. It's pretty cool!So, oddly enough, static typing can actually make your code less noisy rather than more.
The language is in active development, although it is still early days with library support currently lacking (namely, a good plotting library).
I know, it's a terrible language, but it is a great collection of capabilities. Quintessential example of love/hate.
I'm not certain what sort of graphics work you are considering, but if it involves a lot of nonlinearity or higher level functions, it may be a good fit.
Do you avoid Ruby and Perl because each other exist?
Avoid JS because jQuery and mootools both exist?
tl;dr: I don't think the Python2/3 : jQuery/mootools analogy is that fitting.
On the other hand, I still happily use Python(2) for some things, but that is primarily because of familiarity and awesome tools like what was mentioned in the article (namely Theano and pandas). Once the Rust language becomes more stable, I hope to move most of my development to that.
Your analogy breaks on a number of levels.
And I believe Python 2.7 still receives security updates when needed.
Obviously this depends on what libraries are useful to you, but of the common libraries that people talk about, most of them do support Python 3. Numpy, matplotlib, Django & PyQt, for example, all do.
But Python 2.x IS frozen at 2.7 because of 3.
I'm using 2.7, and I can't recall the existence of 3 ever causing a problem.
But I am wary, despite being a python bigot myself, of usin one language for all these things. At a very early point down this road it is simply better to pay the cost of adding a new platform and using clojure to make my DSL
In the end, for production, there is a fine line between bending and breaking.