Plenty of other languages "get shit done".
> And anyway, from my experience, Python has never been the cause of the slowness.
Then you're not actually performing performance analysis. Python is S-L-O-W.
I've written large systems in Python. I've profiled them to find the hotspots. I've micro-benchmarked them against implementations in other languages (C, Java). Python lost. Always. By a lot.
So then I had to rewrite critical sections of the Python in C (or accept that we'd need 32 servers instead of 2). Eventually I stopped bothering, and used a different language, because the aggregate cost of a thousand small performance issues (and a few big ones) is actually quite high.
You totally miss the point. There are large classes of systems where CPU time in your app is not a significant factor. E.g. my go-to example of this is the first Ruby project I deployed in a production environment:
A queueing middleware server. If we served the queues entirely out of memory we could serve a few million messages a day out of 10% of a single 2005-era Xeon core (and while Ruby doesn't scale across cores, this system could easily spread load by simply running multiple queue processes).But of of that 10%, 90% was spent in kernel space. So even if switching to C (or whatever) were to speed up the user-space part of the code 90%, it would only cut CPU time spent from 10% to 9.1%.
This is fairly typical. If your code is CPU-bound, and it's not CPU bound because of stupid algorithm choices, then things can look very different. The first thing to do is find out if you're CPU or IO bound.
This is only true if you ignore response latency as a comparative metric entirely.
> But of of that 10%, 90% was spent in kernel space. So even if switching to C (or whatever) were to speed up the user-space part of the code 90%, it would only cut CPU time spent from 10% to 9.1%.
I'm really curious what you were doing in Ruby that spent 90% of your runtime in the kernel. Having implemented similar middleware, I found that nearly all my runtime was spent just performing basic decoding/encoding of framed packet data, and as a kernel developer, I'm having a hard time imagining what an in memory queue broker could have been doing to incur that overhead
Response latency makes scripting language unacceptable for some types of problems. But in my experience very few latency problems I've come across are down to language choice vs. badly written database queries, lack of caching etc., that'd still be unacceptable regardless of language. Basically, the moment a database is involved, that's where you're most likely to see your low latency going out the window.
Of course there are situations where language choice definitively matters in terms of latency. Despite using mostly Ruby now, there are certainly places where I'd use C. E.g. I'd never try to replace Sphinx search with pure Ruby code, for example.
> I'm really curious what you were doing in Ruby that spent 90% of your runtime in the kernel. Having implemented similar middleware, I found that nearly all my runtime was spent just performing basic decoding/encoding of framed packet data, and as a kernel developer, I'm having a hard time imagining what an in memory queue broker could have been doing to incur that overhead
The moment you find yourself "decoding/encoding framed packet data", you have lost if your goal is a really efficient message queue if said decoding/encoding consists of more than check a length field and possibly verify a magic boundary value.
If the userspace part of your queueing system does more than wait on sockets to be available and read/write large-ish blocks, you'll be slow.
Of course there are plenty of situations where you don't get to choose the protocol, and you're stuck doing things the slow way, and sometimes that may make a language like Ruby impractical.
But there are also plenty of situations where in-memory queues are also not acceptable, and the moment you hit disk, the entire performance profile shifts so dramatically that you suddenly often have much more flexibility.
5 ms you save in your runtime -- for free, without any work on your part -- is 5 ms longer you can now afford to spend waiting on the database.
> But there are also plenty of situations where in-memory queues are also not acceptable, and the moment you hit disk, the entire performance profile shifts so dramatically that you suddenly often have much more flexibility.
I can only disagree here. I view CPU cycles and memory as a bank account. If you waste them on an inefficient runtime, you can't spend them elsewhere, and there are always places where they can be spent.
Excluding huge sections of the problem domain -- such as "decoding/encoding framed packet data" -- is a cop-out. There's no valid reason for it, unless you're going to go so far as to claim that Ruby/Python et al increase developer efficiency over other language's baselines so dramatically that it outweighs the gross inefficiency of the languages.
It doesn't matter if C is 100x faster than Python if your app is running 49/50 seconds running SQL queries.
The exact opposite, I would like to be able to use Python for more stuff but I can't because it's slow.
Nothing about doing it "prematurely" here. I have tried it for specific problems and it was slow. If anything, having built in in Python first is the reverse of "prematurely". And I have, as many programmers have, a very good general idea of where it's slow and what it can be used adequately for without resorting to C.
And what I argued is that people constrain themselves unconsciously, based on that very knowledge, to the kind of projects they do with pure Python.
Nobody writes a AAA game in Python. Nobody tries to build a browser in Python. Nobody writes a performant Apache/Nginx class web-server in Python. Nobody writes a database engine in Python. Etc etc.
(Note for the "Vulcan"-style literal readers: by "nobody" here I mean "not many" / "not a sane person". I'm sure you can find some experimental, bug ridden, half-developed example of all of those done in Python, with maybe 100 users at most. Oh, and I specifically wrote: "pure python").
Almost nobody does any of those things regularly. I've written perhaps two ISAPI modules (in C++), one Apache module (in C) and in none of those cases the code I wrote replaced the bottleneck. It was just that that was the "right" place to put the code. I've done performance analysis countless times (the first time was, indeed, to demonstrate moving off VB3's p-code and into Turbo Pascal's native code was going to have a negligible effect on the performance of a client-server application) and I've seen, over and over again, the bottleneck being the network and the database.
For the exceedingly rare case it's not, I agree, Python/Ruby/Perl/Smalltalk is not the best choice.
No, but the scripting/UI/mod toolkit for the game might be in Python
>Nobody writes a performant Apache/Nginx class web-server in Python
No, but Python can do quite well serving WSGI to Nginx
>Nobody writes a database engine in Python
No, but Python ORMs work pretty well.
The point I'm trying to make is that I don't think that it's a bad thing that Python sometimes gets relegated to a "support" language.
I think the time of the single language programmer is over. If you need something faster than Python, use something else.
Of course it is bad. It would only be good if the inverse (being able to write the whole thing in Python and get the same speed) would be worse.
But it clearly is not -- hence, it's bad. It might be pragmatic and necessary compromise at the moment, but that doesn't make it "not bad".
>I think the time of the single language programmer is over. If you need something faster than Python, use something else.
On the contrary, multi language projects are a kludge born out of (pragmatic) necessity.
Nothing inherently inevitable (or good) about them.
> Nothing inherently inevitable (or good) about them.
Um, except that coherent planes/tiers of abstraction is the foundation of robust software engineering. You write languages to a spec; compiler authors emit bytecodes for another spec; hardware engineers optimize transistors and lay traces on silicon to meet yet another spec. Etc.
Multi-language is the only thing that makes the world work today. The closest single-language runtime that exists is FORTH and things like that.
Every other "modern" language where you get to pretend your process has a contiguous address space, or that there even exists a thing such as "process", is s pragmatic compromise between compiler and runtime technology.
Except that we talk about stuff that should be in the same level of abstraction in the first place. It's not about writing a compiler or OS in Python, or using Python for CPU microcode.
It's about writing the numerical routines in a scientific/statistics package in the same language that you write the core business logic, etc.
But most developers do none of those things in the first place, regardless of language.
I think that's kinda the point; rewriting critical sections in C is pretty much Working As Intended. That's why it's a scripting language. The point is that you didn't have to write the whole thing in C.
No, it's pretty much not "Working As Intended".
It's just "Working as we have been used to, because, well, what can you do?".
As V8, IonMonkey, Dart, Smalltalk VMs, LuaJIT, even PyPy etc have showed, nothing prevents Python from being an order of magnitude faster (and sometimes up to two), while remaining just as dynamic.
Rewriting parts of a Python program in C is not a technical necessity due to some (unsurmountable) theoretical limit. It's a technical necessity because of a non optimal interpreter implementation.
How would you write PyPy to be "an order of magnitude faster"?
That's because you constrain your use of Python in I/O bound problems, like webpages and such. If you had any CPU bound problems, then you would have found Python is quite slow -- and you can get 2-20-100 times the performance in another language.
Python is very slow (compared to what it can be), which is what necessitates tons of workarounds like: C extensions, SWIG, NumPy and SciPy (C, Fortran, etc), Cython, PyPy, etc. A whole cottage industry around Python's slowness. Which is nice to exist, but not an ideal solution: suddenly you have 2 codebases to manage in 2 different languages, makefiles, portability issues, etc etc. All of those are barriers to entry.
However, implementation aside, I think it's invalid to assert that the existence of a cottage industry is inherently a sign of weakness; it's also a sign of just how pervasively Python is used. Find me another language that's used as widely, as popularly (i.e. not just two random guys using it that way), and in as many different kind of settings as Python. (Java might be a contender, except that it's not really used for the really high-performance scientific workloads like Python.)
Integration is hell. This is a universal fact.
I'd rather pick a tool which doesn't add integration requirements!
I made my case explicitly, and it's about technical limitations of CPython that can (and, at some points WILL) be overcome. PyPy is a step in that direction, though not mature yet.
Python, in the form of the standard and most used implementation for scientific, technical etc programming, CPython, DOES add integration requirements.
It necessitates the use of C/Swig/Cython etc extensions for performance critical stuff. And it necessitates it, not because it puts a gun in your head or says so in some contract clause, but in the pragmatic way, that people find it's not fast enough for their needs in pure form.
That's why all the popular related Python projects are not written in pure Python, but in C/Fortran/et al (NumPy etc).
Only if you're talking about the typical CRUD, website, etc program.
Why should we? People use Python for numerical analysis, scientific computing, etc. Why not let them be able to do that, while still writing in Python (instead of using libraries that delegate to C/Fortran/etc)?
And even a website can do intensive computation work. Why HAVE to write it in 2 languages?
The sad fact is you can't always do this. Then you are forced to drop down to C/C++/Fortran.
It would be great if you didn't always have to do this.
https://www.wakari.io/nb/travis/Using_Numba https://www.wakari.io/nb/travis/numba
And the slides from PyCon Talk: https://speakerdeck.com/pyconslides/numba-a-dynamic-python-c...
For example, consider the problem here:
http://arxiv.org/abs/0903.1817
Ultimately I needed to run a math formula on some surfels and connect an edge in a graph if the formula came out True. Following that I had to prune a bunch of edges in the graph. In 2009 (when I did it), I had to drop down to c++. (Admittedly, this was also partly due to a lack of a good python graph lib.)
I'd be very surprised if Numba is helpful with something like that - the branching and complex data structures don't seem like a great fit for it.
1) Yes I do (occasionally).
2) Because it applies to everybody, not just me. You surely heard of the plural form of speaking before.
>If not, why are you holding forth on Python's unsuitability for these applications without knowing about the available tools?
I never said Python is "unsuitable for these applications".
On the contrary, I acknowledged that Python IS USED for these applications.
What I said is that the tools you mention (the ones that make it suitable for these applications) are written in other languages, which is kludgy. I'd rather be able to have them written (and modified etc) in Python too.
So, Python + NumPy + etc is suitable for these applications, but Python the interpreted language without external C/Fortran help is not.
Python and the faster extension languages used are the same kind of tool (well, obviously, both are languages).
So, let's compare them to two hammers.
The difference is Python is made from soft plastic, and while it is light, it only works for softer nails. For harder stuff you need a metallic hammer, which is heavier.
And the sad thing is, the Python-hammer can also be almost as hard as the other hammer while remaining light, it just needs some coating (faster interpreter innards).
So nothing (except besides the man-years of work needed on the interpreter) forbids you to have a hammer both light AND strong (== dynamic and expressive as Python WHILE many times faster).
If anything, I work hard to put more stuff in RAM to avoid hitting disk, as even though we have SSD's in all our new servers, disk IO is still inevitably where we hit limits first, followed by CPU. Running out of memory is a rare event, and when it is it is consistently an indication of a bug.
I don't understand. Concurrent sessions? Do you mean open connections? Worker threads? What exactly?
The maximum speed of the internet, the maximum speed of your database, the maximum speed of your programming language are all basically constants that you can't improve by throwing money at the problem.
As a programmer, available memory is absolutely your domain or else you're increasing hardware/VM costs for yourself or your employer needlessly. Then you've transferred the cost of increased hardware to host your app/site/whatever off to users.
I can see today bit pinching isn't fashionable as it was long ago when memory wasn't cheap (it's not really "cheap" today either since you're paying proportionally greater per VM than the actual cost), but when a basic tenet like efficiency is ignored because DB all you're doing is compounding the problem.
BTW. Most people already deal with the DB issues with aggressive caching and/or reverse proxy so that still leaves the core application.
There are times when the benefits are greater, for example when your software is running a million instances, or you're working on a hardware-intensive game, but those are certainly not common case.
Once you get past a certain level, though, the cost of the next 1GB isn't $20, it's $20 plus the cost of another computer plus the cost of exploiting multiple machines rather than just running on a single one.
Then it's $20/GB for a while again, then $20 plus the cost of adding another machine, and at a certain point you need to add the cost of dealing with the fact that performance isn't scaling linearly with amount of hardware any more.
That last bit might be a concern only in fairly rare cases. But the first, where you make the transition from needing one machine to get the job done to needing more than one, isn't so rare. And that can be a big, big cost.
(Very similar considerations apply to CPU time, of course. Typically more so.)
This casual disregard for resources would explain why a lot of startups run into infrastructure issues so quickly and settle on terrible solutions (or outright snake oil) to mitigate problems that shouldn't exist in the first place.
People need to start realizing that The Cloud is only a metaphor. Hardware isn't intangible and neither is their cost.
Servers don't grow on trees; they live on a rack and grow on electricity and NAND.
It's easy to think everyone has their acts together like facebook or google, but most companies i've dealt with have hardware upgrade processes measured in months or years, not hours or days. You absolutely have to take responsibility for your work as a programmer and make stuff run fast instead of labeling it somebody else's problem.
Well... For those companies, obviously, dynamic languages are not an option. ;-)
In general if your server is making you money, you'll make more money improving the service than reducing the amount of RAM the service takes to run.
In summary, test the conversion rate of a page/feature/etc, not how much RAM it uses. If you're going to performance optimize, attach a profiler and take the easy wins, don't spend too much time on it unless it's a ridiculous amount of resources or your margins on razor thin.
That's not a law of physics. The number of servers you need depends on the number of users you have. The functionality you have to build usually doesn't. So the more users you have, the less true your statement becomes.
I don't think that has anything to do with the languages. I think that has everything to do with the quality of the memory manager implementation, and there is at least one memory manager that does deliver in this respect, namely C4.
I remember trying to optimize some financial software a couple jobs ago and hitting a brick wall because that's the speed the disk rotates at. We ended up buying an iRAM (battery-backed RAM disk) and sticking the DB on it. You can get this a lot cheaper by avoiding the DB and using a RAM-based architecture if you're willing to sacrifice fault-tolerance under power outages (or if you have some other architectural solution for fault-tolerance, like writing to multiple computers).
For vertical progress, i.e. making complex things possible, these micro-bottlenecks are important. These same DBs with their query processing and other domains with extremely high inherent complexity would greatly benefit from high-level languages yet still the most difficult parts have to be written in the already difficult C(++).
Its just another trade off. Beyond the hype, it has interesting choices, and some very poor ones. (No algebraic datatypes, conservative exploitable will-break-your-program garbage collector, no tail recursion)
But the trade off between performance, safety and compile time is quite unique, and should be considered its main contribution to the language design space. Its just too bad it ignores every other contribution from the last thirty years. Like how to do a type system, or how to write a garbage collection that survives long run proccesses.
Anyone know of a human read/write-ability comparison of programming languages?
Perhaps the metric you're more interested in is the difficulty in learning to read/write decent code in a particular language. Of course, this difficulty can vary based on what other languages the learner is already familiar with.
This is a similar case with languages like C++, Ruby and Python. The languages are terrifyingly complex, but the subset that most people use is reasonably usable. If you stray outside this subset there be dragons.
Even off the shelf functionality doesnt have the same time/complexity properties as say c#, java or go resulting in a fairly basic algorithm reaching ridiculous time constraints very quickly.
As for CRUD type apps, you are probably right but its the bits that aren't CRUD which add the most value to your product.
I'm sympathetic to your point, but I don't think this is necessarily true. What adds the most value to a product is determined by economics. It has nothing to do with whether an app is doing basic crud or nlp wizardry underneath, just whether it does something people find useful. There are tons of basic crud apps out there making people millionaires.
Name three.
Economics drives both demand and supply. CRUD apps are easy to make; the market for most CRUD-only apps is heavily saturated with competitors, so the differentiating value of your application is not likely to be your ability to perform CRUD operations. Put it this way - if you come up with a CRUD app that does something staggeringly valuable, expect strong competitors in days not years. Your chance to make millions is small.
What distinguishes your product is not how it stores or presents the data but what value it derives or extracts from that data.
Most CRUD apps are just a form of lock in which provide no value past storing your stuff in one place. That covers pretty much the entire enterprise software space and most trendy startups these days.
Your mentions of maintenance and reliability in the opening sentence are unfounded. What makes a Python app less so than the same app written in any other language? Surely it comes down to the developer. A well-factored app written in Python (with tests) is going to be far more maintainable than a big ball of mud written in C. The choice of language has little to do with it.
I wouldn't drop down to C ever at application level. There is virtually no need to do this these days without introducing more risks from a security and memory perspective and the inevitable situation of "integration hell". I'd rather throw it in a JVM than drop to C. Then again if I started with a JVM I wouldn't have the problem to start with as it solves both ends of the problem. The CLR is equally applicable here. Go hits the mark too.
From my experience here dealing with a massive 100kloc python project. Python maintenance:
a) Horrible to refactor. I mean really horrible. You don't get everything right the first time and when it comes to fixing that it's hell.
b) no type safety or contracts which is hell when working with other developers.
c) indentation issues which are totally horrible at merge time as it introduces a shit ton more work to do.
Reliability:
a) Some algorithms don't scale the same way as they do in other languages due to the points the article raises. It gets to the point that for something in Java is O(N), it can approach O(N^2) due to all the dict fudging and copying which is not good.
b) It's slow anyway - I mean really slow. If something feels slow, it is slow. PyPy may fix this but (c) below causes an issue as well.
c) Optimisations aren't consistent across python versions (consider how crappy IO was in the first py3k drop).
I wouldn't use Python for a new project now. The only utility I see is quick disposable bits of glue.
>> Horrible to refactor. I mean really horrible. You don't get everything right the first time and when it comes to fixing that it's hell
I think dynamic languages like Python and Ruby are an absolute joy to refactor. There is generally less code-junk to be worrying about when making refactorings. I realise this is subjective however, I think it's difficult to make a serious case that any one language is easier to refactor than another. Granted, the refactoring tools in IDEs for Ruby/Python etc. are much more limited than say C#.
I hope you won't mind if I don't get into the indentation or type safety arguments. They've been done to death! :)
I think we're in agreement that if you need something to be really fast you should use a different language. I just don't think that's a reason to not use Python for the rest of your app. And of course as others have pointed out the vast majority of us are not writing apps that require super fast algorithms, we're all far too busy doing the same CRUD stuff over and over right :)
I appreciate your insight on a 100kloc Python project though, I have to admit I've not encountered anything of that size. At work we have a much smaller codebase that has already gotten out-of-hand but I believe that's entirely due to the developer responsible, not the language.
For example, what functions take class Foo that should now expect class Bar? What functions called method Baz, that now should call Bat?
Static typing makes all of that trivial. You know with certainty, when you've hit all the right spots, because the code won't compile otherwise. Also, a lot of that refactoring can be done automatically through tooling. "Replace all calls of Foo.Baz() with Bar.Bat()". Trivial, done... and it's often impossible to do the same thing in dynamic languages. You have to rely on tests catching everything... and how many of those tests need to be updated now, too? What's testing your tests?
I love dynamic languages, don't get me wrong... but refactoring large code based is way easier to get correct in statically typed languages.
Yes, you need testing discipline, but you need that with static languages too - if you think you're ok just because the compiler didn't complain.... Well, that's just a false sense of security.