The CPython crowd apparently made some change between CPython 3.4 and CPython 3.5 which changed how Windows DLLs are called. This broke Py2exe, which no longer works for the current release of CPython. [1]
CPython's approach to C extensions usually requires having the same Microsoft C compiler version used to build Python. This is now a problem for Python 2.7, because the Microsoft compiler version used to build it is no longer available. Microsoft recently changed their API for the C runtime, and CPython has been changed accordingly.
CPython's "C interoperability story" isn't really that good on Windows.
There is a CMake build system for 2.7 that works great with VS2013-2015:
https://github.com/python-cmake-buildsystem/python-cmake-bui...
Note that the problem you mentioned w/ python 2.7 is gone: MS has made a .msi available free of charge to build C extensions for 2.7, indefinitely (thanks to Steve Dower): https://www.microsoft.com/en-gb/download/details.aspx?id=442...
When that is not enough, you can still recompile manually if need be, so long as source for the module is available (which it is in most cases).
On the other hand, native modules are such a big part of the Python ecosystem, that not supporting the API and ABI at all confines any alternative implementations to a very small niche.
This is less problematic on other OSes where GCC dominates and clang is compatible, but it can bite on HPC systems that use PGI compilers.
But on the web there is a point. Unless you can use HTML alone, you have to use JavaScript -- and you can't just plug in a module that you wrote in C.
But you can not buy time.
Developing python applications in many cases is much quicker than developing comparable applications in C++, C or Assembler. Therefore it makes a lot of sense for many companies to choose tools with quick development time to get the product running and care for performance later.
Since there may be so much code be written in python, it cannot simply be converted into a faster language since most critical parts are already accelerated using the proper libraries provided by the language ecosystem. Porting a complex application to another language in many cases is a really hard problem and might even not be possible to be done in smaller steps.
Therefore we need every bit of performance we can squeeze out of python. The whole world will gain from that. If we manage great strides like JavaScript it would be glorious!
Go's biggest problem is a lack of generics, which means that in some cases, you have to fall back to duck typing. So the worst case is no worse than Python's best case.
The next-step isn't Go. It's Elixir. Or Rust if you need a systems language.
This is especially true in the last 5 years, which contrary to your post has seen a shift away from multicore task-parallel hardware (core counts are not increasing) and toward data-parallel hardware (wider SIMD lanes, better on-board GPUs, APU architectures, etc.)
For me, explicit pthreading should never be done except with systems software. In which case you're going to be using a systems language like C, C++ or Rust anyway. Not the oddball middletier stuff that isn't high or low level like Go, Java and C#.
> In the case of an I/O bound Web server, I don't think that Go is going to appreciably result in "leaving money on the table". Python releases the GIL on I/O, which is where your time is largely spent.
GIL doesn't matter, the OP said he was writing single threaded application. This means he's doing async (unlikely) or blocking the entire application process (not just the current request) on I/O.
> Linux is very fast at spawning threads. When you're doing I/O, it doesn't matter.
Not as fast as goroutines, and spawn speed isn't the only (or even the most interesting) measure of efficiency--you can run several orders of magnitude more goroutines on a Linux box than you can threads. Of course, none of this is relevant to this conversation, since the OP is talking about running a single-threaded Python process per core.
> Go's concurrency is not "fundamentally different" from any of the other languages. It's just an implementation of thread-per-connection in userspace. This is an implementation detail, and for a typical Web server I think it doesn't matter much.
You're conflating Go's concurrency model with its webserver implementation. Go's concurrency model (movable M:N coroutines with implicit yielding and a synchronous abstraction over async IO) is fundamentally different than the other approaches (threads and vanilla coroutines with some blend of sync and unabstracted async I/O).
Since you mentioned thread-per-connection, Go's "threads" are much lighter than even Linux threads, so you can service many more simultaneous connections (by many orders of magnitude in extreme cases). This is especially true if you're not using async I/O, and blocking I/O is the default for most of these other languages.
> this is more efficient than spinning up an OS thread per request
Linux is very fast at spawning threads. When you're doing I/O, it doesn't matter.
Go's concurrency is not "fundamentally different" from any of the other languages. It's just an implementation of thread-per-connection in userspace. This is an implementation detail, and for a typical Web server I think it doesn't matter much.
Everything leaves some performance on the table. Everything also leaves some productivity on the table too.
Python is in my estimation 20-50% more productive as a language than Go. I run my web stuff on PyPy/gunicorn/Nginx which within reason, can't be leaving much to be wanted performance-wise.
I've given Go a good-go, but I haven't found anything that strikes that balance of ~25% more productive than Go, but still grant similar performance like PyPy does for me.