See Python, See Python Go, Go Python Go (2016)
blog.heroku.com
blog.heroku.com
It makes calling C# as easy as:
import clr
import System
uri = System.Uri('http://python.org')
And also works the other way around. In both cases you have to mind GIL though.Development on IP3 is, however, still very alive and active so I wouldn't consider the entire project dead:
In native Go, Goroutines are very lightweight and cooperatively scheduled. In a CGo env I believe they each have an OS thread and a full stack. Source: https://www.cockroachlabs.com/blog/the-cost-and-complexity-o...
Similarly, Docker locks goroutines to the OS thread when using the unshare system call to spawn containers. These are of course later discarded. It has to be locked because any goroutine might be stopped at any (automatically inserted by the compiler) checkpoint and resumed on a different OS thread.
Unsharing the network interfaces from half your OS threads is a fun way to chaos test networking in Go.
I wish I could call python code from within a go web server with some ease and safety.
Especially considering that there's quite a few sites which run entirely on Python.
If we go with what I would assume is your definition -- speed of execution versus resources used -- it is certainly possible to build fast Python applications that are efficient.
But use what makes you and your team happy, life is too short for anything else.
This sounds like an issue with the design of your program, not Python.
Python is not fundamentally single threaded - it just has a lock that stops it from taking advantage of threads in cpu bound scenarios.
Python is used in data science because of the C bindings that make it not slow. Also, when in C, you can take advantage of threads since they live outside the GIL. e.g. Dask.
Yes it is. It is not designed to run fast on multi-core CPUs, because there were mostly single-core CPUs when Guido made the language. It has been always a problem since multi-core CPUs are more frequent and it's a front where Python is losing the battle (against Go for example because of way easier concurrency support).
Tomato tomahto
> Python is used in data science because of the C bindings that make it not slow. Also, when in C, you can take advantage of threads since they live outside the GIL. e.g. Dask.
Correct. Python is fast when you aren’t running Python. Of course using C (or anything else) only works in certain situations—there is a cost to crossing the language boundary and very often that cost is greater than what you save by using C. Never mind the added build/package complexity, the security issues, the maintainability issues, etc.
Python is a neat language, but it’s really expensive if your project ever might have tight performance requirements (where “tight” is laughably easy for most other languages). Python can often be made to meet them with enough shenanigans, it’s just costly to implement and maintain said shenanigans.
There is no way around the high memory usage, but a large number of the problems with Python concurrency is not loading nginx (or load balancer) in front and not switching to gevent from PreFork which uses a considerably higher amount of memory per “node” for higher concurrency. That said, gevent is only "performant" if what you’re doing is IO bound. Same thing with any AsyncIO based server.
"data science" sounds like DB or NoSQL heavy so should fit this case, but of course all of this is just general advice and depends on the app/code like others said.
It ranks pretty high on performance in framework benchmarks
That said, it sounds like you’re serving a large model. No amount of async/await or goroutines can solve this problem. A non-blocking web server is a godsend for I/O-bound tasks, but a large model is just a deep call stack - lots of multiply, nonlinear function like RELu, then add, times a billion. This would still block, even if you had perfect async/await code.
I made some assumptions here, but if I’m right, the answer is “shrink your model” and/or “buy more compute”. Neither of which are easy. But if you’re trying to shrink a model, check out Distiller https://github.com/NervanaSystems/distiller
Edit: the restriction I talk about is for event-loop based servers using something like uvloop or asyncio under the hood. Maybe this restriction doesn’t hold for other concurrency modes.
Presumably he’s not serving the model but running it, which is cpu bound, in which case Goroutines would solve the problem.
In the past, we were told that threads were cheap and to use them heavily, especially to achieve parallelism. Now with the advent of async models, we're being told that threads are expensive, and often that a single processor/thread async model is better than a multi-threaded blocking one.
I'm not a luddite, I do agree that async is often better. But I wonder how we got tricked into thinking more and more threads were the answer and how we avoid such trickery again.
The model can do around 10k predictions/s and does it with async, which allows Node to respond to web requests in the meanwhile.
I guess it's a matter of using the right tool for the task, whenever possible, Python for data science, Nodejs for a web backend.
Also take a look at comparison of various frameworks in Python including go http server [3].
People are most productive in the language/framework they are most familiar and will defend it. Every language/framework has their own strength and weakness and I believe over a period of time ideas flow from one language to another.
Indeed today many people in Python community will be moving towards Rust or Go because they think it can solve all their problems they face with dynamic typing and performance, which might not be entirely true. It's upto an individual to decide if they want to go that path.
As an example werkzeug/flask framework developer Armin moved to Rust in spite of it being a complex language with very large syntax surface area and a steep learning curve. In my opinion Rust's complexity, difficulty to learn and understand, and probably a promise of type and memory safety (which is not 100% true given it needs to interface with C in unsafe mode and will only be as safe as the underlying C implementation), makes people adopt it to make them feel better programmers (Personally I would have chosen Haskell, if needed to do the same).
Now all the Python projects he worked on is mostly maintained by volunteers and David. But being a responsible open source developer and contributor before moving in that direction Armin did create pallets projecs [4]. So his decision to move is right for him given his preference and learning priorities.
[1] https://www.nexedi.com/NXD-Blog.Multicore.Python.HTTP.Server
[2] https://lwan.ws/
[3] https://www.freecodecamp.org/news/million-requests-per-secon...
> Keep in mind that this is with 10 concurrent requests, so werkzeug-flask probably chokes more on the concurrency than the response time being slow.
I am not sure though. I’d imagine Go can beat Python performance enough to make up for the (clearly not very egregious) CGo penalties.
(although it looks like it's not close to supporting the needed functionality to import a typical python library)
There is an active fork here: https://github.com/grumpyhome/grumpy
Or you mean to use Python as the client facing API server? I meant how to architect if you use some other platform for the publicly exposed API. Would you build an internal Python API at the lambdas that run the actual data science computation, and invoke that from the client facing API/app?
[1] http://statifier.sourceforge.net/statifier/background.html
https://github.com/google/starlark-go
https://github.com/starlight-go/starlight
It seems to be used by the Delve debugger for example:
https://github.com/go-delve/delve/blob/master/Documentation/...