Show HN: Pump, a dead simple Pythonic abstraction of HTTP.
adeel.github.com
adeel.github.com
Well, not at all. Pump is what werkzeug, webob, and other more friendly wrappers on WSGI already are. Basically pointless duplication of work, without understanding why WSGI can't be this simple.
One basic reason that WSGI can't be as simple as just returning a dictionary is that you don't necessarily want the entire body of your response pre-computed before starting to return data to the client. What about long running connections? What if you want to return the head of the response immediately, so the client can start pulling css and js while you compute the body of the response? What if you want to do chunked encoding to support long polling connections, or responses where you don't know the response size beforehand?
Pump is basically doing what lots of other things already do, except without quite understanding HTTP quite as well.
In the meantime, don't bash WSGI until you understand why it is the way it is.
I'm reposting a reply I made earlier to irahul:
Pump aims to replace WSGI entirely. That is, I believe it does a better job of what WSGI was intended to do.
I understand your point that web developers don't necessarily work with WSGI on a day-to-day basis. But if you look at the Ruby web community, Rack middlewares are much more prevalent than WSGI middlewares. Application developers (as opposed to framework developers) often add functionality as a Rack middleware, so that it can be reused in different applications, even using different frameworks. Why isn't that happening with Python as much? In the Python world, instead of writing even simple middlewares to for basic functionality like https://github.com/adeel/pump/blob/master/pump/middleware/pa... or https://github.com/adeel/pump/blob/master/pump/middleware/co..., every framework ends up reimplementing it. I believe this is because the WSGI API is ugly and not as easy to understand as it could be (just look at the average WSGI middleware).
http://dirtsimple.org/2007/02/wsgi-middleware-considered-har...
I think your approach is harmful and will put developers who use this library in a very bad situation. You are putting a simplifying abstraction on top of HTTP. As with all simplifications, things which do fit perfectly in your tool become very easy, and things that don't fit perfectly become completely impossible.
Finally, I stand by my argument. Werkzeug and WebOb let you put the exact kind of simple interface on top of WSGI that you are attempting to, but with the benefits of not restricting you to a subset of HTTP, and interoperability with awesome tools like mod_wsgi.
Documenting at this level is challenging & rewarding. The hard part about specification work, beyond explaining the design and providing adequate justification, is getting consensus. By doing so, you'll learn quite a bit, open yourself up for critique, and in the end perhaps provide a credible alternative.
</speaking-from-experience>
You should also allow multiple headers with the same field-name, since that's in the spec.
I like the idea of Pump, and am tired of frameworks protecting me from HTTP.
EDIT: Looking at the WebOb code. It does have quite a few conveniences for working with HTTP messages. I'm not sure if copying bodies into temp files in order to make them seekable is a "completely fantastic way" of doing things, but if I were you I'd definitely read through WebOb to see what kind of problems you might be up against.
if you don't like being "protected from http" why not write raw wsgi applications?
WSGI protects you from HTTP. CGI protects you from HTTP. mod_python and mod_perl protect you from HTTP. If you're unable to read and parse the complete HTTP request yourself -- perhaps incrementally, there's an idea -- you're protected from HTTP. Something is imposing policy like how many headers to accept, what the longest header should be, how to fold multiple headers with the same field-name, that it's okay to consume memory buffering all the headers, and so on.
In my ideal world, a web app server has access to the full HTTP request stream, calls an incremental HTTP parser [2] [3], and does whatever it wants along the way. If the typical use case is to accumulate a full request object and call a handler, fine, that can be made convenient. But the web app gets to decide.
Perhaps my issue is not with frameworks (in the sense of Django, Ruby, etc), but with web servers. Except, I view the infrastructure for hosting a web app inside a web server as yet another framework. The common use case is optimized for at the expense of the less common use cases, which become more painful than they should be. Or sometimes outright impossible.
TL;DR -- Libraries over frameworks. In Soviet Framework Russia, you don't call code...code call YOU.
[1] http://www.python.org/dev/peps/pep-0333/#environ-variables
[2] https://github.com/ry/http-parser
[3] https://github.com/mongrel/mongrel/tree/master/ext/http11
The Rack style api is NOT a great representation of the HTTP protocol. Specifically anything with a streaming or chunked response.
start_response/yield may not be the absolute best API for this, I haven't thought to much about that. But if you go look into what was done in Rails 3.1 for chunked responses you may realize it's not actually too great.
When you write your own copy of WSGI to change how some words are spelled, you don't gain much, but you lose the whole WSGI community. This seems rather pointless to me.
Take advantage of existing WSGI tools. Pump comes with adapters for serving Pump apps with WSGI servers and converting WSGI middleware to Pump middleware.
Or if you are really keen on working low level, you use a nice wsgi library viz. werkzeug.
WSGI is a low-level protocol that provides a minimal interface to an HTTP Server ala CGI. It is purposefully not an application-level HTTP toolkit. For example, a WSGI component takes an input stream and returns an iterable which could yield output chunks... of perhaps in infinite data stream. These edge cases are sometimes very important and why the interface is designed as it is: inconvenient as it may be for simple apps.
Also, there's no clearly defined format for middleware added by keys - the included middleware just adds plain old keys as it pleases.
But more generally, I don't really see the purpose of this. WSGI is obviously not ideal (thanks to start_response and CGI environment variables), but it's also quite firmly in place in the Python world. Not to mention that it would take the Pump library quite a while to catch up to Werkzeug or WebOb in terms of having all the necessary HTTP primitives implemented. (Multipart parsing, anyone?) Unless the server makers get on board, Pump is pretty much just an added layer of complexity on top of WSGI. Instead of "server | WSGI | WSGI library | framework", you have "server | WSGI | Pump adapter | Pump library | frameworK".
your solution is going to be slower, more memory intensive and will not be able to be http 1.1 compatible. there is a reason why WSGI was designed the way it is
Why do you find `start_response` unpythonic?
> We could remove start_response and the writer that it implies.
I searched for `wsgi start_response issues` and didn't get anything useful. Care to point out what's the fuss with start_response and why it's unpythonic?
"Note that in the WSGI 2 calling protocol, you would simply modify your return values, rather than needing to create a function and pass it down the call chain."
http://mail.python.org/pipermail/web-sig/2009-November/00424...
In the Ruby world if I want to use the awesome library Sass all I have to do is:
gem install sass
Because it functions as a Rack plugin it automatically works with any Ruby web framework & Ruby web server combo I choose. No special setup required.-----
# WSGI
def app(environ, start_response):
start_response('200 OK', [('Content-Type', 'text/plain')])
yield 'Hello World\n'
# Rack
app = proc do |env|
[ 200, {'Content-Type' => 'text/plain'}, "a" ]
end
WSGI is the Rack for Python. In fact, WSGI predates Rack, and Rack is WSGI inspired. # Pump
def app(request):
return {
"status": 200,
"headers": {"content_type": "text/plain"},
"body": "Hello World"}I understand your point that web developers don't necessarily work with WSGI on a day-to-day basis. But if you look at the Ruby web community, Rack middlewares are much more prevalent than WSGI middlewares. Application developers (as opposed to framework developers) often add functionality as a Rack middleware, so that it can be reused in different applications, even using different frameworks. Why isn't that happening with Python as much? In the Python world, instead of writing even simple middlewares to for basic functionality like https://github.com/adeel/pump/blob/master/pump/middleware/pa... or https://github.com/adeel/pump/blob/master/pump/middleware/co..., every framework ends up reimplementing it. I believe this is because the WSGI API is ugly and not as easy to understand as it could be (just look at the average WSGI middleware).
"Pump aims to replace WSGI entirely." <- this is very ambitious. :)
public class App extends HttpServlet {
public void doGet(HttpServletRequest request, HttpServletResponse response) throws ServletException, IOException {
response.setContentType("text/plain");
response.getWriter().println("Hello world");
}
} # hello.cgi
print "Hello World\n";
I wasn't trying to claim WSGI is the first web server to application server interface. I was just showing Rack and WSGI are similar, and Rack was WSGI inspired.Since both are web server to application server interfaces, they aren't fundamentally different - the difference is cgi was language independent and hence defined for the common minimum. cgi couldn't have been defined in request and response objects - it would have caused trouble for languages which doesn't have objects.
> where your application code is given a request as a parameter to a function, and must return a response?
cgi is not given a request parameter - the request parameters are passed in the environment. And cgi doesn't return a response object - whatever it writes to stdout constitutes the response. cgi had to cater to all sorts of implementation - assuming request/response objects wasn't a possibility.
Servlets and cgi aren't fundamentally different, but I guess we can agree they are sufficiently different.
I don't think anyone other than wsgi library implementors code to WSGI. WSGI would be a problem if that's how python web programming was to be done - but that's not the case.
Also because WSGI is designed to be low-level, and if you want a request object you should really be using a library or framework.