Types for Python HTTP APIs
instagram-engineering.com
instagram-engineering.com
In place where we need client to be able to define the validation the request/response will consist of data, validation json schema. In this case use jsonschema validator [2] [3].
So for every json request/request data, it's passed through a marshmallow model which validates the json. In some cases we use jsonschema validator when we have dynamic json data with schema definition for validations.
For response the function return value pass through a marshmallow model. We are moving towards making all of internal functions/methods with type annotations to generate documentation using Sphinx [4] plugin.
We are not Instagram but are very happy with it and can replace flask library with bottle or other wsgi compatible framework and it will still work.
[1] https://marshmallow.readthedocs.io/en/stable/
[2] https://python-jsonschema.readthedocs.io/
I've wanted to explore using the Typing module to replace Marshmallow since it started making the rounds to see if it results in better performance, but haven't had a chance. I would have liked to see Instagram release a library to go with this blog post so I don't have to do as much legwork.
As first step we are exploring marshmallow with data classes like the way we are dealing with jsonschema using pydantic.
In our product we prefer to use as little 3rd party packages as possible and rely on standard library. When we want to use 3rd party package we look at the code and if my team can support and enhance it then only we use, except in some case where package is better than standard library like requests.
FastAPI has a lot of advantages at this point for someone looking to do the same: https://github.com/tiangolo/fastapi
hug has some catching up to do, some of it just because it got in so early (It needs to be updated to be compatible with mypy and the types defined typing.py!) in any case using typing on API endpoints - independent of how you feel about Dynamic vs Static typing in general, just makes a lot of sense IMHO.
What else does it offers besides performance (switching to async is no free lunch)?
- Integrates nicely with some existing libraries (Starlette, Pydantic)
- Well documented
- Auto-validation of endpoints from data models
- Auto-generation of OpenAPI schemas from those models
- Auto-serves live API docs from that schema
- Easy definition of sync and async endpoints
Again, it was just a tiny proof-of-concept prototype, but I'm sold on using it on future projects.
Is there a package that can do a reverse? Can I give it an OpenAPI spec and get a code stump going?
Solves most of the problems for that.
Connexion [1] is built on top on Flask and does routing and validation based on an OpenAPI spec.
I've recently started developing Pyotr [2], which does the same only based on ASGI and Starlette. It also includes a client module.
[0] https://openapi-generator.tech/
class CreatePostInput(InputClass):
text: str
user_id: str
...
@handler.post('/post')
def create_post(input: CreatePostInput):
...Plus the fact it is built on top of Starlette and Uvicorn and is frequently among the top of benchmarks of python based framework performance.
Fairly complete starter projects.
https://github.com/tiangolo/full-stack-fastapi-postgresql/bl...
@annotate(foo=int, bar=str)
def view_func(foo=0, bar='hi there'):
...
The types could also be arbitrary callables to parse things like datetime and what not. I'd parse the params from either the get params or post data (json, urlencoded, etc).But now I'm using graphene and graphql to handle all this -- it's a better way to do all of it imho. Of course this all came along after instagram, so you didn't have that choice back then.
In that framework your serializers are by default auto-generated from your model classes; this is convenient to get started, just like Django itself.
But if you have a relatively clear idea of what your API should look like, there are great benefits to be gained by providing the specification first. This way, you don't need for the API to be implemented to start developing the clients, even by completely unrelated teams. Second, your spec will already include your routing and validation rules, and there is no need to manually specify e.g. Pydantic models.
I recently wrote a PoC of a framework [0] that uses the OpenAPI spec to easily implement a REST(ful) API; in a nutshell, you need to implement endpoint functions that correspond to the spec's `operationId` names, and it will automatically route a request to the right endpoint. It is fully ASGI compliant and as has a bonus client module, which allows you to do the requests.
Related: are there any good Python libs for doing request/response validation based on OpenAPI v3 schemas?
I would dare anyone come up with even 100 for the Instagram app.
It also handles Union types reasonably well and lets you put in hooks to handle ugly cases. It's used in production on a system with a big complicated payload, and I designed it to be easily extensible if the standard rules don't work for you.
Secondly I find it really impressive that such a large company with so many smart people can produce an application so mediocre and make the experience extra terrible and me wonder what the absolute fuck is up with that company by trying to block desktop browsers from the perfectly useable on desktop web app (the one in which you can upload things).
Especially for a photo centric application (that has since begun to be used for original video production, of course hampered by the insane lack of any options, starting with the aspect ratio) one could expect a normal work flow to include transferring photos from a camera to a desktop computer. Making your browser pretend to be a mobile phone seems like a step that could maybe, if you really tried (to take out the arbitrary restriction that you must have explicitly added in the first place), be made unnecessary.
So that's my (condensed) rant about Instagram as a whole.
And wouldn't a 10% hardware cost decrease be worth a lot at a company with such a high load, or is it all rendered so seldom because of the caching?
Besides most of those problems are UI-related, this article is about the backend and really python is perfectly fine for a backend that has a huge focus on data (for obvious reasons)
Usually, that's exactly the story: there are certain hot spots that account for the vast majority of your processing time.
I work on an optimization engine that evaluates financial plans.
We have a ton of business logic that is not performance sensitive, and we decompose those objects into flat, regular primitives so they can run in a tight loop that does the actual simulation.
We optimized that using numba and get quite acceptable performance.
But let me revisit what you asked:
> I find it really impressive that such a large organization
In a large organization, you need to find people with the right skillset, and if someone is a subject matter expert (mathematician, statistician, etc.) they often know Python. If they can read and understand your code, they can directly check it for correctness.
Or they can load modules into Jupyter and work with them. Have a lingua franca is itself very powerful.
They articles says it is mostly a monolith with several millions line of code.
def f(x: int):
x = "str"
f(42)You can cast things or subvert the type system and the code will still run.
At least where I work, my build process prevents me from running Python code with invalid type annotations, so it's exactly the same as for Java is cpp or any other statically typed language.
You can't grep for dynamic type errors. You can grep for a cast to see where you've made a mistake and find a way to do it without a cast. This of course goes for the named casts in C++. It is harder to grep for casts in C.
In c you have void ptr everywhere anyway.
Examples of untyped languages would be B, assembly, Forth(?) - where the system doesn't distinguish what type of objects are in memory or storage locations.
What the "whole Type thing" adds is type hints. That is to say they tell the programmer what types are expected. These hints can also be used in static analysis to try to restrict the types that can be bound to any particular variable. This can be useful to detect bugs in large code bases.