Django 3.2 – News on compressed fixtures and fixtures compression
paulox.net
paulox.net
Django, RoR - they're almost 2 orders of magnitude behind the fastest frameworks on the composite test and they're an order of magnitude behind the majority of Java / Go - i was going to say C# here but i see asp.net is much faster than the surprisingly homogenous performing "usual suspects" of enterprise web frameworks.
And yet, if i had to launch a site tomorrow, it'd be Django i'd reach for - and with zero hesitation. Instagram, Youtube, Github, Facebook, Dropbox etc. etc. all launched on these "slowest of the slow" stacks and then optimised. Did anyone start out on a lithium (c++) or actix (rust) or other incredibly high performance stack?
This is all a big diversion anyway, I think these benchmarks measure the wrong thing entirely. Instead - show me some measure of how easy & cheap it will be for me to change or add a behaviour to my site 3 years down the road if i've stuck to the prescribed idiomatic approach.
I am B2B. My server costs are like 1% of my revenue using Django. Okay, I could shrink this by 95% by using raw C.. but who cares?
Now if I were an image hosting CDN and Django bill was 50% of my revenue.. it might be a different story.
I think it's the classic right tool for the right job story..
Django is fast enough. You'll be spending at most a fraction of an FTE salary in hosting even with the best of success, and if it comes to more than that, you'll be likely paying for a DBA or for an improved caching strategy for performance.
That's kinda the point.
If your business is making webapps for clients, the fact that you can pip install django-half-the-feature is awesome. What you trade for that is a big dependency tree whose packages sometimes snake their way through your whole codebase, since it's all Django anyway.
FastAPI is its own thing, but it's also heavily based on some other big libraries that also have their own support.
And anyway, what I was saying about performance and decoupling are probably just rationalizations. I focus on writing APIs, and that looks nicer to do in FastAPI than how I've done it in Django. I also really like type hinting, and that's a big part of things. Right now, I'd just like to write more Python and less Django.
Great, you need an ORM, what do you choose? Now you need a serializer, what do you choose? What about validation, what do you choose? Authentication, what do you choose?
A lot of these questions are boring and are not worth spending time thinking about. Almost all choices will leave you with some amount of headaches though (as will choosing Django).
The question to me boils down to... what questions am I interested in solving?
> Great, you need an ORM, what do you choose?
SQLAlchemy. It's very popular.
> Now you need a serializer, what do you choose?
Pydantic, which is included. (Django doesn't include serializers).
> What about validation, what do you choose?
Pydantic.
> Authentication, what do you choose?
The included auth library.
My point isn't that there are solutions, my point is that you have to make decisions, all of which are likely to have rough edges.
FastAPI is a tool made for making APIs. Django is not. So in that specific case, Django actually requires more decisions overall.
And if you like DRF, the author (Tom Christie) wrote Starlette. FastAPI is based on Starlette.
(Vitaly - if you're watching. Hi!)
We have a single Postgres instance and serve thousands of requests per second. You're going to be scaling Django way before that.
Spring Boot basically attempts to turn Java from its statically compiled origins to an annotation-ladled mess full of runtime reflections and compiler pluginitis.
If you're doing that, why not just go the whole hog and use an actually dynamic language and a framework that plays up the language's strengths instead of working _around_ it?
SpringBoot epitomizes the worst of the marriage between an OO fetishist and an enterprise "architect."
Django is stupidly quick to do "standard" things. And Django's "standard" extends a long way. And Django understands a lot more about "standard" than you do until you've been coding your app for a while.
Yes, Django and Python are a bit "boring" today. For those of us with a job to do rather than resumes to build, that's what we need.
https://github.com/TechEmpower/FrameworkBenchmarks/blob/mast...
Django ends up getting benchmarked against "uvicorn" which just uses raw sql, so yeah, of course raw sql is faster. Django has a raw sql option too, but they're not using it for the benchmark.
https://github.com/TechEmpower/FrameworkBenchmarks/blob/mast...
Recently I've been tempted to revisit them and build out a manage.py command to run the domain scenarios that generate these fixtures (to make it easier to keep them up to date), since when testing DB migrations you often actually do want to have stale/old-schema DB state, instead of the default of using the post-migration code to generate your test fixtures. But I'm still not sure I'd prefer them in many tests outside of the migration-testing usecase.
I don't think we'd miss them a lot if we didn't have them for this, but they are handy for this. We use setUpTestData and factories or models directly for most of our testing.
That said, I've just written a new tool for exporting a subset of production data for local use (we have a 2TB database but want a <5GB realistic looking dev database, we export a subset, anonymise, maintain FK integrity, install a bunch of Factories for special test cases). I initially started with loaddada/dumpdata but found them inflexible so I'm using the same de/serialisers with custom code around them. That has been somewhat useful, and if that's the level that this compression code is written at I'll probably use it.
I've found that I really like Model Bakery (https://model-bakery.readthedocs.io/en/latest/) a lot more than Factory Boy for this reason - no factories to maintain. It's very intuitive and easy to use.
``` make dev/init ## runs compose in dev mode make dev/populate ## initializes data ```
would give you an "initialized" data, ready to start working on. When the project has some data in it, I consider it a closer reflection of prod. If you can (data is not sensible), you can update the fixtures with data from prod.
Just make sure all devs use the same days from prod, then. I once was consulting for a dysfunctional team and everyone having different test data drawn from prod made reproducing even easy bugs hard and impossible for more complicated bugs.
But this is like... < 1% of my total tests.
I have a series of E2E tests with Splinter [0] that go through setting up a variety of things from scratch and testing core features. Instead of having really long, complicated test that goes through the steps, I instead wrote a little helper function that "freezes" the DB after each step as a JSON fixture (which is then loaded by the next test), to allow me to break it down into many smaller tests that "start" from various reasonable database states. Works pretty well!
I don't use them however for unit tests or any non-E2E tests. I just use helper factory functions with hardcoded test data.
gzip is basically never worth using anymore except for backwards compatibility; zstd is both faster and more space efficient.
lz4 similarly is worth adding for utmost speed but I tend to think zstd is better in practice (particularly when the control plane is Python where perf is somewhat limited anyway).
The article is about offline creation of compressed data dumps and not about real-time application.
The Django code here use only compression algorithms already present in the Python standard library to avoid unnecessary dependencies.