In recent years Django had multiple major releases, I still remember it as being in 1.x forever. Does somebody know what changed within the Django Community that they break backward compatibility more often?
In recent years Django had multiple major releases, I still remember it as being in 1.x forever. Does somebody know what changed within the Django Community that they break backward compatibility more often?
The deprecation policy[2] is taken very seriously and Django doesn't opt to break things if it can.
Recently there was a very interesting discussion[3] between the Fellows as to whether the version numbering is confusing as this doesn't follow the same pattern as other libraries.
1: https://docs.djangoproject.com/en/dev/internals/release-proc...
2: https://docs.djangoproject.com/en/dev/internals/release-proc...
Idk. I have to grant that Django ORM likes to make your life easy, but lazy loading on property calls is a dark pit full of shap punji sticks. Just overlook one instance where this is happening in a loop and say goodbye to performance and hello to timeouts left and right...
FWIW they do give you assertNumQueries in the testing tools, which makes it relatively easy to catch this as long as you have tests.
But I have no idea if there are database interfaces that make this problem simplistic. In my experience with Django, anything but the most simplistic page will be so noticeable slow that one has to go through the queries and use things like select related. Occasionally it is also better to just grab everything into memory and do operations in Python, rather than force the data manipulation to be done as a single database query.
It is a good tool for its purpose, but it is no replacement for SQL knowledge when working with complex relational databases.
And as already written above, a slow but correct page is preferable to a wrong page because you ORM is omitting related data.
Try using: https://github.com/har777/pellet
Disclaimer: I built it :p
You can easily see N+1 queries on the console itself or write a callback function to track such issues on your monitoring stack of choice on production.
I don't really have a pitch but here is why this was made:
1. we had a production DRF app with ~1000 endpoints but the react app consuming it was dead slow because the api's had slowed down massively.
2. we knew N+1 was a big problem but the codebase was large and we didn't even know where to start.
3. we enabled this middleware on production and added a small callback function to write endpoint=>query_count, endpoint=>query_time metrics to datadog.
4. once this was done it was quite trivial to find the hot endpoints sorted by num of queries and total time spent on queries.
5. pick the most hot endpoint with large number of queries, enable debug mode locally, fix N+1 code, add assertNumQueries assertions to integration tests to make sure this doesn't happen again and push to prod.
6. monitor production metrics dashboard just to double check.
7. rinse and repeat.
For me this ability to continuously run on prod -> find issues and send to your monitoring stack -> alert -> fix locally workflow is the main selling point. Or of course you can just have it running locally on debug mode and check your console before pushing your changes but sometimes its just hard to expect that from every single engineer at your company. Then again your local data might not cause an issue so production N+1 monitoring is always nice.
Unless it adds a bunch of overhead, seems like a no-brainer to enable.
The header feature can also be useful. If you have a client on-call who's complaining about super slow page loads. Just check their network tab and see which response has a query count/time header which seems unnatural.
Also we often had something that was more like 1+3N, basically a 1+N problem but it was looping through the same 3 queries over and over.
If maximum performance, 100% of the time was the end-goal, I would not be writing Python.
Guess who had to refactor this mess and steer away from catastrophy.
I’m haunted to this day.
POC whipped together without good architecture. Having proven itself, usage increases until the application starts to burst at the seams. Program must be redesigned, avoiding performance gotchas.
I still think Django + the ORM give you a lot of runway before performance should be a concern.
I also realized mid project that the default rest framework doesn’t support good api model generation from OpenApi or vice versa. And most devs didn’t understand why using untyped dicts everywhere was bad.
In all transparency, I am openly biased for static typing and have a background in java, c# and Swift.
So what we ended up with looked similar to a statically typed language but without the tooling or performance.
for book in select * from books:
author = select * from authors
where id == book.id
print book.title author.name
In the real world that nested select could be hidden many levels down in method or function calls or even in code run from the templating language. Example: Django templatetags can query the database from the HTML template and any framework or non framework I worked with can do that too.If I'm remembering correctly, that is a fundamentally poor approach. Instead of telling the database to do some work, the database is doing more work, and the application is doing work. One mitigation is to batch the work[0]. Another is to special-case updates[1], which bypasses all the ActiveRecord pre/post-save logic. In either case you aren't holding the tool wrong. The tool is wrong.
[0] https://apidock.com/rails/ActiveRecord/Batches/find_each - you'll note that the batches are "subject to race conditions" - i.e. each batch is its own transaction! And you're still loading the records into your application pointlessly. You're just limiting how many do it at once.
[1] https://docs.djangoproject.com/en/dev/topics/db/queries/#upd... - note:
> Be aware that the update() method is converted directly to an SQL statement. It is a bulk operation for direct updates. It doesn’t run any save() methods on your models, or emit the pre_save or post_save signals (which are a consequence of calling save()), or honor the auto_now field option. If you want to save every item in a QuerySet and make sure that the save() method is called on each instance, you don’t need any special function to handle that. Loop over them and call save()
I like to describe it as Django's ORM will satisfy 80% of your needs immediately, 90% if you invest and sweat, 95% if you're quite knowledgeable in the underlying SQL. But there are still some rather common query shapes that are inexpressible or terribly awkward with Django's ORM.
SQLAlchemy on the other hand never tries to hide its complexity (although to be fair they've become much better at communicating it in 2.0). On Day 1 you'll know maybe 20% of what you need to. You might not even have a working application until the end of week 1. But at the end of, idk, month 6? You're a wizard.
The long-term value ceiling of SQLAlchemy/Alembic is higher than Django's, but Django compensates for it with their comparatively richer plugin ecosystem, so it's not so easy to compare the two.
I've got a customer with two Django apps. One was developed enthusiastically in the object orientation way. The database is a nightmare. The other one was developed thinking table by table and creating models that mimic the tables we want to have. That's a database we can also deal with from other apps and tools.
Where Django’s ORM shines is in the modeling stage: classes in models.py are your single source of truth, relationships only need to be defined once, no separate “schema” file, and most of the time migrations can be generated automatically. It’s truly top-notch.
https://docs.djangoproject.com/en/dev/internals/release-proc...