Data Oriented Django Deux

This video features Adam Johnson at DjangoCon Europe 2024 in Vigo, Spain.

Data Oriented Django Deux
0:20:12
Published July 11, 2024
703 views

Talk: Data Oriented Django Deux by Adam Johnson

https://pretalx.evolutio.pt/djangocon-europe-2024/talk/FTQEBD/

Summary

Adam Johnson argues that Python code is repeatedly transformed and optimized at runtime, while application data layouts largely remain under the programmer’s control. He uses three principles—data layout, batching, and statistical distribution—to show how Django developers can improve performance: use column-oriented structures and databases when appropriate, let the database do more work, right-size fields, and choose suitable Python data structures. He also recommends designing APIs for batch operations, using Django’s bulk ORM methods, pluralizing relationships early, measuring real data distributions to find repetition, applying caching, and returning early when no work is needed.

Key takeaways

  • Data layout often matters more than code structure because processors work best with compact, sequential data.
  • Column-oriented libraries and databases can make analytical operations dramatically faster than row-oriented Python objects.
  • Batch operations such as bulk creation, bulk updates, and multi-recipient notifications avoid repeated database and connection overhead.
  • Design interfaces to handle collections and consider plural relationships early, even when they currently contain one item.
  • Measure the values flowing through real systems to discover repetition that makes caching effective.
  • Use correctly sized fields and early returns to avoid wasting storage, memory, rendering time, and database work.

Summarised automatically from the transcript.

Transcript

3,433 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:02

This is kind of a sequel to my 2022 talk, Data Oriented Django, hence the D, French for two. But it's kind of a reimagining as well. So we're going to approach the idea from a different angle and away with some different takeaways. So friends Pythonistas and DjangoNots, I have some news for you. Most of your code does not exist. With some caveats here in that I mean at runtime. And I mean, well, at least in the same form. Uh here's a simple example from the road of Python. Take this function. We have written if, elif, and else for three different branches of our code. But when we look at how Python Python compiled it using the dis module for disassemble.

0:49

We find there is only one instruction under the hood that represents these three different constructs. This pop jump if false bytecode represents three different things that we're Typing. And that's because we use the higher level syntax to structure our code, but at the runtime, the lower level, more powerful operation lets the computer get on with the work in a faster way. In general, this happens quite a lot with our Python code. We go through many different layers. The Python code we write is compiled to bytecode. That bytecode is interpreted by a C program, C Python. C itself is a compiled language with a lot of different things going on and has many different tools and things, and even down in the processor, instructions it's handed get modified.

1:35

So the things that don't exist in a processor but do exist in our code, things like functions, classes, objects. Even the idea of an individual separate program does not exist. What does exist in a processor then? One and zero. Sorry, I've gone way too low level. We have instructions and data. Everything the processor Is doing is reading some instructions and using those to act on data. As I said, there's many of these instruction optimizations going on in our programs. Here is just a handful of the techniques going on in compilers. and processes. And the point is not to understand these, but it's a very deep topic and I'd encourage you to look at it on the Wikipedia optimizing compiler page.

2:23

The point is here, what we write in code gets optimized so many times by the layers in our execution environment. But when we look at the data side, there's not much going on. I put like the shruggy emoji and struct packing, which is something like the Rust compiler does. But generally compilers Do not and cannot rearrange the data structures we're working with because they're tied to the input data into our program, the output data, and the way we've written our classes, which might be used by multiple different parts of our code. So Most of our code does not exist in runtime at the same form. But our data does. The data layouts and structures and input and output that we use, it still exists inside the processor, and there's not much that can be automatically

3:15

Done about it, that makes us approach the idea of data-oriented design. We focus on the data because it's the one thing we have more control over, and everything else is secondary. This really does mean everything else. This is kind of an anti paradigm paradigm. We shouldn't worry about like layered architecture, classes versus function based views and very various other concepts sold to us to improve any our plain code before we consider The data layouts and structures that we're using. It's okay to talk about that stuff, but we must always be thinking first about the data that we're using. So I'd like to illustrate this with three kind of broad considerations through the

4:00

Talk and give you some takeaways on how to apply that. Layout, batching, and statistical distribution, which I'll cover shortly. So first talking about data layout, how we organize the data in memory. Here's a problem. We're gonna find the mean of 10,000 2D points. Here's maybe a solution that could pop into your head. We take a class point which has an X and a Y value. We generate some random data, 10,000 random points. To generate the mean, we use the traditional mathematical formula. We take the sum of all of the x values, divide that by the number, the sum of all the y values, divide that by the number. If you ran this code, yes, it would run in a blink of an eye, way less than

4:45

499 microseconds on my computer, half a millisecond. But this is way slower than the processor can actually do. If we instead pick what I'll call a column oriented approach, uh which I'll try and explain bit more later. We put the X and Y values in separate lists, generate them randomly again, and then we do the summation and length formula again. That has made the code ten times faster already. Um that's just a better data layout for the processor. But wait, we can do more. We can take a column-oriented array library like Polas or Pandas or NumPy and put our values in that. We use the mean function, so that's doing the same sum

5:32

and lengthing but over in C code, so a little bit faster. And now we're 200 times faster than the original example. So to summarize, we've gone from 500 micro. Mr. Down to three microseconds. 200 times faster. If we use the log formula and think of it backwards, that's the same as 15 years of hardware improvements. That's like taking our code that was running at iPhone 1 speed. And upgrading it to iPhone 15. When we talk about performance numbers, sometimes these very large numbers can become uh abstract, but I like this idea of hardware improvements as a way to think about it. And you know, software chip companies working incredibly hard they have thousands of people on this problem to make computers a little bit faster each year, keep up with Moore's law.

6:21

Um and if we to uh structure our code poorly or our data, we throw that all away and we've undone a lot of work. Why are these slower? Let's look at the row-oriented structure, um, data structure. So we had a list of classes, and then each point has it inside a Python dictionary, which is how classes work, that point to the individual flow. Values. Uh we can call each of these points a row, just like a row in a database table. And they're laid out just kind of randomly in memory. Whenever Python needs to create an object, it finds some free memory and put it There. So we have a list, and then somewhere else we have all the individual point objects, they could be scattered, and then those in turn

7:07

point at different float objects which could be scattered around their memory. The column-oriented format took away one of those. Layers. We had a list and in that had individual pointers to the float objects which might still be scattered around in memory, but it meant a lot less jumping around to find that data. And then when we use a column oriented array library, like polars, It just packs all the float values right next to each other. That's the actually the preferred format for a CPU to read through data. Um it could just do one at a time, it will pre-fetch the next ones because it thinks you're gonna keep working with the same data and so on. So, yeah. This talk isn't to say you should definitely just switch everything to colmorid. But there is some layout considerations, especially when we're doing more heavy computation.

7:55

So here are some techniques that I think apply to our Django programs. is to use column-oriented tools when appropriate. That doesn't just mean these data frame libraries like Polos, Pandas, NumPy, etc. but also databases that are column-oriented. DuckDB is a local one, Parquet and Snowflake are a larger scale scales, and it's it's something that a lot of companies do actually is load their transactional data, which is typically what we're working with in Django, from a row-oriented format into a column-oriented analysis database to make those analytics queries faster. So that's definitely something to consider Consider whenever we need that performance. Another way is to think about it is just to do more in your database, because every time we pull more data over into Django, lay it out randomly in memory, and then work with it, we're a bit slower than if we

8:44

have a little bit of a little bit of a little bit of a little bit of a little bit of a little bit of a little bit of a little bit of a just get the database to work with it. SQL also tends to be faster than Python code. So here's another implementation using an imaginary Django model. We could get the mean point by just asking the database for the average X and Y values and hopefully that will be a bit Custom. Another thing to think about is in your models, we need to right size the fields. Every time we use a very unnecessarily large field type, we are wasting space on the database disk , memory. And in our local server. A classic one is choice fields. You might think of using a text choice field so you get a nice representation, but then you're using, you're storing very large strings, and if you have, you know, a few tens of states, that's wasting a lot of memory.

9:30

We might have strings that are 80 bytes when they're stored in UTF in a four-but byte per character format. Whereas we can use get away with one byte in most cases, uh fewer than 256 case uh states. And then I'd say the last thing in the layout idea is to learn the different data structures that are available in Python. You can't use it if you don't know it exists, so it's worth grabbing them. You've got the classic date set and list, but here are some other Others that exist in Python's standard library, I think we should all have just in the back of our mind for when they come up. Frozen set for frozen, in fact, I won't talk too long, just take a screenshot. The second consideration I have for you is backing.

10:16

And we're gonna go with a with a concrete ticket from inside Django where I optimised the create permissions function. This is the ticket number if you want to look in more detail. I made this function about twice as fast. from five milliseconds to about two and a half in uh one example benchmark. Or more accurately when I was running Django's model tests I found it took eight and a half percent of the runtime. I got that down to about four point seven percent. So when you run managepy migrate, that triggers this piece of code that emits a signal called the post migrate signal. That's run for each app. Over in ContribAuth, that is connected. To this function create permissions, which goes and creates the permission objects for each of your models.

11:06

That code looked like this initially. It would create this C-type set and then loop over all of the app config. uh all of the models in the app config, fetch the content type for that, store it in the set, and then later do some other work. Well, it it just so happens there is also a batch getForModel function. would get for models, which it could have been using, and doing that provided the bulk of the speed up, because then it would not be fetching each one individually, it would just go to the database once. And that's if it needed to go to the database. It does have caching. So that's what the ticket did, but uh yeah, we're probably gonna need a bigger batch to be honest. Uh

11:52

this may be something we could add later. There's still the looping over each app config, so it still has to fetch for each app that you have installed the content types. That could be done in a single batch, but we'd need a new signal. So I didn't propose that yet. But we can see that like optimizing this function was limited because it was being already fed like smaller size batches Than we could have possibly been doing. So, what are some techniques we can use to apply batching in our own code? The first is to avoid these per instance methods, which are also common when we create these fat models um in Django. So for example, try and avoid writing a send notification function on your user.

12:39

Because if you have any chance that somewhere in your system you need to send a notification to multiple users. You're probably losing out on any efficiency gains. So if we write a standalone function send notification that can send to a number of users, we can instantly take advantage of that batching. For example, Django's mail system can send an individual mail or it can send Many mails at once over the same connection. We probably want to use that if we're sending multiple notifications. This also implies that we should generally prefer larger functions, which is kind of antithetical to the clean code idea. that a function should be as minimal and do one thing because if we are doing several things in a function, we can spot the opportunities for doing things in batch. We can maybe see that oh there's the loop over the users here and the loop over the users there.

13:27

We could combine those into one operation. Or something like that. We should be thinking of efficient code first and then maybe uh clean code later. We've also got a lot of batch ORM methods in Django that beginners tend not to have been introduced to. These are all worth learning about. You've got bulk create, which lets you create many model instances in one query. Bulk update, which does the same with a limited set of fields. So if you're loading a bunch of objects, changing one field on all of them, this is what you want. And then we also have the opportunity now in uh bulk create to use update conflicts to do a bulk create or update operation. So you can get some objects, store them in the database, but if they already existed, only

14:15

Only update certain fields that you know maybe you fetched from a third party. And a general design principle as well is to preemptively pluralize. A lot of systems have you know been created with With the idea that object X is only related to one of the other object Y, and then there's a whole migration process to be like, oh now everything should be related to more than one. So for example, here we have a a company that had a single bust, but then turns out it's Legal for a company to have multiple buses, so we have to migrate over to a many to many field and touch every piece of code that did that. So perhaps we should be preemptively pluralizing all such relations in our system. We can still limit them to one later with a check constraint.

15:00

and uh expand the code a minimal amount later. The final consideration I'd like to talk about is statistical statistical distribution of our data. This is another concrete Django Ticket. Well uh I optimized the root to regex function. That's the ticket number if you want to take a look. So when you write in your URL patterns some path and you use Django's new USH from two point zero uh part syntax. Under the hood, Django converts that to a regex using this function root to regex, which returns the regex string plus this converters dictionary.

15:47

Whilst I was looking into optimizing this, because it was taking a significant amount of time, I came up with a surprising optimization, which was this, to add a LRU cache decorator. Which will cache the results of the function so that when it's cooled with the same arguments again, it will just instantly return that result. I found this surprising, uh but I reached this through looking at the dis distribution of the data. Uh and yes, it made repeat calls a hundred times faster. The way I looked at the data that was flowing through this function was to add a little bit of data logging using uh Python's two modules at exit and pprint. So at exit registers a function that will

16:32

run at the end of the process when it exits. And pprint is a a nice pretty printer. So I I logged into a list every time the function was called the value of the root variable. And then I just printed that list out at the end. And this is what I saw: huge amount of repetition. And it took me a second, but then I thought, that looks a lot like the admin. And indeed, if we look at model admin. get. URLs, it calls path over and over again with the same path uh string because it's on a per model admin basis. So there was repetition in the data that I didn't predict, but it uh maybe if I'd seen This I would have, but it was better to look at the actual values.

17:20

And so that's what the cache really helps with these kind of constructs. They also probably exist in REST framework and other API frameworks that repeatedly use the same suffix URL. So here are some other techniques I think we can apply in our Django code. The first is to collect this data about our data, this metadata, looking at what's flowing through our system. It might be as simple as I just with print or pprint, uh or we can look at database metrics like what is the average number of bosses in a company or things like that. And we also have production logs and APM tools which can provide insight in our production environment. Even use a spreadsheet. Maybe you've got some data that's there and you know how to use Excel well, it's all fine.

18:06

Whatever gets you the insights. Caching, as I showed you, is a powerful technique for when there is repetition, that's likely. And we have many layers in Django to think about caching. We have HTTP caching where things get cached on the user on the web browser. We have Django's cache framework, which is for storing values in a separate cache server next to our server. And then this neat library cache tools has ways of caching inside of Python processes with a timeout, which is particularly useful. When we're caching data from our database, that may change. It might be very unlikely it'll change, but we still want to have it. timeout in the in memory so just in case it does come through eventually something like a list of countries.

18:53

And we'd like to check common conditions early in our code. So That we can avoid doing work. Here's another ticket I worked on where the admin email handler was taking significant percentage of unit test time. And this was formatting emails. And we added this early return block at the top. Because previously it would render the email and then try send it to the list in settings. appens. But that list is empty by default and I'm guessing most projects use sentry or something these days to capture errors, not email. So every logged exception was running through this code, throwing away the rendered email, and taking uh two percent of my friend's test suite

19:38

time. Patching it to early return when there's nothing going to happen, it mak it makes makes a lot of sense. Here are some further resources on this concept of data oriented design. These authors haven't necessarily used the term data oriented, but they're all great. Um take a screenshot and Google Google for these. And thank you. I've been Adam Johnson and this is where you get the slides. And these are my books.

Questions this talk answers

What is data-oriented design, and why should Django developers care about data layout?

Data-oriented design puts data layouts and structures ahead of abstractions such as classes, functions, or layered architecture, because data is what remains directly relevant at runtime. Choosing layouts that processors can read efficiently can produce major performance gains.

Discussed at 3:15

How much faster can column-oriented data structures make Python calculations?

For calculating the mean of 10,000 two-dimensional points, separating the columns made the code about ten times faster, while using a column-oriented library such as NumPy, pandas, or Polars made it about 200 times faster—from roughly 500 microseconds to three microseconds.

Discussed at 4:45

How can Django applications use databases to improve data-processing performance?

Use column-oriented tools or analytical databases such as DuckDB, Parquet, or Snowflake when they fit the workload, and do more computation in SQL instead of pulling data into Django. Models should also use appropriately sized field types so they do not waste storage and memory.

Discussed at 7:55

How can batching speed up Django code?

Batch operations avoid repeated per-object work and database queries. In Django’s permission creation code, replacing individual content-type lookups with `get_for_models()` made the function about twice as fast by fetching the data in one database operation.

Discussed at 11:06

Which Django ORM methods should I use to create or update many objects efficiently?

Use `bulk_create()` to insert many model instances in one query and `bulk_update()` to change selected fields across many objects. Django also supports conflict handling in bulk creation, allowing existing rows to have specified fields updated.

Discussed at 13:27

How can I find useful caching opportunities in a Django application?

Inspect the actual values flowing through a function using logging or simple data collection, because unexpected repetition may not be obvious from the code. For example, repeated arguments to Django’s `route_to_regex()` function made an LRU cache effective, speeding repeat calls by about 100 times.

Discussed at 16:32

Why should Django code check for no-op conditions early?

An early return prevents expensive work when the result will be discarded anyway. Adam’s example avoided rendering admin error emails when the recipient list was empty, reducing unnecessary test-suite work.

Discussed at 18:53

Presenters

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Adam Johnson

More videos from DjangoCon Europe