Data-Oriented Django

This video features Adam Johnson at DjangoCon Europe 2022 in Porto, Portugal.

Data-Oriented Django
0:34:58
Published October 17, 2022
3,246 views

Data-Oriented Django by Adam Johnson

Data-Oriented Design focuses on how software transforms specific inputs to specific outputs, on specific hardware. This talk will cover how this way of thinking can inform writing Django code.

Summary

Adam Johnson explains data-oriented design as choosing algorithms and data structures according to the characteristics of the data, the required output, and the hardware running the software. He argues that modern CPUs are much faster than memory and network connections, so performance depends on using compact representations, arranging data for efficient access, and reducing unnecessary movement—especially across the database boundary. Applied to Django, this means optimizing the request/response path, using caching, compression and CDNs, avoiding N+1 queries with `select_related`, `prefetch_related` or Django Auto Prefetch, splitting rarely used fields into related models, and combining database aggregates where possible.

Key takeaways

  • Data-oriented design starts by examining input and output data, including its volume, latency, distribution, and acceptable accuracy.
  • Modern CPUs are often stalled waiting for memory or network data, so smaller data representations and cache-friendly layouts can improve performance.
  • Python objects use substantially more memory than packed numerical representations, making NumPy, Pandas, or other compact structures useful for numerical data.
  • Web performance can be improved with minimal HTML, HTTP caching, HTTP/2 or HTTP/3, compression, minification, and CDNs.
  • Django applications should avoid N+1 queries, choose between `select_related` and `prefetch_related` based on data shape, and use tools such as Django Debug Toolbar to inspect actual queries.
  • Keeping infrequently used fields in separate one-to-one models and combining filtered aggregates can reduce database and memory work.

Summarised automatically from the transcript.

Chapters

  1. 0:00 Introduction to Data-Oriented Design Adam Johnson introduces data-oriented design through the gap between modern computer speed and slow Django applications.
  2. 2:12 The Data Transformation Model The talk defines data-oriented design as transforming input data into output data and emphasizes the importance of context.
  3. 4:00 Data-Dependent Algorithms A number-matching example shows how data volume, distribution, accuracy requirements, and access patterns should shape implementation choices.
  4. 7:11 Modern CPU Architecture The discussion turns to general-purpose CPUs, memory latency, and the hardware assumptions behind software performance.
  5. 10:19 Cache-Aware Python Data The talk explains cache-friendly data layout and compares Python objects with packed arrays, NumPy, and Pandas.
  6. 13:35 Data-Oriented Web Architecture The speaker maps user interactions, browsers, Django, databases, and HTML into a data transformation pipeline.
  7. 17:26 Web Request Performance This chapter covers response latency targets, efficient HTML, HTTP caching, compression, HTTP/2 and HTTP/3, and CDNs.
  8. 20:31 Database Performance Costs The talk examines the high cost of database round trips and introduces strategies for reducing the amount of data transferred and processed.
  9. 22:54 Django Query Optimization Adam Johnson explains the N+1 query problem and compares select_related, prefetch_related, and Django Auto Prefetch.
  10. 28:25 Model Decomposition and Aggregate Queries The speaker shows how splitting models and combining filtered aggregates can reduce database and memory overhead.
  11. 33:02 Further Resources and Takeaways The talk concludes with resources on data-oriented design and a final example of large performance gains from context-specific optimization.

Transcript

5,436 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:00

Friends, Pythonistas and Django Nauts, lend me your brains and let me add a mental model on insight about programming called data-oriented design and how we can apply it to Django Computers are friggin' fast, okay? They're freaking fast. It's something you have to realize. This laptop here, each core can do 3. 2 billion operations per second. And it has 10 cores. 32 billion operations per second total. I tried to compare with Grug-Brained Adam. This computer 64-bit. I generated two random 64-bit numbers and tried to add them together. That counts as one operation.

0:46

It took me 55 seconds. I came up with this, and it was the wrong answer. One of the sixes should be a five. So if we compare the computer with me at maths, I can do. 018 operations per second at 0% accuracy. And the computer can do thirty-two billion operations per second at one hundred percent accuracy. Where does all that speed go? My compu my Django view, I think it's doing hundreds of things, and yet it should be instant and it's taking seconds. What's happening? This is what data-oriented design tries to address.

1:32

But first, what is not data-oriented design? It's not domain-driven design, which we heard about earlier, and uh may not be a piece of crap, I don't know. It's not data-driven design, which is a term I've heard for building layered architectures, and it's not data-oriented programming. Just to confuse you. There's a book out there. That's a kind of adjacent term, but it's not exactly that. So what is data-oriented design? First, we need to adopt Shoshin, the Zambrism term for beginner's mind. Beginner's mind tells you that in any line of inquiry you should try and empty your mind as if you're a beginner, no matter how much you think you

2:22

know already. So, let us return to my high school computing class. I was shown this diagram probably in the first lesson. What is software? Software takes input data, it runs on some hardware to transform that input data and turns it into output data. That's data oriented design, thank you. That's the insight. Software's only job is to transform data. Anything that we add to software, any kind of Thinking about architectures, do I use this technology or that? It's all in honor of this one goal and it's all secondary to this goal

3:12

Users of your software only care about getting that output data and in a timely fashion, accurately, whatever they care about that output data, but they only care about getting it. And whatever you do in the middle As long as it gets the output data, it's it's kind of okay. It's kind of fine. So that leads us to ask what are the characteristics of the data that we're transforming? And Every single problem is slightly different. The data may be in different formats, it may come at a different volume, it may be loads and loads of data at once, or maybe a tiny amount. It may come with different latency characteristics. Your input data might be coming sporadically or in very many chunks very quickly, and then users will only care about their output data with a certain latency as well.

4:00

And that can give you your throughput. And then we can also ask questions about what is the statistical distribution of the data, because that can influence the way we program. Context is everything. If you have different input data and different hardware, that should lead you to build a different algorithm to get different output data Let's look at an example so you can maybe see a bit about what I'm talking about. So you've been given the problem to check if a number exists in a given set of numbers. Perhaps your Python brain is already going, oh, you can use the in operator with a set. So you might come up with this function. There's some set of numbers that's stored already, and then we implement our function

4:49

is this number a match? We just say return number and numbers. Seems pretty reasonable What if there was just a single number in that set? Building a set would be to contain a single number and using the n operator would be slightly wasteful, so we could just use the double equals operator. What if the set of numbers we're checking against is a billion? Putting a billion numbers in memory in a Python set is probably going to exhaust the memory, maybe even on this computer. So uh that's probably not a good idea. So you might end up putting in a database. SQLite would be fairly reasonable. So you could do it like this What though if we were matching many numbers against a few?

5:36

Many functions in our programs aren't called just once. They're called repeatedly with the data that we're matching against So it might be more reasonable to write our function to take in a whole set of numbers to match at once and then do that more efficiently. So in Python you could use the intersection operator. We can take in a set of numbers to search for against another set. Well, that's also built into the set class for us. We could go on. What if the set is all numbers, all odd numbers between one and a million? Well, we could write just a bit of maths instead of storing all the numbers. What if false positives were acceptable? In that case, a Bloom filter might be a better data structure.

6:23

Bloom filters are cool, I don't have time to talk about them here, but do look them up. And so on and so forth. But the important thing here is that the data that we're using as inputs to our function will influence what function we write and also the context because if this function is maybe being called in a loop and we can do it efficiently in one pass that's also better Yeah, so the implementation that we come up with depends on the characteristics of the data we're working with. That is data-oriented design. The other side to it is that our job of the job of software is to transform data But only using specific hardware. When we write a program, we're only going to run that on certain computers.

7:11

The world of computing is infinite. We're targeting rather rather specific computers whenever we write a program. That makes us ask, okay, what is the hardware we're writing for? Are we writing our code to run on a TI-85 calculator? Probably not. At least not in this room very often. Are we writing something to run on a quantum computer? Probably not here. What about the Turing tumble? This is a marble-based computer. It's a huge amount of fun. But probably not. It only supports four-bit numbers. Most of the time, everyone in this here in this room is writing software to be running on the general purpose

7:57

CPUs of the day. That's what's in my laptop. That's what's in servers that I deploy to, that's what's in my phone. Um there's other chips in there as well. But our Python code, our Django stuff is mostly running on general purpose CPUs Well that leads us to ask, what are the characteristics of modern general-purpose CPUs? Here's a picture. There you go. All of our CPUs are built on something called the von Neumann architecture, which is basically the CPU executes instructions one at a time. It fetches the the instruction and potentially data from memory, runs the instruction, and maybe writes something back to memory depending on the instruction.

8:42

So far so good. But this diagram is a vast oversimplification. This is the history of general process general purpose CPUs since 1980. If we peg the relative performance of the processor and the memory transfer at one in 1980 Fast forward to the modern day, and there's a gap of uh around uh uh a thousandfold difference in speeds. So your processor can do a heck of a lot, but it can't get the stuff from memory. This is a diagram I got from a book whose name I forget, but it will be on the blog post So CPU manufacturers introduced a cache between the CPU and the memory.

9:32

So the data is transferred into the cache, which is You know, way faster, but way smaller. So the working data can be fetched and stored from the CPU rather quickly and then when it needs to make its way back to memory. It which um it does, but hopefully that's less often than it's used because most programs they focus on a small amount of data at a time and then they are done with it. That is a vast oversimplification The modern sleep use have three levels of cache, sometimes four. Each layer of cache is about three times faster than the next one, but about three times smaller, give or take. So when you're running something on your computer and you know you're writing your Django view with levels of caching and you're thinking, oh this is a bit hard.

10:19

Your computer's doing it way harder. Uh yeah, so for this computer, uh I looked at the stats. I think these are about right. Uh a fetch from the level one cache takes three cycles. That's three of those operations per second. But going to the level two cache takes ten cycles. Going to the level three cache takes thirty cycles, and going to memory takes at least a hundred cycles. So yeah, we have to work with this reality. This is the hardware we're writing for. What are the implications of this then? Well, put simply, use smaller representations. The smaller it is, the more of it it can fit you can fit into that bottom-level cache, and the faster you can go.

11:11

And then the second implication is to lay out the data in access order. One extra thing the caches is doing is it's pre-fetching data. So you access one piece of memory, it assumes you access the next one. So if you're going through memory in a line, it's actually gonna be almost as fast as going through cache all the time. That leads us to ask about Python. Here's just a quick example of how Python is kind of wasteful. This is a sixty-four-bit computer, so we divide that by eight, that's how many bytes there are in a full s full width integer. But if we import Python's sys module, we can use sys. getsizeov to return the size in bytes of here, the number

11:56

9001. So that's an integer 9001 and it's 28 bytes. What's going on here is that Python has the number, but it also has 20 bytes of other stuff about the number. So it has an object ID. Everything in Python is an object. So the overhead here, you know, it 's it's nearly four times the size of the actual number. Um that's the reality. If we build a list of numbers in Python and we say, okay, it there's a thousand numbers here, multiply by that by eight, that's how many bytes of memory it should take up. It's eight thousand bytes. But if we add together the size of the list, which is the container, and then the size of all the numbers within it, it's 36,052 bytes.

12:47

So again, it's about four times the size bit of wastage. But Python has a solution. You can build uh this object called an array that packs all the numbers together and treats it as one line of memory. And if we build an array of the first uh thousand numbers, then we end up with 8,320 bytes. There's a little bit of overhead, but it's it's nowhere near as much. So it's a bit more reasonable. That said, the array type isn't isn't that useful because it doesn't have many operations. So if you are using arrays of numerical data, much better to use NumPy or Pandas which do the same kind of packed format, but they also have a lot of routines for for transforming the data in an efficient manner.

13:35

Okay. So now we can ask about applying the data-oriented design model to the web. What is the input data, the output data, how are we gonna work with it? Here's my approximate diagram of any website ever. We've got the user and they send clicks or indeed any other interaction like keyboards or perhaps they're using voice control to the browser. The browser sends requests to our something, which we'll define later. And then we need a database because users of most websites expect their data to stay around in between interactions.

14:22

And then it goes back from the database through something that needs to be transformed into HTML for the browser. And then the browser turns the HTML into pixels on the screen and the loop can continue. Note that we're sending HTML to the browser. That's the only thing that browsers accept for rendering stuff on the screen. minus a few small JavaScript APIs where you can directly put pixels on the screen. But 99. 9 % of the time we're using HTML to tell the browser what to render. It's up to us as web developers to choose what we to put in that something cloud. Just like to note, I'm not gonna put the user-browser interaction here. We don't have control

15:07

um of this loop and it's also already hella fast. Like browser super optimized written in C and Rust and whatever and um a well-run web page, there's no noticeable lag for the user. So I'm gonna skip that on the future diagrams. So yeah, we're just gonna look at this. What are we gonna put in the something? Well, one option is to put Django in the middle and then use uh Postgres as our database. So the browser sends a request to Django, Django turns that into one or more queries to the database, getting results in order to construct HTML, give it to the browser. Job done. You could also build an architecture like this, which is not unrealistic in the modern uh

15:54

era. Your first request, you send a bundle of JavaScript to the browser. Then the JavaScript, which we could think of as a program separate to the browser, just running in a sandbox Makes some more requests to an API gateway. The API gateway forwards the requests to service one, two, or three, each of which has their own database. A lot of work going on in order to get the gateway back, and then the JavaScript eventually constructs all the results and turns it into HTML Bit of a mouthful. One question. How fast does our data transfer need to be? How fast do we need to get the HTML to the browser? This is uh the key data characteristic for performance. If we can do it in less than 100 milliseconds, you get a double OK sign.

16:39

Very good. Less than one second. That's pretty good. Um yeah, I put one hundred milliseconds because that's the threshold of perception. So if you can update the screen in less than a hundred milliseconds it'll feel instant to the user. Um anything around that as well is pretty good. If you get to about three seconds, I saw something that, you know You'll lose fifty percent of visitors on mobile devices um who are doing a first page load if it takes more than three seconds. So you're already a bit on the edge there. If you get past 10 seconds, I'm sure you've tried to load a web page, waited 10 seconds, and you just close the tab, you hit refresh, or you're gone. So that's our that's our kind of limit here.

17:26

It's not much actually. Still if we're using like C and stuff, you've still got a lot of uh network interactions between you and uh the browser, so that doesn't leave much time for a perfect experience. Alright, we're gonna go look at the two sides of the diagram now. We're gonna first dive into how we can improve the flow from requests to Django and then Django to HTML back to the browser. Speeding up the request for response cycle. There's quite a lot on this topic, and so I've just put a few bullet points on this slide, and I'll refer you to other resources The first thing we can do is write minimal performance HTML. There are good ways to write HTML

18:13

if you get your style tags and your first JavaScript tags. In the first kilobyte, the browser will start requesting those early on, uh even whilst it's downloading the other HTML. Um don't bloat your HTML with loads and loads of classes. That's some CSS frameworks like to add, etc. Um Introduce HTTP caching. You can tell the browser which things it can fetch once and then store forever, so you can have your images loaded once only. Very good to add. Get on to the latest version of HTTP. There's HTTP3 now and there's HTTP2 way more widely available. Make sure you're using at least the HTTP2. Look into response compression.

19:00

Django has gzip middleware built in, which will gzip your requests. That's a compression algorithm. For text formats like HTML, you'll see like pretty good savings. There's also an alternative called Brottly that's available only on HTTP2, or is it only on HTTPS? One of them. But it could grow that. Um it's slightly better than GZip. Um you can also look at the HTML modification. This is when you're really trying to save the bytes Um I found a Rust-based HTML minifier that um it would save you about one to three percent of your request size. Still something, and I made a middleware in the package Django minify HTML.

19:45

And I think maybe the most important is to deliver with a CDN, because if you're putting your content through a content delivery network, that's a bunch of servers around the world, and the data for your website will go to and from their servers which are closer to users, saving a lot of network time. That's the primary time that we can affect on this request response cycle. I'm going to refer you to some resources here. If you look at MDN, the Mozilla Developer Network, they've got a page called Web Performance. That's a very good section in the docs. Google Chrome team have web. devs, and so there's web dev learn and then I've got two measurement tools, one is web devmeasure and one is web page test.

20:31

You can run either of these tools for freely on your website and you can see what's taking time in a request-response cycle What is the browser having to do from what you've told it in your HTML, how long a connection is taking, you can time them from other side of the world, etc. Um, so you can really dig in and optimize that part of the flow Alright, time to look at the other side, somewhere where our advice gets a little bit more specific to Jenkett. We're making queries to Postgres and Postgres is returning us results. And we're probably doing that a number of times for every HTTP request we get in because we have to assemble different pieces of data on the page.

21:23

So yeah, that's what we're looking at speeding up. Returning to the CPU memory diagram, you might remember it took a hundred cycles to get something from memory. Well, unfortunately the story is pretty bad when talking to a database over a network. It's gonna take 1. 6 million cycles at least to get there. This is assuming a half a millisecond round trip time to Postgres and a pretty much an instant query. If the query starts taking a few milliseconds, this will go up by a factor of ten, you know So that's kind of the answer to the question, where does all the speed go? And anything we can do to reduce the number of times we do this trip, kind of the better.

22:08

So let's look at three ways to speed up queries and results. The first one: avoiding the n plus one query problem. This is a classic. I'm going to dig into it. The second one is to look at splitting your models into smaller logical units. So both less data is fetched over the network, less data is moved around in memory, and less data is also in each in individual table in Postgres so it has to do less work to look through your tables. And and then I'll show you kind of a slightly more advanced technique uh about how you can batch counts into one pass. I hope this This set of things is very small to cover on database optimization, but I hope it inspires you to go look at some other things that I'll point you to.

22:54

Alright, what is the N plus 1 query problem? Who's encountered the N plus 1 query problem? Okay, so reasonable amount of experience. I'll go over it again. Hopefully you learn a little bit of something. Look at this this code. It's it's looking at some books and then it's printing them out by the book name and the name of the author. How many queries does this do? It does n plus 1, if you might have guessed that. The first query comes when we iterate the book's query set. That it turns into a select query for give me all the books database and it gets them in one query. And then within our loop For each of them, when we touch the property, the attribute

23:43

book. author, the day the Django goes, oh, I don't have the author, I didn't fetch that. I'll go fetch it for you But it does that on every single iteration of the loop. How can we fix this? Here's one tool. It's called Select Related. So here we do a single query and we ask for the database, give me all the books, along with all the author data. What happens here? We get a single query where all the book data, author data is joined in. When receiving that data, Django splits it up again into in-memory model objects. So all of the data for the books goes into the individual book objects, all the data for the authors goes into separate author objects that are linked.

24:36

So we've done it. M plus one queries has become one query. This will go way faster. We avoid that back and forth to the database many times. There's a little bit of a drawback to select related and this table hopefully demonstrates that. SQL databases only talk in tables. So when you do a select-related query, it's coming back as a single table of data. We've got all our books. And they uh I've put their names down the left column here. There'll be other columns. Um and then we've got all of our authors joined in. And they're on the right hand side. I've put only the auth the the name column as well, but you can imagine there'd be other columns. The problem

25:21

we've got here is extreme repetition because one author was incredibly prolific. Imagine if we had way more rows. Alpha Cronen Doyle is appearing many, many times in the dataset. At some point, uh the volume of data from passing over the ek the authors repeatedly, the same author over and over again is going to outweigh the savings of using select related. It kind of depends on your data, but you do need to think about it, measure. For this, Django provides uh to solve this problem Django provides a different tool called prefetch related. Prefetch related says when you do this query, also prefetch and do a second query. So we'll get two queries out of this.

26:07

The first query fetches the books when we start iterating on them, and then as soon as Django's done that fetch, it sees that there's also a pre-fetch to do And so it looks through the list of books and it fetches the authors for the books it fetched. So if there were five authors for a thousand books, it's going to do a query that pulls back just the five authors once. No repetition in the transfer data. This can be faster than using select related. It um it does depend, as I say, on the data volumes. You have to do a bit of measurement or guessing. I think prefetch related is generally a better default to reach for than select related because Yes, it does one extra query, but it's uh

26:53

it's limited still. It's gonna be two instead of n With that, I'd like to talk to you a little bit about a project I've been maintaining for a while that my old boss wrote called Django Auto Prefetch. If you swap your models and query sets to the versions provided by Jenkit Auto Prefetch, it has like a slight patch in the ORM. to change Django's behavior. So remember that the N plus one problem was caused by us touching the book. author attribute and Django fetching one of them? Django auto prefetch changes it, changes this logic and says I see you access the author of a single book, but that book was fetched as a larger group, as part of a larger group.

27:38

I imagine you're going to do it again on the second iteration of this loop, probably. So it just fetches the authors for all the books in one prefetch. So we end up automatically with the two queries done rather than N. Um so Django Auto Proof X useful project. It solves your N plus one query problems for foreign keys and one-to-one fields. It doesn't help you with many-to-many is where you're doing a new whole query set within the loop. I don't know if there's a way to do that. But yeah, check it out. And yep. Alright. Second uh technique for um uh uh reducing your data volumes and make things more performant is splitting models.

28:25

Imagine we have a user model It's inheriting from Django's abstract user, so it's pulling in a bunch of fields like username, email, password. And we've also got our own custom fields here like the user's avatar. And then we're given the task We need to store users ACME access tokens and refresh tokens. We're talking to the ACME API and we're going to get access tokens and refresh tokens from their open and from their OAuth thing. We need to store them so we can keep making requests. What you might feel like doing is adding a few fields on user. These are the ACME things for the user, so they should go on user, right? This is not great. And the problem

29:10

is that it's going to slow down every place where users are queried. Every single query set of a user is gonna automatically be fetching those fields from now on. Django's gonna have to do work to deserialize them. They're gonna take up space inside Postgres next to all our other data. One thing you might not realize about Postgres, every time you update a row, it's creating a whole new copy. So everything you add in a single row is having to do a copy of that data every time. So what's an alternative here? Well, we can make a second model. We can just create a model called the UserACME token. related one-to-one with the user model. And on that we'll just put the three fields we need and we'll use the user as the primary key for this model.

29:58

So every user access, user ACME token is related to exactly one user. And if we don't have uh ACME tokens for a given user, that model is simply not going to exist for that user. So we can still determine that easily. This is much better, it gets three thumbs up. Alright, now onto the third slightly more advanced technique. We can do multiple counts in a single pass. Have you ever written code maybe for a dashboard like this where you want to show, say, the count of publisher uh verified books for an author? I called it published in the variable. And the uh un the unverified count as well.

30:44

So maybe two big numbers on the dashboard. These are doing two separate queries right now. Instead, what we can do is use a little bit more advanced feature in the RM and uh send a single query to the database that asks it to count both parts at the same time. And what Postgres is likely to do here is make a single pass of the data, counting each kind separately on the way. So this saves a whole round trip to the database, it saves a whole passover of the data, and you know if an author had many books, we might see a fairly reasonable performance improvement here. So I included this as a slightly more advanced technique that you might not be aware of, which is this filter argument

31:31

on all aggregates. You can do this not just with count, but with sum and average, etc. We have some fantastic resources for optimizing your use of the database, and it's a key thing in data-oriented design. Limit the amount of data you're touching. The first is in the documentation. There's this page, database access optimization. It's a litany of all the things you can do inside Django to improve your queries. Everyone should have a read of list until they know every single part of it. I'd also refer to this book, The Temple of Django Database Performance. This is a very good first

32:16

pass over things like indexes and joins and what the database is doing and features that you can take advantage of. And it's Postgres specific, but you might find most of the things apply to other databases as well. If you want to look a bit more at the N plus one queries problem, I wrote a blog post uh a few years ago Sponsored by Scout APM, just search for that and you'll find it. And then it's very important for us to see the queries that are actually happening in our code as we're using it So I'd recommend you use a tool like Django Debug Toolbar, which is universal, or indeed our sponsors, Colo, if you use VS Code. So, data-oriented

33:02

design. Remember, software's only job is to transform data. And users only care about getting their output data in a timely manner. Here are some resources where I learned about data-oriented design. It's a term that's come out of games programming where every cycle counts and rearranging things in memory really helps. I think the original talk I could find was by this person, Mike Acton, who now works on the Unity game engine, and he gave this talk, Data Oriented Design and C<unk>. It says C<unk>, it's not so uh C<unk> focused, I think. Um

33:48

The thing that got me onto this term was Andrew Kelly, who's creating a new programming language called Zig. And his talk, Practical Data Oriented Design, shows how he applied it inside the Zig compiler for some great speedups. And then this final talk here, Andreas Frederickson, talking about context is everything. This is a very good optimization talk. He wrote something that parses JSON in C<unk> and then continually optimized it. At the start, it could parse the JSON and get the output answer at a rate of about 40 megabytes per second, which is quite a lot of speed for a single call. But as he changed it to use a specialized JSON parser and so on, he managed to get a single core

34:35

speed of about 1250 megabytes per second. So like a 30 times speed up way faster than anything we could ever write in Python. Um but yeah, very impressive. Computers are friggin' fast. All right, thank you. I've been out of

Questions this talk answers

What is data-oriented design?

It treats software primarily as a transformation from input data to output data. The characteristics of the data—such as its format, volume, latency, and distribution—should guide the implementation and architecture.

Discussed at 2:22

How do CPU caches affect software performance, and how should data be laid out?

Modern CPUs are much faster than main memory, so cache misses are expensive. Use compact data representations and arrange data in the order it will be accessed so more of it fits in cache and hardware prefetching can help.

Discussed at 10:19

How can I speed up the request-response cycle of a Django website?

Use minimal HTML, browser caching, HTTP/2 or HTTP/3, response compression, HTML minification where worthwhile, and a CDN. Tools such as WebPageTest and web.dev can show which parts of the cycle are taking time.

Discussed at 18:13

What is the N+1 query problem in Django, and how do I fix it?

It happens when Django fetches a collection in one query and then performs an additional query for a related object during every loop iteration. Use `select_related()` for a joined query, or `prefetch_related()` to fetch related objects in a separate bounded query; Django Auto Prefetch can automate this for some relationships.

Discussed at 22:54

Why should I split rarely used fields into a separate Django model?

Putting fields such as third-party OAuth tokens directly on a frequently queried user model makes every user query fetch and deserialize them, and increases row-update and storage costs. A one-to-one token model keeps the main model smaller and only creates token data for users who need it.

Discussed at 28:25

How can I calculate multiple conditional counts in one Django database query?

Use the `filter` argument on aggregate expressions such as `Count` to request several conditional counts together. The database can then make one pass over the data and avoid an extra network round trip.

Discussed at 30:44

Presenters

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Adam Johnson

More videos from DjangoCon Europe