Cutting latency in half: What actually worked—and what didn’t with Timothy Mccurrach

This video features Timothy McCurrach at DjangoCon US 2025 in Chicago, Illinois, USA.

Cutting latency in half: What actually worked—and what didn’t with Timothy Mccurrach
0:44:42
Published October 23, 2025
223 views

This talk was presented at: https://2025.djangocon.us/talks/cutting-latency-in-half-what-actually-worked-and-what-didnt/

LINKS:
Follow Timothy Mccurrach 👇

Follow DjangoCon US 👇
https://fosstodon.org/@djangocon
https://x.com/djangocon

Follow DEFNA 👇
https://www.defna.org/

Video production by the presenter and DjangoCon US 2025 volunteers.

Summary

Timothy McCurrach argues that performance work should begin with proactive, regular profiling rather than waiting for users to report problems or relying on intuition. Broad production measurements reveal patterns that isolated deep dives miss, while local profiling helps test changes; together they guide trade-offs and led his team to cut average Django template response time from 395 ms to 218 ms. He explains practical ORM techniques such as avoiding N+1 queries, choosing between `select_related` and `prefetch_related` based on the resulting data shape, and inspecting SQL and query plans. He also shows that the biggest gains may come from questioning requirements, removing obsolete work, changing pagination, fixing data, and optimizing code shared by every request—not merely adding indexes or tweaking queries.

Key takeaways

  • Profile production broadly and regularly to understand which pages are slow, volatile, poorly scaled, or widely used.
  • Use averages, percentiles, totals, and local experiments according to the performance question you are trying to answer.
  • Choose `select_related` or `prefetch_related` by considering the size and shape of the joined data, not only the relationship type.
  • Treat N+1 queries, lazy querysets, query logging, SQL inspection, and `EXPLAIN` as practical ORM debugging tools.
  • Before optimizing code, confirm that the feature is still needed and that its current behavior matches user requirements.
  • Optimize the hot path shared by every request, including middleware, context processors, templates, and cache access patterns.

Summarised automatically from the transcript.

Chapters

  1. 0:00 Introduction Timothy McCurrach introduces his performance work and frames the talk as a personal journey.
  2. 1:51 Reactive Profiling The talk examines the limits of “profile first” advice and why teams often investigate performance only after problems appear.
  3. 6:32 Proactive Profiling McCurrach defines proactive profiling as measuring broadly to understand performance patterns before trying to fix a specific issue.
  4. 10:24 Performance Patterns The talk shows how to use slowest-page lists, percentile metrics, and production data to identify recurring performance trends.
  5. 11:55 Benefits of Measurement Systematic profiling reveals performance topology, informs engineering trade-offs, and produces substantial improvements.
  6. 18:08 Deep-Dive Techniques McCurrach covers visualization, repeated samples, user variation, and the complementary roles of local and production profiling.
  7. 21:12 Django ORM Optimization The talk explains query latency, lazy querysets, N+1 queries, and when to use select_related versus prefetch_related.
  8. 30:35 ORM Debugging Tools Useful techniques include query logging, SQL inspection, explain plans, cache inspection, and regression tests.
  9. 32:52 Low-Cost Performance Wins Several case studies show how clarifying requirements and changing functionality can outperform complex ORM or database tuning.
  10. 39:04 The Hot Path The talk focuses on optimizing code shared by every request, including lazy context processors and more effective caching.
  11. 42:15 Conclusion McCurrach summarizes the unexpected sources of performance gains and notes that frontend and network costs matter too.

Transcript

7,259 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:16

Speaker 1: Thank you everyone. So my name's Tim. That's me. I chose that particular picture. It's nearly six years old today because it was taken at DjangoCon US. 2019. I was a very young developer there. And DjangoCon US was kind of a real catalyst in terms of my interest and understanding in Django. And it gave me the confidence that It was something I could contribute towards and it was a community I could be a part of. So it's quite special to be back at DjangoCon US. speaking this time. So I'd like to start by saying thank you, thank you to the organisers for giving me this opportunity and just for organizing what's been a fantastic conference so far.

1:01

Speaker 1: So I'm a developer. I work for a company called Unojuno. We're a freelance management platform. So if you are a freelancer or if you hire freelancers, please do check us out And I've been working there for a while and I've s I sort of got the reputation as the the performance guy. And as a result of that, um over the past year, last past nine months especially, I've been given uh lots of performance work. And over that time, my um ideas about kind of what is performance work is has shifted quite a lot, I feel. which is the kind of inspiration um for this talk. So um it won't be comprehensive. I think it'd be pretty difficult to do a performance talk that's comprehensive anyway in 45 minutes.

1:51

Speaker 1: It's more a, I guess, personal journey revelation. But I hope it will be useful for you. So performance. If you have been to a talk on performance before or read an article or anything like that, the first thing you will see is profile first. And there'll probably be a a quote from Donald Knuth, hope I'm pronouncing that correctly, about premature optimization being the root of all evil. Arguably a misquote taken in isolation like that. And just so we're all on the same page, by profiling, I'm talking about measuring things. There's lots of things you can measure. I'm going to focus mainly on kind of time spent in this talk.

2:36

Speaker 1: But we're told to profile first, and I think we can all agree that's good advice. I mean, for one thing, it's just good common sense, right? There's no point spending hours and hours squeezing every last drop of performance um out of something if in fact the code it took you 30 seconds to write is already fast enough um or maybe kind of really fast for for your use case. But supposing you do realize that it's not fast enough, well you need to know what to optimize. You need to know what are the slow bits, what are the bottlenecks. And sure, you might have some intuition about that, but our intuition is often wrong. And so profiling is just sensible to make sure you're not wasting time

3:23

Speaker 1: optimizing the wrong bits. Also, you want to know if you've made progress, right? If you're doing lots of work trying to make something faster, it's nice to have a before and an after so that you can tell if you've done a good job or not. So yeah, profile first is good advice. But there's a but, and it's quite a big but, and that's that good advice is only really useful if you follow it Um I'm perhaps a bit like Alice in Wonderland who very seldom follows it. Uh I don't know about you. I think um if I look back at lots of the work I've done, um profile first has turned into profile hardly ever.

4:09

Speaker 1: Or maybe profile as a last resort when all the other things I've already tried don't work. Even in performance work, I like to think, oh yeah, I can look at that code and I can spot the N plus one issue here. I can notice there's an index missing there. I suspect I'm not alone in this. I think profiling probably fits into that category of things that everybody talks about, but far fewer people are actually doing. And of that subset, probably a lot of those aren't doing it right. I went online. There are a few companies that kind of do surveys on how developers use um APM tools. And one kind of common thread that came up in kind of lots of the things that I was reading was that we tend to use them reactively.

5:00

Speaker 1: So this survey said that the prime driver for adopting APM is to fix problems. And another survey said organizations are still more likely to be reactive rather than systematic. And you know, in some sense this this might be reasonable, but I think we can do better for a few reasons. For one thing, it's not the best look to have your performance issues being raised by your users. Ideally, you know, you want to address things before they become an issue. rather than scrambling to fix them after it's a real problem. You want to to optimize things for lots of reasons. It might not just be kind of user experience.

5:47

Speaker 1: It might be things like how much resources you're using or environmental reasons. And that reactive style of profiling isn't going to cut the mustard there. Likewise, SEO, so one of the big reasons I did lots of optimization work is our page load times were affecting our SEO results. And without kind of some systematic profiling, that thing's never going to be flagged. So there's some good reasons why we can do better than reactive profiling. But there's a more important reason, and I think it's that the alternative is just so much better, which is what I'm about to talk about. So I don't know if you've ever tracked something

6:32

Speaker 1: like a bad habit. Things like screen time and social media usage are common ones Um for me I drink a lot of tea um and I spend a lot of money on tea and if I'm honest cakes and pastries too. Um and I knew this to be the case um and I had a vague sense of you know I was spending this much money on it. But um this this by the way over there is the the coffee shop where you'll find me virtually every morning programming away. Yeah, so um I I thought I had a a good idea of how much I was spending, uh but a couple of years back, for a month, I recorded all of my coffee shop expenditure. And the results were horrifying.

7:20

Speaker 1: My experience when I started profiling in a more systematic way on the site that I'm working on were very similar. I was just completely shocked at um what I thought I knew being completely wrong. So bear in mind this is a company I've been working for for about five years. Um I thought I had a pretty good idea of where our slow pages were, what our issues were. Um and there were just so many things I had no idea about. So what is proactive profiling? Proactive profiling, by the way, is just a term I've come up with, it's nothing official. I'm just using it for the purpose of this talk. Profiling to understand rather than to fix.

8:06

Speaker 1: It needs to be broad, so rather than just diving deep into the particular view or the particular endpoint that's not performing well, you're taking a broad look at some very basic statistics so you get an idea of the distribution of what's fast and what isn't. And then you combine that with the drill downs. So once you know which bits are slow, then you know which bits to focus on Ideally, this should be on production data because it's not just about what are the slow views, it's about um what are people using, right? Um if you're trying to improve your environmental footprint There's no point identifying all the resource intensive endpoints if in fact nobody's using them.

8:53

Speaker 1: Local profiling is still useful, and I'll talk a little bit more about this later. So there's lots of third-party services that you can use to do this. Sentry and Uelic Data Dog. I was going to kind of compare them, but kind of looking into them a bit, I found they all offered mainly the same features. And I think what's more important is having fluency in what whichever APM tools you're using rather than the particular tool. I'm also aware some of these tools are quite expensive and not all companies can afford them. So some alternatives. are to use open source libraries or just roll your own to gather in the low fidelity stats. It's easy to write some middleware that captures how long

9:38

Speaker 1: each request is taking. And then maybe combining that with something like Django debug toolbar, which you can use on production for your admin users. And that way you can get the broad spectrum as well as the drill downs on production. The third-party services offer a lot of convenience, but that there are ways around it. And finally, profiling regularly. So I don't know how regularly is regular, but I I suddenly found doing this um repeatedly um gave me more benefit as as time went on. So what does it actually look like? Well it's it's it's very simple. What I tend to do is I tend to um find what are the ten slowest pages um on my site

10:24

Speaker 1: by average. If I were to do that right now, I could tell you straight away six of them will be from the Django admin page. And I can also tell you straight away that there won't be very many hits on those pages and they're only accessed by devs anyway. So then I would exclude all my admin pages from my results. And then I would look at the the top ten again. And then maybe I'd look into them and I notice lots of them are list views. And I see actually all of these views have a thing in common, and that's they're all spending lots of time populating drop-down filters. And they're all spending lots of time on the pagination component at the bottom, say. And so you begin to develop these kind of broad patterns.

11:09

Speaker 1: about what areas of your code base are going to give you kind of the most bang for your buck. And depending upon what it is you're trying to do, you might want to look at different things. So um one problem we had is that actually our site was pretty fine for most users, um, but we had a small proportion of users that were really heavy users and had lots of data. And actually it was that our views weren't scaling properly. Looking at the P95, so that's the top 5% slowest pages, actually a much more useful metric. than than the average for that, because it's identifying where do things not scale. Likewise, you know, if you're wanting to profile to improve your environmental footprint, you're probably more

11:55

Speaker 1: interested in that total column. You're actually probably more interested in kind of not time at all, but some other metric. But it would be the it'd be the total amount that that you're interested in. So, you know, it's it's nothing mind-blowing, but I have found it to be um just a massive game changer in a way that I I couldn't imagine. Not just in terms of performance work, but it's affected, I think, my whole programming. So I'll try and kind of convince you why I think everyone should be doing this. So, for one thing, you notice the kind of broad trends, what I'm going to call performance topology. And It's not just kind of, oh, we're we're using paginators that are slow.

12:42

Speaker 1: It changes the way you view uh kind of performance. I was I was forced to stop thinking about views as fast and slow and rather to think of them as well this one's quite fast but it's volatile. Or this one is fast but doesn't scale well. Or this one's actually just slow all the time. sort of time just looking and exploring, just having a bit of a play with the data, you begin to develop much more nuanced view of what's going on with your site. And because your focus is different, you notice different things. So I'll give you an example.

13:27

Speaker 1: Here's some profile from the site. I've changed the names of some of the code, but but the it's it's it's a screenshot taken from our APM tool The way it works is time goes horizontally and each bar is a layer of the stack. So those top two bars, the last bit of middleware and then the view. And then the the yellow and the purple that's um hitting the database. Now if I was profiling this um in a reactive way I would straight away look at that massive query that's taking five seconds and say, right, we've got to sort that out. And then the next thing I would do is I would look at the following five, six, seven seconds um in the view. And I'd be like, well what's going on there?

14:14

Speaker 1: What was going on there was massive overhead developing the creating the query sets with all the prefetched data. But what I definitely wouldn't do is look at that final one second. And there's actually loads of really interesting things going on in that final one second that actually turned out to be really important. for improving the overall performance of the site. They weren't bottlenecks for any given view, but because they were everywhere, they had a massive impact. And this is how profiling to understand rather than like trying to hectically fix a problem, it allows you to notice different things. You gain more concrete insights.

15:01

Speaker 1: So lots of the advice you will see on performance is quite woolly. That's not a criticism of the advice. It has to be woolly. Because it's dependent on so many things, right? It's dependent on your infrastructure, um, the nature of your site, what you consider good performance. Um but when you spend time looking at data, um it kind of puts flesh on the bones of that advice. Um and you can say actually no, I know exactly when I need to use only and defer, and it's for these particular models or in these particular models. particular cases for my situation. But you only really get that flesh by spending time seeing what is fast and what isn't. Increase confidence considering trade-offs.

15:47

Speaker 1: So programming is full of trade-offs. It's trade-offs everywhere And in the past, when doing work, I would see a view and I'd think, okay, we're doing a couple of extra queries there. I could speed that up. But then I would make the changes and all of a sudden the the code was less readable and I was thinking would I understand this if I came to it fresh or if I looked back at it and see six months time. And it felt like a choice I had to make between performance and readability or whatever the trade-off is. But once you have a bit more concrete n understanding of what things are, it makes those trade-offs easy. You can say, okay, yeah, that query that I'm saving, I know that's only going to be one or two milliseconds.

16:34

Speaker 1: And this view is going to be roughly 150 milliseconds. Actually, in this particular case, it's a really easy choice. Readability wins. Not for the sake of two milliseconds. And so in day-to-day programming, it just makes all these little decisions so much easier, even when you're not doing performance work. As an aside, you gained some quite interesting insights about kind of how your users are using your site. If that was your primary aim, I wouldn't advise using APM tools. There are much better tools for doing that. But um yeah, that that that's an aside. Um the next one I think will um be important and that's just it's really really effective. There were all kinds of improvements we made after we started profiling systematically

17:21

Speaker 1: that resulted in a really, really big change. So that statistic there, 395 milliseconds. To 218 milliseconds. That's the average server response time of our template views. And it's actually filtered down to 200 responses, 200 status responses, because the other ones were a lot faster and I felt were kind of made the statistic less honest. So that that that's about a doubling in the page load times. Really quite significant. Um also just spending time in whichever APM tools you use, it's a really useful skill to have. On the rare occasion that, you know, um there is a genuine incident, um

18:08

Speaker 1: being able to get to the root of the problem really quickly um is is is just useful. Lots of the kind of platforms you have this massive array of options and they're difficult to get your kind of head round. So time spent um in them is often not time wasted. So, um hopefully I've convinced you that um you all need to be profiling on a wider range, but you do also need to to zoom in and profile um kind of do do the deep dives. And there are a few things I found that have been useful when doing this. One is visualization. Visualizations are helpful. Obviously we're we're all different and we all kind of um learn and see things in different ways, but some kind of visualization where you're able to kind of see where the bottlenecks are

18:53

Speaker 1: is useful. Don't just look at a single event, especially on production, you get a certain amount of noise. You do a deep dive into one particular request and it looks like one thing's a bottleneck. You do it again and it looks like a completely different thing as a bottleneck. And you need to do this several times to build up a picture of kind of the true nature of what's going on in a particular view. Likewise, don't just look at a single user or a single whatever the thing is that varies across your site. Because sometimes a bottleneck might be restricted to certain profile types. Try and look at slow and fast examples. They each give you different insights.

19:39

Speaker 1: And kind of as you look at these different things, just for a given view, you build up a really nuanced picture of what's going on. I mentioned earlier local versus production. I do think it's important that you're able to profile in production, but local is really useful too. You get different issues on local in production So you might look at something locally and spot an N plus one issue and you think, okay, yeah, that that's obviously the problem. But then when you look um on production, because um it's a scaling issue, actually there's a count somewhere. that's taking a much larger time than the N plus one issue. As your data changes, so where the bottlenecks are

20:25

Speaker 1: will change. So try and try and profile on both. Production metrics are obviously a better reflection of real-world usage. But they tend to be more volatile. You get issues like the noisy neighbor problem. Depends what time of the day you're working at, how kind of much load your site is under. And for experimentation, actually profiling locally I found to be a lot more helpful. I get much more consistent results. Uh and you don't need to deploy too. One way around this um is using the production shell um for testing queries. And if you're on Django 5. 2, a top tip

21:12

Speaker 1: is to write some profiling helpers and put them in the auto import to reduce that friction. Both give different insights. So you've done your profiling, um, very likely Um with a web app like Django, it's going to be something to do with the ORM that is going to be an issue. And I can think of no better place to start than the Django docs. This data access optimization page is a treasure trove of tips and techniques. If you haven't read it, Read it. If you have read it, read it again. I think it's like your favorite book or your favorite movie, and every time you come back to see it, you notice something different.

21:58

Speaker 1: In fact, I was going to find a quote to include in my slides yesterday. I didn't in the end. But I read something and I was like, oh wow, I never realized that. And it's kind of changed my view on a couple of things. So I can't go through that, but I can talk through some kind of foundational principles that I think most of the things on that page flow from. One is say for quick simple database queries, a large proportion of that time is spent in the network lab. So just doing some manual testing, I found that for a query taking 1. 5 milliseconds, about 65 % of that time was not spent in the database, but was spent um transferring back and forth. That that number may vary depending upon your infrastructure sorry and other things but um it's a significant amount of time.

22:46

Speaker 1: Um one way to think about this is to imagine yourself sorry, and the the corollary of that is you don't want to make repeated Calls, right? If you're wasting all this time in the network work clear, one big um call to the database is going to be more efficient than lots of little ones. And the way to think about this is to imagine yourself in a large warehouse. You're not going to, if you've got a list of 20 things to get, you're not going to walk to the back, get the first thing, walk back. Then do it again. Um spend five minutes walking to the back of the the warehouse, get your second item, go back to where you're from. Um that's just not what you would do. You would bring your list with you, you would go to the back, you would collect everything and walk back again. That's I think the way to think about it.

23:32

Speaker 1: Query sets are lazy. So what that means is they're not going to fetch that data from the database until they need to. Normally that means when you evaluate them or you start iterating through them, if you do if query set. If you convert it to another type, like turning it into a list, these are the things that are going to cause the query set to actually hit the database. But before that they won't. query sets internally cache results once evaluated. So once you've got those results, it's not going to keep on doing that long database hit. And your related entities aren't fetched by default. These are the kind of important concepts to understand. So a common thing is

24:17

Speaker 1: an N plus one issue. So let's imagine you've got a problem with coffee shops, and so you create a Django app to help you with that and you have a coffee shop model and you want to iterate through your coffee shops, well what's going to happen? That first line, nothing is going to happen because query sets are lazy. The second line when you start to iterate, we're going to get, can everybody read that by the way? Do I need to zoom in? Can you give it a thumbs up? Is that good? Great. Yeah, it's going to say select these fields from the coffee shop and we're going to get a table back like that. It's only using one database hit, everything's good. But let's suppose I want to support local business and so I record which of these coffee shops are chains.

25:04

Speaker 1: I had a chain for Unkey. And so now, instead of just looping through and printing off all the names, I'm going to loop through and I'm going to print off if there's a chain or not. This time First part of that iteration, I'm gonna do one hit to the database, I'm gonna fetch all those things. Um then once Actually, the first iteration, it's not going to do another hit. Django will know there's no related chain entity and it won't even try and fetch one. But then I get another query sent off to the database. um and I get some data back. And I do a third query off of the database and I get some data back. And you can imagine if I had um a hundred coffee shops lots of with chains, that's gonna be lots of um hits back and forwards to the database.

25:49

Speaker 1: It's like being in that warehouse and getting items one by one. We don't want to do that. And so Django has this thing called select related. And instead of uh doing the repeated things, what Select Related does is when we iterate we get a slightly different query. we get um what's called a join um and Django joins the coffee shop table and the um chain table. And we get one table that's returned. And so it's only one trip to fetch all of the data we need. That's great. But then let's imagine I start adding a review model. And my review model is to review different coffee shops.

26:36

Speaker 1: I noticed multiple reviews can be read for a single coffee shop. And perhaps um I want to list all the reviews. The same thing's going to happen as before. We fetch the coffee shop list But then when we hit that line which says for review in coffeeshop. reviews. all, Django doesn't fetch related entities up front we're going to have to hit the database again. We're going to say select this review where the review. coffe shop id equals one. And then for the next one, and then for the next one. And we've got another one of these n plus one queries. So why not use select related again?

27:22

Speaker 1: That seemed to work last time. Well, imagine if you did. Think about what that joined table would look like. It might look something like this. So I've joined the coffee shop table and the review table. So those the first kind of five so columns are my coffee shops and then I've got the review. But if we look at lines sort of two, three, four, we can see we've got lots of repeated data there. We're fetching the data on the coffee shop multiple times And if I had a coffee shop with hundreds of reviews, I'd be transferring all of this data about one coffee shop hundreds of times. It's not very efficient. And so instead we have this thing called prefetch related.

28:10

Speaker 1: And what prefetch related does is it does two queries. It does one query just to fetch the relevant coffee shops. Once it's got that, it knows exactly what reviews you need, and it then does a second query to fetch the reviews. And um That way we're only passing back the data for each coffee shop once rather than hundreds of times. So we've got prefretch related and select related. And maybe you've come across something like this. So we use select related for one-to-one and foreign keys, and we use prefect related for the kind of the reverse direction of that foreign key or the many-to-manys. And that's that's a useful rule of thumb, but I try to stop thinking in those terms

28:58

Speaker 1: these days and prefer to think about what is the join? What's the joined table going to look like? So let's imagine I want to print off all of my reviews. Well, each review is going to give rise to one coffee shop. So select related is the forward direction of the foreign key rights. So select related seems like the sensible option. But maybe I've only got five, ten coffee shops and I've got thousands of reviews. What's the joined table going to look like? Well I'm gonna have all of this repeated data for my coffee shops. There's only five coffee shops, but for each one I'm requesting it 200 times. That's not efficient actually pre-fretch related would be the better one to opt for in that particular case.

29:48

Speaker 1: Or maybe you've got a more equal ratio of coffee shop to reviews, but maybe your coffee shop table is really really wide and you've got lots and lots of columns, it's going to be lots of data, and all you really need is the name of the coffee shop If you do select related, you're going to pull in all of that unnecessary data. And actually you can use the prefetch object. So rather than thinking select related is for these things, um It's better to think, well, what's the the join going to look like? Um and actually once you start profiling, um

30:35

Speaker 1: you be this is another one of these things I've just noticed you get a better instinct for. So ORM optimization. It can be difficult. I'm gonna quickly go through the next few bits. There are a few useful tools that are useful for debugging. One is to log your queries You can do that via settings. You have this connection object which you can use to get your queries. And you can use that to write things like middleware. I have this is the thing I load into my shell automatically. and it tells me how many queries are on a particular thing. A few other useful tools. Every query set has a query object and the string form of that gives you the SQL

31:21

Speaker 1: And it has an explain object which tells you how the database is going to deal with that particular query. These are useful things. As well as that, when you're getting started, it's sometimes difficult to know, has this query set been evaluated? One thing I always used to get confused about was if I call dot values, is there a cached result of everything? And so this result cache, which is I guess technically a private method, can be really useful for just debugging things. Maybe you think the thing should be select related, but it's not. Likewise for an individual model instance, you can look at the underscore state fields cache attribute and you can see what has been prefetched or selected related.

32:07

Speaker 1: Those tools are also really, really useful if you want to write your own kind of auto-optimization features, things that automatically do the pre-fetching and selecting for you. The other things I find those really useful for Is writing regression tests. So if you've got a function and it's really important that function returns something that's lazy, these things can be really useful for writing tests to make sure that stays that way. Low-code solutions. Maybe you've done um uh some OM opti optimization and you've as as developers I'm especially developers who are interested in performance. We have all these tools at our disposal. We've got kind of caching and in-depth knowledge about indexing. And we see a performance issue. um

32:52

Speaker 1: and we get out the performance hammer and we start hitting it. But actually sometimes taking a step back can be more efficient in terms of developer time and give much better wins. And a couple of stories that I think illustrate this quite nicely. So this was a uh a list view. It was an internal page used by our finance department, uh, and it was taking longer than 30 seconds. And at that point the server was timing out, and so our finance team couldn't access the invoice pages. There are a few things that were slowing it down. The search at the top was you could type in anything, but it would search the invoice field, the client field, hire a first name, hire a last name, freelance the first name, freelance the last name. lots of things. There was loads of data being fetched.

33:38

Speaker 1: It was 17 joins. To make things worse, we had this notion of batched invoices. And to make sure a batched invoice wasn't split between two pages of pagination. It was actually doing this query three times. wants to get the main result and wants to check the invoices a batch wasn't split between two pages at the beginning and the end. So it was a really slow page. And so I started kind of looking into it and I made a few improvements. And I got it down to about 17 seconds, which was still kind of not great, but it was better. So I sort of went along cap in hand to the finance team and I said, you know, I'm I'm really sorry, um it's still so slow, but you know, at least it's usable. And I started explaining to them some of the reasons what

34:25

Speaker 1: which was making it slow. And I talked about the batch invoice And they said, oh, um we we we stopped doing batch invoices for this particular group two years ago. Um and then I spoke about the the search thing and they said oh yeah no we're only ever going to use this um view to type in an invoice number. That's what we care about. Like we you know we don't deal with names and things. And so I'd spent all this time trying to optimize this thing when in actual fact it wasn't a performance. problem, it was a communication problem. So understand your requirements and your users. Question functionality. I think performance shouldn't dictate to functionality. That's the kind of definition of premature optimization.

35:13

Speaker 1: If you do that, you end up with a blazingly fast site that does nothing. But when um when Performance has become an issue. Um I think it's then right to question functionality. Umtice changing functionality doesn't necessarily mean making it less useful. So a really good example of this is I mentioned pagination earlier. So we have a lot of list views in our site. They're all paginated. Um and they all do a count. Now in Postgres, as tables get big, doing a count is really, really expensive. And so a lot of these pages were getting really slow. It hadn't been a problem historically, but as we grew as a company, it was getting worse and worse. So we had a timesheets page that was taking 11 seconds.

35:59

Speaker 1: And our pagination components looked like this. And I spent a bit of time thinking about different things I could do to speed up the counts. maybe some database tuning. Maybe I could use the explain analyze to get an approximation of the row count rather than doing an exact count. That would have been a lot quicker. But actually there was a far simpler solution. Just use a previous next. Nobody cared about the total amount. All they wanted to do was look through the first few pages. Um and so taking a step back and thinking slightly outside the box. was A a lot easier in terms of developer time and B actually a lot more efficient than whatever my clever solution would have done anyway. Here's another example. So we had this view,

36:45

Speaker 1: it took 22 seconds, and as you can see, 21 of those seconds are doing this one query. This is the query. It's not important what it is. I spent a bit of time playing about with things, trying to make some improvements. I made some pretty minor improvements, but we've been talking like 17 seconds, not 22 seconds. And this was a user-facing page. And the problem was that there was a missing in it was missing an index. This table is recording every single search. So if you think thousands of daily users, each of whom are doing lots of searches, this table grows pretty big pretty quickly and it's doing a sequential scan through the whole table. The problem was it was third-party package.

37:31

Speaker 1: And I was thinking, okay, I could fork the package, could maybe do a pull request. I could add a kind of standalone migration to add the index and I was weighing up my options and I thought whilst I'm doing this I might as well ship the improvements I've made. got something to show for the day. I thought before I do that I better do some click testing, make sure I haven't broken anything. So I went to the page and I couldn't see I should have said this is for loading search suggestions. And I couldn't see the search suggestions. And I thought, oh, that's weird. I hope we're not broken anything. So I I dig dug a bit deeper. It turns out we'd actually removed this functionality a year earlier. But when we removed it, what we did was we just got rid of the UI and we were still passing the data back. So again, I'd spent all this time trying to optimize something

38:18

Speaker 1: that didn't need optimizing. And kind of the idea is un understand the full picture, don't reach straight for for the hammer. Especially with legacy code. There have been a few situations where legacy code have been have been an issue with performance. I'm going to skip this next one. The two-second summary is fix the data, not the code. But we made some big speed-ups there. So low-cost solutions require a step back, but often provide a much bigger speed-up for far less developer effort. And you can still use the hammer. You know, you can do these thinking outside the box ways of speeding things up, and then use all your other performance tools to maybe get that last 10%. But think

39:04

Speaker 1: about other options before you reach for the indexes and the ORM tweaking and things like that. The hot path. So this is the really critical bit that all of your requests go through So for us it's the middleware, the context processors, and our base templates. If you're serving up JSON via GraphQL, it might be different. It would be, you know. all the extensions you have wrapped around your GraphQL endpoint. But whatever the path is that all of your requests go through, I think is an area where it's worth really optimizing things. Um a good metric of this is what's the minimum five percent of all your requests? Because by doing that, you get views that are basically spending no time and you get a measure of how much time are you

39:55

Speaker 1: spending in that critical section um of the of the path. So back at the beginning of the year you can see we were spending 100 milliseconds um on every single request. And we were able to make some pretty big speed-ups. One thing we did was we made all of our context processors lazy. So instead of doing this, we used simple lazy object. And what that does, it means the expensive calculation is only executed when you access the thing So there were lots of things we need for lots of templates, but not all templates that were expensive to calculate, and that that made a difference. Another thing, um better caching.

40:41

Speaker 1: So um I think in the past we've tended to think of caching, so each of these um hits here is a a hit to the cache This was a load of assets we had at the top of every HTML page and we included an SRI hash. So that's just a code that lets the browser know the asset hasn't been tampered with. Um you do want to cache those, but we're catch we're getting the same ten SRI caches hashes, sorry, every time. So instead of doing um uh ten separate hits to the cache Just do it in one. Cache the whole partial rather than the individual SRIs. The lesson there is to think about your cache in a similar way to what you'd think about the database.

41:29

Speaker 1: And so this big drop, um it wasn't the things that I would have expected it to be It was 10 milliseconds here, 20 milliseconds here, that we identified through this more systematic profiling, that we wouldn't have noticed by looking at kind of individual. kind of problematic views. And whilst each of these things on their own was small, it added up to a really dramatic effect. And actually that graph does go down by another 10 milliseconds later on, which is what this one was about. So yeah, where does that leave us?

42:15

Speaker 1: Well, I think um if you were to ask me um a year ago, what are the issues with our site, I would have said it's definitely um lots of M plus one issues, lots of indexes missing And it turned out to be all of the things that I didn't know that that made a big difference. I think it's worth saying also that this is only half the picture. I've focused very much on the back end because this is Django. um but the kind of network layer is a is another whole a area and there's no point serving your requests super super fast um if you then have Two seconds with a massive JavaScript framework waiting to render those things.

43:01

Speaker 1: So yeah, that that's something that I just thought was worth mentioning. But um I hope that's been of some interest to you um and thank you for listening.

43:18

Speaker 2: All right, thanks Tim. I think we have time for one question. Um

43:22

Speaker 3: do you have any thoughts about which of the uh profiling tools are worth paying for? Because they all give you like some tools that are free and then charge you a lot for more. When you're when you're getting this like continuous, this proactive profiling going.

43:35

Speaker 1: Yeah. I feel like after this talk I should work as a tech evangelist for one of these firms. So like the one I've used most is Datadog. I know that's really expensive. I find it really, really useful. I don't I I've also used um Sentry , so we um used that primarily for exception handling, but kind of when they launched their APM stuff, um we used it for that as well. I don't think I'd like to comment just because I've not. I kind of took a quick look at some of the other tools in preparation for this talk, but because I've not really used them in anger. I feel it would be unfair to recommend a particular tool. I do think that very likely

44:20

Speaker 1: fluency in the tool you choose is probably more important than the particular tool. Thank you.

44:28

Speaker 2: Okay, thanks. Yeah. I'm sure Tim will be around to answer questions in the hall.

44:32

Speaker 1: Yes, do you come say hi, I think bye. Thank you.

Questions this talk answers

What is proactive profiling, and how should I use it on a Django site?

Proactive profiling means measuring performance to understand the broad distribution of fast and slow requests, rather than only investigating an incident. Tim recommends combining production-wide statistics with drill-downs into the slow or high-impact views, and repeating the process regularly.

Discussed at 8:06

How much can systematic performance profiling improve Django response times?

At Tim’s company, systematic profiling helped reduce average server response time for successful template views from 395 milliseconds to 218 milliseconds. The biggest gains came from many small improvements that broad profiling revealed, not only from obvious per-view bottlenecks.

Discussed at 17:21

Should I profile Django locally or in production?

Both are useful for different reasons: production reflects real usage and scaling behavior, while local profiling is more consistent and better for experiments. Production data is more volatile, so comparing both environments gives a fuller picture.

Discussed at 19:39

Can changing product functionality be a better performance fix than optimizing the code?

Yes. Tim describes replacing expensive exact-count pagination with simple Previous/Next navigation, after realizing users did not need the total count. He also recommends checking requirements and removing obsolete work before spending time on ORM or database tuning.

Discussed at 32:52

What is the hot path in a Django application, and how can I optimize it?

The hot path is the part every request passes through, such as middleware, context processors, and base templates. Tim recommends minimizing work there—for example, making context processors lazy and batching cache accesses—because small savings on every request add up substantially.

Discussed at 39:04

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Timothy McCurrach

More videos from DjangoCon US