Closing session
Published June 13, 2025
This video features Rudi Giesler at DjangoCon Europe 2023 in Edinburgh, Scotland.
Squeezing Django performance for 14.9 million users on WhatsApp
by Rudi Giesler
https://pretalx.com/djangocon-europe-2023/talk/PYFUGF/
At the start of the pandemic, there was a large need for accurate information to combat misinformation. This is how we used Django as part of our South African WhatsApp service "ContactNDoH"
At the start of the pandemic, there was a large need for accurate information to combat misinformation. For this, we developed a South African WhatsApp service, "ContactNDoH", which disseminated accurate information, and later managed registration and bookings for vaccinations.
With over 14.9 million users, and over 850 million messages, we needed to squeeze as much performance from Django as we could.
This talk will go over:
It will focus mostly on backend performance, as the role Django played here was as a REST API.
Rudi Giesler explains how REACH uses Django and existing channels such as WhatsApp, SMS and USSD to deliver healthcare services, including South Africa’s Contact NDOH COVID-19 information service, which reached 14.9 million users. When the service grew to include a symptom-based HealthCheck workflow and large data exports, database scale exposed Django admin’s expensive COUNT queries and unsuitable pagination defaults. He shows how to use PostgreSQL’s estimated counts and statement timeouts, custom admin pagination, and Django REST Framework cursor pagination, then describes a practical performance-testing process using Faker, Locust, Django Debug Toolbar, EXPLAIN ANALYZE and query-count assertions. He also stresses measuring before changing anything, testing with realistic data in a separate environment, and considering runtime, database, memory and CPU bottlenecks while accepting the security and reliability trade-offs of messaging platforms.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: Right. Thank you very much. Thank you for having me here at Django Con. My name is Rudy. I'm a principal engineer at uh organization called uh REACH. I've been at Reach for past eight years writing open source software using Python and Django. I want to give you a quick overview of who who we are, uh what kind of projects we do to give you kind of some uh background information, some context uh around the more technical things, which I'll dive into. in a bit. So REACH is a nonprofit organization based in South Africa that uses existing technology to improve healthcare and to create social impact.
Speaker 1: So what do we mean by existing technology? These are things like uh this is uh technology that our users already have in their hands, things that they're already using and already familiar with. We're talking about things like SMS, USSD, WhatsApp, and sometimes on occasion web as well. Uh but we still use Django, even though web is a small part of what we do. By saying improving healthcare, we focusing on lowering the barriers uh to access to healthcare uh as well as supporting the health systems uh that are uh providing these health services to our users. When we talk about innovation, we talk uh we are using you know
Speaker 1: old existing technologies, but we innovate in the way that we use these technologies in new and interesting ways. Um I'm gonna go over some examples of projects we use to give you some uh more concrete examples. Um In South Africa we have a service called Mom Connect. It started off over USSD and SMS, later on expanded to WhatsApp. Uh the way the service works is uh pregnant mothers uh would go to when they register at a clinic, they can get registered to the service. It then sends them messaging with relevant information to their pregnancy according to the stage of pregnancy that they are in. So it's specific uh
Speaker 1: tailor content relevant to their stage of pregnancy. Um this is things like uh you know the clinic should be giving you iron pills to take at this stage or uh your baby is now the size of an apple. Um the interesting or your next clinic appointment is due soon It also gives them access to a help desk staffed by nurses where they can ask any questions about the content or any questions they have in general about their pregnancy. Another project we worked on in South Africa was Young Africa Live. This was a WhatsApp service. Um it's designed to empower young people to explore and embrace their sexuality and make informed decisions on their Health and well-being.
Speaker 1: This is one of the first services where we looked at doing tailoring and segmentation, having a look at how we can customize the service. To better serve uh their specific uh circumstances. Uh another project was MenConnect. This was also on USSD, SMS, and WhatsApp This was for uh to deal with uh antiretroviral treatment initiation and adherence for men living with HIV. So how do we what can we do over just text messaging in order to help people initiate drugs and also then adhere and keep taking those drugs
Speaker 1: throughout uh their lives. And yes, so another project, the one I'm going to be focusing a bit more on this talk, is Contact NDOH. To set the scene, this was uh near the start of the pandemic, this is when we were having lockdowns uh come into place. Um there was a lot of misinformation uh going around uh a lot of this was happening on WhatsApp. So we wanted uh to create to do something to help um get uh proper information out there. Um we partnership with our sister company turn. io uh to create a uh national
Speaker 1: South African National Department of Health WhatsApp service. So this is uh their service which we uh implemented. Um it sold off being a technologically very simple service. It was keyword-based. You send in myths, we send the thing. So technologically it was fairly simple. We did this so it could be quick to get up and running and easily handle large traffic volumes. Um and this is the service where we ended up with 14. 9 million users. Um so after this, after we had this up and running, we were getting this uh evidence-based up-to-date information to people. Um we started to think about what can we do that's uh
Speaker 1: that's a bit more advanced. You know, the we there there's quite a big limitation on the system that we're putting in there. Um so what can we do uh to now that we have a bit more time, a bit more flexibility. So at the time not everyone could get tested for COVID-19. There weren't enough tests. So there were guidelines about who can get tested, but these were often quite complicated, confusing, they were changing often and no one really knew what was happening. We created a service called HealthCheck on the Contact NDH platform. We worked together with the National Institute of Communicable Diseases in South Africa.
Speaker 1: to develop and then later on updated as new information came in, as things changed on uh what you should do, you know Let's ask you some questions about your symptoms, how are you feeling, and using that create a uh risk classification algorithm to uh to let you know, you know, okay, you need to go in for a test now or you should you need to isolate for a week uh or no you're actually fine. Uh carry on as you were. Uh so as part of this the National Department of Health wanted to integrate this into their data lake Um this is to help them detect uh hot any hotspots and to redirect resources in any areas as well as to gain any sort of insights that
Speaker 1: uh they wanted to to know about symptoms and how they were um spreading across the country. So as you can see this is a lot more uh complex uh system to implement. We need somewhere to be able to store all these results, all this data, and also be uh be able to surface this for the NDH datalake team to be able to pull this. into their system and then run analysis to this. So of course, Django to the rescue. This seems like a very Simple thing, right? Uh we take Django and put a Postgres backend to solve the data. Uh we use Django Rest framework. Uh to create the APIs that we need.
Speaker 1: Um and everything works well, right? Well, until you hit lots and lots of rows in your database. um and things start becoming a bit unhappy. This was a very interesting project for me. It was one morning I was looking at the packaging on the loaf of bread. eating for breakfast and I saw an advert for my service. Uh the service I was working on there. Um that's quite cool. Uh it's never really had a project like that. Most of our things are just advertised inside clinics. So to see it being advertised on TV, even on the loaf of bread I was eating, this was a big service across the country and we had lots of users trying to interact with it So
Speaker 1: what is the one thing that you should never do on a really large Postgres table? Well, count star on the whole table. That is uh there's a whole page of that in the Postgres documentation. Um it will lead to a very slow query and uh be very unhappy So of course your Django admin page when you're viewing all these things, of course it does a count star, but it doesn't only do a count star, it does two count stars. uh in order to generate that page. So how can we get rid of these? Well the first one is fairly easy. The first one is due to when you select all the objects, it shows a little thing up there saying you've selected a hundred
Speaker 1: objects, there are a thousand objects total. Do you want to select all of them It uses a count to get how many total objects you have in there. There's an easy config in your admin model there. You can do the show full result count. equals to false and then it just says show all instead of show a certain amount of number. So that one was easy to fix, but we still got one more. That one is a bit more difficult. This one is related to the pagination, which is unfortunately not something that we can easily change in the Django admin. Django admin uh uses page-based pagination, it puts the page numbers at the bottom. Uh we don't really want to change the whole thing.
Speaker 1: Uh that's a lot of work. So how can we get around this? Uh well we can create a custom pagination class. Unfortunately we That class doesn't allow us to completely change the way pagination works, uh, but we can uh Postgres offers a method, a way to get uh uh an approximate count instead of doing the full count on the table uh using um a statistics that it gathers on there So for these really large database uh these really large tables, we don't really care to get an exact count. We're not gonna want to navigate to exactly the 103 ,000th page of results. Um we do we want to use the Django admin in other ways, but um
Speaker 1: yeah, so but we're unfortunately stuck with this pagination. So we can use this count estimate. um to really uh give Django something that's close-ish um so it can generate its page numbers and things uh but still um allow us to not have to do the the count on the entire table. So So this is how we could implement something like this in our custom pagination class using Postgres. Um yeah, this will just give us at approximate count. Um but what if we don't want to do this on the t on every table, what if we want to make a generic thing uh that
Speaker 1: we can use everywhere? Because we can do better than just get an approximation. What if this table starts with small, gets large, what if in some cases we don't Um we don't want to uh you know always have this approximation, which can be quite wrong for small numbers. Um well what we can do is we can set a timeout on the uh SQL query in Postgres. What this allows us to do is we try to do a count on the whole table. If that starts taking too long, we can then cancel that. um and carry on with our approximation. So we can put these two things together inside our custom uh pagination class. We go through, we try to do a count of the whole table.
Speaker 1: If that takes too long, we can then return the approximate count of the table instead. And that solves the two counts that uh Django admin does and makes our admin page responsive again. We can use it. Great Next, and we'll talk about how we design the API with Django REST Framework Now we can't return all the results uh in a single request. That would be a massive request, uh, take forever. We don't want that. What we want to do is have some sort of pagination, return a page of results and a way to get the other pages of results Django Rest framework has a few built-in options. It does the page number pagination, similar to what Django admin does.
Speaker 1: Obviously, you can't use that. There's a limit offset, which also uses count, so we can't use that. Finally, there is cursor pagination, which it offers. This is the one you want for really large tables. It has some downsides, you can't navigate to an arbitrary position. In the dataset You make your query, you get the first page, you get a way to get to the next page, and to get to the tenth page, you have to go through ten pages. But for our use case, that's not really uh an issue. We have an index on the timestamp field of when these things were inserted. So if you want Part of the data set, you just say give me everything from this timestamp onwards.
Speaker 1: Um very important to have an index on the ordering field you're using, um, because that is how this uh this pagination works uh in order to create the pages it's important to be able to very efficiently um order and go to specific part of whatever your ordering field is there So uh following on you might have noticed that I've missed a part of this. Um I've gone straight to the solution here of what the problem was and how we fixed it. But how did we find that? How do you find where the problem is in your system and how to improve that and how to fix it? So the first important thing is you need to measure before making any changes.
Speaker 1: You need to have a setup where you can go through, test And make sure you know that the changes you're doing has the desired effect that you want. It's making things better, not worse. You're not not doing extra things and has no result. So the first step of that is that you want to be able to have a test set of data in your database. If you do running these tests on an empty database, uh you're not going to get the same results as in production. So for this, if you use a uh If you can use a copy of your production base, great. That works. But oftentimes you can't. Maybe this is before you have any data in your production database. You want to make sure things will hold up
Speaker 1: or uh it contains confidential information that you can't put in your testing environment. For this I find using a library called Faker works really well. This allows us to generate realistic looking but fake data to put in our uh uh database services It allows us to define things just simply in Python and it can generate as a wide range of different data types. um that we can use to to generate uh the realistic looking data without actually using our production data. The next step is now we need to put some sort of load on the system. We need to have we have the fake data. Now we need fake users using our system and so we can see how that works.
Speaker 1: Um tool that's great for this that reviews is called Locust. Um this allows us to do quite nice um load testing uh on the system, you can do fairly complicated user flows with it. Um and then you can see exactly what endpoints are slow, where do you need to focus your your uh time and energy at improving the system? What are the the bottlenecks that you have there? Similarly to Faker, you can define it just using um just using normal Python where you can interact. with the APIs to make requests against yours. And you just describe your user journeys, you describe how often uh how likely each of the user journeys will happen, and it will go and execute those and
Speaker 1: run around with your fake users. Now that you've determined where the problem um where the problem might be, uh this view or or this endpoint uh is very slow under load. Uh now you want to know why is the slow under load. Dig deep into that. One of the most useful tools I've found is quite a popular tool. Uh it's called uh Django debug toolbar. Uh you navigate to a page, it will give you so much information about um everything that was used to um To uh create this page. Um it and the most important thing is it chose all the SQL queries needed uh to generate this page.
Speaker 1: So this is and the time it took each of those SQL queries. So this is where you can see, oh, my Jenga admin is doing two count stars and that's why it's slow. Or Um, you know, this this query is very slow, we should look at putting an index on this field to help it out. Um So this uh this really helps you to delve deep into those things. Um and now uh nearing to the end, I just want to uh go over a few more odds and ends, a few more tools that um kind of really Helped us. So you can use Django debug toolbar or if you're in uh debug mode, you can use the connection. queries. That will give you the SQL uh queries that are running.
Speaker 1: So you can see what is the actual SQL queries that are running. Um if you put an explain analyze in front of that and send that to your Postgres database, it will tell you exactly what the query planner will do, what it's doing. uh what parts of the query are taking long. Uh that really helps you to find out which fields uh might need indexes or where you can really uh change your queries, change how you're making your queries uh to improve them. Uh but of course always make sure you're testing you're running this on a database that is full of data to get realistic results. Um if you have the query set you can just call dot explain on it as well, which does the same thing. Um much easier Uh
Speaker 1: then as part of the uh Django kind of testing framework, there's the assert num queries assertion. This is really useful. You can put this in your unit tests after when you run your view, and it will assert that you only make a certain number of database queries. This can help you catch any of those annoying N plus one issues on your views. It can also ensure that if someone comes along changes a critical view which makes it make another query or more queries you can this test will then fail and you can go and have a look and make sure that uh you're not gonna slow run this performance test make sure this change is not gonna slow things down. um and this release uh yeah uh make your service
Speaker 1: worse. Um another one is PyPy. Um Uh it's of it's very compatible nowadays with a lot of uh libraries and compatible Django and simply by running your things on PyPy you can get a nice performance boost. Uh not always So check your uh check with your benchmarks that you've set up now. Uh make sure that it actually does. But often or not it can just give you free performance. So that's nice. The next one is logmin duration statement. So config in Postgres. You set this to a certain value and Postgres will log any queries that take longer than this So you can put this in your production and then any queries that are taking a long time, you'll immediately be able to see those interlogs
Speaker 1: Hopefully somehow trace that back to the application, what is making those queries, and uh figure out if not, it gives you a good idea of where you might need some indexes on your tables. Uh last one, I want to give a shout out to pythonspeed. com. Uh I've spoken a lot here about uh SQL and optimizing SQL queries for your service. But uh what if you have a uh what if you have a CPU bottleneck or memory bottleneck? Um uh mostly Jengaps are the the database is your limiting factor, but Maybe it is a CPU or memory bottleneck, how do you measure that? How do you find out where in your code is is causing that slow? There are many great articles on there to kind of help you um
Speaker 1: with that. And there's also a great article on configuring GUICorn if you're using it inside of Docker. So if you're doing that, I can highly recommend that for performance analysis and reasons. Right, thank you very much.
Speaker 2: This is such a fascinating talk on multiple levels. For me, what I really enjoyed was that this is very practical, very hands-on. And there's another thing going on through through my mind, like when you mentioned that you saw the advertising for the service on on the bread, I get the goosebumps because I realized that by making this app available, by making sure it stays stays up and running. During the pandemic, but also during during other events. You know, if if you see someone who who saved a person from fire or another dangerous situation, you know, it's it's it's like it's it's a front page material. But you know if you even if you are very conservative with estimates, even if one in ten thousand users uh was saved because of this application
Speaker 2: Uh there there's hundreds of people whom your application might have saved. And that's why I realized that in this room, there's a number of us, a number of you working on uh on applications like this uh that you know when they work during the pandemic when they work during other times that actually saved people's lives or affected them very positively And um I don't know if you think of yourselves like heroes, but but you are true heroes, you know, saving hundreds of people. You're not rampage material, but this is huge.
Speaker 1: We have some time for questions?
Speaker 2: Yes, we have some time for questions. And um there are some um microphones on the left or on the right. And if you'd like to ask a question, you can also raise your hand if you're if you're uh in the audience.
Speaker 3: Um thank you Rudin. That was really interesting. I I'd like to know how you felt knowing that you were delivering this um critical healthcare information and interaction to millions of people urgently in a crisis over a platform owned by one of the most unreliable and dangerous companies that's out Facebook out there on on the internet in effect. That it's it's it's not It's not really a safe infrast infrastructure. You're at the whi-wim of or wait were you at the whim of WhatsApp and Facebook to do the right thing
Speaker 1: Yes, so this is a obviously a a difficult um thing to kind of think about because uh we want to get it in the hands of people, want it to be easy to use, and WhatsApp is such a great tool for that And we've been working with WhatsApp since before they were acquired by Facebook and working with them to better test and develop their API and allow people to kind of run services over it Um but yes there there is a bit of uh of difficulty there, you know we're dealing with sense of health information, but also we're dealing with WhatsApp and Facebook. And there is a bit of security there in that WhatsApp is end-to-end encrypted. Uh we the message is decrypted on the user's phone and then on our side is
Speaker 1: decrypted um only by our sister company, turn. io, which is a trusted entity to us. So they run the decryption kind of uh side of the of Facebook's infrastructure there. But yes, uh there is there you do need to have a a level of of of trust to that because these are black boxes that Facebook kind of handles hands you to run and says these will decrypt your messages and send it to you. You um there's some trust that You know, the app isn't doing anything strange once the um data is decrypted on the phone. But I think um it's It 's difficult but you have to kind of weigh out the advantages of that and advantages of using these platforms. Um this was also available um
Speaker 1: Uh some parts of it was available over USSD and SMS, uh which is even less reliable. Now you're trusting your mobile network operators, you're sending these things over plain text, it's not end-to-end encrypted So, you know, what what is what is better there's no matter what medium using, unless I guess web is you can trust a bit more. Um but yeah, it with these convenient uh kind of uh channels to communicate with the users there's some trade-offs and you have to trust at some point and kind of do your best to keep their data and information safe.
Speaker 2: Do we have more questions in the audience? Please.
Speaker 4: Um I just want to thank you for the talking. Um How do you manage the lockers with foreign keys and complicated patterns? I know it's a specific question, but uh that was very interesting. Thanks
Speaker 1: Uh so yes with with Locust um uh for for us it was it using uh foreign keys wasn't that difficult. Uh but What you can do is define your user journey such that the user creates this thing, and then upon that you'll get a response which has the foreign key, and then you can use that for the next steps of your user's interaction Um so you don't need to hard code these foreign keys in. Sometimes there will be some hard coding if you say needed to test a specific um you know a specific article or page um or maybe selecting from randomly from a list and you kind of hope that All of these are relatively similar to have a relatively similar cost to generate.
Speaker 1: But yeah, it can be a a bit difficult. But Just try and simulate what a user would do in that. And a user wouldn't know what the foreign keys are, what the primary keys are of the thing.
Speaker 2: Two more questions, please.
Speaker 5: Um you showed the Django debug toolbar and uh in my experience it's a great tool, but it tends to be a little bit slow. Do you have a solution for this or do you just use other ways of finding the problem when we have s slow view or something like this.
Speaker 1: Yeah so Django debug tool is not something run in production. This is something run separately an isolated test environment with the fake data. So for us having it slow is is not an issue. Um but uh I do know we didn't go over there's a lot of tools for kind of uh sampling and profiling um things in production, that Postgres query, uh Postgres configuration option to show slow queries is an example of one of those. And there's a lot of kind of tracing frameworks and tool sets that if you want to do this sort of thing without on in production without affecting performance. That is a way to do it. Um but yes, this the we stuck to using it in places where the the speed didn't matter. Um
Speaker 1: we just wanted to Uh it doesn't matter it takes ten times longer to generate than in production as long as it gives us the relative ratios of this is the part of the view that's slow and this is the part that's giving us issues. Thanks.
Speaker 2: And the final question please.
Speaker 6: Yeah, you mentioned uh low cost and load testing. Um how do you make sure that you don't test the limits of your testing framework or the how do you generate this big of load because you have a high availability and a a very big system to handle those requests. Yeah.
Speaker 1: Yes, so um if you go in depth into Locust it does have some things around that creating a I think they call it a locust swarm, locus locust, um uh where you can have multiple machines making requests. Uh but yes, that's uh kind of an important uh thing to make sure that um you know you have a target of you want to load test so many users making so many requests per second make sure you actually um Make sure you're actually reaching that and make sure that uh you know you're not being your you the the server running the request is not being bogged down. That being said, Locust um I found scales up pretty far before you have to do a like multiple machine setup. Um it can take you pretty far with quite high loads
Speaker 1: and most of the time it will be your your servers, but obviously as you said if you testing it on a um on a a large clustered setup where you can handle a lot of a lot of load you're not just looking for where those pain points are Yes, there are tools within Locust to kind of help you generate more load, but that's a good point. Make sure that you're actually reaching the targets that you want to reach on the loads.
Speaker 6: Small follow-up then. Do you test against your production environment or do you set up a separate one?
Speaker 1: Set up a separate one. I guess you could, if you want to, test uh during a quiet time against your production environment. Um I prefer to keep things separate, uh test on a separate setup with a copy of the production database or generated fake data. Um I find that it's But uh I find that's worth it just for the kind of safety of um of not taking down your production system.
Speaker 6: Thank you very much. Great talk. Thank you.
Disable Django admin’s full result count with `show_full_result_count = False`, then use a custom paginator that first tries an exact count with a statement timeout and falls back to PostgreSQL’s approximate count when necessary.
Discussed at 8:38Use cursor pagination rather than page-number or limit-offset pagination, since those approaches can require expensive counts. Put an index on the field used for ordering, such as an insertion timestamp.
Discussed at 13:18Populate a test database with realistic fake data using Faker, then use Locust to simulate user journeys and load. This reveals which endpoints are slow and where the system’s bottlenecks are.
Discussed at 14:54Use Django Debug Toolbar to inspect the SQL queries and their execution times, then use PostgreSQL’s `EXPLAIN ANALYZE` or Django’s `QuerySet.explain()` to examine the query plan and identify missing indexes or inefficient queries.
Discussed at 17:12Use Django’s `assertNumQueries` assertion around a view or operation. It can detect unexpectedly high query counts and prevent later changes from silently making critical views slower.
Discussed at 19:30Often it can provide a performance boost because it is compatible with Django and many Python libraries, but the result should be verified with benchmarks rather than assumed.
Discussed at 20:16Set PostgreSQL’s `log_min_duration_statement` to a suitable threshold so queries exceeding it are logged. Those logs can point you toward the application code or database fields that need optimization.
Discussed at 20:16Define the user journey so one request creates the related object and returns its foreign key, then use that returned key in subsequent requests. Avoid hard-coding keys unless you specifically need to test a known object.
Discussed at 27:03No. The talk uses it in an isolated test environment with fake or copied production data, where its overhead does not matter. Production issues should instead be investigated with slow-query logging and tracing or profiling tools.
Discussed at 28:26Set a target number of users and requests per second, then verify that the target is actually being reached and that the load generator is not saturated. Locust can scale substantially on one machine, and it also supports distributing a swarm across multiple machines.
Discussed at 29:46The speaker recommends a separate environment, using a copy of the production database or generated fake data, to avoid risking an outage. Testing production during a quiet period is possible but not preferred.
Discussed at 31:03Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025