High Performance Django at Ten: Old Tricks & New Picks with Peter Baumgartner
Published October 23, 2025
This video features Peter Baumgartner at DjangoCon US 2014 in Portland, Oregon, USA.
By Peter Baumgartner
Django makes it easy to build a site and get it running on your laptop, but how do you go from there to a site that can gracefully handle millions of page views per day? This talk will show you the modifications and supporting services needed to make your site scale. Topics will include caching, uWSGI, Varnish, and load balancing.
Help us caption & translate this video!
Peter Baumgartner argues that Django’s ability to handle large traffic depends less on Django itself than on the servers and infrastructure around it. In a live, deliberately rough benchmark, he compares Django’s development server with uWSGI, Nginx load balancing, and Varnish caching: replacing `runserver` with six-process uWSGI raises throughput from about 28 to 72 requests per second, while two application servers behind Nginx nearly double capacity. Varnish produces the biggest improvement by serving cached pages in milliseconds, protecting backends from cache stampedes, and continuing to serve cached content even when both application servers fail.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Cool. Thanks everybody. Um yeah, so today we're gonna talk Uh high performance Django, uh basically taking your site from run server on your laptop to being able to get uh uh on the front page of Reddit. So my name is Peter Baumgreiner. Like you said, I'm the founder at Lincoln Loop. We are a Django web agency, so uh we build Django sites. We help people with Django problems and we help people learn how to scale and build new sites that are built for large high-scale traffic. So we've been around since 2007 and uh in those seven years we have learned uh a lot about Django
and how to make it. run fast and handle lots of traffic. We've learned a lot of those lessons the hard way. So we actually recently wrote a book by the same name, High Performance Django. uh that kind of bundles up all those lessons we've learned and uh packages them uh in a nice little ebook now and probably uh a print edition later. Actually uh I have a couple here. Um maybe we can give some out to people asking questions. So if you're a reader of Hacker News or uh any any other uh web publications, you may have seen this before. Django doesn't scale. Uh how many people here uh believe that does Django doesn't scale?
Okay. I'm gonna be a little controversial and say that's true. Uh Yeah. Yeah, Django does not scale. Uh you may say what about Instagram and Pinterest and discuss all these people are using Django and they're you know at at massive uh massive numbers of users and traffic, how are they doing it? Well I'd say they're not actually using Django uh as much as uh supporting Casta players. So You know, it's their database, uh, Postgres or MySQL, it's uh Memcache and Redis, uh doing load balancing with Nginx, uh, maybe doing caching with Varnish. None of these are Django, but they're really what makes
a Django site scale. You could use them just as easily with PHP. So all those servers, uh how do they work? Um how do you plug them all together? That's that's going to be the uh the focus of the talk today. Um We're going to do a lab demo. I um basically gonna take a site from running and run server and we're gonna blast it with a lot of traffic, see how it performs. and then scale up. I'm not going to talk at all about how to optimize your code today. I'm not going to talk at all about tuning your databases or how to do caching. We're basically going to work higher in the stack than that. We'll be talking about how you serve your Django application, how you do load balancing, uh, and things like that.
Um I can't do this on my laptop because I I literally need like multiple servers. Um so uh if any of you are doing like uh massive BitTorrent downloads now or like watching cat videos or something, um I would appreciate it if you shut it off because it's going to be a really lame talk if uh I can't get out to the internet. So um like I said, we're gonna be throwing a lot of traffic at these servers. It may look like uh we're benchmarking Django, but um really what we're gonna be doing would be a terrible benchmark. Um I'm gonna set up a fake Django application. It's going to have fake data in it. I'm going to use EC2's network. Who knows what's going on there? And I'm going to use Docker containers inside virtual machines on shared infrastructure.
Again, you know Neighbors could be doing anything. So don't take the exact numbers we're gonna we're gonna see here to heart, but more um of the difference. the the difference between uh each different uh setup we're gonna have. Okay, so uh first up we'll take a look at Django. So let's see here. This is a um EC2 instance I have set up. It's got a Django application on it. It's a M3 extra large So that is uh has four cores in it and about 15 gigs of RAM. So uh a decent sized box, it's not massive, but um you know
for for what we're doing. uh it'll work well. I also have a RDS Postgres database server. Uh it's a T2 medium. Um I think that is two CPUs and four gigs of RAM, probably a lot less than what you would use on a big production site, but uh for our purposes it'll work well. So Uh I'm gonna use a tool called Fig. Um fig, let's see, is a way to um kind of manage uh Docker containers. So what we're going to do here is we're going to spin up a memcache D instance that is going to handle cache sessions. And then we're going to spin up our web server running run server. Uh
so that looks like this fig, and we'll call our fig file and spin it up. So That creates our web container and our memcache D container, and we're up and running. So I can show you what the site looks like right now. This is another box running in EC2 , and it's going to be what we're throwing to throw all the traffic at it. So This is just a app. I threw in a bunch of fake data. It's user profiles and the users have a couple foreign keys to a company and their job title. And they have a profile.
That profile has a decent sized text blob in it and some links to the people they share a birthday with. It's not a totally trivial Hello World app, but probably not as crazy as what you'd be doing in production. Again, just kind of a demo here. What you'll see in the real world is going to be a little different. Okay, so this is JMeter. So um you may be familiar with uh like Apache Bench. uh A B or Siege, they're they're good at uh kind of blasting uh specific page with lots of traffic. JMeter uh is kind of like those on steroids. You can do uh really complex
uh test plans and um run them against your site. So what what I'm going to do what what's going to happen is we've got this requests object and it's going to go loop through everything that's under it. So first it's going to hit the home page of the site as an anonymous user Next it's going to hit what I'm calling a hot profile page. So in uh real world site, um typically you'll see If you have a site with lots and lots of pages, there's usually a small subset of pages that really get hammered and then kind of a long tail after that that don't see quite as much traffic. So we're going to try to simulate that. So this is going to pick a random profile page between 1 and 50 and hit that Next up we're gonna uh
on 10% of our loops we're gonna log in. So uh that login will um basically will hit the admin login page that'll uh give us a the C Surf token. We'll use that token and credentials to authenticate against the server. And then we'll post to create a new profile. And then we'll hit the home page as an authenticated user. So kind of stimulating a, you know, a site that's got some logged in traffic and those users are doing something. Another 10% are just going to hit a random profile page. So the database has about half a million profiles in it. This is going to kind of simulate that long tail of web traffic. So I'm all set up here to uh hit my web server.
Um I'm gonna do uh 50 concurrent users and we're gonna go through this loop ten times and see what happens. So fire off my test here and we can see our response times. So uh what's happening is uh we started off pretty good. We were below like 200 milliseconds But really quickly, we uh are are getting really bad response times. We're now up to a second and a half. Um that that's not really uh a great uh um response time uh to serve requests out of Django. So basically what's happening is is we're we're overwhelming the server. And if we go back and look This is Htop.
So we can see our, well, you can't quite see, but um, those are run server instances in uh Python there. And we're not really utilizing the whole server. There's you know at most maybe you know 30-40% of this, any of the CPUs are getting used. So um But but requests are queuing up. Basically this is because uh because run server is just a single process. Um as of recently it's multi-threaded, but uh still um we're only using uh a single process. So not really great performance. Basically the server's not fully utilized. I can go back here and uh I'm gonna pull and and we're gonna keep track of the the the data we're seeing here.
So um This is going to be oh shoot. Make this font a little smaller so it fits on the screen. Uh So that's one server. Oops. Uh fifty concurrent connections and our requests per second were twenty-eight point two. And our average response time was 1353.
So about 1. 3 seconds, that's uh that that's not gonna fly um if you if you're on the homepage or Reddit. It's a little too uh too slow. So um Let's go back and uh see what else we have. So normally you're you're not going to put a site into production and run serve with run server. That's uh a bad idea. Usually you're going to use a production uh whiskey server. Um we're gonna use UWISGI today. Uh you might be familiar with GUnicorn, Apache Bod Whiskey. Um any of those uh really uh are gonna get you by the same This is just the one uh we prefer. So I'm gonna go back and kill off uh our run server container. And while that's happening, I'll show you
our UWISGI configuration. So you can't quite see the edge of the screen there, but these are processes. So instead of running one process like run server, we're going to run six processes. And uh again we have um threading on, so uh multi-threaded, multi-threaded, uh six processes. uh we we should expect some better performance here. Um the rest of this stuff is kind of boilerplate you don't need to really uh worry about. So uh I'll show you as well So this is how we are gonna start off uh our UWISGI process here. Fig.
Alright, UWISGI's up and running. And we'll uh erase our previous test results. Um UWSGI's gonna go a little faster, so we'll loop over this 20 times here. And uh start off our tests. So You can see here performance is much better. We were pretty quickly over a second with run server. Here we're you know under half a second on almost all our requests. You can see the authentication takes a little longer. That's uh the password hashing in action. Um we actually want it to take a long time. So Uh that's good. And we can see we're serving a ton of requests.
Let's see how our processor's doing. So uh we're utilizing a lot more of the machine. Uh there we go. Um so it's uh you know it's it's definitely uh we may have just finished the test run there, but uh as you can see we you know we were uh hitting 90% uh CPU usage. And it does look like we finished. So let's see our results. So uh We're gonna go back here, UWISGI, again, 50 concurrent users, and our requests per second are up to 72. 4. So uh 150% better. That's that's a pretty good improvement. Our average response time is down to 279.
So uh much better performance, exact same server. Uh all we did was was swap out run server for basically a real WISGI server. Um Let's see what happens if we take that same server and instead of throwing 50 concurrent connections at it, we throw 100 concurrent connections at it. Uh and I think this is gonna take a little longer, so for uh for short on time I'm gonna bump that down to 10. uh loops through. We'll erase our old results and fire it up again. So we should see here uh we we're pretty much uh maxing out the server Um it's under a lot of load and the and the the load average is just going to keep creeping up here.
And if we look at our response times we can see they're also jumping up. So last time we were hovering around you know half a second, now we're hovering around a second. So in the real world, if you saw this on a server, what you would be saying is we're maxing out this server. If we throw much more load at it, we're going to start, our requests are going to start timing out. uh we're gonna start dropping requests uh so um this this isn't gonna get us uh what we need so anybody know uh what what the next step is What's that? Uh that might work. Anybody else?
Caching? Caching, we could do caching. Another server? There you go. So so this server, you know, uh we maxed it out. Um we we could we could uh get a slightly better optimized uh whiskey server that might buy us a little bit. Um we probably would be a really good idea to look back at our application and see if it's uh you know if there are places we can optimize it with caching and uh improve the situation. But um let's just uh you know throw more money at the problem. Uh you know, one server's not enough, let's try two. So uh that looks like this. We're going to use Nginx as a load balancer and put uh two servers behind it.
Uh instead of Nginx, you can use uh something like HAProxy. uh Amazon ELB uh there there's lots of options here so I'm gonna go back to my web server kill it off I'm gonna bring up another one uh Uh so this is uh web two and web two oh There we go. So Web2 uh looks just the same as the other box. They're identical. This time we're going to bring up uh UWISGI. You notice the other one said UWSGE HTTP. This is uh UWISGI um using the USGI protocol. So uh It saves us a little bit of overhead,
basically converting HTTP into what Nginx wants to use, and then back to HTTP and then down to UWISGI. So um we should get uh a slight slightly better performance by using uh UWISGI's internal protocol, which Nginx can speak. So Here's our load balancer. I'll show you. This is our Nginx configuration. Sorry. Nothing exciting here. These are some settings that are known to boost the performance a little bit, kind of just boilerplate. Here's our server that we've defined. We're going to pass back to a UWISGI cluster that we'll define when the container spins up.
And this include UWISGI params does everything we want it to do. So that looks like Fig Nginx. Okay, so there's our our UWISGI cluster we defined. And let's see if our looks like our web server is running. Okay. So we're going to go back to our test plan here. We're gonna loop over it 20 times and instead of pointing to our web server, we're gonna point to our load balancer, which is called LB.
Oh, I haven't put in our uh our other U Whiskey here. So that we did run it with a hundred and we got uh A little better throughput there, 101 requests per second, but our average uh response time jumped up to 750 milliseconds. So uh yeah we we kind of decided that we overloaded the uh server there, so let's kind of forget about that one. That's not good performance. Um uh overall. So I'm gonna erase this, start up against our load balancer, and we can see we're already doing uh Pretty close to to double the requests here, let's see how
our response time is. Our response times way back down So kind of what we'd expect. You know, we we served a certain amount of requests with one server, we double the servers, and uh we're getting close to double the requests. At the same time, we're probably throwing twice as much traffic at our database, so uh you want to make sure that your database can withstand uh all this extra stuff. Um Yeah, I think there's some kind of funky stuff going on with the network here with these uh those those gaps, but Hopefully we still get a decent result out of this. So that was 100, let's see. So we have Nginx, 100 concurrent users.
And we did four hundred and three average response time and hundred and thirty-seven point four So we pre came pretty close to to doubling uh our our our first option here. The response time that's higher than it should be. I think we kind of had some anomalies. If we ran the test again, I think we would see it's it would be really close to that uh initial um UWISCI instance. uh maybe a little bit of overhead, but um that uh Nginx uh is pretty efficient in proxying. So uh that's that's Nginx, a hundred concurrent uh
users. What if we have two hundred concurrent users? Let's see what happens then. So I'm gonna erase these results, fire it back up And let's see. Our response times are starting to go up. Let's see what our servers look like. So uh this is Web 1, it's pretty maxed out there. This is web two. It's probably also gonna be pretty maxed out. Yeah. And let's take a look at what our load balancer is doing. That's nothing.
So load balancers are super efficient. Really all you need to give Nginx is a big fat network pipe and it can handle lots and lots of traffic on a on a small machine. So Let's see how we're doing overall. Kind of like when we bumped up UWISG uh uh to um you know more than it could handle. Uh we're seeing about the same thing now. I'm I'm guessing our average Response time is probably gonna get close to a second here. And uh request per second, we did better, 166. So a few more requests per second
at two hundred And but our response time was eight hundred and five milliseconds. Um That means uh basically we saw with with very little load um we should expect response times around 200. So if we're at 800, um that means we're we're basically overloading our server, processes are waiting. So uh let's let's strike this one out as well. So next up, um We could keep adding app servers, right? Uh so if two didn't work, then we could add three. And if we can't handle it with three, we can do four. But maybe we can get a little smarter here.
If you keep adding app servers, you're basically pushing the problem down your stack and having load issues. on your database is uh is not fun. Um you you you know you can throw hardware at that for a certain amount of time but uh once you run out of hardware options um that problem gets a lot trickier. So maybe we can get smarter. And uh this is gonna be the last one that we're gonna benchmark. So uh instead of using Nginx as a load balancer, we're gonna use varnish. Varnish uh does the load balancing just like Nginx, uh, but uh it can also do caching. So uh Then when those requests come back from Varnish or come back from our backend
through the load balancer, Varnish can grab a copy of it and and serve that uh to other users. So With varnish, I'm going to bump up the number of times we're going to loop through this here and erase the previous results. And fire it up. So pretty quickly we should see the request per second jumping well above what we were at before. Uh and and let's take a look at our response times. Um our response times are actually a little oh. Okay. I didn't switch to uh to varnish here, so that explains why we're seeing the same thing.
Uh so we're gonna stop their web servers. I'm gonna stop uh Nginx. Varnish uh does not speak the USGI protocol like Nginx does, so uh I'm gonna uh start those up uh again um in the HTTP. uh with the HTTP protocol. So there's our first web server. Here's our second web server. And um Varnish is is really amazing uh and I don't think it gets enough love in the Django community. Uh I don't hear a lot of people um talking about using it, so Let's take a look at at the varnish config and
kind of walk through that really quick so you can see what's happening. Just like Nginx, we're gonna uh define our backends uh when our containers spin up and include that file. Um varnish is uses a configuration language called VCL. It sort of maybe looks like what you would use to configure Nginx, but it's a lot different. So uh you define these functions. Um VCL received is what happens when a new request comes in. So uh it sees a new request and what we tell it is If that request is to the admin URL, or if uh the request has a cookie called session ID We want to bypass the cache. We want that person to always go through to the back end and get fresh content.
They're authenticated or they're trying to access the admin and log in or something like that. If they're not in one of those, we want to unset the cookies. So Varnish will look at a request and uh basically um determine whether or not it's unique. One of the ways it does that is by looking at the URL. Another way is by checking the cookies. So Your Google Analytics package is going to set cookies for a user, and you may have other reasons that there's cookies set for anonymous users. The backend doesn't care about those, so we we wipe all those out. Next up is the VCL hit method. The VCL hit method, that's what happens when it finds something in the cache.
So first uh what we do is we check the TTL, uh the time to live. If it's um still basically uh still valid, uh we're gonna deliver that right from cache. Uh What this next part does is uh let's see here we go. Uh this next part um we can also define a grace period on our cache. So there's there's basically two timeout values. One we if we're in the first timeout, we just deliver it. If we had passed the first timeout but are still within the second grace period, what we're going to do is serve that stale content to the user. And fetch new content from the back end in the background. So the user doesn't have to wait for Django to return the response, but any future users are going to get
a new copy of that data. or that page. So that's really nice. It can prevent uh if you have a really hot page like you're on the front page of Reddit and and you have one page that's just getting hammered and hammered and hammered and then your cache expires What's going to happen is you're going to have a hundred users all flood through to your back end and require request that same page before the cache refreshes, refreshes. So that's uh they might call it a cache stampede or dog piling. So this is basically protection against that. And then the the the last um method here is where we actually said uh the DCL backend response. We're going to set the grace period and the time to live on the requests that are coming back out.
Varnish also respects cache headers, so this is something you can define in Django. But for for our purposes, we're just going to keep it simple. And we're setting a five-second time to live and a five-minute grace period in production depending on the you know type of site you have, you you maybe could run those uh much higher. Um but when you're on uh you know when you're on Reddit um even having those set really low, uh five seconds might be enough um for you to to withstand that. Basically that means If you if you've got a hundred users, a hundred concurrent users sustainably hitting your site, only one request every five seconds is going back to your backend. So that can be a huge load off of
your servers. So all right, now let's uh spin up varnish. Whoops. Up uh Uh I did switch those over. Okay. So let's try this again. Erase our previous results. That's what I'm expecting. Okay, so if you can see here uh our throughput skyrocketed. We're uh 550 requests per second. Uh and if you look at our average response time, it it's uh 178 milliseconds. So and this is the this is the the really interesting um one
to me is uh if you look at the average and median response times on our hot pages, uh two milliseconds. So you're you never ever will get Django running this fast. Um Varnish uh is is very fast, you know it's a cache, uh so um it does what you'd expect. No matter how much optimization you do in your code, your database, anything, it's never going to do this. You can have the full page caching on. So this is where Varnish really shines. So uh that test is already done. Let's add that to our list here. So we're running varnish, uh 200 concurrent requests. And we did four hundred and fifty-six requests per second.
Oops. And our average response time was two hundred and three. So uh compared to Nginx running on the exact same servers, uh we we more than doubled uh our the request per second we're handling and we cut our response time in half. Uh all the exact same hardware, all we did was change uh change Nginx with varnish. So that's a huge win. Um let's see what uh varnish can do if we were to throw 400 concurrent users at it. So this is a lot of traffic. This is 400 people simultaneously uh hitting your site. Most sites will never see this much traffic.
So we're going to run that again. And then uh while that's running, let's take a look at what our servers are doing. So this is web two, this is web one. Uh load is A little bit lower you'll see than than what it has been in the past. You know, before when we were overloading the servers, we were, you know, really just uh totally spiked. Um it's it's They're getting used pretty heavily, but uh still not a ton. It looks like we might have we finished already Yep, uh so we're already done. Uh that's uh we'll add that to our list here varnish uh that was 400
oh wait I think I we might have had a uh another one of these kind of blips in the network. Yeah That's not normal, that doesn't uh happen. This is why this is a terrible benchmark. Um usually that would be smooth, so uh just pretend you don't see that giant spike there. Anyhow, oh and then while that's running, we can also look at uh varnish. Just like Nginx, it's it's not even breaking a sweat. Um you put a bunch of RAM in your varnish machine and you make sure it's got lots of network uh capability and uh it'll handle a ton of traffic. So looks like it's wrapping up now.
Uh So that was 400. These results aren't great, but we'll add them in. I've run this like testing it a million times, uh, so just take my word on it that uh Normally it's about the same as far as the uh response time and slightly better request per second. So three sixty wait a minute, four seventy-six. I got these backwards here. And three is sixty-four point eight. So uh while we have this up, let's um let's look at some other kind of cool things about Varnish.
Uh I've got a few minutes left here. So I'm gonna Go in and uh instead of looping over this 30 times, we're gonna loop over it a hundred times uh to give me some time to show you what's happening on the servers. So I'm gonna start off our tests here and I'm gonna go back to our load balancer. And this will let me go into our varnish server, uh our varnish container that's running. So one kind of cool thing uh varnish comes with is this varnish hist varnish histogram. Uh this is showing us uh in real time the requests that are hitting the server. The dashes or the pipes, the vertical lines are ones that are hitting the cache.
Uh you see that 1e negative 5, that's 1 millisecond. Uh and then the the hash marks are the ones that are uh cache misses and hitting our back end. So those are getting returned around in the neighborhood of a second. It also has varnish uh top. Uh no, that's not what I want. Varnish stat. Uh so here you can see hitness ratios, uh number of uh connections and all that. So A big performance win is basically just uh letting Varnish serve more of your cash and you can track that uh you know it does a really good job of showing you um what your uh hit and miss ratios are.
And this is the really awesome thing. So what I'm going to do now is I'm going to kill web two there, and I'm going to kill web one. And let's see what's happening. So take a look at this. Our hot pages are not errors. We're still serving content on all those pages. So You've totally screwed up. You've deployed a massive breaking change to your live servers in the middle of being on the front page of Reddit. And Varnish is still chugging along serving your content. 10% of your users are getting errors. Yeah, that's bad, but uh you know, your important content is still up and running. So you you know you can see the the error ratio shooting back up on all these, but uh
Varnish is still serving our homepage and those 50 profile pages in two milliseconds. So uh You know, we can spin back up our web servers and Varnish is going to reconnect to them and uh you know basically fix all that stuff. So Uh use varnish, it's really great. That's that's the uh lesson of this talk. So we did about 450 requests per second with Varnish. If you were to uh you know do that sustained for a day, that's 40 million requests in a day. You know, th there's people that do a lot more than that uh in a day, but they do it a lot on a lot more than three servers.
And uh, you know, if you were to be on the front page of Reddit or something like that, uh that'll get you by just fine uh you could probably do it on a lot less. Uh I wouldn't be surprised if you could do it on one server running uh Varnish and and your application. So um pretty good results from from where we started uh with uh hang on so we started There we go. With run server at 28 requests per second. So like I said, Django doesn't scale. Don't use run server in production. Varnish scales. Use varnish. And uh yeah, that's all I have. So uh like I said, we wrote a book. Uh it's called High Performance Django.
Um you can check it out at High Performance Django I also have uh a few copies here that are loaded on these nifty little USB keys. Um if anybody wants to ask questions, I'm uh you can win a copy. Thank you.
Under load, runserver uses only a single process, so requests queue up even when the machine still has unused CPU capacity. In the demo, 50 concurrent users pushed average response time to about 1.35 seconds.
Discussed at 9:32Replacing runserver with a multi-process, multi-threaded uWSGI setup raised throughput from about 28 to 72 requests per second and reduced average response time from roughly 1.35 seconds to 279 milliseconds on the same server.
Discussed at 13:35Put a reverse proxy such as Nginx or HAProxy in front of multiple uWSGI application servers and load-balance requests between them. This nearly doubled throughput in the demo, although the database must also handle the additional load.
Discussed at 16:00Varnish can serve stale content during a grace period while fetching a fresh copy in the background, so many users do not simultaneously overwhelm the Django backend when a hot page’s cache expires.
Discussed at 27:48Varnish caches complete responses and serves cacheable pages without sending every request to Django or the database. In the demo it increased throughput to roughly 450–550 requests per second, with hot cached pages responding in about two milliseconds.
Discussed at 30:08Yes. Cached pages can continue to be served even after all backend servers fail, so cached home and profile pages remained available in the demo while uncached requests produced errors.
Discussed at 36:24Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026