Just enough ops for developers with Peter Baumgartner

This video features Peter Baumgartner at DjangoCon US 2022 in San Diego, California, USA.

Just enough ops for developers with Peter Baumgartner
0:43:18
Published November 3, 2022
613 views

Most developers don't want to think about operations (aka Ops, DevOps…). PaaS providers (Heroku, Fly, Render, etc.) do an awesome job of getting your app live on the internet, but there's a lot to ops beyond just deployment.

This talk will answer the following questions:

  • How much CPU and memory should I give my services?
  • How do I know if the app is overallocated (costing too much money) or under-allocated (slow/overloaded)?
  • How can I make my app faster?
  • Will autoscaling help me save costs?
  • What about serverless?

If you're a developer and want to run your applications successfully without deep DevOps knowledge, this talk is for you. It will help if you have some basic Django developer experience, but other than that, no specific knowledge is necessary!

This talk was presented at: https://2022.djangocon.us/talks/just-enough-ops-for-developers/

LINKS:
Follow Peter Baumgartner 👇
On Twitter: https://twitter.com/ipmb

Follow DjangCon US 👇
https://twitter.com/djangocon

Follow DEFNA 👇
https://twitter.com/defnado
https://www.defna.org/

Summary

Peter Baumgartner explains the minimum operations knowledge Django developers need when deploying through platforms such as Heroku, Fly.io, or Render and managed database services. He covers how CPU, memory, disk, and I/O affect performance, cost, and failure modes; how to configure workers, avoid excessive memory use, handle files and queries, and scale applications and databases. He argues that managed services remove much of the operational burden, but developers still need to understand resource usage and monitor CPU, memory, latency, uptime, and error rates. He also compares horizontal and vertical scaling with serverless, including cold starts, unpredictable costs, database connection limits, and platform constraints.

Key takeaways

  • Use platform-as-a-service and managed databases to avoid unnecessary infrastructure work, but understand the resources your application consumes.
  • Disk access is slow and often ephemeral, so store user uploads in object storage such as Amazon S3 rather than on the application server.
  • CPU capacity depends mainly on concurrent work and request duration; improving application performance can be cheaper and more effective than adding CPUs.
  • Each Gunicorn worker consumes memory, and exceeding the memory limit can kill the application, so monitor usage and avoid loading large files or querysets entirely into memory.
  • Stateless applications are easy to scale horizontally, while databases are usually scaled vertically until read replicas or more complex approaches become necessary.
  • Autoscaling and serverless handle some traffic patterns well but introduce issues such as startup delays, uncertain costs, database connection overload, request limits, and the loss of shell access.
  • Monitor more than a homepage or average latency: track resource saturation, high-percentile response times, uptime, and application and infrastructure-level errors.

Summarised automatically from the transcript.

Chapters

  1. 0:00 Just Enough Ops Introduction to the talk’s pragmatic approach to operations for developers using managed platforms and services.
  2. 2:44 Compute Resources Overview of disks, CPU, memory, and why understanding basic server resources matters when choosing a hosting plan.
  3. 4:17 Persistent File Storage Discussion of ephemeral disks, deployment behavior, and using object storage such as Amazon S3 for uploads.
  4. 5:55 CPU and Web Processes Explanation of CPU usage, concurrent requests, Gunicorn workers, blocking, and application performance.
  5. 10:39 I/O and Asynchronous Work How database queries, network calls, and file access block Python applications, plus the tradeoffs of async programming and gevent.
  6. 16:19 CPU Capacity and Backlogs How to recognize CPU saturation, request backlogs, timeouts, and when to optimize or add CPU.
  7. 17:06 Memory Management How Django processes consume memory, how to avoid excessive allocations, and how to diagnose memory leaks and out-of-memory kills.
  8. 24:56 Database Resources Applying the same performance fundamentals to databases, with an emphasis on managed services and keeping frequently used data in RAM.
  9. 27:17 Horizontal and Vertical Scaling Comparison of scaling out application instances versus scaling up resources, including stateless apps, databases, and autoscaling.
  10. 31:59 Serverless Tradeoffs How serverless changes scaling and pricing, along with cold starts, database connections, request limits, and other constraints.
  11. 37:22 Observability Recommendations for monitoring CPU, memory, response times, uptime, and error rates in production.
  12. 39:49 Questions Audience questions about process scheduling, load testing tools, and open-source monitoring options.

Transcript

7,310 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:20

Speaker 1: Thank you. Um, thanks, Brent. Really quick about me, I'm the founder at Lincoln Loop. We are a full-stack Django web agency, so we help clients uh build sites in Django, uh fix Django problems, all that stuff. A few years ago I co-authored High Performance Django, which is all about scaling large Django sites and currently building AppPack, which is like Heroku for your own AWS account. So it simplifies a lot of setting all the resources and doing the AWS type things. And it was kind of the motivation for this talk. As we brought developers on uh to AtPack to use it to deploy their own sites, we found um there were actually a lot of kind of gaps in knowledge even when you take away all the traditional ops

1:10

Speaker 1: type things. So That's what we're going to talk about today. A good preface to this talk is my talk from the last in-person DjangoCon here. uh which is prepping your project for production. Uh it's on YouTube and uh yeah, it's online, you can find it And it is all about the things you should do in your code base to make it ready to move off of your laptop and to run out on the internet somewhere. Um but we're not going to talk about that today. Today we're going to talk about uh just enough ops. So the reason it's just enough is what I've found is developers typically fall into two groups when it comes to ops. They're either really interested in it, they want to learn more, they want to dive in with both, you know, jump in with both feet,

1:58

Speaker 1: or they don't want to ever think about it, they hate it, they want to touch as little of it as possible This talk is for the second group. So we're going to try to do our best to kind of skim over things and distill them down to the things that are really important. The reason we can get away with this is because uh today platform as a service and managed services are really good. If you're in that second group, This is where you should be deploying your sites. Don't go get a droplet on DigitalOcean and try to figure out how to set it up yourself. Uh you're gonna end up doing a lot more work. So uh what I'm talking about here are things like Heroku, Fly. io, render, uh sponsors platform here have a service around that.

2:44

Speaker 1: So um and and databases too. Uh if they your provider there doesn't provide one, there's um crunchy data and CITUS and all these people that will provide great databases for you. So use those. Um they're a huge help. When you use those, you get to forget about a massive amount of problems you would normally have to think about with ops. So you know hardening systems, making sure they're secure, uh how the heck do you do deployments and get your code up? secrets management, all this stuff. Forget about it. You don't have to worry about it. But you do have to understand some basics about how your site gets out on the internet and to do that you really have to understand kind of the fundamentals of a computer.

3:30

Speaker 1: And so today we're going to talk about CPU, RAM, and I. O. among a few other things. And the reason for this is this is like the first decision you have to make when you go to one of these service providers. How much CPU and how much RAM do you want to give to your service? This is flat. io 's pricing page. This is Railways pricing page. So that's generally how you how much your app costs to run is based on the answers to these, you know, how much CPU and RAM you use. So we're going to go back and do some computer science 101 here. This is a server. It's a one-use server, and these things down here are the disks. And the green things up there are

4:17

Speaker 1: the memory and the two gray boxes are are the CPUs. You don't ever have to know that, but let's fun to look at a server every once in a while. So we're going to talk about those three things. So disks, you may also hear them referred to as HDD or SSD. That's a hard disk drive or solid state drive. This is your persistent file storage on the server. It holds the operating system, it holds your application code, it holds all your dependencies. Disk access is super slow compared to the rest of the server. So generally you want to avoid accessing the disk for any kind of critical path in your application. And it's usually ephemeral, which means when you do a new deployment, you're often going to end up with a totally new disk that doesn't have any of the stuff on it from the disk you were previously

5:07

Speaker 1: operating on. There are services that kind of hook up a network file system for you where you actually can maintain files on that disk. I would say in general, try to avoid that stuff. It's just going to complicate things and it's going to limit your abil of your possibilities of kind of where you can deploy your app. When you need to handle file uploads from a user, they should go into something like Amazon S3 or any other file storage. Any cloud provider is going to offer you something like that. Next up, uh so really that's all we're going to talk about with disks. Like generally ignore them. If you need to use it for a temporary file or something, that's fine. But uh otherwise I would recommend avoiding it. Next up is the CPU, also known as the processor.

5:55

Speaker 1: This is the thing that's actually responsible for code execution. One of the first questions you're going to get asked is how many CPUs do you need to run your app? And unfortunately, the answer is it depends, as with most things in software engineering, right? So If you look at the pipeline of how an application goes through the system, uh or sorry, how a web request goes through the system, you're gonna have Uh that web request come in, your Python application is running, maybe you're using GUnicorn or something like that for an application server, and it's going to hit that. And that's going to need to do some computation, and that computation happens on the CPU. So this is kind of the simplest way you can set up an application.

6:42

Speaker 1: And the FastAPI folks have really great analogy on their site of kind of how to think about this in real world terms. Uh if you think of the CPU as the cook out there flipping burgers, uh the cashier is taking orders, and the customer is the request coming in. So I said it depends how many CPUs you need. Uh just like it depends how many cooks you need. If somebody came to you and said, hey, I have a restaurant, how many cooks do I need? You'd probably start asking a lot of questions. How long does it take to cook an order? How many customers are you expecting? The general rule of thumb is the more users and specifically concurrent users, users coming in at roughly the same time, the more CPU you're going to need. So if we take our simple example of

7:28

Speaker 1: one process and one CPU, if you have two requests come in, one of those requests is going to sit there and wait in line. So this is called blocking. Blocking is generally bad. It slows the system down and uh users aren't happy about that. So a naive approach here is just add more CPUs, right? Add more processes, add more CPUs, and we can handle all these requests. That's great, but uh the CPU is one of the pricing things on the scale here. So if you add a CPU for every concurrent user, yeah, your app's gonna probably be fast, but uh it's gonna be really expensive to run. So you want to minimize your costs, and to minimize your costs, you want to maximize your CPU usage.

8:15

Speaker 1: The first way we can do that is via multiple processes per CPU. So instead of running just one Python process, we can run two Python processes. So this in uh in our analogy we've got two cashiers now taking orders. This is great. We get to uh we get to benefit from from both these processes handling requests coming in. But in reality, you're going to have always one request is going to be waiting. The CPU can only do one thing at a time. So it's going to be basically interleaving these requests, handling some of one, handling some of the other. And this works because as we're going to talk about in a little bit, the CPU isn't necessarily always doing work.

9:02

Speaker 1: when a process is basically being executed by the CPU. So uh you can get a little more performance by running two of these. If you're using GUnicorn, this is the worker's flag, so uh you can just set that and that will set the number of processes that it creates to serve these requests. The general rule of thumb for this is to use double your CPU count. And it's not a hard, fast rule. Sometimes you'll get away with you 'll have better performance with more of those. Sometimes you might even need to run fewer. It really depends on kind of how much uh how much CPU work your each request is doing. Um you'll know when you've added too many uh workers when either you run out of memory, which we'll talk about in a second, or

9:52

Speaker 1: uh your performance starts to degrade. And this can be a tricky one because you really don't know until your application is under load. When you've got lots and lots of users coming in, that's kind of the only way to check this. So this is why people do load testing. There's tools out there for that, but that's beyond the scope of this. The biggest one that I harp on with folks is the best way to get more out of your CPUs is to improve the performance of your application. So if you look at the restaurant analogy, if everything you're cooking, every meal you cook takes an hour to cook, you're never going to get a meal out faster than an hour. So if you've shortened that to five minutes, you know, you can do a lot more work with just one cook.

10:39

Speaker 1: So here's kind of what that looks like in terms of web requests. If you have uh 100 millisecond response time, one CPU can process 6,000 requests in a minute, roughly. If you have one second response time, that goes down to just 60 requests. So a big takeaway here, slow responses mean you're using more CPU time. And it means it's a higher cost for you and your users probably aren't going to be as happy either because they're going to be waiting longer. So the next way we can optimize CPU, we kind of have to dive back into our computer science 101 and talk about I. O. So I. O. is input and output. And essentially it's anytime you are accessing a file or doing something over the network on your system.

11:29

Speaker 1: So You're waiting for some external thing to give you data or to accept data from you. Database queries fall into this, third-party API calls, reading files, all that falls into I. O. Everything stops and waits for I. O. and kind of the default Python uh setup. So your Python process stops. Your CPU process stops, your request backlog stops, everything just sits there and waits for this response from the network or whatever it is In the analogy, this would be like if you went into a restaurant, place your order, and the cashier sends it back to the cook And the cook sits there and stands there watching the burger cook for five minutes. And you sit there and stand in front of the cashier and don't let anybody else come to the cashier until you get your food back.

12:19

Speaker 1: So uh obviously not very efficient. The techie way to say this is synchronous I. O. is blocking. But synchronous I. O. isn't the only option we have. We also have asynchronous I. O. Asynchronous I. O. is you probably heard about. There's async IO in Python. In recent versions of Django, we can now use some of these async tools. Uh they're really great for some use cases, but in general, I would say asynchronous programming is hard. Like um People might tell you it's easy, but it's it's 100% more complicated than just looking at a program and reading it line by line. So it's certainly doable. In some use cases it makes a lot of sense, but be careful about kind of going too too deep into this rabbit hole

13:09

Speaker 1: because uh hardware is cheap and programmer time is generally expensive. So one of the reasons we use Python and not Rust or Go or C to write web applications is because we have great productivity with it. Sure, we could use something that's faster, but uh it's not worth the extra cost of our time. So keep that in mind as you kind of go through this optimization journey. The next way you can maximize CPU and uh kind of utilize asynchronous I. O. is via G-Event. So G-Event is it's magic. It goes through your synchronous program and it will monkey

13:54

Speaker 1: patch all the places where you Access files or access the network and turn those synchronous calls into asynchronous calls. Like most things, magic, when it works, it's amazing. When it doesn't work, it often fails in like really spectacular ways where you might have one request getting data from a different request and potentially you know different customers reading different things. If you've programmed your application really well, that doesn't happen. If you've accidentally set a global variable somewhere, that's really easy to do and a tough one to debug. When you're running gunicorn with gevent. your application is going to look more like this. So you've got your, you still have your Python processes, but each of those Python processes can spawn

14:43

Speaker 1: what are called greenlets. They're kind of like threads in programming. And they're these really lightweight, um, it's called coroutines that can uh they know when to suspend themselves. to when they're actually doing some sort of I. O. and another uh process can come in and do real CPU execution. So uh this can be a great way to kind of remove that async or that I. O. stall that you might have in a CPU without going through the depths of uh actually making your whole program async. When you use this, these are the flags you would use with GUnicorn, worker class G event, and then by default it will spawn up to a thousand of those greenlets

15:29

Speaker 1: per process. If you have four processes running and you run a thousand greenlets on every process, you're going to have like 4,000 connections to your database and your database is probably going to fall over trying to handle all those connections. So a good thing to do is to limit that so you're not spawning thousands and thousands of database connections. The other thing people do here is uh use a database pooler, database connection pooler, but um this is probably enough for you. So if we optimize and we we're maxing out using as much CPU as we can and kind of saving our costs that way, the the flip side is what it looks like when you don't have enough CPU. So any of these providers you use are going to probably give you some nice dashboard that's going to tell you how much CPU you're using

16:19

Speaker 1: for your application. Either of these graphs is probably cause for concern. Obviously the one on the left, sorry, the one on the right is you know it's 100% maxed out all the time. That's a problem. But even if you're hitting peaks at 100% that and you there are other requests coming in, that probably means those requests are waiting. Could mean that those requests are timing out So this is a good sign that uh you might need to either optimize your application more or use less or use more CPU. This is what it looks like when those those requests start backing up. Essentially in the restaurant, your your line's getting longer and longer. Eventually, when you kind of hit this scenario, you'll start seeing these types of error codes.

17:06

Speaker 1: This means either your application is accepting requests but not responding in time, or your application has crashed completely and isn't even accepting requests. or you know kind of some something in between there. What you get depends a little bit on the uh proxy your provider's running in front. So that's all about CPU. Next up's memory. Memory's a little bit easier. So memory is an ephemeral cache It is when your application starts up, all the code in your application gets loaded into memory, your dependencies get loaded into memory. Any variables or Python objects that you're creating during the course of your application running. are going into memory. Um if you're doing any sort of file processing, if you're reading a a file that is all going into memory.

17:57

Speaker 1: And then there's pointers to kind of external things you're working with, like open files or network connections, and those are tracked in memory. Memory is really fast, but it's limited. Disks you usually have lots and lots of space. Memory is another thing you you pay for it. And so generally you want to use as little memory as possible to save on costs. Memory usage for a typical Django app, I would say, is probably between 128 and 512 megs. I've seen things outside of this, but uh I'd say on average this is what you're seeing. And keep in mind that's per process. So if you're running GUInicorn with four workers, that's going to be 512 megs of memory per process. So your baseline is two

18:42

Speaker 1: gigabytes of memory. Unlike CPU, if you exceed the memory you have available, your application is just going to shut off. CPU, you can be at 100% memory and it's going to be sl or sorry, 100% CPU. And it's gonna be slower, things are gonna wait, but they're gonna keep processing. They can kind of queue up. Memory, uh generally if you exceed it, uh you're just gonna get the axe. And this is called um on Linux machine you'll see OOM killer. So there's a process that actually goes through and kills processes that are kind of exceeding that memory limit. If you hit that situation where your process is dying and you look at your graphs that your

19:28

Speaker 1: provider is offering you, and you can see, wow, yeah, I'm using 100% CPU, there's really just two options here. You increase the memory that your provider is giving you or you reduce the number of workers you're using. Either one of those, you basically raise the ceiling or lower your usage. When you look at those graphs, your memory usage should be stable. Should in a perfect world it looks something like this. It's like flat. But this is an application that isn't really doing much processing at all. In reality, it might look more like this, where you've got some peaks, but the the thing that I like on this graph is it keeps coming back to that same baseline over time. time. So there's some processing happing, some things are getting loaded in memory, but they're all getting freed up.

20:17

Speaker 1: I this graph is a little spikier than I would like to see. I don't mind the bumps, but the the spikes here are a little bit of cause for concern, but you can see we're still not hitting that 100% mark, so it's okay. With Django , a few tips on memory and something to keep in mind here, your laptop is probably incredibly faster than anything you're going to deploy your application to. And it has way more memory. So I hear a lot like, well, it worked fine on my laptop. Like, great. Yeah, it works fine on your laptop, but your laptop would cost like $5,000 a month on a cloud somewhere. So keep that in mind. So one thing is if you're working with files and you're working with big files, don't read the entire file into a string or a bytes

21:05

Speaker 1: object. That whole file is now in memory. So if you have a two gig file and you read it into some Python object, now you're using two extra gigs of memory. And then the other one you see pretty frequently is really big query sets. So Django model instances actually use a decent amount of memory when they are instantiated. And when you evaluate a Django query set, it's going to create a model instance for every object there. So if you're creating a thousand model instances, that's going to be a decent bump in the amount of memory you're using. Using an iterator instead of basically appending the iterator method to your query set will help prevent um creating all those objects in memory. Another option is using values.

21:50

Speaker 1: If you don't actually need the model and all its methods and everything in it, you can just pull out a couple fields from that model and that will use significantly less memory. And then if you have uh really big text blobs in your database like json or whatever uh you can use dot only method to only pull certain fields out uh there's dot there's dot defer as well Uh so um consider that uh if you've got a really you know an object with text that you're not gonna need, uh you can use that to avoid pulling that into memory. So the the worst kind of memory looks like this. This is a memory leak. Memory keeps going up and up and up. Eventually it'll hit 100%. It'll come crashing down because your process gets killed, and then it'll just start climbing the mountain again like the price is right

22:37

Speaker 1: guy. So Python has a garbage collector in it, and the garbage collector's job is to free up any memory that's not being used. So when you enter a function, you create some variables or Python objects. When you leave that function, Python knows, okay, this memory is okay to free. This is one of the great things about Python. You don't have to think about manually managing your memory. As soon as you drop down into a C extension, which might be your own, might be uh something you're pulling in from a dependency, all that goes out the window. It's risk it's responsible for managing its own memory. So uh Kind of a C extension with a bug in it can be a source of a memory leak. Opening files and not closing them can cause memory leaks because you're keeping all those pointers to the files open.

23:24

Speaker 1: And then uh global objects, which generally we're not doing in Python or in Django if we're kind of doing it right, but sometimes you can accidentally create these if you're using a good Python linter like Flake 8. it it probably is going to catch errors like that where you're accidentally creating a global object and writing to it. But that that can cause a memory leak as well. There are memory profilers out there to debug this, but it tends to be a kind of a challenging problem, and it's more common than you might think. So most application servers actually have a way you can just say, hey, after you run a bunch of requests, just restart on your own, because I know you're going to run out of memory eventually. So uh the in Geunicorn this is max requests. Uh

24:09

Speaker 1: one thing you have to be careful with this is you can essentially duplicate the scenario that the memory killer is gonna do where uh All your all your workers, all your processes are gonna hit that max request number at about the same time because they all started about the same time and they're all getting requests at about the same rate. So uh the jitter here will put a random amount on the number that it kills those uh processes off and will make it kind of stagger over time. You want this to be a big number, so like in this case, this is over, I think, a couple of weeks. So I don't want to be killing processes off every hour. That actually takes work. The process needs to get loaded back into memory. So, you know, I've I in this scenario I would have lots of time.

24:56

Speaker 1: If I if I if I uh rotated those workers every couple of days, that would be more than enough to stop this scenario So that's a lot so far. Um we just talked about the application and now we have to talk about the database. The good news is We've hit the fundamentals and they're basically the same when it comes to a database if you kind of squint at it the right way. So um Instead of a web request coming in, we have a database query coming in. Instead of your application server handling it, the database server is going to handle it and it's going to hit the CPU just like your application. One good thing about databases is they're tuned like so well.

25:41

Speaker 1: Like they are just like amazing feats of engineering in general. So you, if you're using a managed provider, somebody that says, you know, sells you a Postgres database or a MySQL database, and they're doing more than just like spinning up a empty Docker container for you, they're going to have this kind of tuned out of the box to the point where you probably don't have to think about it for quite a while. So you don't need to worry about, you know, how many workers am I running or, you know, uh types of IO or anything like that. One thing you do think about with databases that you don't think about with applications is your database should fit in RAM. So uh with any database you use there are ways to go in and look at how much uh Space on the disc those tables all take, and you should have

26:30

Speaker 1: about as much memory uh as those those tables take on the disk. And the reason for this is just like we said before, the disk is super slow. Like disk is like lava. Don't touch it. So you want it to all operate in RAM, and then that's fast. And again, this is a rule of thumb. There are times where you may have some tables that aren't heavily used and you don't need those in memory. And the database manages what's in memory and what's not. So stuff that isn't accessed isn't going to get loaded into memory. And there may be times where application performance isn't super important and it's okay if you have to wait a little bit extra for a query to come off the disk. Scaling is the next topic. So we've optimized our CPU.

27:17

Speaker 1: We're you know running great and we're getting more users, more web requests to our application. and now we're starting to hit you know 100% CPU usage. So to do this you use scaling. Instead of running on one CPU, we're going to run on two CPUs or 10 CPUs or 100 CPUs, whatever There's two flavors of scaling. One is horizontal and the other is vertical. So when you sign up for your platform as a service provider, You're gonna run your application, your application's gonna run in something, and they're gonna call it probably something different all the way across the board. They're gonna call it a container or a dyno or a virtual machine or your firecracker virtual machine. or your Lambda, but it it it's all like an instance of this thing running and you've got resources dedicated to that instance.

28:07

Speaker 1: And scaling horizontally means more of those instances running. So you're just kind of stamping out more of each one. Scaling vertically means I'm going to keep my one instance and I'm just going to keep adding more resources to that one instance. With our application, since we have kind of removed the disk from the picture like I talked about before, we're not storing files there. Our application is completely stateless. And that's a good thing. That means scaling horizontally is really easy for us to do. We don't have to worry about sharing some sort of state between all those application servers. They can happily run independently of each other. Your database, on the other hand, is stateful. It has all your data in it. So

28:53

Speaker 1: that is one where you generally are scaling that vertically. And you can scale those vertically pretty far. It's uh not often that you see folks having, you know, requiring multiple databases to serve traffic. So generally that's my recommendation is if if your CPU is using 100% and you've optimized your queries, just keep adding CPUs. At some point you're gonna see a Oh wow, the the limit's not that far away. Uh and if you get to that point, then you can look at scaling horizontally. That usually means read replicas. Uh it might mean sharding. If somebody talks about sharding, figure out some way that you don't need sharding. It's like a 99 or you know 0. 001% problem. And sometimes people like to jump to it, but it's

29:39

Speaker 1: it's hard and try to avoid it if you can. So some of these providers are going to offer auto-scaling. Autoscaling is really great. So if you have traffic that looks like this, where say most of your users are in the United States and they mostly use your application during business hours in the United States. You're going to have this kind of roller coaster of CPU usage for your application. It's going to have predictable peaks and predictable valleys. And you can turn on auto scaling and generally these are operating based on your CPU usage And as your traffic ramps up, your CPU usage is going to ramp up, and the autoscaler is going to kick in and say, hey,

30:25

Speaker 1: we need to scale horizontally. We need more instances running. And it will stamp some of those out. And as the traffic keeps rising, it will get distributed to those new instances. It'll run at that higher point. at the peak and then as it scales back down it's going to tear down those instances. So uh you're only paying for the extra resources during your peak hours when you need it and not all the time. That's really great. But auto-scaling is not magic. So if you have a traffic pattern that looks like this, where say, hey, at 8 a. m. we're gonna start a class and everybody's gonna log in simultaneously. uh and we're gonna go from zero users to two thousand users. Um autoscaling does not handle these types of situations well because

31:13

Speaker 1: It needs to detect that CPU usage has risen. It needs to spin up new instances. That might take a few seconds. It might take a couple minutes, depending on the provider you use. And it has none of the time to do this in this scenario because it's just jumped up immediately. And there are things that we can do here. One is if if you have predictable traffic like this and you know it's coming. You often you can schedule scaling events. So say, yeah, I I know a bunch of people are going to log in at 8 a. m. Spin my servers up at that time. The other option is uh you just run at peak capacity and you know you're gonna be over provisioned during the off hours. Like I said, uh in the grand scheme of things, hardware is cheaper than programmer time.

31:59

Speaker 1: It might not be the end of the world to have those extra resources out there. There is one thing that can almost handle this kind of traffic and that's serverless. So serverless is kind of a totally different paradigm here. Instead of us running these containers or virtual machines or whatever that our platform as a service provider gives us. We're basically delivering our code to them and we're saying you pass these web requests to our code. You don't have to run a server or anything like that. Server in the sense of an application server like GUInicorn. The request just goes right into your code and it gets processed. And we're kind of back to this model where we have one CPU, one application, and uh you know

32:45

Speaker 1: one web request. For serverless, the pricing is per request, and then per milliseconds, generally per milliseconds, you're responding to that request or processing that request. So it's definitely different than just saying, hey, I want two CPUs and four gigs of RAM for a month. And I know exactly what that cost is going to be. This one's a little more challenging to try to figure out. With serverless, you get to basically forget about having to deal with scaling. Um a serverless application can go from running zero uh instances of your application to a thousand instances of your application in

33:30

Speaker 1: probably a couple of seconds. You don't have to think about CPU allocation. Generally, you're just going to be running one request at a time, and it's going to kind of scale that out to CPUs as needed. Um so you don't have to think about that. And like I said, you don't have to think about you know Gavin or UWISGI or any application server you're using. But uh serverless is not a magic bullet. So um yeah you get to forget about some stuff, but you get to worry about some new stuff So one is costs. If you have a hobby app or some personal website and you don't end up on the front page of Hacker News by accident, you don't have to think about it. Like it's probably free to run. run. But if you're running a real world application that's getting lots of traffic, good luck figuring out what your costs are going to be.

34:19

Speaker 1: I I don't I don't know how to do this well. Like you can probably maybe get some rough estimates, but until you're actually running it, it's um it's challenging. Cold starts are a problem on serverless. So a cold start is just like that ramp up time we talked about where uh When you s when it when you're auto-scaling would start a new instance, it might take a couple minutes. With serverless, it happens much faster, but it might take Five seconds could take ten seconds. Depends on a lot of things, the size of your application, how long it takes your application to kind of start up from a uh state where it's not running. Um if you know it does it's not much, but if you have users that have to wait 15 seconds for your their application or for the to get a web request.

35:04

Speaker 1: uh responded to that's generally a problem. It's not uncommon to see these just time out completely on the first request. So there are ways to deal with this, but definitely something that you have to take into consideration. There's also there's been improvements in this area, but it's it's still something that exists out there. Next up is database connections, just like when we were talking about G Unicorn and running all these greenlets. Scaling your application to thousands of instances and then having all those instances try to connect to your database at the same time is often going to topple your database over. So often you need to put some sort of proxy in between uh your serverless application and your database that can um like pgpooler or most of the

35:51

Speaker 1: you know amazon or google will probably offer something like this for you uh but you still have to plug it in and configure it and all that There are limitations with serverless. So the size of a request is generally a limitation and the maximum duration, the uh any kind of request can take is a is a limitation that you don't have with with the other options. Upload size is often, you know, could be smaller than how your hand how large the uploads you're handling are right now. Um so you might have to do you know uploads via JavaScript or something that don't go through the server, which are great. Like they save your server CPU time and that I. O. time. But Uh probably extra code you're gonna have to write. Um

36:36

Speaker 1: the max duration is generally not a problem for web requests. On AWS it's 15 minutes, but if you have background processing you need to do and those background processes take more than 15 minutes. You have to solve for that. You split up your process or you might need to run it on some other infrastructure. Keep that in mind. And then finally, remote shell access. So most of these providers are going to give you a way. where you can you know SSH uh into some container instance somewhere and operate with your remote in your remote environment, the database and all that stuff. Generally on serverless you don't have something like that. So if that's important to you, consider that and see if there's any options there.

37:22

Speaker 1: So that's it. Final thoughts. Observability, which is the fancy word for knowing what your application is doing. is really important here. Like that if you kind of take away one thing here, you can't just toss your app onto Heroku or Fly or one of these providers and just be like, yeah, they got under control. I'm never going to think about it again. I mean you can, but if like you care about uptime or you care about how much it costs or you care if your you know application's crashing and users can't get to it, don't do that. Track your CPU usage, uh, you know, look at it periodically. If you can set up alerts that tell you, hey, you know, you're using 100% CPU usage. Uh, same thing with memory.

38:08

Speaker 1: Watch your response times and not just your average response time, even like a 99th percentile response time is something that Likely most of your users are going to hit. If you have users that are going to make more than 100 requests to your application, they could be hitting like the very far spectrum of your worst of the worst response time. So keep an eye on that if your you know average response time is 250 milliseconds, but your maximum response time is 30 seconds. somebody's probably having a bad time uh with with those pages. And that's also potentially blocking those CPUs and you know you know could could be causing timeouts and all those other things we talked about.

38:55

Speaker 1: And then uptime monitoring and error rates. Uptime monitoring is great. People generally like point something at their homepage and say, my homepage is up. Site's good, but if you once you have a bigger application, you can have lots and lots of errors and your homepage loads loads fine. So error rate is a good one to track. It's basically tracking, you know, for every good request, are there any errors? And you may find there's some threshold there that You know, it's okay if I have 0. 05% of my responses be errors, but if you're having 20% of your responses being errors, that's probably something you want to know about. Um that's all I've got. So thanks so much. Um I think I might have a couple minutes for questions

39:41

Speaker 1: if there's any. Um yeah.

39:49

Speaker 2: Yeah, thank you Peter. Does anybody have a question they'd like to ask? Hi

39:53

Speaker 3: Pete, thank you. Um so uh when if I'm running multiple processes with say um GUInicorn and one is blocked on um Synchronous I. O. with the st um the OS run the other one.

40:07

Speaker 1: It's dumb. schedule between those um I don't know exactly what the algorithm is but it might happen when it's on IO and it might happen when it's actually doing processing.

40:26

Speaker 2: Anyone else have a question? Um

40:37

Speaker 4: you said it was outside the scope of the talk, but do you have any recommendations for load testing?

40:42

Speaker 1: Uh you know, like for A site where you just need to test like, hey, can a thousand people get this page? Um really basic load tests like uh there's one called um you know A B Apache Bench. Yan, what did they rename Hei or is Hei still Hey? Do you know? Okay. Yeah, there's one called Hei, H-E-Y. Um if you're doing more advanced stuff. Uh J meter is out there, which is um clumsy, but it works. There's another one I've used called K6, and and those are for places where it's like Okay, I don't need to just get a page. I need to get a page, submit this form, read the response, go to this other page afterwards, and you can kind of script some of those.

41:30

Speaker 1: And then there's uh another one, locust, uh, is is of interest. That's a Python one.

41:42

Speaker 2: This is probably the last question.

41:47

Speaker 5: Hi. Well in the past I have used like for the error rare monitoring and the CPU and stuff. uh new relic and data doc but uh do you have any recommendation on some open source or free alternatives to those?

42:03

Speaker 1: Yeah. I love Sentry and Sentry you can host on your own, but you can also host with them. That's probably my my biggest recommendation. And one kind of caveat here is The stuff that goes into that error monitoring system is happens when your request completes So you can end up in a situation where you're dropping requests because the request can't even get to your application because it's overloaded. or your application is crashing because it's running out of memory, your error reporting system probably isn't ever going to see those because those uh you know it never gets you your application never completes the request. So uh in addition to those, I I still recommend keeping track of something where um

42:50

Speaker 1: any of these systems uh platform as a service they're gonna have some sort of proxy or load balancer out in front of your application And that that's what they're going to report on for how your application is responding. So I would do both.

43:05

Speaker 2: All right, thank you.

Questions this talk answers

Where should developers deploy Django apps if they want to avoid most operations work?

Use a platform-as-a-service provider such as Heroku, Fly.io, or Render, along with a managed database. These services handle deployment, security, secrets, and much of the infrastructure for you.

Discussed at 1:58

Where should Django file uploads be stored?

Store user uploads in object storage such as Amazon S3 or an equivalent cloud file-storage service, rather than relying on the application server's disk.

Discussed at 5:07

How many CPUs does a Django application need?

It depends mainly on the number of concurrent users and how much computation each request requires. More concurrent requests generally require more CPU, but improving request performance is often a cheaper way to handle more traffic.

Discussed at 6:42

How many Gunicorn workers should I run per CPU?

A common starting point is roughly two workers per CPU, though the right number depends on the application's CPU usage. Too many workers can exhaust memory or make performance worse, so validate the choice under load.

Discussed at 9:02

How does slow response time affect CPU usage and application capacity?

A slower request occupies CPU time for longer, reducing the number of requests one CPU can process and increasing costs. For example, the talk contrasts about 6,000 requests per minute at 100 ms with about 60 at one second.

Discussed at 10:39

How much memory does a typical Django app need?

A typical Django application may use roughly 128–512 MB per process. With four Gunicorn workers, for example, that baseline is multiplied across the workers, so it could require about 2 GB.

Discussed at 17:57

What happens when a Django process runs out of memory?

Unlike CPU saturation, exceeding the available memory generally causes the process to be killed by the operating system's OOM killer. You can respond by increasing available memory or reducing the number of worker processes.

Discussed at 18:42

How can I reduce memory usage in Django?

Avoid reading large files entirely into memory and avoid materializing huge querysets. Use queryset iteration, `values()`, `only()`, or `defer()` when you do not need complete model objects or large fields.

Discussed at 21:05

What is the difference between horizontal and vertical scaling?

Horizontal scaling adds more application instances, while vertical scaling adds more resources to an existing instance. Stateless Django applications are usually easy to scale horizontally, whereas databases are commonly scaled vertically first.

Discussed at 27:17

Does autoscaling handle sudden traffic spikes well?

Not necessarily: autoscaling needs time to detect increased CPU usage and start new instances, so an immediate jump from zero to thousands of users can overwhelm the existing capacity. Scheduled scaling or maintaining peak capacity may work better for predictable spikes.

Discussed at 30:45

What are the main drawbacks of serverless for Django applications?

Serverless can scale rapidly and removes much of the need to manage application servers, but costs can be difficult to predict and cold starts can delay requests. It also introduces database-connection, request-size, duration, upload, and remote-shell limitations.

Discussed at 32:45

What should developers monitor after deploying a Django application?

Monitor CPU and memory usage, response times—especially high-percentile times—uptime, and error rates. A homepage check alone is insufficient because other endpoints can be failing while the homepage still works.

Discussed at 37:42

What tools can I use to load-test a web application?

For simple tests, Apache Bench or Hey can check whether a page handles a target number of users. For scripted multi-step flows, the speaker recommends tools such as JMeter, K6, or Locust.

Discussed at 40:42

What is a free or open-source alternative for application error monitoring?

Sentry is the speaker's main recommendation and can be self-hosted or used as a hosted service. It should be combined with monitoring at the proxy or load-balancer level, because requests that never complete may not reach the application error-monitoring system.

Discussed at 42:03

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Peter Baumgartner

More videos from DjangoCon US