Hosting and DevOps for Django with Benjamin "Zags" Zagorsky

This video features Benjamin "Zags" Zagorsky at DjangoCon US 2023 in Durham, North Carolina, USA.

Hosting and DevOps for Django with Benjamin "Zags" Zagorsky
0:43:19
Published November 22, 2023
1,597 views

Production server infrastructure is a complicated beast that requires configuring and coordinating dozens of tools and services. You have a new Django application and you're ready to deploy it; what next? You have an existing Django application and you set up the servers yourself; what can you do better?

I’m the co-founder CTO of Zagaran, Inc., a software consulting company. Over the past 10 years, we’ve built, maintained, and deployed dozens of Django websites, and have an extensive playbook for how to do that well. In this talk, we'll draw from that playbook and go through the main issues for a creating a robust and secure Django deployment. For each issue, we'll look at the technologies and techniques to solve it. We'll focus on AWS as a hosting platform, but the techniques at play will work on any major cloud provider. This talk will cover the following:

  • Application hosting
  • Resilience to server failures
  • Automating deployment
  • Secrets management
  • File storage
  • Error monitoring
  • Server maintenance
  • (and more!)

This talk was presented at: https://2023.djangocon.us/talks/hosting-and-devops-for-django/

LINKS:
Follow Benjamin "Zags" Zagorsky 👇

Follow DjangCon US 👇
https://fosstodon.org/@djangocon
https://twitter.com/djangocon

Follow DEFNA 👇
https://www.defna.org/

Video production by the presenter and DjangoCon US 2023 volunteers.

Summary

Benjamin “Zags” Zagorsky explains how to run Django services that tolerate failures, scale with demand, support testing, deploy automatically, protect data, and make errors easier to diagnose. He recommends managed hosting as the default: Elastic Beanstalk for straightforward Django applications, ECS for containerized and more varied workloads, and Kubernetes when multi-cloud or on-premises portability is essential, while stressing that flexibility brings additional operational work. He covers load balancers, HTTPS, Gunicorn, background tasks, environment-specific deployments, infrastructure as code, system updates, managed databases, S3 file storage, environment configuration, and secrets management, with particular emphasis on stateless application servers, automated replacement, backups, encryption, restricted network access, and backward-compatible database migrations.

Key takeaways

  • Use managed hosting rather than maintaining individual servers: Elastic Beanstalk is a sensible default for Django, ECS adds container flexibility, and Kubernetes suits genuine multi-cloud or on-premises requirements.
  • Put a managed load balancer in front of stateless application servers to support scaling, avoid single points of failure, terminate HTTPS, and improve protection against unwanted traffic.
  • Treat deployments as creating a tested new running copy of the application, then switching traffic to it; automate code, migrations, infrastructure changes, and system updates.
  • Maintain separate staging and production-like environments, including a copy of production data when testing significant migrations, and pin dependencies so tested images are reused consistently.
  • Use managed, encrypted, private data services with backups and failover, store files in S3 through Django’s storage API, and keep configuration and secrets outside the code repository.

Summarised automatically from the transcript.

Transcript

8,151 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:21

Hello, Django Con. You've spent all this time writing Django code. Make sure you know how to run it, right? So we're gonna talk about how to run it. So I'm Zags. That is short for Benjamin, uh, as you can clearly see. Uh and I am The co-founder and CTO of Zagoran. We're a Boston-based consulting firm and we do full stack web and mobile. And people say full stack, sometimes they mean front-end and back end. No, no, no, no, no. You gotta do DevOps too to really be full stack. I do, and so that's what I'm gonna focus on. But there's a lot of blur between DevOps and the backend. In my 10 years of doing this, I've worked on a lot of Django projects, I've worked on a lot of DevOps and a lot of the in-between. So I'm going to talk to you today about how to, if you're setting up servers for the first time, what do you even do?

1:11

Or if you've set up servers already, how do you do it better? But first, what does better mean? Let's define good. So six goals here for a good hosting. We want to be resilient to server failures. Even if you're working with virtual servers and the cloud and all that stuff, servers fail. I have seen it happen, both from hardware failures and also from software failures, engineer failures. Humans are fallible, everything we touch is fallible. That's the nature of the reality. Let's embrace it and let's insulate ourselves from it. We also want resources to be scalable. Unless you are going to have constant user load Always? You're not? Okay, so we want to be able to scale our resources up and down. Down also matters. And that's both changing how

1:56

much Resources we have for each server. We potentially want to be able to size individual servers up as well as the number of servers. We want to be able to add more servers or potentially reduce servers. We want to be able to test our code. That's good advice in general, but also we want our infrastructure to support that. So think not just production environment, but also staging environment, maybe some others. Maybe a prod copy environment. Hmm We want automated deployments because manual is slow and manual is error-prone. We want our data and secrets to be backed up, we want to protect against accidents and adversaries. And we want our errors to be easy to notice and diagnose. The other thing I want to talk about before we start is what does a good server setup look like? Now they all look a little different, but this is actually a decent template.

2:42

If you're working in Django. The setup for your service is going to look something like this. Over here on the right, we've got your API service that's running your Django project. Whether you're hosting a website or a REST API. In front of that, you're gonna want a load balancer. I'll tell you why soon. Over on the left, you might have a task framework, can be running Celery or something else, it'll be doing asynchronous tasks. Those are gonna be pulling from a task queue. We want that middle layer to be stateless. And again, I'll tell you why soon. So then now down at the bottom, that's where we have a persistence layer. That's where we're gonna have our relational database. That's where we're gonna have our file storage, maybe some other things. And then finally, there's there's probably a code repository in there too, uh, where you can update your code. And we want to talk about how we can push updates to this setup.

3:27

The four things I'm gonna cover here is how to set this up in the first place. That's hosting, how to update it with new code that's deploying, how to really manage that bottom layer because when anytime we have persistence. That's a little tricky, so data and secrets, and then finally, how do we tell when stuff goes wrong? I'm going to focus mainly on AWS because that's what I have used the most, but that is not to say that if you are not on AWS, this talk is irrelevant. Azure and Google Cloud have analyzed services for nearly everything that I'm going to cover, and I have included in the slides The equivalent services on both of those. The configuration is gonna be different. Some of the details may be different, but the big picture, the methods, how to think about it, that will be the same And by the way, the slides are in Slack already. So you don't have to be taking pictures or hurriedly typing everything that's on the slides.

4:16

Hosting. How do we set it up? So here we're gonna focus on being resilient to individual server failures and scalability. Those are the two things that I want to really focus on for hosting. And to start, I want to take just this slice. Perhaps the most important slice, given this is DjangoCon, because that's where you run your Django project. So let's just let's start with the most important thing. I want to give you three options for hosting. And one of those options is not setting up your own EC2 server, that is to say, a virtual server, installing your code base on it And managing everything yourself. The reason that I am saying don't do that is because it doesn't meet the two criteria that we're trying to hit for hosting. If that server fails What do you do?

5:02

You have to spin a new one up yourself, reconfigure it. If you need another one, again, you have to get in there, set it up yourself. And yes, you could write scripts to do that. But at that point, congratulations, you have reinvented platform as a service. Please go use platform as a service. So the three options I want to give you are platform as a service. On Amazon, this is Elastic Beanstalk. Managed container services on Amazon, this is ECS, and Kubernetes, which Amazon has a hosted version of called EKS. Now why these three is because these three work a level more abstract than a server. These three technologies are working on a running instance of your application. That's the unit of work. And they're gonna manage the servers for you. In order to do this, you need to tell these technologies how to set up a server.

5:49

You need to give them all of the information for how a server gets built to run your application. But once you do that, magic. Except not, because I'm gonna explain how it works. What they can do with that information is they can replace dead servers automatically. You can scale your servers up and down, either number of servers or size of servers, just by changing a configuration parameter. And Beyond that, you can set up auto-scaling. All three of these technologies have support for setting up rules under which you're gonna add or remove servers automatically, like if you have 75% CPU across all of your servers, just add another one. Let's start with platform as a service. Platform as a service, which on Amazon is Elastic Beanstalk, is a purpose-built hosting platform for

6:37

Websites and HTTP-based services like Django. So if you have a Django project, platform as a service is a great starting point and a really sensible default. On this diagram and the next two, orange is what you have to do and green is what you get for free. And I will also mention it just because the colors don't necessarily come through. The thing I love about platform as a service is there's so much green. So you set up with Elastic Beanstalk an application and an environment. That environment that is that running copy of your application, your code And Elastic Beanstalk sets up for you everything else. It sets up the load balancer, it sets up the servers, it sets up a reverse proxy, a web server logging, sets up your code on that, and is gonna manage those servers for you.

7:22

It's going to go spin up servers, replace them when either they die or you're deploying something new, all the kinds of stuff. And it will also set up security groups for you. Security groups, this is Amazon parlance for firewalls. And what Elastic Beanstalk will set up is what you want your security groups to be. It's gonna have that load balancer open to the internet or to other services. Depends if this is a private service. you could restrict who can send traffic to that load balancer. But that load balancer is going to be open to a broad range of consumers. And Your servers are only going to be accessible by the load balancer. The way you do this in Amazon is you say this security group accepts traffic on this port from members of this other security group. So for example, your application server security group only accepts traffic from the load balancer security group.

8:12

And therefore your application servers can't, no one can access those from just the outside world. So great from a security perspective, that's the same model we're gonna want to set up with the other technologies. One other great thing about Elastic Beanstalk is that it lets us host Django just natively on the servers. So we don't need Docker. You can use it, it does support it if you want, but you don't need it. Elastic Beanstalk has native Python support and can go set up a server for you with just your code running on it. Sounds great. Why do we even have other options, right? Well, the downside to platform as a service is that in being so purpose-built for web and APIs It has limited flexibility. So if you want more flexibility, we've got other options.

9:00

That's where a managed container service is going to come in. So on Amazon, this is ECS, Elastic Container Service. With a managed container service, you go set up a Docker image to do whatever you want with that Docker run command and it'll go run it for you Now already, not just from the more orange on this diagram, but already just from that statement, you are going to have to put in more work. You have to configure your own Docker container to say, here's what I want to run. More flexibility, more DevOps work. That's a trade-off you have to make. I'm just showing you the options. In addition, you're gonna have to set up a bunch of other things that platform as a service will do for you. You have to set up your own load balancer, your own security groups. You need hosting for your Docker image. That's what this Elastic Container Registry is over here.

9:46

You need to set up your own logging. That's what CloudWatch is here. And you need to set up runtime configuration for your Docker images. So that's the ECS cluster and task. And once you've done all that, then you can make a service in ECS. That is ECS's name for that running instance of your application. And that will go run containers for you. Now, in the old days, and old is a few years ago, but not as many as you'd think. You had to go set up a pool of servers yourself for an ECS cluster in order to run containers on them. That's really the point of a cluster was to have a pool of servers that it could source containers to. No more. You can now use Fargate. This is Amazon's serverless compute for containers to just totally eliminate that headache from this.

10:31

So this has gotten easier over time. But it's still going to be more complex than platform as a service just because you can do different things with it. So if you're running not just Django, but Django and Celery or things like that, this is gonna be a good option. So far though, we've only looked at options that are running within Amazon specifically. What if you need to run potentially in more than one cloud? That's where Kubernetes comes in. So Kubernetes is also a container management service, but rather than being a cloud native one, it's an open source one. And the reason that this matters is because Kubernetes can run on any set of machines, not just Amazon's Google, or Microsoft, but even your own or one of your customers that are sitting in their data center. So if you need to support multiple different clouds or on-prem deployments,

11:21

That's where Kubernetes is really gonna shine. You can even run a light version of it just on your computer to actually run a fairly high fidelity copy of your production infrastructure But as is the theme of this section, with more flexibility comes more DevOps work. So with Kubernetes, in addition to configuring your application to run Inside Kubernetes, inside the cluster that it's managing, you also need to set up that cluster. Now, Amazon and also Google and Microsoft will give you a fairly large leg up. If you use Amazon's elastic Kubernetes service with Fargate, that will go set up a Kubernetes cluster for you. It will go source those containers onto Amazon's serverless compute infrastructure for containers, but

12:07

you still need to bridge the gaps between your Amazon account and Kubernetes. So Kubernetes is managing its own networking, so's Amazon. You need to bridge those. So that's where we've got Traffic, we need to route it not just from the load balancer straight to our containers, but through an ingress and a service to actually get to our containers. Kubernetes calls them pods. You also need to do a similar thing with authentic authentication to bridge Amazon's users and roles with Kubernetes users and roles. There's a couple of those Let's call them friction points where in order to get that cloud independence within a cloud environment, you're gonna need to bridge those gaps. This is also still Docker-based, we're gonna have to host our images somewhere, so you could still use ECR, you can still use CloudWatch. Um

12:52

And having set all of this up, we can get a Kubernetes deployment, which is what Kubernetes calls that running copy of our application. So just to step back and look at these three Against each other, platform as a service, given this is DjangoCon, I can say that is a really sensible default for Django applications Great starting place. If you need that extra flexibility of arbitrary workloads, managed container services will do that a bit better. Also, managed container services like ECS will play a little nicer with infrastructure as a service, things like Terraform. Just because platform as a service is trying to do some of what Terraform is doing for you, you can make it work with Terraform and I have, but There's gonna be a little friction there. And then Kubernetes, that's what's really gonna give you that multi-cloud or on-prem deployment, that just true cloud flexibility.

13:40

Don't just say we might need cloud flexibility and like say great let's go there. Every piece of flexibility comes with a cost. So If you really are going for cloud independence, it's not just about running on Kubernetes. You also are gonna have to give up on some of my other recommendations that I'm gonna make soon about managed databases, managed file storage, to like really have that run anywhere, you actually need to push those into your Kubernetes cluster and find Docker images that are going to go do those for you. Let's talk about the other aspect of hosting. How do we get traffic in from the internet or from our other services into our Django project? High level traffic is going to come in, it's going to hit a load balancer. That load balancer is going to have your public SSL certificates.

14:28

It's gonna terminate public HTTPS. You can optionally use a self-signed certificate on your servers to re-encrypt that traffic going to your servers. You can then also optionally have a reverse proxy there, can be helpful for serving static files. Then you need a web server, something like Unicorn, to serve your Django project. Now in this model here, this is serving your static files from the same place as your Django project, which I love doing because it makes it really easy to coordinate those deploys That's not the only way to do it, but that is the way I'm going to show you in these next few slides. So first off, application load balancers. These are Amazon's hosted load balancer. The reason to have an application load balancer, namely a hosted one, is it lets you have multiple servers without a single point of failure.

15:13

If you set up your own load balancer, hey, you have multiple servers, so if one of them fails, you're still good, but what if your load balancer fails? Amazon is sourcing load balancers from a pool of load balancers when you use their hosted version, so you have no single point of failure anywhere in that diagram I just showed you. And even if you don't have multiple servers all the time, if you have multiple servers some of the time, if you're doing auto scaling or even manual scaling, having a load balancer is essential to enable those times when you do have multiple servers And even if you're not convinced by that, um having application load balancers will do more than just that for you. They do protect against some kinds of DDoS attacks. just by preventing that traffic from even hitting your servers, which is great. Adding in a reverse proxy can help

15:59

if that is something you care about and this is a highly public website, have more layers of protection. Um and on the other side, if you're worried about cost, Amazon will let you share an application load balancer between multiple environments. So you can use the same one for staging, production, prod copy, and so on, and only have to pay for one. Beyond that, they can do HTTPS. Isn't that great? Um HTTPS is not that hard to set up these days. But why bother? Why take that extra work? The whole point of this is actually I always say DevOps done well is you do less work managing your server so you can focus on writing your Django code, right? Isn't that what you all want? So if you use an application load balancer along with Amazon Certificate Manager, Amazon Certificate Manager, you set up a couple

16:46

C names in your DNS, and it gives you auto-renewing SSL certificates. You plug that into your load balancer and your load balancer terminates SSL for you. And it can even also it can even do the HTTP to HTTPS redirect. So if people type HCP colon slash slash in the in the browser bar, it'll do the redirect for you if you set that up at the load balancer. Django can do this too, but Why bother your application servers with that traffic? Do it as far out as possible. Spare yourself the load. Two other security bonuses. Amazon has actually a weird statement on re-encrypting traffic from the load balancer to your servers. By default, that is going to be unencrypted. And Amazon says it's It's fine, it's within the data center and it's within your virtual private cloud.

17:31

So like it's kind of hard to eavesdrop on that. But maybe you have compliance reasons you need to do it. And maybe you don't trust Amazon. So If it's one of those or both of those, just add a self-signed certificate to your application servers and tell your load balancer to send traffic to them via HTTPS. One other thing to think about, uh application load balancers by default accept a broad range of versions of SSL. This is because The default settings support legacy browsers. If you don't care about legacy browsers or you care about them less than you care about security, go disable those old versions of SSL. In fact, all versions of SSL are insecure. TLS is the secure stuff Uh and if you want the up-to-date information, check out SSL Labs.

18:17

They've got a great self-test tool. The last thing in this stack is the very poorly named technology of a web server. I really hate this name. Somebody should come up with a better one. GUInicorn is an example. It's not the only one. There's other ways to do it. Fundamentally, what this does is let Django receive HTTP requests. So GUNicorn or your web server, that is the technology that is sitting there actually calling your Django code. In case you ever wondered how that happened. It does multiplexing too, so for G Unicorn, a a reasonable setting for workers, that's how many processes it's got how many how many requests it's gonna be able to handle in parallel. Reasonable default for that is two to four times however many cores are running on that or on that machine that you have. But that's not to say that's how many requests can actually be waiting.

19:05

Unicorn will also have a queue of requests and as long as you're handling them pretty quickly, you can actually get through thousands of requests very fast, even though you're only processing, say, four or eight at a time. Unicorn can also do static hosting if you pair it with the white noise library. It's not gonna be great from a performance perspective, but the reason you would wanna do that is if you're running in a Docker-based environment like ECS or EKS. It's very hard to shove both GUnicorn and a reverse proxy into the same container. It's also a hassle to set up two containers for every running service. So if you don't care a ton about static. Hosting performance. White noise is just a great way to get those static assets hosted with Gunicorn. If you do care about performance, adding in a reverse proxy like Apache or Nginx.

19:51

is going to give you better performance on hosting those static assets. It'll also give you that better DDoS protection. But if you really care about performance for hosting your static assets, go get a CDM. The other thing that I want to touch on very briefly, not because it's not interesting or complicated, but because it's too interesting and too complicated and really could use its own 45-minute talk is task servers. So very high level. How these work in general, you're gonna have some queue, you're gonna put tasks on that queue. And the task servers Don't re accept any incoming requests. Their firewall rules are closed to all traffic. Instead, what they do is they go check the task queue either when they're done with a task or when they have nothing to do saying, got anything for me to do

20:36

And then they go do it. And then they go put the result, hopefully down in that persistence layer. In order for this to work, you need some task framework, something running on those task servers to go. check the queue, figure out what data on the queue means in terms of code that needs to be run and run it. I have three this slide is a little AWS centric, but three options. The first two are AWS specific. The first is if you're using Elastic Beanstalk, there's a second type of environment called Elastic Beanstalk Workers, and these use what I can only describe as an incredible hack. They take tasks off the queue and turn them into HTTP requests that they send to your application. So you can implement a task framework just as a REST API. Isn't that great from a Django perspective?

21:22

They're really good for, I mean it's just a great way to get background tasks like long-running scheduled jobs. It's not they're not as good for on-demand tasks. They can do on-demand tasks, but if you start having thousands a day or many different types of on-demand tasks, it's not going to be a great fit for that. If that is your model, using AWS Lambda, this is Amazon's serverless compute with SQS, is gonna be great for that. It lambda with SQS can do magic autoscaling based on how many jobs are in your queue, but has a 15-minute time limit on any job it's running. So not a great fit for those long nightly or weekly jobs. If you need both, that's what salary is for. You're working in Python, Celery is just the de facto does everything task framework, but

22:08

downside there is complex setup. If you were hanging out in the blissful land of platform as a service. Celery is one of the main things that is going to kick you out of that and into needing a managed container service. All right. That's task frameworks. Really needs its own talk. Now you've set up servers. How do we update them? Unless you're planning to never change your code. Okay, if you if you're never changing your code again, you can sleep for five minutes. Goals for deployment. We want our code to be testable on other environments before we push it to production. So we need to have other environments. So let's let's watch for that as I talk about how to deploy. And we want deployments to be automated. And we're gonna be taking code from our repository and pushing it out to all this infrastructure.

22:57

Now, a deploy has potentially many steps. In addition to obviously updating your code, you potentially need to install old package updates, both system and Python, JavaScript, compile CSS, collect static. Maybe do other infrastructure changes. Migrate is a very common thing you need to need to do. And then restart your web server. This is the wrong way to think about it. The right way to think about it is take all of the hosting technologies that I just mentioned, assuming you're on one of them, all of these hosting technologies are thinking in terms of a running copy of your application. And they all need to already know how to set up one of those, which is steps one through six at a minimum. So instead, what we're gonna talk about for a deploy is how to make a new running copy of your application

23:47

Then we're gonna need to do peripheral infrastructure changes. Migrate is actually the hardest thing to do with deploying, so that's gonna be the focus of my next three examples. We're gonna then put that new running copy of our application into rotation, stick it in with the load balancer, and then once it's good, we're gonna throw out the old ones. And remember how I said those those servers in the middle should be stateless? This is why. also to insulate you from individual server failure. So deploying Elastic Beanstalk, the other reason I love it is You run a single command in the Elastic BeanSock CLI, you run eB deploy, you give it an environment name. Look at that multi-environment support right there. And what that does is it goes and it takes your code, it takes the current commit that you're on in the repository where you run this command. It bundles it, it uploads it to the cloud, it gives it to Elastic Beanstock and says, go deploy this.

24:38

And then Elastic Beanstock will either spin up new servers or actually do it in place on your current servers. There's a setting for it. It will go spin up those so that new running copy of your application. In the process, it's going to use all of the configuration for how to set up that copy of your application. And that all actually just lives in your code base. There's a couple of configuration files in either. eb extensions or.platform. That tells Elastic Beanstalk how to set up that copy of your application. So it's gonna do that. And as part of that, you can even run peripheral infrastructure changes like migrations. Elastic MeanStock is the easiest platform to do migrations on And then once that's good, it's gonna switch that on, it's gonna chuck the old copies, you're done ECS is going to be different because Docker

25:24

is a substantial piece of it. So the key for ECS is Docker build. That is where we're going to prep that new image, that new Set of instructions for what a new running copy of our application is. So that's where we're gonna have our code updates, our system installs, our static files, all of that stuff. We need to tag it and push it. Peripheral infrastructure changes are a little complicated here. For ECS, running migrations, best way I found to do it is you can run a copy of your service. with an overridden run command to do migrate, and this way you can run migration with the context of all of your code and and uh and Python installs. And then finally, once we've done that, then we can do update service. And what update service will do is it'll, based on your new Docker image, it'll make new containers, put them in with the load balancer, throw out the old ones.

26:15

And Kubernetes is the same model, right? Docker build is gonna do all of that, creating that new copy of our application, and then Kubectl, the Kubernetes CLI, is going to put those new containers into rotation throughout the old ones. And Kubernetes does have a way that you can configure here are particular commands, temporary containers I want to run when deploying. Now, both Kubernetes and ECS, there's a lot of steps involved here. That's not just a single command, there's multiple steps, so you want to make it a script. Make it a Python script. Please don't do bash scripting. Python is way better at scripting. You can use the subprocess module. It's really great. And control flow is better. Right. It's just you error handling is better. So many things are better. So do a Python script. You've got the Bodo library at your disposal.

27:00

It just makes it makes things so much nicer. All right, so just to give you the side by side, if you're using Elastic Beanstalk, EB deploy, that's gonna be doing most of the work for you. And then just the rest of it is all in your configuration files, either in eB extensions or in dot platform. Meanwhile, on ECS or EKS, Docker is doing most of the work. That's what's doing your code, your dependencies, your static files. You have to find some way to do those peripheral infrastructure changes, and then it's just down to whatever CLI. Updates your technology to get those new containers. So what about other environments? We've talked certainly a lot about staging and production. I would highly recommend having a prod copy environment in there as well. The reason for that is because your

27:46

code depends on your data. And in particular, if you're testing major migrations, you should really be running those on the copy of production data to make sure they work. And you may also want to have other environments too, something like demo or training that you deploy after production. that you can use to show off your product without having real data in it. Now, all three of the methods that I just said will work on any environment. Elastic Beanstalk explicitly takes an environment as a parameter when you deploy. But the other two as well, you're creating a Docker image based on code on a particular branch, and then you are updating one of your environments, one of those running copies of your application. Really, the only thing that I would suggest if you are using a Docker-based solution would be skipping the Docker build on post-staging.

28:32

You already have a Docker image that you know works on staging. Just have that Docker image be what you deploy to these other ones so you're not potentially deploying something different. Pinning dependencies is also very helpful for making sure whatever you tested on staging is going to work on these other environments. Now, the other thing at the point at which you have four, maybe five environments of your service that we should talk about is not just automating deployments of your code, but automating deployments of your infrastructure changes. Whoa, infrastructure as code. So Terraform is my personal favorite here. All of these are preferences. There's so many technology choices. Go use whatever you love. But use something like this. So what Terraform does is

29:18

it describes your infrastructure in a set of configuration files. You run Terraform apply and it goes and sets up that infrastructure for you. You're like, okay, but How have I saved any time? Well, you've saved time if you use modules. With the Terraform module, you can say here is the infrastructure for one copy of my service. Here are the inputs, here are the things that vary between environments, and then with a mere five or six lines of code, you can make five different copies of that set of infrastructure very easily. So already some time savings, and then the real kicker is when you need to change it. You go change that module definition in one place, you run Terraform apply. Terraform goes and looks at what do you have set up already on the cloud, what's in your new configuration, computes a set of changes for you.

30:05

Tells you here's what I'm gonna do. It does confirm before running. And then you you say yes, and it goes and changes your infrastructure for you in a coordinated fashion across all of your environments. So automated deploys of infrastructure changes too, not just code. Now we've talked about automation. The pinnacle of automation is the zero-click deployment, continuous deployment. Zero click is a lie. Uh you still click something, it's just merge rather than deploy. And so how this works is you set up a mapping of git branches to environments, and when code is deployed to a particular git branch. It gets deployed, well it's merged to a Git branch, it gets deployed to an environment. And all of these, there's a bunch of paid services, they're they all follow that same model. You pick one, I mean really the important thing to do is to get to

30:52

A single command deployment, get to that automated script, and then if you want, you can set up zero click continuous deployment. A last thing to think about on deployments is system updates. Please do these. Uh Equifax lost half of America's social security numbers by not doing this. So you need to update them regularly. Platform as a service, it'll do it for you. Love platform as a service. You just set up manage platform updates on Elastic Beanstalk, give it a maintenance window, does it for you? But with Docker, your system installs are part of your Docker image. So if you're using Docker, you need to periodically rebuild your Docker image and do, you know, after yum update and upgrade, and then redeploy that. Wouldn't it be convenient if you already had automated deployments and maybe even a CD system to do it for you?

31:38

Hmm. It all comes together. All right. We've focused a lot on the application servers. How about data? These don't really do anything without the data. So here obviously we're gonna talk about data and secrets need to be secured and backed up, but also we're gonna come back to resilience to server failures. Because it matters here too. So we're down here in the data storage layer, the murky depths of persistence. Using a managed database like Amazon's Relational Database Service, will take care of most of these goals for you. They will give you managed database and operating system version upgrades. They'll give you automated backups. Really great. They give you storage auto-scaling. You can say give me 30 gigs to start and just add more storage when I run out.

32:26

So you don't have to pay for more storage than you're using. There's also settings when you set up a database, you check a box, you get encryption at rest. Do it. Now there's one other consideration that people often skip because it costs money. Uh you can also check a box and say, give me a live read replica with automatic failover. It costs twice as much because you're running two databases. But the reason you want it, at least for production environments, Amazon calls this a multi-AZ deployment, by the way. The reason you want this is if your server, your database server, fails. You will lose all of your data since your last backup. So if your backups are daily, you could lose up to 24 hours of data with a live read replica You reduce that to seconds. So you have way less potential for data loss and you also have

33:11

much Better uptime guarantees because with that automatic failover, it's substantially faster to get that database back up. So we said encryption at rest. How about in transit? You just tell Django please please use SSL. So this is for Postgres. You can either say SSL mode require or uh in your database options or in the database URL, depending on how you're configuring it. Wouldn't be very good to encrypt our data at rest and in transit if someone can just steal our database password and suck all that data out of our database. So let's talk about security groups for a sec. We also want our database to only be accessible from our application servers. Same model as the application server should only be accessible from the load balancer.

33:56

Because those are the only things that should be talking to our data. And application servers also includes task servers here, by the way. So usually what this will look like is the database security group will be open to two different security groups, the application servers and the task servers. And as an added layer of security, at least on Amazon, you can have your database not even have a public IP address. As long as your database, your application servers, and your load balancer are all in the same VPC virtual private cloud They can talk to each other on private IP addresses, and your your database can just not even exist from the perspective of the public internet. Um your application servers can too as long as you use a NAT gateway so that they can still make requests out to go do things like pip install. One other thing to mention on databases and deployments comes together in other ways.

34:42

There's a race condition when deploying migrations You should deploy migrations first, application code second. All of the examples I showed you do that. However, even if you do that. If a request comes in in between you doing migrations and you updating your code, you're gonna be running an incompatible version of your database schema and your application code Two options here. You could either use a maintenance mode when deploying, that is to say downtime, or you can use backwards compatible migrations. And if you want to know more, this is a really deep topic. I talked about this at DjangoCon last year, so see my talk on Migration pitfalls and solutions. That's schema data. What about files? For some reason we call this block storage, it's file storage. Um on Amazon, this is S3. S3 gives you unlimited file storage.

35:29

And I do mean it. Instagram is built on this. If you have more files than Instagram, talk to me afterwards. Incredible durability guarantee. That is 11 nines. Very cheap. Three things for S3. There's a checkbox for encryption at rest. Check it. There's a checkbox for mail versioning. Check it. This will prevent you from losing data from accidentally overriding your files, which I have seen happen. So just check that box. And every object on S3 has its own permissions. Which is usually not what you want, so there is a way to configure on a bucket block public access for everything in this bucket. Check that, unless you actually do want the bucket to be public. Now what about S3 and Django? Django storages. It makes it so easy. Uh

36:14

Django has a pluggable storage system which lets you just put files in a file field and whatever storage engine you have configured, that's where it will stick the file. So with Django storages, you truly only need five settings to connect this to S3. And then any file that you put in a file field will get uploaded to S3. And so you can have that file on S3 's incredibly durable, incredibly scalable file storage, but also still have it be accessible through your relational data. It's beautiful. And you can even have this be. Swappable so you can have this be you know on for servers but off for local development if you want those files to just be stored locally. Speaking of pluggable configuration Django Environ makes it really easy to do. This is a great tool for letting you access environment variables for pluggable settings on servers.

37:01

You can read from a. en file for local development. Just put that in your gitignore. And then it does type coercion, defaults, many other things that are super helpful. Um great for configurable settings like debug or secret key or your databases. Now, one thing worth noting here, your secret key, your databases, those might be kind of sensitive data. How do we store that? That's what Amazon Secrets Manager is for. Purpose built technology for storing literally configuration That is sensitive. A secret in Amazon Secrets Manager is not a secret key. It's a JSON blob that should be all of the secrets for one of your environments. And it's very, I mean, all of these technologies, with not that much work, you can plug an Amazon Secrets Manager secret into

37:46

a running copy of your application as an environment variable. It takes about three lines with Django Environment to read that in as environment variables that you can then reference in your settings. And even if you can't do that mounting, just replace that OS environ get with a Bodo call to Amazon Secrets Manager. All right, we've set everything up, we've deployed our code, we've stored our data. What could possibly go wrong? Well, your code, of course. So, if we have errors, we want to know. And we want to know not just that we had an error, but also what. I've got one very easy recommendation for you if you're working in Django. Sentry is great. It is not the only product here. There's actually several other options out in the hallway exhibition booths.

38:31

Sentry is one that I've used a lot. This is a paid product. All of the ones in this section are paid products. Um but what Sentry does is with three lines of config, plugs into your Django project, watches for any unhandled exceptions, grabs them, sends them to Sentry, dedupes them by stack trace, get notified once if you get 100 copies of the same error, and then gives you an explorer for all the layers of that stack trace variables that are in scope. Fundamentally the point of a good error monitoring system, century or otherwise, is so that you do not have to look at logs. That's how you know you're succeeding. I still look at logs if there's something that goes wrong with the it with a deploy, and otherwise I do not. Now, code errors is not necessarily the only thing you might want to fix. If you want to know, is my application up? Uptime robot?

39:17

Single purpose tool. It polls your application every so many minutes, sends you an email if it's down, sends you an email when it's back up. It's great. Um if you need to dig in more on what is my application doing, for that you want An application performance monitor, something like Datadog or New Relic. If you want to know what are my resources doing, my server resources or my load balancer database, right? That's for that you should dig in all these clouds, have a way of seeing metrics on your cloud resources. So in Amazon that's CloudWatch, you can see CPU or network or whatever on any particular Amazon resource that you're using. And there's many more things that you could use for monitoring. This is just an initial list of my my go -tos when starting a new application. or working on an application that's starting to have issues

40:05

and come under a lot of load. Now, when something goes wrong, what do you do? Well For the most part, you're gonna go deploy new a new version of your code to staging, test it, production, and I already showed you how to do that. So that's easy. But what if you need to do something else? What if you need to get on that server? This is where you would use SSH if there weren't a better option. Now, just to talk about reasons you would want a SSH. At a minimum, getting onto a server or container, you can do whatever you want. You can really just poke at whatever's going on with a stick, really great for debugging. But also, if you have followed my advice, this is the only way to access your database. Remember how your database is only accessible from your application servers? But that's okay because your application servers have your Django project, so you can just go

40:52

pythonManage. py, shell plus and just talk to your database through Django, which is way better than any other way that I've ever done it. But don't SSH, please. If you're on Amazon, there's a way better option. It's SSM. This is a what I would describe as a cloud native SSH. Um this is one Google Cloud has an equivalent. Azure doesn't really. On Azure you you still kind of have to muck around with bastion hosts But with SSM, what you can do is you use your Amazon credentials rather than shared SSH keys, which I would describe as a security menace. SSH key, shared SSH key said it's not not SSM. So you use your Amazon credentials and it just jumps you straight into the Amazon resource. even if that is a Fargate container, a private server with no public IP address, you don't open port 22, and it even logs your S

41:40

your SSH sessions for you. So just better than SSH in basically every way. Um and if you're on Elastic Beanstalk, I wrote a great command line tool called EBSSM. Check that out, it makes it even easier. So I wanted to end with. The big picture, or shall I say just the complete picture? I've gone through a lot of technologies here. Here's the list. This is a slightly different view of them. On the right is what I would describe as the pretty much everyone should be using these. And on the left, these are the conditionally useful. So you need to host your application somehow. And I would really push you to use one of platform as a service, a managed container service or Kubernetes to do it. You need a load balancer.

42:26

You need a managed database, block storage, error monitoring, uptime monitoring. Some way of storing secrets securely and some way of automating your deploys. Everything else you might need, I showed you those in case those are relevant to your situation. Some people need a task framework, other people don't. Who knows? So these are just these are the options that I have shown you today. Go check them out. This is my email. I'm going to be in the hallway afterwards for questions. My slides, as I mentioned, are on Slack, but please reach out if you have questions about anything I've said or if you need consulting help, DevOps related or otherwise, drop me a line. Thank you.

Questions this talk answers

Why should I use a platform-as-a-service instead of managing my own EC2 server?

A platform-as-a-service manages the servers for you, replacing failed instances and making scaling largely a configuration change. Managing EC2 yourself requires you to build and maintain those automation capabilities, effectively reinventing a platform-as-a-service.

Discussed at 4:02

What are the best hosting options for a Django application on AWS?

The speaker recommends three options: Elastic Beanstalk for a sensible Django default, ECS for more flexible container workloads, and Kubernetes/EKS when multi-cloud or on-premises portability is important. More flexibility also means more DevOps work.

Discussed at 4:32

How should I route internet traffic to a Django application?

Traffic should reach a hosted application load balancer first, where public HTTPS is terminated, and then proceed to the application servers and a web server such as Gunicorn. A reverse proxy can optionally serve static files and add another protection layer.

Discussed at 13:28

Why does a Django deployment need a load balancer?

A hosted load balancer lets multiple application servers share traffic without making the load balancer itself a single point of failure. It also supports scaling, can absorb some DDoS traffic, and can handle HTTPS termination and HTTP-to-HTTPS redirects.

Discussed at 15:13

How do background tasks work with Django and Celery?

Task workers do not accept incoming requests; they pull jobs from a queue, execute them, and usually store results in the persistence layer. Elastic Beanstalk Workers suit long-running scheduled jobs, Lambda with SQS suits short on-demand jobs, and Celery is the flexible general-purpose choice when both are needed, though it requires more setup.

Discussed at 20:36

How should I deploy Django code without causing unnecessary downtime?

Build a new running copy of the application, apply required infrastructure changes such as migrations, put the new copy behind the load balancer, and remove the old copy once it is healthy. Stateless application servers make this replacement-based deployment model possible.

Discussed at 22:57

How should I manage staging, production, and other Django environments?

Use separate running copies for environments such as staging, production, and a production-data copy; the latter is useful for testing major migrations against realistic data. With Docker, promote the exact image tested in staging rather than rebuilding it for later environments, and pin dependencies.

Discussed at 27:46

How can I automate infrastructure changes across multiple environments?

Use infrastructure-as-code such as Terraform to describe a reusable service module and vary only its environment-specific inputs. Updating the module and running `terraform apply` lets Terraform calculate and apply coordinated changes across environments.

Discussed at 29:18

How should I run a Django database in production on AWS?

Use a managed database such as Amazon RDS, which provides automated backups, managed upgrades, storage autoscaling, and encryption at rest. For production, a multi-AZ deployment with an automatic failover replica reduces potential data loss and improves recovery time.

Discussed at 31:38

How do I secure a Django production database?

Require SSL for the Django database connection, restrict database access to application and task-server security groups, and avoid assigning the database a public IP address. Keeping it inside the private network prevents direct exposure to the internet.

Discussed at 33:11

How do I deploy Django migrations safely?

Run migrations before deploying the new application code, and make the migrations backwards-compatible so the old code can handle the intermediate schema. The alternative is to use maintenance mode and accept downtime during the deployment.

Discussed at 34:42

How should I store Django media files and uploads?

Use Amazon S3 with `django-storages`, which integrates with Django’s pluggable storage system so files in `FileField`s are uploaded to S3. Enable encryption, versioning, and bucket-level public-access blocking unless public files are explicitly required.

Discussed at 34:50

Where should I store Django secrets and environment-specific settings?

Use environment variables and a tool such as Django Environ for configurable settings, keeping local `.env` files out of version control. Sensitive production configuration should go in a dedicated secret store such as Amazon Secrets Manager.

Discussed at 37:01

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Benjamin "Zags" Zagorsky

More videos from DjangoCon US