Forklifting Django: Migrating A Complex Django App To Kubernetes by Noah Kantrowitz

This video features Noah Kantrowitz at DjangoCon US 2019 in San Diego, California, USA.

Forklifting Django: Migrating A Complex Django App To Kubernetes by Noah Kantrowitz
0:29:02
Published October 25, 2019
1,150 views

DjangoCon 2019 -Forklifting Django: Migrating A Complex Django App To Kubernetes by Noah Kantrowitz

Everyone is talking about Kubernetes, but migrating existing applications is often easier said than done. This talk will cover the tale of migrating our main Django application to Kubernetes, and all the problems and solutions we ran into along the way.

This talk was presented at: https://2019.djangocon.us/talks/forklifting-django-migrating-a-complex/

LINKS:
Follow Noah Kantrowitz 👇
On Twitter: https://twitter.com/kantrn
Official homepage: https://coderanger.net/

Follow DjangCon US 👇
https://twitter.com/djangocon

Follow DEFNA 👇
https://twitter.com/defnado
https://www.defna.org/

Intro music: "This Is How We Quirk It" by Avocado Junkie.
Video production by Confreaks TV.
Captions by White Coat Captioning.

Summary

Noah Kantrowitz explains how his team moved a complex, single-tenant Django monolith from a Terraform-and-Ansible deployment process to Kubernetes without rewriting the application. He covers container images, init containers, Helm’s limitations, and the custom Kubernetes operator they built to manage deployment sequencing, Django migrations, generated secrets, and service state through convergent reconciliation. He argues that Kubernetes works best when teams replace procedural deployment scripts with idempotent, desired-state systems, and describes how GitOps, Argo CD, CLI helpers, and a simple web UI made the resulting platform easier to operate.

Key takeaways

  • A forklift migration keeps the existing Django application largely intact while changing how it is packaged and deployed.
  • Multi-stage container builds can separate Python build dependencies from the smaller runtime image, while shared images can run Gunicorn, Celery, and Channels processes with different commands.
  • Helm improved environment reuse but did not adequately solve migration ordering, error handling, testing, secrets, or permissions.
  • A Kubernetes operator can model deployment as a convergent state machine, safely coordinating infrastructure setup, migrations, application rollout, and error states.
  • GitOps makes Git the authoritative source for application and infrastructure configuration, allowing automation to correct drift and restore the platform from one repository.
  • Operational tooling such as CLI wrappers, environment checks, reports, and a read-only UI helps non-specialists work with a Kubernetes-based deployment system.

Summarised automatically from the transcript.

Chapters

  1. 0:00 Deployment Challenges Introduction to the complexity of deploying a multi-tenant Django application and the limitations of the existing manual and Ansible-based approach.
  2. 2:25 Migration Goals and Constraints The goals of the forklift migration, including adopting Kubernetes while retaining managed AWS database and messaging services.
  3. 3:10 Containerizing the Django Application Building container images, deploying the Django processes on Kubernetes, and addressing migrations, static files, and supporting services.
  4. 6:45 Helm-Based Deployments Using Helm to package deployments and manage multiple environments, along with its limitations around sequencing, testing, secrets, and security.
  5. 8:19 Kubernetes Operators Introducing custom resources, watches, and controllers as the foundation for a more reliable deployment system.
  6. 10:37 Convergent Systems Contrasting procedural and convergent systems and explaining idempotence, Promise Theory, reconciliation, and protection against state drift.
  7. 14:49 Operator Implementation Implementing the Django operator with deployment state machines, controlled migrations, Celery Beat handling, and dynamic credential management.
  8. 19:30 GitOps Workflow Moving the source of truth into Git and using continuous synchronization to reduce configuration drift and improve deployment traceability.
  9. 22:54 Developer Tools and User Interfaces Supporting the Kubernetes platform with a command-line toolkit, environment diagnostics, and a read-only web interface.
  10. 25:08 Questions Audience questions about managed services, learning Kubernetes, and making GitOps accessible to non-engineering teams.

Transcript

6,113 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:15

Speaker 1: Hi, I'm Noah. This is mean all places. I run the GiveTracker team as part of my ride cell and transportation software. And I'm here to talk about deployment being GIFO. It's a difficult process. There's a lot of individual lips. They have to happen in the right order over and over again. It has to be fast, has to be easy to use. In the beginning, at the dawn of time, we all did this manually. We got stage machines, we did it itself, and it's great, exactly, not creative. It's slow, error-prone, and time-intensive. Over time we've got new tools like MapRef, Chef, and Splow. They help them automate things. So the rise of how a big challenge has a flood. Am I live now? Awesome. Um

1:01

Speaker 1: so with the rise of new tools came new challenges. Um so let's set the stage. My main application is a fairly standard Django monolith. We're still on Python 2, but that's not really relevant for deployment. We've got salary for background tasks, channel for web sockets, we're using Postgres for the SQL database, and we use Redis and RabbitMQ. Uh additionally, one sort of complicating thing is that my main application is single tenant. That means we deploy one instance per customer. Mostly this just means that we have all the same problems as everyone else, but much more so because we have to deploy a lot more environments. So whereas most people would have You know, a couple of dev environments, a couple of QA environments, one staging and one prod. We have a dozen or two prods. Um so this leads into what our old deployment system looked like. Uh there were two main pieces, um Terraform for setting up all the stuff on AWS.

1:47

Speaker 1: And then Ansible for configuring the systems and deploying the actual application code along with some Python scripts for things that didn't fit into either of those. Uh this was a good system. Secrets lived in Ansible Vault, access control is managed via SSH key distribution. Pretty solid all around. But there were some downsides. Setting up a new instance took an hour or two of work overall, running a whole bunch of Terraform stuff, making sure that everything was all up to date, dealing with with bugs that had creeped in this last time we did it. And that hour had to be someone on my team. Similarly scaling up and down if we needed to boot a new web server, that took a little while.

2:17

Speaker 2: And we really couldn't even think about doing auto-scaling.

2:20

Speaker 1: So good but not great. We also have a bunch of microservices, but I'm not going to talk about those for this talk.

2:25

Speaker 2: So we had some constraints for what we wanted to build as part of the new system. We knew we wanted to use Kubernetes. We'd already been using it for some other projects, and it's my weapon of choice for container wrangling overall.

2:34

Speaker 1: Additionally, we wanted to keep using uh RDS and Cloud AM QP for databases. I'm probably more in favor of running databases on Kubernetes than most people, but I do have a very small infrastructure team, and so I will basically always trade money for not having to worry about databases. Uh also that would make the production migrations easier because we didn't have to move any data around. What do we mean by a forklift migration? This means I can make changes to the application But not like a Greenfields rewrite everything and go and make everything glorious. Uh we had an application. We wanted to pick up the entire application, basically the way it was, and just drop it into Kubernetes. Before I start talking specifics about my application, let's define some jargon for

3:10

Speaker 2: that haven't heard it before. Two generic terms. A container is a cool way to run a process and an image of the larval form of a container. Uh Kubernetes is many things to many people, but for the purposes of this talk, Kubernetes is an API for making containers jump through hoops to do useful work. And pod is a container-specific term that every time you hear it, you can basically just pretend that I said the word container. So with all that laid out, I was ready to roll up my sleeves and jump into porting my application to Kubernetes. First step in containerizing any application is to make a container image.

3:38

Speaker 1: Multistage builds are super useful for Python applications because Python apps usually have much bigger build time dependencies than runtime dependencies. One thing of note here is the fact that I'm not using a virtual env. Some people will disagree with me, but I don't really see the point in using a virtual env when it's a single-purpose container in the first place. But so here's our build image. And then our run image. So we ran the build phase first, and then we copy over just the files we need into a much smaller run image, along with setting up a non-root user for the application to run as and anything else that's sort of runtime specific. My first attempt to get things set up in Kubernetes was very, very simple, just static Kubernetes manifest YAML, no automation beyond KubeCuddle apply. Uh fitting an entire Kubernetes YAML file on screen is really not gonna work. So here's just sort of the core of it.

4:23

Speaker 1: Um but it gives you a general idea. Start up the container, mount up a configuration file, we put all our configs in YAML and then read them in the settings. py. Run some migrations and then launch to Unicorn. Pretty simple, like these are all the steps that you'll probably have seen a million times just done inside Kubernetes. A quick improvement from there was to move the migrations from being run just sort of as part of the main container into a thing called an init container, which only runs the first time a pod is launched. Um and uh for some bonus points, this sort of sort of shows a very easy way to do some dynamic configuration injection, again, using init containers. Um but this still had a lot of issues with ordering. Init containers run every time a pod starts, including every time you launch another replica of the pods. If we wanted four web servers, it would try to to run the migrations four times, which isn't really what we want.

5:10

Speaker 1: Django migrations are generally safe to run more than once, but you can get some weird locking issues and things like that. It was just not a good scene. We also had issues where like because we weren't running the migrations first when we were upgrading Celery D, it would roll out the new code before the migrations had finished and everything would get really confused. So good start. This proved that we could run our application under Kubernetes, but we clearly needed a more uh organized deployment system underneath it. A few things I glossed over just so we're keeping track of them. Um the other three main daemons work exactly like the Gunicorn pods. They just run a different command instead of G Unicorn, it's running Celery D or Daphne or the channel worker thing. Um We build all of the daemons into the same container image as opposed to building separate images for each one. This was mostly because the shared uh dependencies are so much bigger than anything that's specific to web

5:56

Speaker 1: or cellarity or channels. That it made release management a lot easier just to have one release artifact instead of four. Um I'm also ignoring CeleryBeat for right now. We will get to it later. Um my application is mostly an API that's used by mobile applications. So we do have web interfaces in it, but they're mostly either used rarely or used by administrative staff. So I decided to keep things simple and try to avoid using S3 for the static files. For things that are dynamic like user uploaded uh media files, that still has to go in S3 because you need a shared system of some kind. But for the like actual static files, we build that into the image and we serve it using a cut-down version of the CADI web server. So that just means that we don't have to separately deal with versioning the static assets and the deployed image. Uh and for the initial phase I was uh going to sort of start with RDS, because we knew that's what we wanted to use in the end, but launching and destroying RDS cluster takes about 30 minutes.

6:45

Speaker 1: Uh and that was getting really annoying for rapid development. So we did most of the initial phases of prototyping using Zolando's Postgres operator So by this point I had a working proof of concept. I could deploy my entire stack on Kubernetes and use the application. The application itself worked like normal, but deploying it was still not really great. We needed something that was more repeatable and reliable. So, next tool in the standard Kubernetes quiver is a thing called Helm. It calls itself the Kubernetes package manager. This is overall fairly similar to what we had before. Um the idea is you just take the existing static manifests, you dump them into a Helm chart, uh, and then you can install that chart as many times as you need with the different inputs. I'm not going to go over the chart itself because it's basically the same thing we saw before, just the more curly braces. But overall Helm did improve some things. It made deploying multiple environments way easier. We could store just the values that had to vary between the environments in YAML files and feed them into Helm.

7:33

Speaker 1: That part was great. It's also widely supported by basically everything in Kubernetes ecosystem because it's such a standard tool. And we had some operational experience with Helm from other projects. But it also brought some problems. Um it basically doesn't help us with the ordering and sequencing issues we had with migrations and things running out of order. It does have a thing called hooks that let you kind of do what you would think, like before I deploy, please run the migrations. But we ran into no end of problems with it. Error handling wasn't great, debugging wasn't great. And in the end we just decided that it was too complicated. To work with for that uh function. There's also some major gaps in the Helm ecosystem. Um unit integration testing of Helm charts is basically impossible. Um, and secrets management exists but is basically delegated to plugins. Uh and then the big problem with Helm, Tiller, I'm not gonna talk a ton about this

8:19

Speaker 1: because they have already released a beta of Helm 3 which removes Tiller, but the rough version is it's a security nightmare. Um It's a single point that all requests have to go through, and everyone gets access to whatever permissions Tiller itself runs with, which is usually the cluster give me all permissions admin. role. They are removing this, so it's going to get better, but it's not there yet and all of the workarounds for not running with Tiller have various downsides that we didn't love. So okay, Helm worked, we could use it, but it wasn't quite as smooth as we wanted it to be. We had a lot of problems with what happens if there's an error during deployment, how do we deal with that? Uh and it made uh self-service fairly tricky to set up because of the permissions issues I mentioned with Tiller. And so with that, we have our current and hopefully final approach: a Kubernetes operator. Um, this is gonna be the bulk of the rest of the talk.

9:06

Speaker 1: Uh Kubernetes operators tie together three main concepts: CRDs, watches, and controllers. Let's go over each of those. Uh a CRD or custom resource definition is a way to add new object types into Kubernetes. Just like pods and services exist in Kubernetes, we can make our own. In this case we call it Summon Platform, because that's the name of my Django monolith. Um it lets us add a new object type and we can add whatever fields and parameters we want on that just like any other object in the Kubernetes API. In this case, you can see we have a version field which takes a string of what version we want to deploy, and we have a config thing that takes a dict of some Django settings, in this case just turning debug true. This is a bit more verbose than a Helm values file for anyone that's worked with those, but it does allow for some nice things like you can have type validation on your variables. which is a thing that Helm itself does not natively support.

9:52

Speaker 1: And you can use tools like KubeCuddle Apply with this directly as opposed to having to go through multiple intermediary tools. So the CRD gives us a place to put all of our configuration. It holds the data for what we want to deploy, what we want our configuration to be, all that kind of stuff. But now we need to actually do something with it. And the driver for those are a thing called an API watch. It's basically like push notifications for the Kubernetes API. Whenever an instance of my custom object changes or is created or is deleted, I want to get a notification in my code so that I can do something with that new data. Or lack of data if it's a delete. The heart of an operator is a controller. Each controller sets up some API watches, waits for a change, and then does its best to make sure that the state of the world reflects the new config. Repeat indefinitely. That's what an operator or what a controller is,

10:37

Speaker 1: sort of deep down in the heart of an operator. Before we dive into talk more of the specifics of our Django controller, uh let's pull back a bit and talk about building convergent systems. And before we do that, let's pull even further back and talk about the opposite, which is procedural systems. Um a procedural system is built by giving the system all of the steps that you want to run in the order you want to run them. Like say a bash script. Um that is the steps that you want the the code to take in the order you want it to take them. Real simple, easy to write. As opposed to a convergent system where we don't tell it what steps we want to take, we just say, this is the state that I want the system to be in when you are done. Figure it out. So if you've worked with Ansible, for instance, it is a convergent system. You don't tell it what steps you want it to take. You say this is what state I want the system to be in when you are done, and it figures out what it has to do to get there.

11:25

Speaker 1: Very briefly, two other important concepts that I'll be mentioning throughout this. Itempotence, uh, or a system being itempotent means that it take it only takes action when it needs to. To correct the state of the system. So going back to the Ansible example, if you tell it you want a package installed and the package is already installed, it doesn't try to install it again. That's all it means. Long word, simple.

11:44

Speaker 2: Promise theory is a mathematical framework for designing convergent systems by breaking them down into smaller subsystems that adhere to a specific contract of you give me this data and I will try to make the state like that. Promise theory systems can work at several levels, so we can start with small promise theory actors and compose those into bigger actors and compose those into bigger actors. In practical terms, this means taking the thing that you want to do, in our case deploying a Django app, and trying to break it down into multiple smaller convergent bits that can each be built and tested independently, which helps to reduce the surface area of your actors and keep keeps the code a lot more manageable. So we don't write a single giant function that does everything. Instead we write a small object that creates a rabbit mqv host, a small object that creates an S3 bucket, and then we use those in building our deployment.

12:31

Speaker 2: Just like, again, I keep harping on Ansible because most likely used. When you're writing a role, you can compose that out of other roles. Why does all of this design and theory matter? Because part of being successful with Kubernetes or any convergent system is to rotate your thinking from procedural to convergent. You might have noticed that my brief description of Promise Theory lines up very well with also how I briefly describe Kubernetes controllers, and that's no accident. Kubernetes uses these patterns because over the last decade or two, we figured out that when you're managing big complex distributed systems, it is far better to work with a convergent system than a procedural one. The real problem deep down, and we'll talk more about this a little bit, is state drift. Failures happen. Like there's going to be errors during deployment. And after that error occurs, you don't really know what state your system is in.

13:18

Speaker 2: You could go and find out and try and correct it, but that's very time consuming. And if you've got, say, just like a big bash deployment script with Like rsync and whole bunch of stuff. A lot of people have written those. I've definitely written those. But unless you account for every possible starting state, if the system is in an error state of some kind and it's in some weird configuration, your script may just break randomly.

13:39

Speaker 1: Like You're trying to rsync and the destination folder doesn't exist. What's it gonna do? It doesn't know how to deal with that. But if you are building a convergence system and you just tell it that folder must exist, no matter what errors kind of go on, it'll make it right and it'll eventually get you to the goal state. So controllers can be used for all kinds of things, but most of them in Kubernetes follow this pattern. On any

13:59

Speaker 2: change to an object, read that object in, generate a whole bunch of other Kubernetes goop, apply that into the Kubernetes API, or sometimes like talk to external. systems, so like the one that makes S3 instead of talking to the Kubernetes API, it talks to the uh AWS API. Um and just loop this forever. So every time there is a change, you get the object, you read what was requested, like I want an S3 bucket called foo and it should have you know permissions public. Talk to the AWS API. If the bucket doesn't exist, create it. If the permissions don't match what was requested, fix them. Repeat forever. This loop ensures that everything gets rebuilt effectively from scratch. So if we are just sort of running and I go and delete that S3 bucket, the next time it reconciles, it'll just recreate it. Don't have to care about what error state things we're in. All it wants to do is know what is the current state of the universe and what is the desired state of the universe

14:49

Speaker 2: Um every custom type will have usually one custom controller that is attached to it. Package up all of those multiple controllers and multiple custom types together, and you end up with an operator. But okay, enough about the ideas behind the system. Where's the code? Unfortunately, writing complex operators in Python is doable, but a bit tricky. Um, there's two main projects that are trying to make it easier to write Kubernetes operators in Python. One is called Cop and one the other is called Metacontroller. They're both, however, aimed at relatively simple use cases and they're pretty early in development. There is also a low-level client library that's auto-generated for Python. But then you have to write all of the custom controller and custom tools, uh, custom type specific stuff yourself, and no one has really written a thing to make that easy. So in the end, we decided to go with the more community standard kubeilder and controller runtime libraries, even though those are in Go.

15:36

Speaker 2: I'm going to summarize the things that we wrote in Python pseudocode from here, but if you go and look at my actual repository, which will be linked at the end, it will all be in Go. So so be it. So the first problem that we tackled as we were writing our custom operator was the big problem we had with Helm, sequencing and ordering.

15:53

Speaker 1: We added a lightweight state machine to figure out sort of where in the deployment process our system was. The first phase, initializing, covers the setup of databases and other underlying infrastructures, sort of like big global stuff that's usually only done once. And then we had migrations run per version, deploying perversion, and then ready. And error when that happens.

16:12

Speaker 2: This is some of this pseudocode specifically around dealing with migrations. Um this is one of the nice things about writing custom operators though, and we'll see this a couple of times throughout the talk, uh, is that because

16:22

Speaker 1: this is real code as opposed to just being a declarative system like Helm

16:26

Speaker 2: or other Yang we control it and we can make whatever behavior we want.

16:31

Speaker 1: So it allows very, very careful customization of exactly how migrations work in Django, say.

16:37

Speaker 2: Which is subtly different than other frameworks and lets us very carefully control the behavior of what happens if say uh there is an error during migrations. What do we do? Like what does that mean? In the case of Django, it doesn't attempt to auto-roll back, for example. So we have to understand that there was an error and flag that to an operator, as with some other frameworks where it'll auto-roll back and you just try again. Um I put aside Celery Beat earlier, so now's a good time to pull it back out. Um there's two problems with Celery Beat for anyone that runs it. Um the first is that if you run two copies of Celery Beat, um it's the thing that schedules all of your uh scheduled tasks. So basically it's like cron for celery. Uh and if you run two copies of it, your schedule tasks just run twice. If your tasks are all perfectly idempotent. That can sometimes be okay, but it's still gonna be a load increase.

17:23

Speaker 2: And usually you'll end up with one or two tasks that if they are run twice, bad things happen. Um it's certainly easy to happen unexpectedly, and people do not understand why. The second problem with Celery Beat is it is a stateful tool. It stores a a bit of information for every task like when it was run last. By default it stores that in a file on the local file system. There's a tool called Django Celery Beat that puts into Django Django SQL database, but that's a lot of write load. Like we have a couple dozen tasks running every, I think, five seconds. And so that would just be an enormous amount of write load on our SQL database that we didn't really want. So this brings me to a tool that I hope everyone will now be able to use called Celery BDEX. BDEX solves both of these problems.

18:06

Speaker 1: The first phase, it does locking between multiple instances. of celery beat so if you have more than one running only one will be active um and it also allows storing the celery beat state into your cache backend so either redis or memcache um which is usually a lot better for live data like that than seek There is a catch though, uh which is it is Python 3 only, and if you remember, my app is not yet on Python 3, so I actually can't use this. So be better than me, use CeleryBDX. Otherwise, uh you gotta deploy it on a stateful set and it's kind of a pain in the ass. Um

18:36

Speaker 2: so again with pseudocode, um, but this is another example of why it is nice to be able to control uh at a code level what is going on during our deployment. This is an example of we auto-generate the passwords for all of our database users as we're creating them automatically. And so we wanted to feed those passwords back into Django. I mentioned before that we put all our config into YAML, so we

18:59

Speaker 1: want to read the dynamic passwords for uh our database users, both SQL and RabbitMQ. We put those into the YAML dynamically, render some YAML, put them back in a secret that gets mounted into our container, and then the settings. py can read it as if it were a file. And because this is in the controller, anytime anything changes, so If we want to rotate all the passwords, we tell the system to update the passwords, and then it'll cascade through the system naturally.

19:24

Speaker 2: So that covers the underlying tech. What about workflow? It's just as important as having good deployment systems, having good deployment systems.

19:30

Speaker 1: Workflow. Um our original was uh fairly simple deploy. sh. It just wrapped running Ansible. Pretty simple, repeatable, but it had some issues. The biggest ones were that there was a source of truth mismatch. Um And uh that because convergence was both partial and manual, uh, things could drift over time. With the source of truth, uh the idea of mapping out sources of truth is to figure out where in your system a given piece of data is authoritative.

19:57

Speaker 2: So for example, with our Ansible code, that was authoritative in Git. Git defined what was the correct way to configure a system. But where did the information about the versions that should live on each server go. In our case, that didn't really go anywhere. It was basically authoritative on the running systems themselves based on what version was git cloned onto them. So that we had no like centralized overview of what versions were where and there was no easy way To like backup and restore that information because it was just what was on the machines. Um this also comes up if you're using Jenkins parameterized build jobs for doing deployments, same problem. The thing that records where your versions live is Jenkins build parameters. Which isn't a thing that you can easily like review and back up and restore.

20:42

Speaker 2: And then the related issue of drift, that because Doing a push becomes a manual action, it is possible for things to be missed and to slowly drift over time.

20:51

Speaker 1: Usually in our case, this meant an old test server sitting in a disused corner of our AWS account that was unpatched for months or years. So GitOps. The central idea of GitOps is that Git is your only source of truth. All information, configuration, code, everything has to live in Git somewhere, somehow. Um and then you have some automation that is reading out reading configuration and infrastructure changes out of Git and affecting them onto the universe. This approach gets you a lot of great things because the full state of your configuration is reflected in Git at all times. the automation can continuously clean up that drift. So if you make a change to your deployment system somehow, it will take effect across every machine simultaneously. That also means that you can take down all the prod simultaneously, so make sure you have good tests. But it is nice that things will never be out of sync.

21:39

Speaker 1: It also means that if there's an emergency, you have all of your stuff in one place. I can push one button and restore my entire infrastructure. because it's all in one place and that is the authoritative source. It also helps just bringing a more code-based workflow to operational changes. You get things like code reviews and commit messages and things like that. It even uh it can be used as a kind of dis count audit log of seeing who made changes when and hopefully why. But it's not perfect.

22:05

Speaker 2: The biggest friction we've had in switching our teams to GitOps

22:08

Speaker 1: is frustration that they have to get what they feel are minor operational changes fully through a code review process. We're currently using the GitHub branch protection rule system, and it's it's got some flexibility in determining who does reviews, but not a lot of flexibility in what should be reviewed. So our goal is to eventually try to replace that with a custom review management system, but we haven't yet built that, so there's been a lot of frustration from some of our team members. Also, we've had some team members that are engineering adjacent. They're not working in code day-to-day, but they do need to interact with the deployment system. And they didn't necessarily have GitHub accounts. They didn't know how to use it very well, so we've had to guide them through that. It's been okay since then, but it was definitely a thing that we sort of forgot to factor in is that there are people that don't sit in GitHub all day.

22:54

Speaker 1: And of course, uh if your team members uh don't fill in good commit messages or don't review things things consistently then that data source is not going to be very useful. We do have a lot of uh update foo. yaml as the commit message, which so it goes. So put it all together and we get a new end-to-end workflow based on GitOps. Um we use a tool called Argo CD to do the critical watch your Git repositories and sync them into Kubernetes. Um but there's a lot of tools around there. Find one that works for you you the general idea is to have GitHub be or Git in general be the source of truth. Alright, some bonus things. we do have a few extra minutes. Second piece of the puzzle that we built was a CLI tool to handle common tasks that our engineers would come up with. The first two of these, for example, are just wrappers around the

23:41

Speaker 1: Cuddle Xec uh command line tool to give you a bash shell or a Python shell on a given instance of the application. The last one is a wrapper around PSQL to pull the right database information, user password, host name, and just feed it into PSQL so they didn't have to go find that every time. We've been adding more debugging assistance tools over our uh to it over time and it's been generally I think well received. We took a queue from Homebrew. And we added a command to our thing to both uh examine environmental issues and potentially auto fix them if it knows how to. Um this has been very helpful in just getting more people set up on the new platform. Um we can just tell them, hey, please run right And paste me the output. And uh this is still mostly in testing. We haven't rolled it out to production yet. Um as a as an ops person, I like command line tools, so I'm cool

24:26

Speaker 1: just like kubectl describe as my

24:28

Speaker 2: interface to a lot of things, but not everyone on all of our teams feels the same way, so we started to build a fairly simple read-only user interface so that they could navigate to things via the web and see nice graphical pretty things. Um so while our Django deployment logic in this thing is definitely fairly tailored to our needs, it's still pretty general Django, uh, and it is all open source, which is nice. It is not particularly well documented, but it is open if you want to go look at it. Steal ideas, steal whole sections of code. That is highly recommended. And yeah, thank you so much for having me. Any questions?

25:08

Speaker 3: Hey, thanks for the talk. It was really insightful. Um one of my biggest fears from moving to a PAAS like Heroku to a More flexible Kubernetes environment is Heroku does so much for me with um managed Redis, Postgres, like they tell me when maintenance needs to be done, they do it. Have you still have you I guess my question is have you still kept that managed or you are you managing that yourself?

25:37

Speaker 2: Um For most of Postgres, we use RDS. For uh RabbitMQ, we use Cloud MQP. Redis in particular, we decided to run ourselves because Uh Amazon does offer Elastichash Redis, but it is remarkably expensive and Redis is pretty easy to run. But yeah, I will point out that you can actually use Heroku Postgres without using Heroku. They don't like super advertise that fact, but if you're on Amazon, it's running in Amazon. You can just keep using Heroku Postgres if you like it.

26:05

Speaker 3: I like that, yeah. Okay, thank you very much.

26:09

Speaker 4: Hi, yeah, thanks for the talk. Um I have a question. You know, as someone who's tinkered with Docker and, you know, kind of been getting into these types of technologies, maybe even using them in production, y wha how would you recommend someone diving into Kubernetes as kind of the next next level.

26:27

Speaker 2: Really just what I showed in the talk of start writing some YAML. Like pick an application. The one I generally recommend is WordPress because it is really like it's got a slick installation process and there's a million and a half guides to every possible flavor of WordPress thing you could want to do. So just like start writing some big YAML files to deploy it, like deploy, I guess, MySQL if you're doing WordPress, deploy a WordPress container. Like it's got a fairly well-polished Docker image already, so you don't have to make that yourself. and just try deploying stuff. Like GKE is really easy to get started with. And if you remember to shut it down when you're done using it, it's pretty cheap. So yeah, that's that's where I begin.

27:09

Speaker 5: Hey uh I was just curious when you were in the section regarding GitOps and you're talking about getting people engaged in that process. You said that there was difficulty with people who weren't engineers being engaged with that process. Who is not that's not an engineer cares enough about your deploys to get involved in that? Anyway.

27:30

Speaker 2: Mostly the product management team and the release management team. The release management team has some engineers on it, but it also has more PME kind of people that do want to be involved. They want to like see when deploys are happening because it's part of their checklists, but they're not like living and breathing and

27:56

Speaker 5: So when you uh as you finish up that uh more like friendly UI is the intent that a lot of those people would just be pushed towards that because it's mostly about communicating to them, not necessarily them contributing.

28:08

Speaker 2: Yeah, probably. If they want to be able to see like system status. And we're also working on sort of custom tools, especially for the risk management team. like they really want to be able to see uh not just what version is in one place, they want a big overview of like which given a you know what versions are in prod and where are they. Um but because everything is just in YAML files, it's pretty easy for us to just just write a little script that like generates them a nightly report or whatever and emails it to them. So a mix of that web UI and just writing little custom integrations that either talk to the Kubernetes API directly or just read and parse YAML files out

28:41

Speaker 5: Cool. Thanks.

28:42

Speaker 2: Awesome. Thank you so much and enjoy your break. No, no, no.

Questions this talk answers

What does a forklift migration to Kubernetes mean for a Django app?

It means taking an existing application largely as-is and deploying it on Kubernetes, rather than rewriting it as a greenfield project. Noah’s team kept RDS and CloudAMQP so the production migration would not require moving data.

Discussed at 2:34

How do you containerize a Python or Django application efficiently?

Use a multistage Docker build: install the larger build-time dependencies in one image, then copy only the required files into a smaller runtime image. The runtime container should also run the application as a non-root user.

Discussed at 3:38

What is a Kubernetes operator and how does it work?

An operator combines custom resource definitions, API watches, and controllers. A controller watches for changes to a custom resource and repeatedly reconciles the actual Kubernetes or external-system state until it matches the requested state.

Discussed at 9:06

Why use a convergent deployment system instead of a procedural deployment script?

A procedural script assumes a known starting state and can fail after an error or partial deployment. A convergent system declares the desired state and keeps reconciling toward it, so it can repair drift or recreate missing resources regardless of the current state.

Discussed at 12:31

How can Kubernetes deployments run Django migrations in the right order?

The team used a lightweight state machine in its custom operator, with stages for infrastructure initialization, migrations for each version, deployment, readiness, and errors. Because the operator is code rather than just declarative templates, it can precisely handle migration failures and prevent new code from rolling out too early.

Discussed at 15:53

How can you safely run Celery Beat in Kubernetes?

Celery Beat needs protection against multiple active instances because duplicate schedulers can run tasks twice, and its state must be stored somewhere durable. Celery BDX addresses both by locking between instances and storing Beat state in Redis or Memcached; otherwise, a stateful deployment is needed.

Discussed at 17:23

What is GitOps, and what problems does it solve?

GitOps makes Git the single source of truth for application versions, configuration, and infrastructure, with automation applying Git’s state to the environment. This provides reviewable operational changes, a history of who changed what, continuous correction of drift, and the ability to restore infrastructure from one authoritative place.

Discussed at 20:51

Should databases and services be managed inside Kubernetes or kept managed externally?

Noah’s team kept most PostgreSQL on RDS and RabbitMQ on CloudAMQP to avoid the operational burden of managing databases, while running Redis themselves because managed Redis was expensive and Redis was relatively easy to operate. He also noted that Heroku Postgres can be used independently of Heroku.

Discussed at 25:37

How should someone get started learning Kubernetes?

Start by deploying a familiar application, such as WordPress, using Kubernetes YAML; deploy its database and application container and experiment with the setup. A managed cluster such as GKE makes initial experimentation easier and relatively inexpensive if shut down when unused.

Discussed at 26:27

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Noah Kantrowitz

More videos from DjangoCon US