Fighting Homelessness with Django with Benjamin "Zags" Zagorsky
Published December 6, 2024
This video features Benjamin "Zags" Zagorsky at DjangoCon US 2022 in San Diego, California, USA.
Django's migration system is one of its greatest strengths as a framework. It can automatically generate migrations based on your changes to your models and can detect which migrations need to be applied to a database. But, as the size of your development team and user base scale, there are pitfalls that you need to be aware of. Not all migrations can be safely reversed, and trying to rewind bad migrations on a production database can cause a data disaster. Not all migrations can be safely deployed without downtime, and trying to deploy them can give your users and your engineers a wall of errors.
This talk will cover the following:
This talk was presented at: https://2022.djangocon.us/talks/django-migrations-pitfalls-and-solutions/
LINKS:
Follow Benjamin "Zags" Zagorsky 👇
Follow DjangCon US 👇
https://twitter.com/djangocon
Follow DEFNA 👇
https://twitter.com/defnado
https://www.defna.org/
Django migrations let model definitions act as the practical source of truth for database schemas: `makemigrations` records schema changes in version-controlled files, while `migrate` applies them to a particular database and tracks applied migrations in the migration recorder. Benjamin Zagorsky explains how divergent branches, migration reversal, destructive changes, and failed production deployments cause trouble, and shows how `RunPython`, staged changes, backups, maintenance mode, and testing against production-like data reduce the risks. He argues that backwards-compatible migrations are the safest default for zero-downtime deployments and branch switching, while complex changes may require multiple releases or scheduled downtime.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Hello Django Con. Who's excited for migrations? Alright. So so I'm Zags. As you can see, that's short for Benjamin. And I am the Chief Technology Officer at Zagaran. What is Zagoran? We're a software consultancy. I'm one of the founders. It's named after its founders. hence the name similarity. And we do full stack, web, mobile, back end, server management, the whole gambit. Django is our primary backend technology, because it's awesome. And One of the great things about software consulting is that I get to see problems across a whole range of industries, and one of the things I see is the same problems over and over again. So part of what I'm going to be talking to you about is, well, how can you learn from my experience of the same problems everybody has
and not necessarily have to do some of these things the hard way. One of the other fascinating things about migrations, one of the reasons that this is a topic near and dear to me, is that it's not just about back-end. It's actually at the intersection of back-end databases and server management. It's about a whole set of parts of the application stack that collide in potentially nasty ways. So what we're going to start with is some migration fundamentals. I mean hopefully all of you have at least run the command migrate before in your lives, but we're gonna dig a little deeper on those, get a little more background on the stuff that you need to know Uh and then we're going to talk about things that can go wrong. We got four big ones here. What happens when you got multiple branches? What happens when you try and undo your migrations? How to get those zero downtime deployments on production.
There's actually migrations things you need to do, not just infrastructure. Um and then what happens when everything goes wrong? So let's let's jump into it. Migrations When you boil down to it, are two management commands. We got make migrations and we got migrate. But taking a step back, the point of migrations And the reason that Django is so awesome as a web framework is because migrations enable you to write your code as if your models are the source of truth for your database schema. And migrations are what make that lie almost true. So when you change your models, you run make migrations, and Django is gonna generate for you migrations that describe it goes and looks at what migrations you have already. Let's add all those up. Let's look at the difference between that.
And your models and generate new migrations to capture that new delta. And most of the time, the auto-generated migrations are right. Most of the time that is actually what you want. And the times when it isn't is basically what we're gonna talk about for the next 20 -something minutes. So those migration files, those live in your code base. You can look at them, you can edit them. And they ship with your code. They should. Those part of why they're in the code base is because those live with those model changes. You can ship them on the same branch. And then when it comes time to actually apply those schema changes to your database, that's where migrate comes into play. And at this point, we're not talking about a particular database. You can have probably do have several. You've got your local development, your staging, your production. When you run migrate, we're talking about a particular database. And just to look under the hood, Django's got a table in which it's tracking your migrations, which ones have been applied on that database.
There has to be a table in the database because otherwise, where else would it be? You can actually look at this. It's called the migration recorder. You can go query it. You can see what's in it. Fundamentally, there's two fields in there you care about. Migration name and applied timestamp. That's what it keeps track of. And when you run migrate, migrate's gonna go look at the migration recorder, see what migrations have already been applied on this database, go look at your migrations, see which ones aren't listed in the migration recorder, and it's gonna go apply those. Looks like this. So you run migrate says, all right, we got 15 and 16, haven't been applied, we're gonna go apply them. That's not all you can do with migrations. You can go backwards to Backwards is a little weird in syntax. You don't say undo this migration.
You say please rewind to this particular migration's state. So if we want to undo 16, we don't say undo 16, we say go back to 15. Um part of the reason for this is because this way you can run a rewind agnostic to which migrations have been applied on that database. So If 16 hasn't been run, you can say please get me to migration state 15, it can say already done. Um means this is an item potent operation, one of my favorite words in programming. One other thing that I want to call out, one of the best tools in migrations for some of these nasty situations, is run Python. You can add your own Python code that's going to be run as part of your migrations And run Python is actually the tool for all the low-level pitfalls.
Let's say you want to add a constraint to a database field. Right? You've got a field, it was previously nullable. Shouldn't be anymore, but you still have null data in there. RunPython is your answer. Run Python is for doing those accompanying data modifications that go along with your schema modifications. RunPython has two arguments. This is going to become relevant very soon. A forwards function and a backwards function. The forwards function, that's the code that runs when you're running the migration normally. Backwards function, that's what happens when you're reversing the migration. What could possibly go wrong? So, first of all, let's say you have multiple branches. And each of your multiple branches is making database changes. So you're over here happily working on feature one. You've made your database changes.
And Your buddy says, hey, I need you to go help me on feature two, do some testing, do some code review, whatever. And so you go and check out feature two, which has also made some database changes. And you go run the app and it crashes because, well, you're you've got some some different models on the feature shoot. I forgot to run migrations. Oops, all right. So you run migrate and it still crashes. Because well, what's happened was you've run your migrations over here on feature one. Okay, those model changes, they're not on feature two. These are currently divergent feature branches. You say, all right, I need to rewind migrations. Let's rewind to the migration state of Maine. So you say migrate back to migration 18, Django says nothing to do. Huh? So, Django, in order to remote rewind a migration, needs to have that code of the migration.
The migration recorder is not tracking what were the operations that you did. It's just tracking the name of the migration. And so if you try and rewind to a state, if there's migrations in the migration recorder that aren't present in the code, and you say rewind to a state before those. Django is going to ignore those migrations as part of the rewind. And this is a great feature when we when you get into squashing migrations, things like that, but for this particular case, it's actually really obnoxious. So the solution, pretty straightforward, you have to reverse your migrations before switching branches. Okay? Sounds Kind of obvious when you lay it out like this, but you gotta reverse migrations, then switch branches, then migrate. That's not the only solution. I got two other options for you. If you have a backup of your database from that common branch, if you've got a backup from main
, You can just restore that back up and run migrations from there. For local development, this is great, especially if you're using like SQLite. That's just a file you can copy around. That's really easy. And also just to tease two sections forwards, if you've got backwards compatible migrations, you don't You don't have to worry. You don't even have to rewind them. It's great. This is why you care about backwards compatible migrations, among other reasons. One other thing I want to call out here is this is not just about local development. If you have a shared staging environment where people are deploying different feature branches too, you have this same problem. Great way to tank your staging environment if you're not paying attention to this. Now, if you're using like CI to deploy your staging environment, that first option is really obnoxious because that first option requires you to actually have some introspection into the states of migrations on different branches
Which is part of why I've got these other two options on here. If you're doing a full CI deployment to staging, you probably want to consider number two or three over here of either having some canonical staging database that you restore to. That you update periodically as you know new releases happen, or be using backwards-compatible migrations. The other thing that happens with branches is, well, what happens when you're done with your branches? The chickens come home to roost. Okay, you merge the branches back in, all good, you run migrate, it crashes. This error is delightful. It tells you how to solve it right there in the error message. Please run make migrations-merge. Now, if Django knows how to solve this, why is it bothering to error? Because that is not always the solution. Now, just to like
unravel the the onion here of of the various merging that we're doing, first Git has merged your code. And Git is gonna check for merge conflicts in your models. But it's not going to check for merge conflicts in your migrations because those are separate files. So from Git's perspective, those are completely mergeable. But Django knows better, or knows maybe better, that you actually need to check to make sure those migrations are compatible. That they can actually be run in an arbitrary order and that they're not trying to do conflicting operations. So this is just a step for you to get the human, the developer in the loop and actually check those. Heuristic If you don't have merge conflicts in your models, the migrations are also probably fine to merge, but that's not always true. So do check them.
That's why this happens. So we've talked about one of the big solutions for branches is reversing your migrations. Wouldn't it be great if you could always do that? Well, why couldn't you? So, first off, some migrations will just straight up error if you try and reverse them. I showed you run Python a few minutes ago. That second argument, that reverse function, is unfortunately optional. If you don't pass it, your migration can't be reversed. It will error. Now, there's this is very simple to solve. Empty function, no op. There's even a no-op that ships with the migration framework that you can fill in. So if you're using runPythons, just fill in a reverse function, even if it's a no-op, so that you can literally reverse the migration.
Um but there's other reasons you actually want to use the backwards functionality of runPython. So if you're removing a constraint, okay, you had a field, it was previously unique. And you say, nah, this field can have duplicate values now. You drop the unique constraint. And then you add data that has duplicate values in that column. And you want to reverse that migration. Well, the opposite of removing a constraint is adding it. And you are now trying to add a unique constraint to a column with duplicate values, that's gonna crash. So the answer here is run Python reverse. This is one of the things it's for. It's for trying to clean up that data mess that may be there so that you can undo some of these my database changes if you're trying to reverse your your migrations. So having a runPython reverse function. Here you could even do runPython no op forwards, just a reverse function, just to clean up this data so you can actually reverse that constraint
drop. One other thing that can go horribly wrong, this is not an exhaustive list, by the way, but one other thing that can crash, if you try and delete a field, it's non-knowable, no default. You can't Because the opposite of that is adding a non-nullable field with no default, which should be obvious. You can't do that. What what data would be in the field? The solution here, you gotta split it into steps. You gotta first change that to be a field that is nullable or has a default. And then you can drop it. And you probably need to run Python in the middle if you're not going to give it a temporary default as part of deleting it, so that you can literally rewind these without erroring. But that's not the only thing that can go wrong. Deletion has other problems. And the problem is that reversing a migration isn't time travel. It's just undoing a schema change.
So if you have this perfectly normal char field, this is not a magic trick, right? Take this char field, it's got a default, you delete it, you migrate, you reverse the migration, your data's gone. This column is gonna have just the default for every row. And that's because what you did was you dropped a column and then you re-added the column with a default. I don't even know how Django migrations could solve this problem other than by keeping around shadow copies. of dropped columns, which would be insane. So how do we fix this? First off, back up your database. This is not going to be the only time I recommend this. The reason backing up your database is great is because database backups are time travel. Now that's not particular you can just restore the backup and be back to before you ran the migrations.
That's not particularly helpful if you also want to keep any other data changes that happened since Since then, if this was like something on production and it's been live for a day and you need to undo this somehow, that's a problem. But a backup will still help because a backup can still let you really deconstruct that column from the backup. You can still get the deleted data because it's over there in your backup. There's other options too Another really good preventative measure is have your fields be knowable for some time before deleting them. You can have that field be knowable, don't talk about it in your code, just leave it in the database for a full release, and then next release you can drop it after you're sure you don't need to roll back that that release. And right, changing a field from knowable back to not knowable, that's a lot easier than restoring a deleted column that you have to reconstruct from a backup. Second that second point that's gonna come back again when we talk about backwards compatibility, so remember that.
Um another thing you can do, run Python backwards functions. They are magic from a reversibility standpoint. This is great if what you've done is you've you've deleted a column because you've moved that data somewhere else. So if you've split it into two columns or put it in a different table, you can use a backwards function to reconstruct that column from where that data got moved to. And at the very least, use the no-op so that you can reverse the migration without it crashing. So we've talked about that's how to run your migrations backwards safely. How about running them forwards safely? And here's where you go, Zags , what are you talking about? Backwards is the weird edge case. Forwards is what migrations are designed for. Right? What could go wrong? Right? Aren't they safe forwards? Race conditions. When you deploy your code to a server, you have to do two things.
You have to update your database schema and you have to update the code in your server. And no matter which order you do this in, a request can come in in between. Now, for the rest of this section, we're gonna talk about updating the database first, server second. This is the way you want to do it. This is the way you want to do it because not only does this mimic how code updates happen, In development and on staging, but it also optimizes for additions, adding fields to your database, which I would argue is the far more common operation. So fundamentally the situation we're talking about here is this middle state. Django is designed for the sides. It's designed for when you've got model state one, database state one, model state two, database state two. Django migrations are all about updating your schema to match your models. But in the middle of a deploy, we're in this weird intermediate where we've updated the schema, but we haven't updated the model.
Migrations aren't guaranteed to work here, but some of them do. And the ones that do are called backwards compatible. Because, well, that is a change to the database that still works with your old code. There's a lot of things that are backwards compatible. Okay, right off the bat, things that don't change the schema. The whole problem with migrations. Not being backwards compatible is a schema code mismatch. If you're not changing the schema, that's you know, that's trivially backwards compatible. And there's actually a lot that falls under this category. Run Python. You're changing your data now, your schema. That's backwards compatible. Changing some of the vanity attributes of models. Things like blank or choices don't actually generate schema generator and any schema changes. Or squashing migration. That's just administration of your migrations themselves if you're not actually changing the schema.
So these guys, you can deploy. No downtime, no problem. But there's more on this list. If you deploy your database changes first, code changes second, you can add nullable fields all day. Because if you add a nullable field, old version of your code is running that doesn't know about that field. It's just gonna make objects and whatever your database is will fill those in as null. It's gonna be querying objects and not talk about this field. And that's fine. You can add a model and an old version of your code that's not talking about that table. Doesn't care. You can remove constraints. You can add constraints as long as all of your code and data are already satisfying it, and you can remove a model if you're not talking about it. There's a great list, but it's also not everything you'd want to do. So what about that other stuff? Well, the question is how can we get other stuff to look like that list?
So if you want to do a rename, you can do that without doing any schema changes. As long as you're willing to tolerate legacy names in your database. So if you use the db column or the db table attributes, you can do renames, rename it everywhere in your app, not actually change your database schema. Pretty cool. You can even move a model between apps without doing any database schema changes. You gotta pull in the separate database and state. But that can be done completely backwards compatible. You can switch branches in development or on staging all day, no problems, zero downtime to production. More complicated stuff, we can get there by decomposing it into two deploys. So if you want to add a constraint, well to add a constraint, what we have to do is first Start acting like we have that constraint already.
Get all of our code to start satisfying that constraint. Okay, you want to make a field not knowable? Stop setting null. And then once that code is everywhere, once that's on production, merged into main, everybody is developing off of that. Then in the next release, we can do a run Python to fill in. Any null values that are already there, and then at that point, adding a constraint, all our code and all of our data already satisfies it, that's backwards compatible. Removing a model, same story. Remove all code references to it in one release, delete the model in the next. And what this looks like here, right, so you do the actual feature, all the code that you you care about, right? That goes in ships, that's in version 1. 8, and then you can do you know a little bit of cleanup, merge some some some database migrations are and actually adding the constraint into 1. 9. And so each step going from 1.
7 to 1. 8, that's a backwards compatible release. 1. 8 to 19, that's also backwards compatible, even though the whole story is not. You couldn't go straight from 1. 7 to 1. 9 with zero downtime, but each step does does work. Removing a field, this one can also be decomposed into two releases with help. There's the Django Deprecate Fields library, gives you the deprecate field. Option, uh uh function, you can put that around a field. And what that does is in that first release it makes that field knowable. Convenient for reversibility too, no? And then in the next really and then you stop talking about that field. And then in your next release you can delete that field. Um and this both gives you the backwards compatible deletion of a field and also much safer from a reversibility standpoint.
While I believe in theory every change can be made backwards compatible, there is definitely a trade-off in terms of development effort. So there are some things that Depends on your use case, but in many use cases are not worth the effort. And this is what scheduled downtime is for maintenance mode. So for some of these complicated changes, if you're going to be splitting or merging fields or models, if you're going to be changing the type of a field, right, things like that. Here it's useful to have scheduled downtime. And what scheduled downtime is is what you want is a maintenance mode that where your site rejects. all requests. I mean it doesn't reject them. It says sites down for maintenance. Those don't hit the database. Okay, if this is a web you know site, if you've got background tasks, you want to make sure that that task queue is not, you know, you aren't don't have workers picking up from the task
you right. So the point is whatever you're doing, have a period where nothing's hitting that database so that you can deploy changes to the database and code. Together without those pesky requests coming in in between that gives you such a headache. But backwards compatibility is not just about zero downtime deploys. It's also about switching branches safely in development and staging. So if you are going to be doing this for a particular change, but you're using backwards compatible migrations everywhere else, make sure to give that branch extra care, shepherding through development and testing. All right, so so so far we've you know we've been dealing with kind of small things that can go wrong. What about the doozy? You're in the middle of deploying to production and your migrations fail. There's some good news, right? Well, migrations are atomic by default.
Please keep them that way. Uh what that means is that if a migration fails, it will get rolled back. But if you have a group of migrations, the whole group is not atomic. So if you're running five migrations, the first two work, the third one fails, you're not two migrations in. And your deploy has hopefully been aborted. I strongly recommend don't deploy the code after the uh failed migrations. Another reason to do migrations first in the deploy, that's the riskiest operation in a deploy. So aborting the deploy on failed migrations, that's uh it's just kind of a good circuit breaker. Um if your migrations are not backwards compatible, you are in a world of hurt right now. So another reason to use backwards compatible migrations. If you think you can deploy non-backwards compatible migrations, you don't need a maintenance mode because you deploy
super fast under a minute, you're not going to get that many requests. And your migrations fail part way through. Well that's not that fast to deploy anymore. You've now got potentially an hour or two while you sort out this mess. And that's a lot of time for requests to come in that are gonna fail. So Have a maintenance mode ready, definitely put it up if you haven't at this point and your migrations aren't backwards compatible. But the other problem is my preferred way when I get into trouble to fix a database. Manage that PyShell, you know, gives me access to my database and all of my other code, all of those constants and helper functions that make it easy to manage my data, part of why Django is great. If your migrations aren't backwards compatible, that's not gonna work because you are potentially in a state where not only is your database schema not compatible with the old version of your code that's on your servers, it's also not compatible with the new version of your code that you're trying to deploy.
You might not have a version of your code where you can just manage that PyShell and talk about the model in question So, how do we fix this? First, don't be here. Um, test your migrations on a copy of production data before you run them. It's not enough to test your migrations on staging because the success of a migration is data dependent. If you're adding a constraint, That will work on a database where all data meets that constraint and fail on a database where it doesn't. Now, this test this is good, but it is not foolproof. If you take a snapshot in the morning, you test your migrations on that, it works, and since then new data comes in that's gonna break your migrations, you can still end up in this situation. So, not foolproof. The only way to make that foolproof is if you turn on downtime, test your migrations, and then run them.
So if you're in this situation, how do you fix it? Manage. py shell, that's still going to be my recommended go-to. If your migrations are backwards compatible, it will work. Even if they aren't. It still might, um, but not necessarily. So if you don't have backwards compatible migrations, next up, uh you can reverse the migrations. Hey, isn't that great? Well, it's assuming your migrations are reversible. Now, quirk here, right, the reason that's that's a fairly complex sentence on the slide is that Well, you need to have the code of the migrations to reverse them. And if you aborted your deploy, that migration code isn't actually on the servers, it's still on the old version. So you gotta get the migration code onto the server somehow to then reverse the migrations, assuming your migrations are reversible. And then you could go jump into Matchup High Shell or something else.
Third option, database backups. They are time travel. If you were doing this under downtime, this is great. Great way to do backups with a big database change. Set downtime, take a backup, then do the deploy. Then you can restore to the backup. There's been no changes happening. Last option here, really if nothing else works, manage. py db shell. Get right into that Postgres terminal and debug it without any of the assistance of any of your Python code constants or any of that stuff. I have that last, not because Postgres isn't great, but because there's a lot of other assistance you really want. Alright, so top five recommendations. Backwards compatible migrations are great. If you're not going to use them, make sure to reverse before switching branches
and use a maintenance mode when deploying Run Python backwards functions. That is how you can get reversible migrations real easy. Make fields nullable, delete them later. Great for both reversibility and backwards compatibility. Test your migrations on a copy of production data before deploying so you don't end up in that disaster scenario. And just take the database backups all the time. You should have regularly scheduled backups. Only back up the data you want to keep. And definitely back up before making major changes. That's all I've got for you. This is my email. If you've got questions for me or need contract software work, I'm also going to be out in the hallway If you want a copy of my slides, I posted them to the Salon F Slack channel right at the beginning of this talk. So thank you very much.
`makemigrations` compares your models with the existing migration history and generates migration files describing the changes. `migrate` applies the unapplied migration files to a particular database, tracking them in Django’s migration recorder table.
Discussed at 1:56Reverse the migrations from the current branch before switching, then run the other branch’s migrations. Alternatively, restore a database backup from the shared branch or use backwards-compatible migrations so the branches can coexist safely.
Discussed at 6:38After merging the branches, inspect the migration files rather than blindly following Django’s suggestion to run `makemigrations --merge`. Confirm that the migrations are compatible and can run in either order; model-level Git conflicts do not guarantee that the migrations are safe.
Discussed at 8:57Give `RunPython` operations a reverse function, even if it is a no-op, and use reverse functions to clean up or reconstruct data when necessary. For destructive changes, split the work into stages—for example, make a field nullable first, clean up its data, and only then remove it.
Discussed at 9:43No. Reversing a migration restores the schema, not the deleted values; re-adding a dropped column generally fills it with its default. Backups can recover the old data, and making fields nullable for a release before deleting them reduces the risk.
Discussed at 12:03Deploy database changes before code changes and make each migration backwards-compatible with the old application version. Add nullable fields, avoid incompatible changes, and split constraints, renames, removals, or other complex changes across multiple releases; use maintenance mode for changes that cannot reasonably be decomposed.
Discussed at 14:26Keep migrations atomic, abort the code deployment, and enable maintenance mode if the migration is not backwards-compatible. Test against a copy of production data beforehand; if failure occurs, use Django’s shell or reverse the migration when possible, and restore a backup or use the database shell as a last resort.
Discussed at 22:15Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026