Django Through the Years with Andrew Godwin
Published November 3, 2022
This video features Andrew Godwin at DjangoCon US 2014 in Portland, Oregon, USA.
By, Andrew Godwin
An in-depth look at Django's new migrations framework, explaining the component architecture, highlighting issues with multiple database backends, and showing how management commands typically get routed through the software.
Help us caption & translate this video!
Andrew Godwin traces migrations from the early SQL-file approach and Django Evolution through South’s development, adoption, and eventual incorporation of a new migrations framework into Django 1.7. He explains that Django migrations were redesigned rather than simply ported from South, separating schema editing from migration execution and introducing deconstructible fields, historical app registries, declarative operations, migration graphs, state tracking, autodetection, optimization, and support for data or custom SQL migrations. The design represents model changes as operations that can update an in-memory state quickly, while the schema editor handles database-specific DDL, including SQLite’s table-rebuild workaround. He argues that migrations are modular, extensible, and manageable once their underlying concepts are understood, though dependency handling and optimization remain technically complex.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Thanks Russ for a very nice introduction. Um hello everyone. As uh Russ mentioned, I am Andrew. Uh I am previously famous, I said still famous. Famous is a wrong word. I am the author of South Famous, yes. Um, and of course of the now released 17 Django Migrations, uh and a senior software engineer at Eventbrite, where I am sort of on our architecture team doing database stuff and sort of general difficult design decisions and things that are sort of very analogous to what I've been doing in Django. So what this talk is about is essentially The history of South, why it came about, how it developed, and then what became of the 1. 7 Django migrations, and more importantly, how those migrations were developed. uh sort of what the design decisions were, how they're implemented and sort of the pitfalls and the issues that I came across.
I want you to come away from this talk with an idea of like how that stuff works and not be scared of migrations. One of the biggest injections of new code into Django in the last couple of years, I think. It's mostly very clean, of which I'm very proud. It's a little bit complicated. There's some stuff in there that is very computer sciencey, which means it's very difficult to understand sometimes, but I'd love to try and explain to you and get some of the ideas across. So as uh Russ mentioned I have been to every single Django Con in the US. This is me at the first one looking very young and slightly scared of a stage. And in fact, I have, I believe, everyone I've been to given a presentation on migrations or nearly or nearly thereabouts. I have put out the first slide deck I ever did um on Speaker Deck for you. Um this will be on my website after things so you can read the URL. Um it's a very short eight slide deck, but it's it's interesting that how that first slide deck and that first introduction of South
is very similar in some ways to we have today and different in others. But after all this time, after the past seven or so years of doing migrations I've concluded that they're pretty good. Um this is my only conclusion, really. Uh there's not much more you can say. Some people complain, some people like them, but You know, in the end, people seem to really like migration. South became the sort of integral part of Django. And people are saying, oh, you know, every tutorial went, you need to install Django, and then you need to install South, and then you can keep going. And for a long time, South is a sort of like You know, it started off being compassor and then it was the only player in the field. And then people were like, well, why isn't it in the docs? Why isn't it officially supported? And then finally, as as you learn, it's now 1. 7. But How did that kind of thing come about? So I need to sort of step through the history of South Fear. So South was initially launched in August of two thousand and eight
It was launched in fact from the second to the right window of this building here. This is the stable block in Cornbury Park in Oxfordshire, which is a very lovely place to launch code from. I was working at the time for an agency called Torchbox. uh who are sort of small development agency in Oxfordshire and uh they have for some reason um offices in a country estate um so this sort of fits the British stereotype very it could be Downton Abbey like it's a British stereotype like a I launched my code from this wonderful country manor, it's fantastic kind of thing. So the the first few versions were launched from there and they were sort of made for a single project. At the time we were making a very simple sort of CMS for a customer. I say simple, moderately complicated. And previously at the company, the method was we have a directory full of SQL files, we run the SQL files, and we just sort of keep the number in our heads.
That then improved to we have a column with a number in it, we run the numbers bigger than that number, then update the number. And then at some point we sat down and me and Nick Birch, my colleague at the time um said that sort of thought we we want some sort of more formal system to do this, some kind of some kind of better system. And out of that um South was born. Curiously, uh Simon Willison, who was also working with Nick at the same company the year previously launched almost exactly the same month I think D Migrations which was sort of his version of that and looks shockingly similar so I think Nick had influenced both us on that. But in the end um D Migrations was MySQL only and that sort of was one of its demises on that front The other competitor at the time was Django Evolution, done by our wonderful value in Russell Keith McGee. It was I select I I used it certainly for a couple of years before I wrote south.
quite popular at the time. Um I'm not sure why it fell out of popularity. Um it did. Uh it was in some ways sad to lose competitor because like as you can see If you look at the times, you know, my releases grew further and further apart as sort of the pressures of of having other things didn't keep up with me. Now that's all of the early years. South at this point, before South 0. 4, there is no auto detection. You call South and you go, hi South, dash-add field this, dash-add model this. dash dash remove model this. And then I believe it was migratory that came in with the idea of a sort of auto detection and storing state. And then I summarily nixed that from them and put it into 0. 4 , of which I I I thank them very much for
coming up with the idea and a decent implementation too that I could sort of heavily borrow from. But so sort of South as you know today probably appears more around 0. 4 than 0. 1. Um 0. 1 migrations I believe still run on 1. 0. There is still full backwards compatibility, which I'm very proud of I think only I have migrations from the first few versions that really exist anymore. But so from 0. 4 or 0. 5 you start seeing things like different operations, the auto detection gets a bit better. it sort of becomes more of a thing where you can simply go change models, run so schema migration, and then have migration. That's fantastic And then sort of as we get to 2009, 2010, um we approached ability. And very optimistic me, with the bear in mind, this is four years ago. Um In the short term, South
1. 0 will be released, says Andrew of 2010. He's very wrong. He doesn't realize it yet, but he's very wrong. This is a period where sort of South becomes pretty much what is required of it. mostly by the community, like there's a lot of issues, a lot of missing holes, but it's good enough. Like the big horrible bug reports stuff coming in. People generally sort of work around the other issues and it's not perfect, but it does it does the job. And then sort of between 2010 and 2013 there are no release there are no major releases. I call it sort of the long gap. Part of this is that you know at that point I left university, I was in the world of work, I was freelancing, that's reasonably high stress. Um and then finally in May 2013 at DjangoCon Europe or thereabouts I released 0. 8, which essentially just
a collection of bug fixes and a few new features that accrued over the years. At the same time, roughly I launched the Kickstarter, which funded the remainder of the work until this day. So I asked for two and a half thousand pounds. I got, as you can see here, almost 18,000 pounds. I'm very grateful to everyone who contributed. many of them here. I've shaken at least a few of them by the hand. If if you gave, please come to me and I would love to congratulate you again. This essentially funded the work in Django. The idea of this was that I I could go down to three days a week at my job at Laniad. I could then work one or two days a week on Django and get the work done and then rather than having it done sort of a mysterious timescale, we could approach that almost like Almost being paid to work on it. Like there's a guaranteed time. You can have this sort of fixed time to not only write code but review bugs, communicate with people.
I experienced a little bit of that in having like a set amount of time per week to do this kind of stuff. It's absolutely invaluable for this. And then sort of Again, there was sort of a long period of me working on Jenga migrations and so I wasn't working on South and then uh a couple of months ago, well I think it's last month uh 1. 0 was released and then of course yesterday 1. 7 was released. And that's sort of the culmination of all this work. Django 1, uh sorry, South 1. 0 is essentially just 0. 8. 4 with one extra change, which is a way of having migrations that can exist alongside Django migrations. So basically South now looks for a South migration directory by default. And then force that back to a migrations directory, meaning you can ship a third-body app with both South
migrations and with Django migrations, which is a big part of our sort of upgrade plan. But of course Why you did this all about? So, you know, in 2010, in the same blog post where I write, oh South 1. 0 is coming real soon now. Um I also write For at least a year now, people have been suggesting to me that South should be in JOCore. And by people I mean like Jacob Haplan Moss and other core developers coming and saying, Andrew, the thing you're doing should really be in core. And I'm going, mm-hmm. And they're just walking away, sort of like, you know, just terrified of the workload that that might involve. So the the idea for for this sort of new version of migrations in core has been has been building since before this point, but certainly since this point. And you know at this point I have Let's say I think five years of of SAF bugs and SAF issues and like as a maintainer
all you ever see from at a project like that is is the issues tracker essentially or the bug report mailing list. And so Your life is pretty much nothing but a continuous stream of issues and bug reports where you can basically become convinced your software is absolutely terrible and then you come to a conference and everyone goes, oh we it's great software, we rely on you and people go, this is fantastic and then suddenly you're like, oh wow, right people actually like this stuff. Conferences almost one of the reasons I come to them is that it's a nice kind of more realistic snapshot of who's using Django and what they're doing than sort of the much more self-selecting thing you find on issue trackers. So at this point I was thinking about you know how do I structure this? What do I what's the kind of thing I do? So the initial plan I had in 2010, in the blog post I'm outlining here. was I would have two different components. So I was very keen that the sort of situation, the environment that South grew up in, that of competition, out of sort of
Different ideas would persist. And so I was like, okay, there are common parts of migrations frameworks, in particular the schema backend, where you abstract things like adding tables, removing tables, changing fields and things. That you should abstract those out across databases. Like Django already does this for queries using the ORM. We should probably extend that and do all the difficult work of like, well, how does Oracle add indexes or how does SQL Server remove unique constraints? And map that into a common part of Django called a semen backend. And then also add some ORM hooks. So the key thing that South does when it's doing different versions of models is that it plugs different migrations, it versions the models So when it does auto-detection, it has this big giant dictionary at the bottom that you'll see in a bit if you're not familiar with it, but you probably are. And then from that it makes brand new, well, brand old versions of models in memory and then runs migrations on those.
It reflects how things looked at the time you made it. And so Django didn't really support this. There's a module called South. hacks, which as its name suggests, is a load of higher-wall multi patches to make this work. And so my goal was to, okay, let's remove that module, let's make, let's have Django have proper sort of first class support for having coexisting module versions and have that kind of stuff in there. And then at the same time, separate of this, we'll have South 2. And South 2 will use those underlying APIs in Django and then build a better user interface and better migration handling on top of it. That sounds good. But of course, at some point the conversation's moved to well, while Django Core at some point we'd like to slim slim things down. So you know, contrib um dot uh local flavor was a good example of this being moved out We also think that things that are essential to web development should be in there.
Like Django is a batteries-intuitive framework at some point. And so the plan kind of changed to, well, let's just roll it all in. And so the two components there, the migration handling and the user interface. Suddenly sort of roll into Django and then so this is the Kickstarter at this point. Um I very optimistically proposed uh that I would backport the same kind of migrations to 1. 4 through 1. 6 The code for this does exist. It is shockingly slightly functional. South 2, at least in its Nathan state, I should probably release it just for interest. is an automated source port of Django source code to 1. 4. There's a whole load of regular expressions, a whole load of mappings that takes the source code from migrations transforms it through certain things
and dumps it in the South2 directory with certain sort of helper functions to make it work and manages to run migrations one by four and I have no idea how it works. Um like it has a whole basically it has a giant monkey patch module that makes 1. 4 look sufficiently like 1. 7 the migration scale, oh it's probably fine, let's try and run. And of course it falls over and all the complex stuff, but it it is fascinating and awful. At some point I'll show it. But of course, as I'm suggesting, that kind of fell on the right side. So the revaluation revised plan was To focus all my work on Django 1. 7 and unfortunately I couldn't deliver South 2. The sort of aforementioned South 1. 0 migration path is the new suggested solution. You can have both proper South migrations on these weird sort of secondary ones that South 2 would have and 1. 7 ones and use the full power of both frameworks.
So if you want to write raw code or raw Python, that is an option to you both of those. It is a higher burden on maintainers of third party apps. I realize this. I'm very sorry. If I can help you write these things, please come and ask me because you know I'm happy to help people port stuff over to 1. 7 if that makes their lives easier. But with this new plan, you have sort of Django containing all these different parts. So it's not really moving south into Django. Like a common narrative is like, oh yes, Andrew is just moving south into Django. This has been a very common misconception over the over the cusp. Last couple of years. Instead I am adding migrations. It is a brand new framework. I say about five percent of the code is from South, if that of the new stuff. A lot of the ideas are from South, but it's entirely rewritten pretty much. And so it's not just a simple source port, it's been a lot of work of redesigning how migrations work, improving some of the designs
Agent South failed on. and then changing sort of the core sort of structure of how migrations work and making them modular, making them reusable, as we'll see in a bit, you can take half these components and do that, use them yourself separately from Django migrations. I did however keep that separation that you saw on that initial slide of my initial plan there. Django still separates the idea of schema changing from migration running. If you want to, you can just use the schema editor and never touch the migrations library. So for example, if we're doing an IATN framework that needs columns for every language, we have you covered now. You have a supported Django interface to changing schema in a reliable fashion in transactions. Like that is there, that is core code, you can rely on that. If you want to, and I suspect you probably do want to use a full migrations framework, we have that too.
You can swap out parts if you want to. Like if you have a very custom big company problem, you can probably swap out one of these parts to do just what you like. And that's sort of the adaptable way it's meant to be working. So I'll go through these different parts. I'll describe them briefly and then sort of more detail in the next few slides about them. So the schema editor is what I called sort of the abstraction earlier, is the piece of code that takes the idea of modifying databases called you know DDL that's called data definition language and adds a sort of schema editor to every backend in Django. You can do connection. scheme editor, you get one of them, and it has methods like add model. It has methods like add field, alter field, remove field, change you need together, all this kind of stuff. And so you can call that and do things. The second cool thing that's sort of separate from that is fields can now be deconstructed.
I'll show this later, but what this means is that you can take a field, uh sort of say that's a char field or a foreign key, and you can ask it, how are you made? You can take that field and get it into a serializable format where you can recreate it again a second time. Before this, you couldn't take a Django field and serialize it. There were model instances. And most of the things they had were stored in instance variables, but some of them weren't. The way South did this, so South, I think point one to point five actually read the source code of your models. py file, parsed the line out of your model, rejects out the little theory bit and took that part which is horrific. Um but worked surprisingly well as many links did in South. They're not well implemented but they worked. And then later on South gained a whole load of rules about well, if you see a char field you need to do this. If you see this other
field you do this And then if you were sort of a custom third-body field, you could implement a method called South Field Triple. And what that is, that's essentially what deconstruct is. It said, okay, you're a field, tell me how I make it for you. And deconstruct is just that formalized. So all the core fields in Django have this. It is now a requirement in 1. 7 for third-party fields if you wish to work with migrations. However, there are plenty of docs on how to do it. There are some examples in those docs. And if you're inheriting from a field inside Django and you haven't made any changes to the keyword arguments. You don't need to do anything at all. So again, if you want help with that, just email me or Django Dev and we can help out. But that's one of the few sort of imposed requirements on third-party apps in 1. 7. Finally, the sort of the ORM hooks I mentioned in that first slide back when I was doing the separation, now model options. apps.
So you can now make models Living in different worlds. And so in particular, the app loading stuff in 1. 7 changed what we used to call the app cache to a sort of more an app registry. It's a much better formalized concept. And so now you can load models into different versions of apps, of app registries. And so what the migration framework does is it makes a different app registry for every point in history, as I'll show you later. And so we can easily have like 40 different versions of a user model and address the right one, have the right foreign keys pointing to the right places every time. Migrations themselves are much more complex, of course, they have much more moving parts. They are roughly separated into operations, which are the abstractions around the individual schema editor methods. Things like add a field, remove a model, those are still operations. I'll show you those in a bit.
The loader and the graph are this sort of way of abstracting out you have these files on disk. What do they mean? Like how do the dependencies work? How do I resolve stuff? How do I plan migrations? That's covered in that section. The executor simply takes a plan from the loader and the graph and just runs it. So it basically takes a list of things to run Loads them up, runs the operation. If it fails, it rolls back the transaction if it can. If it works, it prints okay and keeps going. The auto detector is the most complex part, arguably. This is what takes your current model state and the current migration state and sort of intelligently compares them and writes out full migrations. That's improved a lot since South, but we'll see that in a second. And then the state is the final part, and state is a very clever thing where rather than working directly on versions of Django models, what Migrations
does is it works on a sort of much stripped down version of them called state. And because the way it w because of the way operations work, they can work on state very quickly. We can run through from nothing to your end state in hundreds of migrations in a fraction of a second. Because we're not doing model objects, we're just chasing chain dates. Like state is nothing but like a couple of classes and dates. So you can run all that stuff, and that's a much easier way of doing it. I'll show you here. So the first sort of key sort of cornerstone concept here is operations and state. Now these two kind of go together. Operations are, as it says here, a declarative representation of model changes. If you look inside a new migrations file, you'll see a big list of operations in a list. All it's saying is, hello, as an as a migration, I am this operation, followed by this operation, followed by this operation.
In South you had two methods, you had forwards and backwards. But what operations do is they abstract away the concept of those two ideas. Because you know, obviously the backwards is simply the reverse of the forwards. So operations know how to work both ways. Operations additionally know what changes they represent without doing database calls. An operation you can say, here's a state. What does your new state look like? So you know you can do a quick in-memory change of Well, you know, I end up adding a model to adding a field as a model, so here's the new version of the model you get. And then you can render those out to different versions we'll see in a bit. And so what it sort of looks like is you have this thing where you have you tell a state, an operation takes you to a new state. And in internally in memory, migrations is keeping a state for every individual thing between operations. And in fact, a migration is nothing more than a sequence of operations.
And so you can see here that, you know I've said migration one migration two, that could be one migration. In fact this is how um the new squash migrations command works. It takes all the operations, it just concatenates them, it optimizes them away a bit, and then there's a brand new migration from those ones So you can see that in fact migrations are now more of a generalized framework or closure for operations in a sort of specific concept. You could actually do away with them technically, but they're there for sort of useful reasons so you can address things by name. These are the operations we can't really ship with. There's quite a few of them as you can see. The most important ones that you might not know about are RunSQL and RunPython at the bottom Those are operations that you can manually put into a file. Like migrations are meant to be writable. Please feel free to edit them. They're fantastic.
Good for that And run SQL, you give it some SQL, it runs it. Not very confusing. And runPython looks basically like an old South method. It takes it you get two arguments. You get um a app , which is like the old RM object, and you get a schema register, and you can do whatever you like. So run Python is for data migrations, it's for really complicated changes, it's for things like adding stalled procedures if it's too complicated. Stuff like that. So you have all the power of the old stuff if you want it, but by default we have a much more simplified migration format that's much more understandable and much less prone to manual error of missing certain different things. The schema editor, as I mentioned before, is this sort of abstraction over the DDL in databases. This is pretty easy for most things. Oracle has a few niggles, SingleScale has a few more niggles
Unfortunately, SQLite, while a fantastic database, does not have support for altering tables. SQLite, I think you can drop tables, you can add columns. Um that's it. You can't alter columns, you can't drop columns. And so Django and South both have a full emulation layer where if you ask to do things something simple like can't do, we Make a new copy of the table that looks like you want it. We copy the data over, we delete the old table and rename the new table to the old table. All in one big go. And so if mysteriously your SQL byte migrations are really long, it's probably because we're having to emulate the functionality SQL byte doesn't have In general, you shouldn't be doing heavy migrations on SQLite. It's meant for embedded systems or quick development. Don't do serious, don't run production on it, please. It's a single access locking database.
It's not great. The other thing to mention is that Scheme Editor takes Django models and fields. So South's version of this called DB took table names and sort of more direct it took fields as well, but it took table names and column names. Whereas this takes field names and models or objects. And so it's a bit more high level and this actually gives us more power to know, well, they passed a foreign key, so while it's called something, it's called it's actually got column of something underscore ID. Whereas South has all these sort of things like if it says unto ID, it's probably a foreign key and sort of with exceptions. So sort of try and work around this stuff. Using it's very simple, so this is a very simplified example of how to Do a couple of changes. It's a context manager, so you can just do with context connectional scheme editor add create model author add field
book author foreign key author. So as you as you can see here. The first one is making new model. I just pass in a model instance. The reason we have version models is that I can give them the correct version of author in the migrations framework for this. And then the second one is adding a field. So you pass in the model, you pass in the field name, you pass in a field instance , and that's it. And then again, the field instance here would come from a deconstruction, which is of course this bit. So deconstruction is this new thing where every field is deconstruct and it returns you the arguments you pass to dub unscore init to duplate itself. It doesn't have to return you the ones it was made with, as long as the things it returns you give you the same kind of field. That's the requirement is that basically it's it's a clone without having things like copy or the copy um module and things like that. So using it is very simple if you want to use it.
You just call deconstruction field. You get back a four tuple of four things. The first thing is the name of the field on the model. So this is field. name on most uh field instances. It's none this example because this is a field that's not parented to a model, so but it would normally be something like author or you know height or something like that. Um the second one is a fully qualified path to the field. So as you can see here, this is Django DB models child fields. If you're a custom field, this would probably be like myapp. fields. custom field or so on. The third argument is positional arguments. We encourage you not to return these because they're much less portable across versions of your fields that changes. But if you have to have them, they're there. And the fourth one is keyword arguments. As you can see here, this chart field spat out the one keyword argument I put into it so I can
Take these things I can do init of that class with star args, double star keyword args, and it reconstructs itself And in fact, this is also a really useful way we think we were using it in Michal Petruska's not yet merged composite fields work because he needed to play to clone fields. Django doesn't have one of those, at least not the one that's properly, but this is one of those. And so even his work is using this kind of stuff now. The graph is this sort of generalized idea. If you know graph theory and mathematics, it's a directed graph. It is non-cyclic as well, so psych cycles are bad. And it's sort of an in-memory representation of both the structure of the graph and several methods for traversing through it. Most importantly, the leaf nodes are basically the most recent migration for every app.
If we see more than one leaf on an app, it means that you have merged two branches. So the nice thing is that Django migrations can, because they have explicit dependencies on their pair and inside the same app. They know when they've emerged, because it then says, oh, there's two apps that point to the same parent. Oh, that must be a merge, and then it can immediately quit and tell you, ah. It also has root selection, which gives you the first migration of every app. That's how it decides where to start. It has planning, so you can say, I want to get to these four nodes, and it will give you an exact in-order set of migrations to run, to satisfy dependencies, to get to there. And of course loop detection. If you try and load in a set of dependencies with a circular dependency, it will just go, it's a circular dependency and quit with you and error and showing you the cycle so you can break it. The auto-detector by default will not make cell dependencies.
This has been most of the release blockers for the last three months. It turns out that foreign keys and and proxy objects and custom user models we'll get to later are very difficult. Uh but essentially you know internally it looks like this. Like this is a particularly complex one for a small example, but you've got three apps here and a fourth app that's called My app, other app, app three. So the leaves are all the zero zero ones. So the leaves are the top ones on each each column, and then the roots are the bottom ones. And as you can see here, if you wanted to apply other app to, you've got to apply like a good shit all the migrations in that tree first in a certain order to get to that point. And so the loader and the graph will tell you these things, and more importantly, they'll load them from disk. The loader also loads from the database. It sort of goes into a table we have called Django Migrations and reads out the
applied state of migrations and it uses that to inform this part of the graph. A useful feature that isn't illustrated here is that there's also support for having migrations that replace other sets of migrations. So you can so if you want to shrink down again, so you've got like hundreds of migrations You can merge them all into one and then declare this one migration replaces all of these. And then the loader will intelligently swap in and out. So if you're midway through that set It won't swap it in because you can't be halfway through a migration. If you are below or above the set, it will swap it in and reduce the tree down. So it's m it's very intelligent to sort of giving that stuff and the command squash migrations makes those files for you at least to a limited extent. Hopefully in 1. 8 we can get a better version with better detection, but it's a good start and much better than South's blow everything away, run fake zero and everything and start again.
The autodetector is the most complex part of uh Django migrations and the one that I've rewritten at least once during the beta phase of Django. So the basic thing here is it takes two different states. Your current project state, so state can have a has a thing where it says, hi state, make yourself from Django, from the current models. It's very simple. And then it takes the final state for my migrations branches. So it sort of runs through all them in memory, goes, ah, here's a state, and then it just compares them. And then from that it outputs a set of operations and a set of migrations and dependencies between migrations and all the kind of stuff. One of the problems is that there is a very complex dependency set between migrations. In initial Django and SyncDB and 1. 6 and below
Django just makes all the tables in one big transaction and then declares foreign keys as well, we'll do them at the end. And so m so it can sort of get away with a lot of stuff at that point. Because migrations forces you to have these separate componentized things and that you can run them whenever you like, we can't just leave foreign keys till the end, as it were. Like the end might not be for another couple of weeks in that case. And so we have this very rigid dependency scheme where we go, okay, we have this foreign key points to here, and this foreign key points to here, so you have to run this migration first, then this one second, and this one third. And it gets more complex when you have circular rings of foreign keys. Like say I have two apps that foreign key to each other. How do I do that? In the old stuff they'd be made at the same time, but if they're in different apps one comes before the other. And so The autodetector has to split that into make this model, make this model with a foreign key, add a foreign key to this model.
So it sort of gets a sort of a cycle like a non-cycle thing like this And so that's sort of the difficulty here. The sort of brief way it works is that we we take we take a very dumb diff. We sort of take the two states, we go, okay, here's a set of models that have changed, which is like you know an intersection of ones that still exist. Here's ones that have been removed, a set difference. Here's ones that have been added, another set difference. Here's some fields that have changed. And then we take all those, we dump out a raw set of operations, not any order, just sort of generalized order, onto a big list in in the auto detector. And then the operations themselves have individual dependencies. So things like a foreign key will say, I depend on the other end of me being created. or things like adding a model or say I depend on optionally my same name of me being removed first and things like that.
And a big bit of code goes through and rearranges everything to try and satisfy the constraints and dependencies. It pretty much always works now. There are some very edge cases that um are either bugs or some cases aren't supported because of custom models. And you end up with this sort of nice set of migrations that are in dependency order with defined dependencies that spit out. There is a whole talk on how this works. I will give at some point. If you really want to find some nasty code to well, it's beautiful code, but also hard to understand to work on, that's what they go and look. The optimizer is another complex part of this. The optimizer takes a set of migrations and returns you a smaller set, hopefully, with the same effect It can just return you the same set and go, I can't do anything with this. It errs in sort of caution. But the idea is that when you're squashing migrations and also during auto-detection, which is very verbose, like
auto detection outputs a add field for every foreign key. And then the optimizer comes along and says, well, you've just got add model, add foreign key, add foreign key, add foreign key, and squash them into one add model with foreign keys. So what it does is it sort of takes your list of operations, sort of a as a big list of things in order, and it steps over them pairwise. It goes one, steps over like this, then two, steps over like this, and then if it finds a pair that matches a certain pattern, so for example If it finds an add field and a delete field, it goes, okay, those two I know I could optimize away to nothing if they match if the names match If it finds an add model and an add field, it goes, oh, those two can optimize. Screen along, you have the middle one is like, oh it's its ad fields on the model, we can't
be done in that, and there's oh okay. This is this is two pairs that we could potentially optimize So it says, okay, then then the models match, the field is not already in there, that's good. And then we have to look at all the all the intervening operations to make sure we can't We can pass through them. So for example, if my ad field is add a foreign key to a model and the middle one is create that model, I can't push the ad field through the creation of the model, it won't work afterwards. And so there's all these checks like, well, can I push through the individual operations? And if it can, we can replace those two things with one and then restart the thing again. It keeps looping and looping until it basically has no result and then returns. This is always not perfect. It errs on the side of not reducing because that's a safer way. We know that what was there worked before.
It can fall over on very, very, very complex stuff, but I've not seen any bug reports for a couple of months, so it's probably pretty good on that one. But nonetheless, you know, this is somewhere that could do with an improvement. If you're very good at optimization stuff, please help me with this bit. I would love to have an optimizer be even more intelligent, even better at optimizing. There are some things it just misses that it should be it should be catching. The final part of this in in terms of design is a brand new format. So I'm sure many of you are familiar with this. This is It's I I'm not sure that it's like size size, like size one font. If you can't see it, which you can't, the top section there is the actual content of the migration. There's a forward and backwards method. I think it's just, yeah, it's it's just a create table. Like this migration is making one model. That's all it's doing.
Um and so you have this Tiny tiny bit of um actions at the top and then this giant chunk of frozen around the bottom. And the reason that's there is that what South does is for every single migration, South dumps the entire historical slate of models serialized into a big dit. And so when we want to load that stuff back up again, we go, okay, let's read the dict, load it into memory, and then go from there. This is what the new format looks like. Notice it's much bigger. You can see it from the back of the hall. What we're doing here is we say we'll create model. That's what we're doing. Like it don't repeat yourself. Don't have all the stuff at the bottom. And the key thing we do we can do here, let's go forward there, um is we have in-memory running. What happens here is that We can take a state and we can apply the migration operations in memory to these states and
end up with a result. And we can have models at any point in history. And we can run from those. And so we don't need that big dict anymore. We can just derive that dict from all the previous operations in history. And it's m it's almost as fast too. That's a fantastic way of doing it. So I'll quickly go through how the two commands run. Two main commands are make migrations, which is the old schema migration, data migration commands. The reason it's got a plural is that make migrations will make a migration and all its dependent on the migration So you can pass it an app name. It will try and limit itself, but if you have a migration that needs another migration, say auth or user something, it will make that migration as well because it needs it it knows that needs the dependencies. South would make you give you migration and get every like it's not going to apply, but good luck. Because it didn't have a dependency result Migrations does.
So what it does is it takes your two states, pass them to the autodetector. So basically the number two there is going getting a state from disk. Number three is getting one from the project, and then it takes them, it auto text them as a big list, it optimizes them so it has a much shorter list. And then it passed them to a writer. And the writer is a small class that can take a list of operations and write out a Python file in pretty decent, like indented, nice code that you could understand as a person rather than just like a Big lump of stuff. Migrations are meant to be human readable and human editable as well. Migrate then takes those different things It passes them to it sort of goes to the loader and says hi, load migrations, loader goes to the disk and finds the stuff and brings it up again. And then it gets a big list of migrations, it runs through one at a time. has a transaction, then runs the operations individually.
The operations actually just call the schema editor individually. And then once it's all done, there's a thing called recorder, which I haven't mentioned here, but it's very simple. It just says Yeah, it has methods like app marked this is applied, marked this is unapplied, what's when marked is applied? And so it just says applied, applied, applied, applied. And we store that in the database individually in multi-db2. So if you've got more than one database, you have to run my grate on each individual database. On the plus side, the database routers do now have AllowMigrate, which is on by the operation. So you can individually turn on and off migrations for different databases. So sort of the thing I want to address here at the end is what went wrong. Like the design is fantastic and wonderful, but like the real juicy bits here, what what what was I getting ah ah ah for like three months there? Um so the first thing um Um
the thing that I really don't like, and no no thought of us here who uh may have been involved in this, is swappable models. Now, for those who aren't familiar, I I I imagine you are. Swapable models are when you can replace, in this case only the supported one is auth. user, and you say, okay, all thought user is actually now this other model. And inside Jang all the file keys repoint, all the constraints repoint, or like if you're trying to even get the model, it just turns into your model, which is fantastic. Um except for dependency graphs. So normally this is fantastic. The problem is that has a setting and the setting is outside of migration. So depending on the setting is I can just change to this And then it just goes crazy and then a robotack is and it's just I don't know. And suddenly there's a giant Roblox on Tabridge, which isn't that's quite fun.
Um and so the problem is like The dependency graph with this setting changes at runtime and there's no way of telling what it was before. Like you can't even we can't even tell there was an error because like as far as Jenny is concerned, the the previous history is gone. There's no reference to it. Like we didn't know that it used to be a different setting. And so In the case of auth. user, we have a few workarounds in Django that support this stuff and we generally do a pretty good job and if you point dependencies at it, we try and resolve them. But you can still get loops out of the auto-detector. If you auto-detect with it on one setting and then change it, then there's no guarantee the resulting migrations are cycle-free in independencies. Unmigrated apps are a big issue. So the problem here is that we still run syncdb apps and we the code still exists until 2.
0. And so my management. This is fine unless you've got foreign keys between apps that are in different sections. So the choice was, do we not allow you to have foreign keys from migrated apps into unmigrated apps? meaning that we'd have to run migrations first so that they were there for the um the other way around? Or do we not allow keys foreign keys from unmigrated into migrated, meaning we had to run unmigrated first? The decision we have is that you can't have foreign keys from unmigrated apps to migrated apps. This is a little problem because we moved all of the core contrib apps to have migrations. So if you depend on those, you need migrations, sorry about that. Um but on the plus side this is much better for the way forward. Uh we like if it if it was the other way around then it would be this horrible thing where you had to like
Bringing in one app would just corrupt your project, you couldn't do any changes at all. So it's a little annoying, but adding migrations to an existing app is very, very trivial, even easier than it was in sound. Like there is I think like this much documentation on the page like you add a migration and then Django auto-applies it because it knows a bit like Django reads it says oh there's some trait models in here we've already got those tables and just marks it as applied so it just does that bit for you So you can just ship a migration. In fact, 1. 7 ships migrations for every contrib app. When you run migrate, they'll auto-apply themselves the first time because they know that they're already there. And in fact when you run it in 1. 7, 1. 8 even, which is sort of the current dev trunk, you run migrations and it sort of auto-applies and then it improves the length to your email fields and users. It sort of f it fixes all the old like all these low number bugs we couldn't fix for ages We now ship migrations of that stuff in core.
So that's fantastic. Uh test persistence is another one, and particularly on MySQL. Um I'm not a big fan of MySQL. Uh test persistence in Django. We promise you that between every test you will have the same starting state. In particular, the data you loaded at the start will always be there. That was great with the old stuff where syncDB all you could do was load stuff in initial data. So we just replayed those. With migrations, you have data migrations. So you can Make stuff whenever you like. We can't even tell what happens. And so on things like Postgres, where we can roll transactions back, we go, okay, apply the test, roll that transaction. Apply the test, roll back transaction. But on my ISAM and things where you can't do transa and transaction test cases, we can't do that. There aren't any transactions. We can't roll it back. So what we do is
we after after we migrate the start of your test run, we then we then dump data into memory, copy of all the data from the migrations, and then we load from that for every test after a flush. So it's a little bit slower. Um notably you the tests are three times slower with it turned on, but it turns itself off in Django core for things that don't need it. But if you're having issues with test persistence, there's now a setting in tests you should turn on that says, I actually do need I actually I do want this or I don't want this. So sort of tell Django to either speed up or slow down as as required I perhaps was not terribly professional and didn't read the docs when I was doing migrations. I forgot a few meta -options. I forgot the order with respect to existed and had to add it during I think the beta phase. and a few other ones, but Jang has a lot of sort of like slightly odd meta
options of like, oh, we have this thing where you can just make foreign keys be orderable or this other thing where you know So a lot of the options had to go through and add a second run of like, well, we need to support this, we need to support chaining this stuff, and different different attributes as well And finally, uh proxy models were the last thing. So proxy models are this thing where you can have a model but it's a proxy to another model. And it has no tail of its own. It has no sort of it has no fun it's merely wrapper for extra code methods basically or extra an extra um sort of objects. And so to migrations, they're basically worthless. Like they don't they don't have a table, so we don't have to do anything with them at all. But you can foreign key to them and so we need them around to sort of sort of point foreign keys at and sort of just say point at that empty thing and be on with it. And so I had to go through and had a separate set of things that said, well
If this model being added has proxy true, then skip all this database part and then just leave it lying around. And that's sort of another sort of the gotchas that I forgot about. Like you know, I had assumed during the development of this that proxy models would be Useless because you couldn't they were they weren't in database stuff of them, but like I forgot you'd foreign key to them. So it you know it bears in mind that you could think about all the things Django offers to start with And uh but we're there. Um this this slide was written before we read 1. 7. I'm very glad it's still it's actually in date now I'm presenting it. Uh 1. 7 is out now, it has migrations. Um please use them. Please report any bugs as you find them, but like I am pretty confident in the quality of this release. We have gone through a very painful like three, six months of bug reports and fixing and really sort of hardening that that stuff up.
Like migrations are fantastic. I hope you will use them. I hope that you upgrade the third party apps to use them. If you need any help, I'm always around. Like I'm on IRC. You can email me, you can email any developers. I want to sort of get this painful transition period of going to migrations over and that we can all live in the wonderful future of migrations. Thank you very much. Yeah, I think it's a pain, but I'm not sure.
South began as a formal replacement for manually numbered SQL files, gained autodetection and state tracking, and eventually informed—but was not simply moved into—the brand-new Django 1.7 migrations framework. South 1.0 also provided a transition path for apps supporting both systems.
Discussed at 2:44The schema editor provides a database-independent API for changing schema, with operations such as creating models, adding or altering fields, and removing fields. It can be used independently of the full migrations framework and is implemented for each Django database backend.
Discussed at 14:34Fields implement `deconstruct()` to return enough information to recreate the field: its name, import path, positional arguments, and keyword arguments. Third-party fields need this support to work with migrations, unless they simply inherit a Django field without changing its keyword arguments.
Discussed at 16:05A migration is a sequence of declarative operations, such as adding a model or field, and each operation knows both how to change the database and how to update an in-memory model state. This lets Django reconstruct hundreds of historical states quickly without repeatedly creating full model objects.
Discussed at 19:10Use the `RunPython` operation, which receives a historical app registry and a schema editor. It can perform data transformations or other complicated changes that cannot be expressed through the standard schema operations; `RunSQL` is available for custom SQL.
Discussed at 20:43Django emulates unsupported SQLite alterations by creating a replacement table, copying the data, deleting the original table, and renaming the replacement. The speaker cautions that SQLite is not suitable for heavy production migrations or production workloads in general.
Discussed at 21:29The migration graph records dependencies between migration files, identifies roots and leaves, detects circular dependencies, and produces an ordered plan that satisfies all dependencies. The loader also reads applied migrations from the `django_migrations` database table and can substitute squashed migrations when appropriate.
Discussed at 25:19The autodetector compares the current project state with the state produced by the existing migration branches, identifies added, removed, and changed models and fields, and turns those differences into ordered operations. It then resolves dependencies, including difficult cases such as foreign keys between apps and circular relationships.
Discussed at 27:41The optimizer examines migration operations pairwise and combines or removes operations when it can prove the result is equivalent—for example, folding field additions into model creation or eliminating a matching add-and-delete pair. It deliberately favors safety, leaving operations unchanged when it cannot confidently optimize them.
Discussed at 30:01Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026