Empowering Django with Background Workers
Published July 11, 2024
This video features Jake Howard at Wagtail Space NL 2024 in Arnhem, Netherlands.
Wagtail Space NL 2024 https://nl.wagtail.space
Automatically transcribed, so expect mistakes in names and technical terms.
Okay, so I probably don't need to reintroduce myself. I'm Jake. I know those things. Um and today I'm going to talk to you very quickly about recovering deleted Wagtail pages or in fact this trick works for any Django model. So to set the scene, people tend to use Wagtail as a website or a blog. A Torchbox, as has been said, we actually use it for our internal documentation, which we call our intranet. It's got our processes, company information, links out to other places, various things like that. It has been around for a while and it's slowly grown. Um and sort of In the middle of 2022, we decided to restructure that content in a big information architecture redo. We wanted to make things easier to find, we wanted to remove duplication and generally make it
a easier and nicer place to store content that wouldn't just be forgotten about. Unfortunately, and that's why I'm standing here That didn't quite go to plan. Um one afternoon, sort of while a lot of this restructuring was going on because of the size of it, it was a very long-term effort, um I wanted to look at our internet to reference a process um and I couldn't find it. And it turns out what happened was the entire sysadmin section, which is the team I work under, the entire section of the internet had completely vanished. Um and it was my job to work out what on earth had happened. So the very first step was to lean very hard into Wagtail and one of its features, namely the site history report. Um and
that conveniently has a nice place to show exactly what has happened inside a Wagtail site. And it showed me almost exactly what had happened. Um exactly kind of what I'd thought. Uh one of our staff members went into the intranet's Wagtail site and deleted all of this sysadmin content. Um and that was great. Um and because this they went into the sysadmin page and hit delete on that, it deleted every page underneath it, all 105 of them. And it did this nice and quickly. At the time there were some issues around usability. It didn't prompt you saying, are you actually sure you want to delete 105 pages? It does do that now. Um and there were also issues with how it was surfaced in the site history
report, which made my life a little bit more difficult. Those have now been fixed as a direct result of me pulling my hair out in this. Doesn't need to happen again. And so I messaged the person to better understand what had happened, um, assuming that they didn't mean to delete all 105 pages in Hanlon's razor format. Um I've redacted their name for what I hope are obvious reasons. If you're watching, you know who you are. I know who you are. In the end, what had actually happened was they'd made a new sysadmin section a while ago before Moot Before switching strategy to moving the pages under an existing tree. And what they'd intended to do was delete their temporary one that they'd made, and they didn't, they
deleted the one with all the content in it. And sure, Wags has confirmations for these kind of things, but when you're deleting lots of pages and expecting to delete a lot of pages, you might not read the message perfectly. My line of manager referred to this process as radical reorganization, which I think I quite like. But with all the content gone, I needed to get it back. I need those 105 pages. I can't just use some radical reorganization and forget about it. I need the content. And so naturally I had to do some restoring for backups. Um our intranet is a living document. It gets updated fairly often. Rolling back the entire system almost two days by the time I noticed it. Would have been potentially losing critical changes, not to mention the time that people had spent making said changes. It'd be annoying, but we could do it, but I'd rather find a different solution.
In an ideal world, what I need is a partial restore for backup. I want to restore just the sysadmin pages, leaving all the other pages completely untouched. And using a few tricks from inside Django and inside Wagtail, it is absolutely possible, and I absolutely did it with zero downtime or user impact of any kind. And here's how I did it. Step one, I needed the database. We back up our intranet nightly, so I downloaded the backup, spun it up locally, and it being a Django app, spun the code base up locally, loaded in the data, nice and simple. Step two, I needed to find the actual page models locally that were deleted. Behind the scenes, Wagtail's pages are in a tree-like structure implemented using Django TreeBeard.
When a page is deleted, TreeBeard is the one that finds all the child pages and deletes them. And then Django goes through the database and deals with cascades and everything else. So Getting the page, the system in page had an ID of 91. I get that and then I call get descendants on it, which nicely gives me all of its descendants exactly in the Wagter page tree, which is nice and convenient. Step three, I need to find out what was deleted. And this is where the subtle bit of magic happens. When you delete a page, you delete more than just a page. You're deleting the specific model, you're deleting Wagtail's page model You delete revisions, you delete related models, you delete through tables, you delete everything. Get descendants just gives you the pages. It doesn't give you all these extra things that might be related through relationships. And when you call delete, you can get a number of objects, and there's generally quite a few things that get deleted.
But if you've ever deleted anything through the Django admin, not the Wag2 admin, the Django admin You'll know it's actually capable of going when you delete this, here's everything else that is going to get deleted. And that's implemented with this undocumented but actually really simple to use API called the nested object collector And it really is that simple to find everything that's going to be deleted without actually deleting it. It doesn't delete the models, it just shows you what would have been deleted if we had done a delete. But it's slightly easier than using a transaction and then rolling it back. You can just call this. The next step was to serialize it. Collector. data, in the case of the code snippet from before, now contained all the model instances
which were deleted, but they were in memory on my laptop, and my laptop is conveniently enough not what is running in production So what I needed to do was take that data, serialize it so it then could be loaded into production. And if you're thinking of something like Django's fixture settings, that's exactly what I did. Django 's fixtures create a JSON representation of a model so that they can be saved from one location and loaded into another. It's most useful when you're doing complex testing fixtures, hence the name, but you can use it for many things. Um now this code is s exactly what we use, but there is an extra bit which is slightly more complicated, which is this no M2M serializer, M2M standing for many-to-many. Uh Weng Django serializes a model with a many-to-many
relationship. uh which doesn't use a custom table, so it's just the implicit through table. It inlines that definition when you generate it into fixtures. Which is nice and easy to work with because everything's just there in one thing. The problem is is the nested collector also finds all of these nice things. And so you end up with duplicates and referential integrity issues when it comes to loading them back in Which is a problem. So what this little three lines of code does is it says to the serializer, when you find these things, don't inline them, just ignore them. Generally, that's a bad idea, but because we can rely on the nested collector finding them for us, we can ignore them, so we only end up with one copy in our fixtures, not two. Now, because we have what is functionally a standard Django
fixture, the inverse to load it back into our system is just manage. py load data, the same API that you would use for loading any kind of fixture. And if we combine all this code together, this is somewhat truncated, the version of the script that I used to build everything together. Um there's a file name at the top as well that got truncated as well. Um It's not actually very much code. It's not very complicated. And if you understand what each line is doing as I've walked you through it, it's really easy to understand. Now, for what I hope are obvious reasons, I needed to test this restore before running it on our production intranet. Um so what I did was I Loaded the old back up and I did exactly what the person who will rename
nameless did. Went in, deleted that page. And then I tried to take the restored file, the fixture file that I generated, run it through load data, load it back into our system, and confirm that all the pages were there and they were all the same. And they mostly were, but I am glad I tested it because it did come up with some issues, namely around Wagtail search indexing. So we use the Postgres full text search indexing. Which means that the Nessa collector did actually pick up on those search indexing things. Wagtail or Postgres or something didn't really like when I tried to restore those as is. So there's some extra command line flags you can add the load data to strip those out, and then it really did just work. Step six, showtime, the tense
bit. Once I was happy everything was that everything was tested and working, I ran exactly the same steps on production. Um our internet runs on Heroku, a platform as a service. platform. Um so I had to do a few dances to get the JSON file up there. But once that was done, everything was fine. Because I am a good sysadmin, I took another backup just before doing this because I last thing I want to do is make things worse. And with the data vial in place, I crossed everything and ran load data. And slowly pages starting popping back into the admin as if they'd never left. I ran the check tree management command, everything worked. I ran update index and update reference index to make sure that they were all updated and pages could be found
and once I did that, they could. All the pages now appeared back in the admin. They were searchable both from the admin side and from the internet side, and everything seemed fine. So with just a few hours work, the pages were back. And the benefit of this is there was no downtime, there was no content freeze, and as far as I'm aware, there was no data loss either. We got all the pages back. And most people didn't even know there was an issue. If you'd never needed the sysadmin pages, you'd never known they were gone, which is a thing in itself. But we never needed to tell people, don't make any changes, we're gonna roll a bunch of stuff back. It just happened I've used this trick, sadly, quite a few times in my career, both for restoring of Wagtail pages and also playing Django sites. Ironically, just a few weeks after the blog post that this talk was based on was published, Dan and I had to do a very similar thing on a very similar Django-only
project. use exactly the same trick and it worked absolutely perfectly. And so I'm hoping that through the pain that I've had to go through to discover all of this stuff, you can use it and hopefully it will help you as much as it does me So this is a link to said blog post. It exists on the Wagter. org website. It's exactly what I've just said, but in written form. Hopefully this trick works out for you as well. Thank you. I'll leave it up for just a second so you can Oh it's gone. Come find me if you want a link, I'll put it in Slack.
Restore the backup locally, identify the deleted pages, collect the related objects that were deleted with them, and serialize those objects into a fixture. Load that fixture into production with Django’s `loaddata`, then rebuild and check the relevant indexes; this can restore the pages without downtime or rolling back unrelated changes.
Discussed at 4:01Serialize the collected model instances as a Django fixture, transfer the fixture to the production environment, and load it with `manage.py loaddata`. For implicit many-to-many through tables, the talk’s serializer skips inlining them because the collector finds them separately, avoiding duplicates and referential-integrity problems.
Discussed at 6:24Test the restore against a copy of the backup after reproducing the deletion, then verify that the pages and related data load correctly. In the speaker’s case, PostgreSQL full-text search indexing objects needed to be stripped during loading, followed by rebuilding the search and reference indexes.
Discussed at 8:46Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 27, 2024
Published June 27, 2024
Published June 27, 2024
Published June 27, 2024
Published June 27, 2024
Published June 27, 2024