A New Approach to Multitenant Wagtail, Stephanie C. Smith and Addison Hardy

This video features Addison Hardy and Stephanie C. Smith at Wagtail Space US 2022 in Cleveland, Ohio, USA.

A New Approach to Multitenant Wagtail, Stephanie C. Smith and Addison Hardy
0:28:23
Published March 30, 2022
785 views

Summary

JPL’s Design Lab built a multitenant Wagtail platform to replace nearly 100 separate codebases with a single deployable product. A management backend creates and administers sites, while each tenant gets isolated databases, media, users, and dynamically generated Nginx, uWSGI, and Django configuration; PostgreSQL notifications let running containers adopt changes without redeployment. The shared frontend improves branding consistency, and developers can create local sites, isolate migrations and management commands with the `WCP_ALIAS` environment variable, and work on multiple branches in parallel. The team planned to open-source the platform and potentially separate the management backend from tenant application code through configurable service types.

Key takeaways

  • JPL replaced nearly 100 site-specific codebases with one shared Wagtail platform and consistent frontend components.
  • Each site has isolated databases, media storage, users, and tenant-specific Django settings derived from a static alias.
  • A management backend creates sites and hostnames, while PostgreSQL notifications update running containers without rebuilding or redeploying them.
  • Nginx and uWSGI route requests to tenant-specific Django processes using generated configuration and the `WCP_ALIAS` environment variable.
  • Developers can create disposable local sites, run management commands and migrations per tenant, and work on multiple feature branches simultaneously.
  • The team intended to open-source the platform and evolve it toward independently configured service types.

Summarised automatically from the transcript.

Chapters

  1. 0:00 JPL Web Platform Context Stephanie introduces JPL’s Design Lab, the need for a sustainable CMS, and the goals of consolidating many sites onto one Wagtail platform.
  2. 3:01 Management Backend Concepts Addison explains the shared codebase, deployment environments, site aliases, tenant isolation, and the role of the management backend.
  3. 4:32 AWS Infrastructure and Runtime Updates The talk covers the AWS GovCloud deployment architecture and how PostgreSQL notifications let running containers apply site changes without redeployment.
  4. 6:51 Management Dashboard Demo Addison demonstrates creating sites, assigning hostnames, accessing Wagtail, and viewing API and container health information.
  5. 10:01 Database and Request Routing The speakers explain the database layout and how load balancers, Nginx, and uWSGI route requests to the management backend or individual tenant sites.
  6. 10:47 Nginx and uWSGI Configuration Addison details the generated Nginx server blocks, path routing, uWSGI emperor mode, and per-site application processes.
  7. 16:54 Dynamic Django Settings Stephanie explains how the WCP alias selects each site’s database, cache prefix, Wagtail site name, and S3 storage location through layered settings.
  8. 18:27 Site-Specific Management Commands The platform applies migrations, permissions, and re-indexing per site, using the tenant alias to target individual Django instances.
  9. 20:56 Developer Experience The speakers show how developers create disposable local sites, work across feature branches, and run isolated migrations with dynamic make commands.
  10. 24:17 Open Source Roadmap Addison outlines plans to open source the platform and separate the management backend from configurable service types and application repositories.

Transcript

3,934 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:00

Speaker 1: All right, uh welcome everyone uh and welcome to those of you online. Thank you for being here. Um we're going to talk about a new approach to multi-tenant wagtail. I'm Stephanie Smith and this is Addison Hardy. And we're both here from JPL's Design Lab. And we're going to talk with you today about a platform that our team has been working on for about a year now. Before we dive into the details, we want to give you a little bit more context about what we do at JPL and what the landscape of the web is like there. So JPL or the Jet Propulsion Lab is one of 12 NASA centers and it specializes in robotics, space, and earth science missions.

0:45

Speaker 1: Some of our recent missions include the Mars rover Perseverance, the first ever Mars helicopter ingenuity, along with contributions to the James Webb Space Telescope. And as you can imagine, supporting missions like these requires a lot of communication products. And this is where our team comes in. We're part of Design Lab, which manages hundreds of internal sites for JPL's communications. And we needed to find a sustainable way to support these sites and also provide a CMS. A CMS was very important because we needed our content editors to own their content. Our team is a very small team

1:31

Speaker 1: and we just don't do not have time to help update content. Because of that, another important goal for us was to reduce overall developer overhead. One of the ways that we did that was by making a primary goal of ours to reduce our code bases down from literally almost 100 code bases to one or two. We've been able to now focus on developing a single product. We also reduce developer overhead by being able to spin up a Wagtail site. uh literally at the click of a button without the need for a developer and that's been uh really helpful for onboarding new users

2:17

Speaker 1: Another reason why we wanted to build a platform like this was because over time There's been a lot of branding inconsistency coming up on all of these different sites. Old branding guidelines out of date, maybe slight tweaks to one site, but not on the other sites. And with having all of our sites on one platform, we now have the same front end for all of those sites. And if we need to update the styles or update a component, it's updated on all of them. And now I'm going to hand it off to Addison, who's going to tell us a little bit more about the concept behind the management backend, which is one of the keys to our approach to multi-tenant wagtail.

3:01

Speaker 2: Cool. So the management backend. I guess to start off, I just want to cover a couple of concepts that will come up throughout the presentation. just so that uh it all makes sense. Um so you'll see this prefix uh WCP uh that just comes from the internal project name uh web content platform And uh so all of our WCP containers are built from a single code base. So all of the Django and Blacktail code and then the management backend. That's all packaged together in a single repo. We've got three branches, development stage, and production.

3:47

Speaker 2: And when we push to a branch, it uh triggers the deployment and builds a new container image. from that branch and that gets deployed into an auto-scaling cluster, one for each environment. Sites can be created and managed via this management backend. And sites are identified by their alias. And that's a static identifier. And it's referred to as the WCP underscore alias throughout the platform. And you'll see that variable come up rest. presentation. Each site's data is kept completely isolated from all the other sites. So each site has its own database. If you upload any media, that's all kept separate. All the users are separate. From the end

4:32

Speaker 2: user perspective, it's like they're using a single Wagtail site. And for operations, the management backend provides an API and web dashboard for our team to help administer the cluster. Infrastructure-wise, just a brief overview. We run on AWS and Govell, which is If you're not familiar with GovCloud, there's a few AWS regions that have like heightened security and compliance for government applications. Each environment again runs from a single Docker container image and we store those in uh the Elastic Container Registry ECR. And then we use

5:18

Speaker 2: Elastic Container Service to run the cluster. And in front of each environment's cluster is a network load balancer. And that's that's the request entry point into the environment. So the management backend serves as a wrapper around um the the Django YTEL app And it lets you create new sites, manage existing sites, and it also manages the sort of active container configuration, things like the network configuration. The management backend has its own database and API and those exist outside of the context of Django. And that was done so that we could potentially separate them in the future.

6:04

Speaker 2: And we'll talk some more about that at the end of the presentation, sort of future plans. And yeah. All right, so since any of the containers in the cluster can handle requests for any of the sites And that includes requests to the management API. We needed a way for the containers to let each other know when a change taking place. So let's say you go to the management dashboard and you add a host name to a site. That request is going to get routed to one of the containers in the cluster. It'll update the database and then

6:51

Speaker 2: post a notification using the Postgres notification channel to let the other containers know they need to update their configuration. And that lets us add sites and add host names without building new containers or um you know needing to do a deployment or anything like that. The the currently running containers pick up the change and just reconfigure themselves. So I want to give a quick demo of the the management dashboard Okay, so what you're looking at here, there's two active sites, and this is running in my local environment. So there's the default site, and that's created automatically the first time a container starts up in one of the environments.

7:40

Speaker 2: And the default site is also uh where the the API dashboard live um on that default site host name and each environment sort of has a brute host name um and uh Then there's one additional sort of regular site, it's aliases test. And there's links here to view its homepage. or to sign into its uh Wagtail interface. And then for each site, uh it shows the current host names And then you can add and remove post names using this

8:27

Speaker 2: table right here. And then if you wanted to create a new site, Um let's call it um why count space So if I hit create , what the management backend is doing is it's creating a database for the site, running initial migrations to you know, create the initial database schema, doing some other kind of management tasks, and then adding that site to the management database. And then if I give this site a hostname Maybe uh you know sample. And this uh wcp dev uh jplweb. net, that's just the host name we use for local dev.

9:15

Speaker 2: So I'll add this. And then if I go to the homepage for this site, you can see I've now got a new Wagtail site and I can sign into it. And up here in the top left you can see the site alias. So There's the site I just created and then there's that site with the test alias. Alright. Oh, and one other thing, there's also the management API specification here and monitoring about the health of the containers available in the dashboard. All right, database overview.

10:01

Speaker 2: So yeah, the management backend has its own database, and then each site has its own database and each environment, all those databases live in in one Postgres server for each environment. Here's a sort of just overview diagram of the overall overall architecture. So requests coming through the load balancer. They come to one of the containers. Initially, they're picked up by Nginx and then Nginx decides where to send the request based on the host name and the path. So it could be to the management backend or to one of the Django sites. Uh and this just breaks that out a little bit more. Um

10:47

Speaker 2: so um Based on the hostname and path, Nginx may directly serve a static file or static files. And it uses the we use the send file directly for that sort of an efficient way to uh unless you copy files um directly into the response buffer uh without um sort of duplicating uh that content memory twice. I think Netflix contributed that to Nginx. Um and then dynamic requests are sent upstream to UWISCI. which is a library for managing applications over the WSGI protocol. The management backend automatically generates

11:33

Speaker 2: these two types of network configuration files, so Nginx site files and UWISG VASL config files. And that happens during container startup or when a container receives a Postgres change notification. For the Inginx site files, we generate one for each site and one server block, like you can see in the example for each hostname. And when a request comes in, Nginx looks at these server blocks and looks at the request hostname and finds the one that matches the best. And it has sort of a order of priority for the matching. It looks for an exact match first. Then it does a sort of a fuzzy match or a rejects.

12:22

Speaker 2: And then You can see we're setting this variable WCP alias and that's a Nginx variable and it only exists within the context of this server block. And we set it to the site alias. And then we include routing rules from another file. And these route rules, it's how Nginx matches paths. So the server block is where the hostname is. And then once it's found the server block based on the hostname, Nginx looks at the request path and tries to match it against these location blocks. And originally we were generating these in the site files themselves.

13:10

Speaker 2: But that was causing a lot of Duplicated directives, which wasn't that that big of a deal, but we also wanted to be able to check these route files into the repo and sort of edit them as code files. And the reason we were generating them in the first place is if you look at the first location block, we're passing that request upstream to the WISCI. And the way we have UWISG configured is to listen via a socket file for each site. And each site socket file is its name for the site alias. And so we were generating those directives from Python. But by switching to using this Nginx

13:55

Speaker 2: variable setup, then we can include you know this route file into each server block. You can see this include line at the bottom there. And then that that variable sets it dynamically Umisky, we're running that in emperor mode. And in that mode, UWISKI manages each application as a vassal. UISKI, by the way, has some really interesting native conventions if you never write the documentation. And uh basically what emperor mode is is the emperor is sort of like a master process, and then the vassals are are um processes controlled by by that uh emperor process.

14:40

Speaker 2: And you can configure how many processes you want each vassal to have. And you can give it as many as you want, but it's constrained sort of by the number of CPU cores and RAM and things like that. like that. We give each vessel two at the moment in each of our containers. For Django, we're using USGI's Python 3 plugin and Django's WSGI mode. And then management backend generates a Vasel config file for each site. And one of the key things we're doing in there to support the multi-tenancy is we're using USD's uh E and B option and that lets you set an environment variable In the context of those vassal

15:26

Speaker 2: processes. And the environment variable is called WCP alias. And we set it again to the site's alias, the static identifier. And The way Emperor Mode works by default, the way we're using it, is uh basically you tell you whiskey to watch a directory. And if it sees a new configuration file appear in that directory, or if the contents of one of those files changes, then UWISGI will either create a new vassal or reload the existing one. And when the vassal loads up , if you're using the Python 3 plugin and pointing it at Django, it loads an instance of the Django application in memory.

16:12

Speaker 2: And we read this WCP alias environment variable in our Django settings. And those Django settings get evaluated just once when the Django application gets loaded to memory. And I'm going to hand off to Stephanie to talk about what we do with that setting. Oh yeah, sorry. So this is an example of what the UWISG config file looks like. You can see we're setting the socket file name and that's where Nginx sends the request, Python 3 plugin, and then we set that WCP alias environment variable. No, I'm not really.

16:54

Speaker 1: All right, thanks Addison. So as Addison said, uh when a U Whiskey vessel loads a Django application, uh the Django settings are then processed. And most of them are static, but there are a few key settings that are dynamic and that are tied to that same WCP alias environment variable. And those key settings are the cache key prefix, the database name, the Wagtail site name, and the S3 bucket folder name So if we take a closer look at the way we've structured the settings inheritance, you can see on the left is where it starts. with the base settings file, and that's the file that contains all of these static settings. So these are the settings that are shared across all of our sites.

17:41

Speaker 1: That is then imported into what we call the dynamic settings file. And there's a code snippet just below showing what that file contains. And you can see it imports the base settings and then it gets the environment variable called WCP alias. And then it uses that environment variable to set things like the database name and the Wagtail site name, among other things mentioned in the last slide. And then the last settings file in the inheritance chain is the environment settings file. And this is the one that has settings specific to our environments. So development, staging. production and local environments.

18:27

Speaker 1: And so the key here is per environment, the only thing that we actually need to set is the Django settings module environment variable. And we basically set that variable to the environment settings file that we want to use for that environment, which is importing the dynamic settings file, which is also importing those base settings. And that's how we generate our settings dynamically. All right. So an application instance is loaded into memory, the dynamic settings are processed, and what happens next is during container startup. A task is run to then update the environment. And in that task, which is right here, this code snippet,

19:14

Speaker 1: after the list of sites is retrieved from the environment Functions from the WCP Django class are called for each site, and those functions apply things like new migrations, adding custom permissions, and updating the site index. So this happens for every site in that environment. And if we take a closer look at that WCP Django class, well we use the WCP Django class to call the manage. py action , call manage. py in the application. And if we look at it more closely, you can see there are a few methods defined. The first one is the manage method And essentially what that does is it forms a management command, but it's also passing that same environment variable

20:03

Speaker 1: WCP alias. And if you look at the other methods, the migrate, permissions, and the re-index method, you can see that they are all using that manage method. to pass management commands. And again, that means that those methods are essentially passing the environment variable WCP alias. And that is how each site is addressed with these management commands. So in practice, it is possible to run a management command on just one specific site. And that kind of blends into our next topic, which is developer experience. So as a developer, what is it like to develop on this platform? And for this talk, we wanted to focus on what the key differences were for developers from a standard Wagtail instance to

20:56

Speaker 1: our multi-tenant platform. And covered in Addison's talk about the management backend, we know that the local environment configuration is the same as the production environment. And that's from all the dynamic config files being generated and that dynamic settings file that I just talked about. So again, the only thing that needs to change between environment is setting that environment settings file through that Django settings module environment variable. And another key difference between our multi-platform , multi-site platform is that we can spin up a site

21:41

Speaker 1: quickly, so at the click of a button. So this has been helpful helpful for onboarding users, but it's also really helpful helpful for developers. So that means if you get your database, your local database in a bad state, you can quickly just spin up a new site and destroy the old one. And on top of that, you can spin up as many sites as you want locally, which leads me to my next point, where this then allows you to work on multiple feature branches simultaneously. So I'm going to talk a little bit about how you would go about doing that. So locally, so say I was working on a new feature branch,

22:28

Speaker 1: what I might do is I would spin up a new site by clicking on that create site button that Addison showed in his demo of the dashboard. So I have my new site called Example Created. And maybe I'm making some model changes and I want to create some migration files. So to do that and to only do that for that one site, all I need to do is pass an environment variable with that management command. So in the more verbose version version, if you're using Docker Compose, for example, should look very familiar. With the main difference being that we are passing that same environment variable, WCP alias, in that first line there. Now to make this a little bit more manageable for us as developers, we've also

23:16

Speaker 1: created dynamic make commands that Surprise also pass along that same environment variable of WCP alias. So now it can be as simple as make migrate the alias name of your site, which has been really helpful for us as developers. And we can you can also unapply migration. So you can get really complex with your management commands just as long as you pass that. uh WCP alias. And so you're done with your feature, it's code reviewed, merged in, and once it's deployed, That's when the new migrations are run for every site in that environment. So locally you can isolate your migrations, but at deployment, that's when it's run for everything.

24:05

Speaker 1: And so with that, I'm going to pass it back to Addison, who's going to talk about what we can do next or what we're hoping to do next with our multi-tenant platform.

24:17

Speaker 2: Thanks. Thank you, Stephanie. Yeah, so um we we've been um thinking about sort of the future roadmap for the the management back end uh at gpl and and just in general um we went through uh gpl 's process over the past month um and got permission to open source uh everything. So we're working on doing that over the next um probably two to three months. We need to you know package the code base and write documentation and stuff but we want to get uh the the management back end code open source and um The other thing we've been talking a lot about

25:04

Speaker 2: is potentially separating the management back and out into its own repo and not having any of the application code live in that repo. and introducing a new um a new concept uh for the management backend which would be the service type So a service, when you would create one, like if you went into the dashboard, there'd be a services tab. And each service would have a static identifier. You'd give it a git repo URL and a branch name And when what would get uh built at as the container image would be just the management backend. And when it deployed during container startup, it would get all of the current services and sites.

25:51

Speaker 2: And for each service, it would do like a shallow clone of that git repo and branch. There would be a services folder and it would clone each service into a folder in the services folder named as the service identifier. And then when you created a site, you could choose from one of the service types. And to support that Just wanted to briefly show this. So we're envisioning having a sort of a node configuration file that would live in the root of these service repos. And it might be called wcp. yaml or something like that.

26:36

Speaker 2: And uh we'd have a documented spec for this configuration file. But just as a quick example, it might have a service type like Django. And it might have various action scripts. So a setup script. So after the container points down the repo during startup, it would also point down the repo if a new service got created. It would run the setup script to like install the services, Python, NPM dependencies, stuff like that. And then it might have a tenant script. And the management backend would call that tenant script once for each site of that service type, but again passing that WCT alias environment variable. uh and you know it's to the script and then that script could do site specific actions like run migrations

27:24

Speaker 2: um and there's just uh some quick examples of that here um And uh that's something that we're still you know actively working on the uh the architecture for, uh but that's just kind of a quick look at uh where we're going. Want to say thank you to the rest of our team who's not here today. Uh that's it

27:54

Speaker 3: Thank you so much. It was my fault I forgot to give you the signal to stop us too wrapped up in your presentation. Do you mind taking questions on the Slack and in the hallway as people are?

28:04

Speaker 4: Yeah, please. Yeah.

28:07

Speaker 3: Great presentation. I apologize for not giving you time at the end.

28:10

Speaker 2: No, no worries.

28:12

Speaker 1: Thank you.

28:12

Speaker 3: Awesome. Thank you. Bright in the

Questions this talk answers

Why did JPL build a multitenant Wagtail platform?

JPL needed to support hundreds of communication sites with a small team, while giving editors control of their content and reducing nearly 100 codebases to one or two. A shared platform also makes branding and component updates consistent across all sites.

Discussed at 0:45

How does this Wagtail platform isolate data between tenants?

Each site has its own database, media storage, and users, while all sites run from the same application code and container cluster. Requests are routed by hostname and path through Nginx to the appropriate site application.

Discussed at 3:47

How does the management backend create and manage Wagtail sites?

It provides an API and dashboard where administrators can create sites, add or remove hostnames, view site links, and monitor containers. Creating a site provisions its database, runs the initial migrations and setup tasks, and registers it in the management database.

Discussed at 5:18

How can the platform add sites or hostnames without redeploying containers?

The management backend updates the database and sends a PostgreSQL notification to the running containers. Each container then refreshes its configuration, allowing the change to take effect without rebuilding images or deploying again.

Discussed at 6:04

How does the platform use WCP_ALIAS to configure each tenant?

The alias is passed into each uWSGI application process as an environment variable. Django uses it to derive tenant-specific settings such as the database name, cache key prefix, Wagtail site name, and S3 storage folder.

Discussed at 16:54

How do developers run migrations or management commands for only one Wagtail site?

They pass the site's WCP alias through the environment, either directly with Docker Compose or through the platform's make commands, such as `make migrate <alias>`. This lets developers isolate migrations and other management operations locally; after deployment, migrations run for every site in the environment.

Discussed at 22:28

What are the planned next steps for the multitenant Wagtail platform?

The team plans to open-source the platform and potentially move the management backend into its own repository. They are also designing reusable service types, where sites can be associated with a Git repository and branch and initialized through documented setup and tenant scripts.

Discussed at 24:17

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Addison Hardy and Stephanie C. Smith

More videos from Wagtail Space US