From Dinosaurs to Snakes - Decomposing the Enterprise Monolith
Published October 13, 2024
This video features Michael Nicholson at Django Day Copenhagen 2022 in Copenhagen, Denmark.
"How 12 factor app helped Divio migrate over 10k Postgres instances" by Michael Nicholson at Django Day Copenhagen 2022. Talk description at: https://2022.djangoday.dk/talks/michael/
Michael Nicholson explains how Divio used 12-factor app principles to migrate more than 10,000 customer PostgreSQL databases from version 9.6 to 13. Divio’s platform treats databases as attached resources, keeps configuration outside application builds, and runs maintenance tasks as first-class, stateless processes. These practices allowed an automated runbook to provision replacement databases, copy data, switch applications to the new configuration, scale the migration workers, and remove the old instances with few technical problems. Nicholson also covers the limits of the approach: automated migrations used maintenance mode unless customers followed the zero-downtime procedure, and customer code still needs to support fast shutdowns, idempotency, and background processing.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: I don't think I've come quite as far as Paolo or a couple of others here, but yeah, Kim. Good morning everyone. My name is Mike Nicholson. I work for Divio as a cloud solutions engineer. I'm British from the beginning, but I live in Stockholm, Sweden. So I've traveled Stockholm for today. Today I'm going to talk about how 12-factor app principles helped DVM migrate over 10,000 Postgres instances. without too many problems. That being the important part. No, that's not working now.
Speaker 2: Isn't that?
Speaker 1: No.
Speaker 2: Can you the the
Speaker 1: There we are. Okay, so just briefly I will talk about why we wanted to do this in the first place. uh who Divio is and why it was actually relevant for us. What is a 12-factor app? What are the 12-factor app principles? And then the bulk of the talk will be talking about how the 12 factor actually helps with upgrading the database version and why it made it easy for us. So the Y is quite easy. The final release of Postgres QL nine point six was Scheduled uh back in August 2021 for the coming November
Speaker 1: and in October 2021. AWS followed up on that and announced the end of support for Postgres 9. 6 , sometime early 2022. Uh the date they announced has changed a couple of times. I think now it's actually dead So who's Divio and why did we actually care about this? Why was it a big deal for us? Well, our tagline is multiple cloud management without the headaches. What that means in practice is we put a pause in front of your applications to make it easy for you to deploy in the cloud without caring about what's actually happening in the cloud. So in the instance of a database, you would add a database
Speaker 1: service to your application, to your project. We'd take care of all the setup on AWS or wherever, whichever cloud you want it to deploy it. on. You wouldn't need to care about the implementation details. You just say, I want a Postgres database. So we have um obviously most of our customers have a database. A fairly common requirement. We do have we do support multiple different databases, but Postgres is the most common one. We did MySQL, I think we've got a couple of Mongo ones somewhere and a few others So almost all our customers, yeah, vast majority of the customers use uh Postgres, uh, I think. Not all of them on 9. 6, some of them are on later versions, but
Speaker 1: 9. 6 until relatively recently was our default. So if you asked for a Postgres database, you would get a 9. 6 So that means we had quite a lot of database implementations on currently running on Postgres 9. 6. And since AWS is also our default target, if you don't want Azure or Google or something, it did mean that most of those were also on AWS. So we'll talk a little bit about 12-factor app. 12-factor app is a methodology particularly suitable for building software as a service. applications.
Speaker 1: As the name suggests, it consists of twelve principles. And following these principles can help you build for portability, scalability. and continuous deployment pipelines. As I say, it's especially suitable for deploying applications in the cloud. Now I'm not going to talk in depth about everything on the 12 factor today. Some things were more relevant for this particular use case than others. And also this is only a 20-minute talk, so perhaps don't have time to cover everything So I'm going to focus on uh the six primary ones that I think were of most interest uh just for this particular case. I think probably all of them came into play in some way or another and a lot of them are dependent on other ones.
Speaker 1: But it's these ones in bold that will get get a particular mention. If you're not familiar with the 12-factor app uh principles, um this is the 12factor. net. website, I do highly recommend people uh go and investigate it. It's uh well worth um reading up on So how do we use 12 factor? Well, two ways. We apply 12 factor principles to all our own code. So everything that we build, everything that we write, everything that we deploy follows all the 12-factor principles. That includes the control panel that we provide to our customers for them to set up their own applications for deployment. Deployment in the cloud.
Speaker 1: But possibly more importantly, especially for this case, is that we enforce it on our users. Which means They have to follow the 12-factor principles in order to deploy an application on Divio. That's actually easier than it might sound, and largely it means that you containerize your application. And you ensure that you don't store any configuration or data in the container itself. If you're doing that, you'll 95% of the way there to having follow the requirements that we need. The next slide's going to mention run books. So that was a weird echo.
Speaker 1: So just talk very quickly about what a run book is. Basically, it's a procedure for executing a set of specific test steps that will automate a commonly run task or procedure in order to produce a specific outcome. Obviously for this use case that commonly run task procedure is upgrade a Postgres database from 9. 6 to 13. The runbook wasn't new for this use case. We use runbooks for other things. I think one of the very common examples that we we've implemented previously is CloudShift. So for customer has their application in, for example AWS US East 1 and they want to move it to Frankfurt.
Speaker 1: We have a run book that will essentially pick everything up and drop it down in Frankfurt again and redeploy there. So we have a number of rumbles that we define for different things. So here we wrote a new one specifically for upgrading Postgres. So the first 12 factor principle I'd like to talk about before I actually get to it wasn't the next slide that mentioned one book, it's the one after this. First 12 factor principle I'd like to mention is admin processes. And 12 factor says about this, that you run admin management tasks as one-off processes. And if you dig a little deeper into this, what it actually means is that
Speaker 1: your any maintenance task, any admin task, pretty much any task at all that you're running for your application, be it the main application itself or anything around the outside uh is going to run in an identical environment to your actual long running processes, your task itself or perhaps background workers, that sort of thing. So maintenance tasks, short-lived tasks, they are first-class citizens just as much as the application itself is under the 12-factor principles. Yeah. So our upgrade run book, you can't really, that's the whole run book of this side here. You can't really see that, but that's just to indicate sort of the size of the thing and where we're at on it
Speaker 1: But you can see it's not particularly long. Not lots of steps in there. Implementation details are kind of abstracted away. We have some fairly human readable uh uh method names here to make it easier to show. But that does mean that I can sort of throw the run book up here and it should make it relatively simple for people to see what's actually going on So the first thing we did was um our runbooks consisted of a couple of things here: preconditions and operations. So preconditions in this case was We're only interested in customers in applications and environments that actually have a Postgres 9.
Speaker 1: 6 database. Customs running on my MySQL, not interesting. Customs running on Postgres 11, not interesting, not for this particular process So the next twelve factor principle I talk about is configuration. And the importance thing to think of here is that configuration uh is stored in the environment not in the application um So admin processes, we've already talked about them. They're the first they're first-class citizens. They're built and deployed exactly the same way as the application itself. Here.
Speaker 1: So and a build under 12 factor principles is purely coming from the code base. There's no configuration involved in the build at all. We take the build. We take the code base and we turn that into a build. It doesn't know anything about what it's connecting to at that point, it doesn't know anything about the environment that it's going to be running in at that point. It's purely a build from the code base. The configuration is external to all of this and is injected into the application at runtime. So a build plus a configuration together will then give you a release And this is quite important because it lets you then say the same build can be used for
Speaker 1: uh with different configurations to produce different different releases here It also let us use the same uh configuration for production or test when we were actually building these things, when we're actually building the RUMBOK and testing it out. So it was just a question of testing the configuration rather than changing the build. So there's another 12 factor app principle that I won't go into huge detail on, but number five, their build release run ties quite heavily into this. Which is yeah, you build your code build from your code base, you release combined two steps,
Speaker 1: and then you can sort of run the final result of that. So principle number eight is concurrency. And this is particularly important for scalability because all processes are of equal importance. Everything's um Basically running the same sort of thing, but it's the configuration that's determining what's happening in there. This makes it very easy to scale. Um, and you scale out based on the type of process that you're talking about. So If you get a lot of heavy load on your application, you can have another web process. If you're running a shop application, it's Black Friday.
Speaker 1: You can double your number of web processes. If you've got a heavy worker logo and there's a lot of background tasks running, you can just add more work processes. It's easy. I've put maintenance on here. Our maintenance is actually running as work processes. All our mention that all our code is actually developed using Django and we use RabbitMQ and Celery for background tasks. So I should have mentioned that earlier on. So yes, our maintenance tasks actually run as workers, but I wanted an extra example on the end there for just to make it obvious. But what this does is make it very easy to scale things. When we were trying to upgrade a lot of databases during the same period here, we could just simply throw more work
Speaker 1: processes at it So onto the actual operations for the run book. First thing we have to do is collect the service instances that we want to upgrade. So the service identifier here is referencing backing services and that's another 12-factor out principle that we will get to in a moment. Let's see. Which is here. So 12 factor says about backing services is that we treat these as attached resources. So backing services, it can be a database, Redis cluster,
Speaker 1: object storage, anything else basically that you need to connect to your main application. Doesn't matter if that's a local service or it's a party service. This is treated as an entirely separate resource that is later on connected to your application via configuration. I'm making changes to the backing services. largely is about changing configuration settings, injecting them into your application. So these sort of changes will take effect when you deploy, you then deploy your main application and it receives information about the new configuration. And this is rather critical for what we're about to do now with the databases.
Speaker 1: So First thing we want to do, we've got a 9. 6 database. We want uh a 30 uh Postgres 13 database. So the first thing we do is actually just provision a Postgres 13 database using the same configuration that we had for the 9. 61. And that's it. That's at this point when this is run , our project, each or the project that it's running on at the time, will now have two databases. 9. 6 and a 13. Projects at this point in the application is still running against 9. 6 database. It doesn't even know that the 13 exists at this point. It's just been 13 is there, it's on AWS, you could SSH into it if you wanted and start mucking around with it.
Speaker 1: But the application knows nothing about it yet, because it hasn't been told about this configuration. The next few steps we run, at least during the automated, when we're doing it in the automated side, we do it in maintenance mode. And that's largely because we're going to be copying data around and we cannot guarantee when customer applications are going to be changing or adding or removing data. We did provide all the steps that we're talking about here. There are ways for the customers to do themselves. So we did provide instructions to say if you want to do this with zero downtime, these are the steps to reproduce it. You just need to be sure that for these few steps here
Speaker 1: that you're not changing your database, otherwise there's a risk that you'll you'll lose. uh a number of transactions during that period, uh just the bit where the data is being copied around. So it was perfectly possible to do this with a zero downtime, but we couldn't make that assumption on behalf of anybody when we're doing it in an automated way. Okay, so 12 factor principle number six talks about processes Execute the application as one or more stateless processes. So because the build comes from only code, doesn't know anything about the configuration at that point.
Speaker 1: And because the configuration is then injected into the app at runtime, it means the application itself is stateless. You can't guarantee that anything that you did earlier is still going to be around. You could cache, for example, if you needed to read in a large-ish file in some kind of process, you could cache that, do your read-in, that would all be fine. But you could not guarantee that cash would still be there later. So um And what this means is in practice is that it's very easy to do a graceful shutdown and a relatively quick startup. Which is also quite important
Speaker 1: for what we need to do in a bit. So next stage in the run book was to actually back up the data from the old services, the 9. 6 ones. Restore that same backup to the new services and then to enable the new services. At this point, there's still two databases. They've both got the same data in them, but we haven't deployed the application at this point, so it's still running on the 9. 6 database. It still doesn't know about the 13 database at this point.
Speaker 1: Again, you could SSH into 13 and run SQL and do whatever you wanted. It's there, it's provisioned, it's working, it's got the data in. Uh it's just that the application does not yet know about it. So the next step is to deploy the new service instance. And we do that by essentially deploying the application and telling it about new configuration. We don't deploy obviously if it wasn't already deployed, so we do do a quick check for that. We're not going to force deploy someone's application if they'd taken it off themselves for whatever reasons they might have had. If that's the case, then those settings would take effect the next time they deployed it themselves.
Speaker 1: So But otherwise, yes, we simply trigger the deploy process, which is combining a build with the new configuration. And that will then cause the application to know about the Postgres 13 database. The adopt equals true there is telling it that actually you can adopt the old um build because we haven't changed anything about the build. There's no point in us trying to build a new container here because we haven't changed any customer code. All we've changed is the config. So we'll just reuse the old build. It saves time Toll factor number nine talks about disposability
Speaker 1: and says maximize robustness with fast startup and graceful shutdown. And this is largely possible because , as we talked about a couple of slides back, the processes themselves, they are disposable because they're just because they are stateless. There are some things you need to think about when you're coding for this. This is that you should minimize startup time as much as possible yourself, whatever processes you need to run in the application. Any you should strive to minimize the runtime of uh any requests or anything coming into the application. So if you've got anything that runs for a long time. hand it off to a background process, provide a callback to
Speaker 1: let you grab a status for it whenever you need it and just sort of let that run. Um processes themselves should be either idempotent or um rollback on termination uh to avoid any issues with that sort of thing. So But this largely mirrors the process that we've done here. We changed our configuration because we set up the new database. We've changed the configuration to say we want to point it to the new database We've stopped our application , we've redeployed it, combined the build with the configuration, and restarted it up there to pick the new configuration up.
Speaker 1: And finally, one last step at this point obviously we're running on the Postgres 13. We do still have the 9. 6 floating around. So we want to get rid of them. It's not particularly important for the customer. If we hadn't got rid of them, AWS would have come along a couple of months later probably and tried to force upgrade them to 13 Something we particularly wanted. And also 10,000 database instances on AWS isn't cheap. So they're not renowned for Yes, quite so yes. We did not want ten thousand database instances that weren't really in use.
Speaker 1: So quick summary. Maintenance tasks are first-class citizens, just as important as the application itself or any other long-running tasks that you may have as part of that. Configuration is stored in the environment, so environment variables, secrets manager, whatever, not in the application itself. Backing services, in this case the Postgres database , are attached resources. And the application is executed as one or more stateless processes. All of these help with fast startup and graceful shutdown, though there is some things that you need to think about when you're building things on the fast startup side.
Speaker 1: And if you do everything right, then this becomes then highly scalable and just lets you start and stop new processes as and when you need to do so. So I think that's it for me today. Thanks everyone for listening. Thank you to Benjamin and Emile for uh arranging this and letting me come and bore you all today and flick my talk upon you. Um Divio is at at Divio, I'm at at Mitch J Nich. And uh any questions?
Speaker 2: Thank you so much.
Speaker 3: And uh Like before, we'll start with questions from the audience. And there is a question there.
Speaker 2: in general that you talk?
Speaker 1: Runbook's more of a more of a concept that we've used. Yes, we've built our own implementation for uh for run books here. Um but as you saw it it's Basically it's basically a script. Yeah.
Speaker 2: And I like the declarative syntax in a sense. Yes. So it's similar to Ansible, but
Speaker 1: Yeah, essentially we've taken there's a lot of concepts that have been lifted out of Ansible there to run this. It's the same sort of thing. Um, we've abstracted away all of this deprovision yell stuff. That's you know, that's
Speaker 2: tagging and so on.
Speaker 1: Yeah.
Speaker 2: system code running uh inside that because in Ansible it's YAML basically.
Speaker 1: Yeah. Yeah. No our run books are but everything's Python on our side. So yeah, our Runbooks run as Python.
Speaker 2: Is it open source?
Speaker 1: Uh that isn't no I can talk to our CTO and see what he says.
Speaker 3: Okay, great. So the question was if uh the run book that Mike has shown was internal based on the other thing. Okay. We'll take the next question down there. Yeah.
Speaker 4: I I'm wondering about uh we're talking about uh statelessness and and then uh Multiple processing and stuff like that and it's it's a lot of things going on actually. It's yes the same thing but in a huge scale as you can say.
Speaker 1: Yes.
Speaker 4: Uh how do you uh do surveillance and monitoring uh for these kind of things, some kind of third-party uh product or you created your own stuff to get some kind of bashboard or uh cockpit for it uh to uh I mean if anything goes wrong, just a little thing goes wrong.
Speaker 1: Yeah.
Speaker 4: You may miss it.
Speaker 1: We we do integrate with a number. Sorry, you want to repeat the question? I beg your pardon.
Speaker 3: The question is uh Quite frankly, how do you do monitoring of all these processes that you're running?
Speaker 1: We do use a number of third-party tools for the error and for the error tracking stuff. For example, we have Sentry. We do have some internal metrics that we use as well. That's our purpose, that's our own uh solution that's plugging into this stuff. But um Yeah, and we we have Sentry, we've we datadog, we use um a few other things that we plug into here. So there's a there's a reasonable amount of monitoring uh going on. I mean we run As a matter of fact, we're running the we're upgrading 10,000 databases here. I mean they're all customer databases, so we've got a lot of ongoing monitoring for all of these customers. sites anyway. So we didn't have to do anything extra for this. We already had all this monitoring set up.
Speaker 3: Yes, in the middle.
Speaker 5: You mentioned in in the prelims of the talk uh that it doing it this way fix most of the difficulties with upgrading Postgres so far Would you be prepared to speak to in a high level some of those remaining difficulties and whether they were intrinsic to that grade or the scale you were doing it at or And you're not kind of thing.
Speaker 1: Um
Speaker 3: so yeah, are there any intrinsic issues remaining after the 12 factor?
Speaker 1: From a technical perspective, not really, to be honest. This let us take care of everything pretty smoothly. I think we did have a bug in the run book earlier on when we ran a couple of smaller batches and that was fixed very quickly and we sorted that out. I think any sharp edges on this one were more to do with customers and what they wanted and what they didn't want to do and what they were complaining about. Yes, we had a few of those cases. But Yeah, we got a handle and most and because we provided the I say all the steps that we put here, you could do them manually yourself as well
Speaker 1: via our control panel. So you could go and do the backups and the restore and put it all down without actually taking the application online. That was usually enough to solve the problem for the um for any customers that had any issues with it.
Speaker 3: The internet is trying to uh catch up on the question asking. Unfortunately, the question has already been asked. Just a reminder that uh everyone's very interested in the source code of the internet. Oh the rumbox.
Speaker 1: Okay, yes. I'll yes, I'll have to talk to our CTO when I get back.
Speaker 3: Any more questions? Yeah?
Speaker 6: Uh I'm struggling with disposability with the code base I'm managing at work, so and and I can manipulate that code because That's my job. But but you're running pool on behalf of your clients.
Speaker 1: Yes.
Speaker 6: How do you make sure that uh you You know, we shouldn't run long running streaming responses and long running jobs, but how do you ensure that your clients code bases don't do this? When do you terminate? How do you dispose a service that you are not totally in control of?
Speaker 1: Yeah.
Speaker 3: Yeah. Long running services, how do you dispose of something that you don't have control of?
Speaker 1: Yeah, specifically customer. Um yeah, I mean we In general we don't as a as a as a general rule for customers got some kind of long-running process. it's it's not ideal, but we're not gonna just go along and shut it down for for whenever. Um but uh yeah it it is A lot of it is up to them. They need to build for this sort of stuff in the first place. As I said, we do enforce these principles on our on our customers as well. And in order to deploy The application has to be containerized and they have to follow a certain certain number of guidelines. I don't say that all our customers are following them fully when it comes to maybe especially the I think what you're talking about with long running requests
Speaker 1: maybe. I'm almost certain there's people that are running some uh some longer ones without providing callbacks. And we see we see stuff like that pop up occasionally on sport questions and things and usually the advice is well. we suggest you refactor in this way, essentially. So um yeah, I I think that is usually, yeah, we'll we'll give the customer some guidance some guidance when we see that this stuff uh when the system happens uh basically because not yeah you're right not everybody builds for item potency or rollbacks or that sort of thing so Yeah. The human factor, I think, is the answer. Yeah.
Speaker 3: The thirteenth factor.
Speaker 1: Yes.
Speaker 3: Any other questions? I have a question. So and we do have time for more questions, so don't hesitate. But uh in order to get 12 factors introduced in your project or at your workplace, uh Do you meet any obstacles, any friction? When is it too late to go for 12 factor? Can you always introduce them one factor at a time?
Speaker 1: The first half of that question I can't answer for you, I'm afraid, because they did that before I started there. So I'm not sure what they meant. And to be honest, I think it's always been there for them. I think they've built this this this has been inbuilt from the start so um I don't think they've seen a lot of friction with that sort of thing so um Sorry, now I've forgotten the second part of the question. What was that?
Speaker 3: Yeah, is it ever too late? Uh or how would you introduce it uh when you haven't really followed it?
Speaker 1: I don't think it's ever too late. Um, and I think a lot of people are probably introducing this a a number of these factors anyway, simply by moving in towards Docker. you know containerizing everything because that when you're doing that that does kind of put some restrictions on you as well I think and and you end up doing some of this anyway even if you don't realize you're doing it. There are some rough edges sometimes, for example, uh Django migrations. Um You can't build them into your build process because when you build the code base into a build, it's not doesn't have the configuration, so it's not connected to a database. So your migrations would run, but they're not going to run against anything.
Speaker 1: So we have this concept of uh sort of release commands, which is the final thing that runs once everything's been put together along the 12 factor principles. it will do those remaining couple of things like the migrations once we've finally injected the configuration and and uh sort of uh connect wired the two things together and said the app's connected to this database. I'm not really sure if that answered the question, to be honest, but I'm not sure I yeah.
Speaker 3: Yeah, I mean it's a good point that uh We may have already introduced twelve factors without knowing it, so that uh takes a part of the heat off from uh
Speaker 1: I I think I think probably people uh recently sort of tried to containerize things. and Docker I stuff have probably done a root some of this without considering that that was what they were actually doing. So yeah.
Speaker 3: Let's see. Any more questions from the crowd here? I don't find any questions from the internet. Is that correct? There isn't. All right. But uh let's give a big hand to Mike.
It is a set of twelve principles for building software-as-a-service applications, especially to improve portability, scalability, and continuous deployment. Divio applied the principles both to its own code and enforced the key requirements for customer applications.
Discussed at 3:54Divio automated the migration as a runbook: identify PostgreSQL 9.6 instances, provision PostgreSQL 13 with the same configuration, copy a backup into the new database, update the application’s injected configuration, redeploy, and remove the old instance. Because the application build was separate from configuration, the existing build could be reused without rebuilding customer code.
Discussed at 8:50Under 12-factor principles, the build comes only from the codebase and configuration is injected at runtime. Divio could therefore point the unchanged application build at the new PostgreSQL database simply by changing configuration and redeploying it.
Discussed at 10:06Maintenance tasks ran as worker processes, so Divio could scale the upgrade operation by adding more workers when many databases needed to be processed at once.
Discussed at 11:37The automated process used maintenance mode while data was copied, because ongoing writes could be lost. Zero downtime was still possible if the customer followed the documented steps and avoided changing the database during the copy and cutover period.
Discussed at 15:40It used existing monitoring rather than building a special system for the migration, including Sentry, Datadog, and internal metrics. Since the databases were customer resources already being monitored, the upgrade process fit into that existing monitoring setup.
Discussed at 26:17There were few technical problems; an early runbook bug was fixed quickly. Most remaining difficulties involved customer preferences or objections, which could generally be addressed by letting customers perform the backup, restore, and cutover steps themselves through the control panel.
Discussed at 27:39The speaker does not think it is too late. Containerizing applications with Docker often introduces several of the principles naturally, although some tasks—such as Django migrations—need to run later as release commands after runtime configuration is available.
Discussed at 32:18Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published October 13, 2024
Published October 13, 2024
Published October 13, 2024
Published October 13, 2024
Published October 13, 2024
Published October 13, 2024