From Dinosaurs to Snakes - Decomposing the Enterprise Monolith

This video features Michael Nicholson at Django Day Copenhagen 2024 in Copenhagen, Denmark.

From Dinosaurs to Snakes - Decomposing the Enterprise Monolith
0:35:11
Published October 13, 2024
72 views

Django Day Copenhagen 2024 talk descriptions: https://2024.djangoday.dk/

Transcript

4,782 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:00

Speaker 1: And welcome on stage, Mike. Thanks so much for coming. I suggest that we start by giving you a big applause. Status.

0:18

Speaker 2: Alright, thank you. Um I'm going to talk about sorry, is that working? You hear me okay? Yeah. I'm going to talk about some of the process we're going through at work to um Get rid of our old legacy monolith and move it into a Python Django architecture. Something we're very much in in the early middle of. So So who am I? If you recognise what that is, I know roughly how old you are. That's what I started coding on. I've had most roles

1:04

Speaker 2: over the years and spent a lot of time as a legacy coder, so we'll get to that in a bit. And these days I'm pretty much Python and Django. So I'm on Blue Sky and Mastodon. Technically still on Twitter but haven't been on there for months. So this is our legacy system, the traditional big ball of um we'll say mud. That's a little unfair and it's more really people's perception of the legacy system. It is big. It does an awful lot of stuff. It's been around since I think

1:51

Speaker 2: 92 on our site. So there's an awful lot of convoluted code weaving its way in and out. There's multiple does multiple different things. Um so there is a lot, a lot of co-complexity in there. Um And on a very basic level, it kind of looks something like this. We've got the monolith in the middle. There's a thousand other systems. that it kind of integrates with in various interesting ways. Flat files, SOAPs, APIs. Yeah. Hundreds of users, most of them kind of uh customer care representatives or things for the company, but lots of business people as well.

2:43

Speaker 2: And we have the database which is not a relational database management system. It's essentially a load of flat files. So that's what we've been working with. A little bit more complicated than that. So that's essentially what we've been working with for many, many years. And a lot of these systems that we integrate with are they're also dinosaurs and probably will stay that way for a lot longer because they work. Some of them are really small. Some of them probably don't have anyone actually actively maintaining them because that person left 15 years ago but You know, I think you all know how that is.

3:31

Speaker 2: So uh uh this discussion's been on ongoing for many many years. But the the the cons of the current system, the code is complex, different areas very heavily tied together. The legacy language that it's written in I won't go into that here, but ask me later if you're interested. It's dying. Uh there are very few consultants available for it. Um so it's hard to get hard to get hold of people. Uh in fact I think we've got most of them. And recently the licensing costs shot through the roof, largely because the company that owns it knows that it's dying and are trying to squeeze the last few dollars out of it.

4:18

Speaker 2: However, there are an awful lot of pros with it as well. It is very quick to develop in. quicker than any other of their in-house systems despite the fact that it's probably larger than any of the other in-house systems. The the database, uh the flat file one I was talking about, it's uh it's an ISOM-based database, um CISOM to be specific. Uh it's very, very quick for both read and write. Obviously it's not relational, as I said, you've got to maintain the relationships yourself. And because it's been around for over 30 years. most things are solved problems. An awful lot of new business demands that come along can be met with you know

5:06

Speaker 2: configuration and minor tweaks here and there. I don't want to put that last one into pros. Yeah, that's not really a pro. It it can't can't really be moved into that. Yeah, maybe, but that's not something I'd want to attempt. So what were the options here? They want to swap it out for many reasons. So the options are do nothing, which should always be an option. It should always be the first option. Do we really actually want to do this, to do anything here, or are we okay?

5:52

Speaker 2: is. Do we build a new monolith or do we start looking at a distributed architecture with services and blah blah blah? So excuse me. Uh so the do nothing option. Uh yeah. I mentioned there aren't many consultants left for this language. Most of them are kind of around my age, so in the next five to ten years. But people are looking at me looking at retirement and dropping out of the team and that would leave them in a bit of a bind. And the huge support costs I meant mentioned earlier have led

6:39

Speaker 2: to us going out of support with the language itself. um which isn't isn't a problem from a development point of view, isn't it really a problem for the team because like I say we've most things are solved problems. we can we can sort things out ourselves if we need to but enterprise companies do not like being out of sport they like to have that box ticked um so this this this wasn't an option And we we did kind of know that because they've tried to swap this out three or four times over the years and it hasn't really gone very well. So option two, build a new monolith in a new stack.

7:25

Speaker 2: It doesn't solve the problem of code complexity. That would all still be there. We've got all of that stuff. It doesn't enforce the separation of concerns. We could have gone the modular monolith routes, built the monoliths, separated everything into smaller apps and tried to enforce that. But the yeah, the the risk of leaking across the boundaries there and and just ending up with the same big ball of mud in a different language was very real. I think. So more to the point, um it would still be a single team probably maintaining uh the single system there. Um and I'd say this this is a huge system. It's it's It's um it's billing, invoicing, customer care, general ledger, um

8:13

Speaker 2: inventory stuff and dozens of other things. So all sorts, all things that don't necessarily normally hang together. So uh not an impossible not an impossible idea, uh but perhaps not the ideal one uh from an organization. point of view. So the third one was build a more distributed architecture. The main benefit of that being that we could have multiple teams with smaller areas of responsibility more easily manageable systems. If any of those ones needed swapping out in the future for something, that will be a lot easier to do rather than having to deal with this massive legacy monolith again a second time

8:59

Speaker 2: in Ten, fifteen yes. Um but moving to services or microservices You're swapping code complexity for architectural complexity. The complexity doesn't disappear, it just moves into a different area And microservices themselves should not be seen as a solution to um system complexity in any way. It's they're a solution to organizational complexity. They they make it easier to manage things in teams, smaller teams, more focused teams. So but this was the this was a preferred option.

9:46

Speaker 2: And going into the analysis, we did kind of know this was a preferred option as well. This was largely what was talked about from the beginning but but this is the route that that we're going and probably the route most people would have expected from the start so So the proposed stack, pretty standard for this room probably. Python, Django, Postgres, Reddit, Rabbit. um celery. Initially we're talking about deploying everything onto the same server as the current monolith. because that was our uh server. But with making sure we keep one eye on what we can deploy to

10:31

Speaker 2: AWS, they're very much on a cloud transformation journey. at this point so things are being slowly moved into the cloud at work um but starting this new project straight in the cloud wasn't wasn't quite on the cards there. So initial deployments all to all to the same server, which is which isn't a problem because it's it's it's you a huge huge server. And we're talking about not so much microservices but but sort of macro services um to coin a new phrase. Eights to ten, I think we might be on eleven or twelve actually at this point.

11:16

Speaker 2: that reflect kind of the the core functionality of um the legacy system there. So We're talking eight sort of 10, 11 smaller monoliths essentially, each running as their own service. more would be breaking up into too many, you know, creating too many future teams and making it a little bit pointless, adding unnecessary uh unnecessary complexity to the architecture and all the communication everything so um we thought this was roughly This was a good goal and and manageable from an organizational point of view.

12:02

Speaker 2: And event-driven communication via uh using RabbitMQ. Um as a message broker. Okay, so what did we start with on the knowledge base? We had all the domain knowledge and experience there, because these people we've been working with this legacy system. for a long time. Little to no Python or Django experience and very little experience around distributed architecture. Within a team, that is. There is some within the company, but not so much within the team. because it's just because it's tied to this existing legacy monolith.

12:47

Speaker 2: So at the beginning excuse me Velocity was prioritized and we started building to get a proof of concept out. just to show that we could actually deliver something. Initial functionality was all built around Django Admin. Have you and if you've seen some red flags, I understand that, but um the idea being to make as few changes as possible to the existing monolith um because the goal is obviously to to deconfigure. mission it. Something we're still not sure if we'll actually do long

13:33

Speaker 2: term, if if if it will ever fully go away. We're not sure yet. And they've been we've been uh sharing databases to a certain extent to avoid some certain bits of architectural complexity and speed up the development. And that kind of worked for a while A little later into the process, it's been going for about to what have we known now for two and a half to three years, something, I think. So we got a bit more knowledge in the team now.

14:20

Speaker 2: We've got a couple of new Python Jenga resources on board, including me. Some of the existing team members have had training. A lot of the uh existing legacy coders are still very much used to uh some more procedural based code. Um OO concepts are can be a little confusing sometimes um just yeah just because it's it's they're not used to it so There's a lot of reliance on reading across databases and I'm also talking about to the legacy database here. We've got drivers in to read some data. from there that we haven't yet migrated into the new things. That's that's the less problematic part to be honest.

15:09

Speaker 2: Everything's still being done via Django admin at this point. This isn't now. This is a little while back, with no real front end other than sort of login stuff and things. And we don't have tests at this point. because velocity prioritizing velocity kind of uh yeah uh precludes writing tests because tests take time So next step. I'm hoping you learn from what we did wrong as well as what we did right here, just to be clear. So next step, okay, what what patterns are we looking at here?

15:55

Speaker 2: So what we've done is we've introduced uh domain models for the core information concepts and here we're getting into the domains some domain-driven design stuff. So if you took a photo of Joseph's slide earlier when he's talking about some of this stuff, now it's time to take it out and look at it sneakily. So we've created domain models for the core information concepts. Ours are just data classes for the domain models, not uh we're not using Pythantic. um just because that would have been another new thing to introduce to the team um I think and the data classes stuff that we've got working just fine. I don't see any issues with that so um we're also using this service repository pattern here um the repositories are mostly a wrapper around uh um

16:40

Speaker 2: the the Django ORM stuff um but they should allow us to also put some more complex SQL stuff in there later on if we need to and and hide some of the slightly weirder things that might be going on with the uh with these cross database reads until we can factor that stuff out later on so at least that's all self-contained in one place And the services, as as Joseph mentioned, that they are they're largely the business logic side of things there. We're using Celery for asynchronous tasks and also integrating with the RabbitMQ stuff and introducing PubSub for uh some communication between the services.

17:27

Speaker 2: Um co very do I need to skip across couple of slides possibly. Very quickly the your code would look something like this. You'd create a domain object, you'd get your service. You'd call a method on method on the service which handles business logic. You'd see if there were any errors for that and you'd handle them. Um so your your main code will just be. And if you know the reference for that, you you probably know what the ZX81 was as well. And you're probably British, actually, in fairness. Alright, so this is kind of what the current architecture looks like.

18:13

Speaker 2: We've got the monolith over there, we've got the Postgres database there, we've got our new services marked with a little snake, old ones with a dinosaur, all that sort of stuff. So Everything's reading Postgres. We've got some APIs coming in, but we've also got some legacy database reads going on from the services. Yeah, it's not not ideal. Um but it but it is working at this point, so So the next step is introducing the uh the pub sub architecture. Um so we Get RabbitMQ in place. We use the Outbox pattern at this point, so we we put

19:01

Speaker 2: messages into a specific table and we've got a worker that runs, just picks them up out the table. and shoves and pushes them to RabbitMQ to a rabbit MQ exchange. We are PubSub worker which is subscribing to in each well in each service. This is only in one service at this point. We're trying to roll it out to the others now. So we've got PubSubworker, which, or more than one probably, which will subscribe to various Q using RabbitMQ that it needs all integrated with Celery. Celery is calling in some places into APIs, into the legacy database. We've still got the legacy database reads going on, we've still got the Postgres , some Postgres sharing going on, and we haven't done anything with any of this stuff down here.

19:53

Speaker 2: The final goal looks a little bit more like this. Um Which is we've still got our pub sub stuff with all the services on the left hand side, but now we've very much got this kind of separation. We've done away with these with these um uh this database sharing and the legacy reads. And we've put this, we've kind of wrapped um the legacy system in um our own broker service which allows the legacy system to throw things into an outbox um and sort of create their own uh and actually interact with the rabbit m queue and and be part of this pub sub architecture, something it it can't do natively.

20:39

Speaker 2: So we're having to um uh write our own broker um fill that to handle that stuff. That will also, the broker, it's also intended that that will s sort of act as a an API gateway a little bit into the legacy system, which will also allow us better to control some of these incoming APIs and perhaps apply things like the Strangler pattern later on to start moving things around when we need to, because then we can redo direct stuff uh a little bit better. And of course that's actually the goal. We don't want all that stuff um still communicating directly without VOC with with the legacy system. We want that coming in

21:25

Speaker 2: calling the APIs or or integrating with RabbitMQ. So all that stuff needs to move to one of those two other things on the left. But that's probably a lot further down the line. And they're not systems I'm dealing with, so I'm just sort of going la la la la la la la for that So the biggest risk of this for and and this is very subjectively in my opinion. It's not something everyone agrees with, but I I think the biggest risk that I see is that we We partially deliver the new architecture. We lack some of these service boundaries that are sitting in place, but it's all deemed good enough by the business who see everything hanging

22:12

Speaker 2: together and actually working without really seeing what's going on underneath, which means what we've end up doing is kind of extending the existing monolith so it's now exponentially more complicated and now includes Python, Django, Rabbit MQ, Celery. Redis, yeah. Lots of other things, which is, you know, not not ideal. We still have all if if that happens, we still have all the problems from the earlier slide about yeah um licensing and you know cons available consultants all the other stuff so all the organizational problems would still be there but um I mean it depends a little bit on how much we actually get out and into the new stuff

23:01

Speaker 2: So for anyone going through this, um some thoughts about how to sort of maximize your chances of success uh if you are doing it. Get the buy-in. Don't try and do this sort of part-time with with a uh with a group of people that that you know know the b know the business side well but uh we suffer a little bit from people getting pulled off the the new stuff to go and work on the leg system all the time because obviously again not that many consultants available with that knowledge so there's this constant pull uh to take resources off so But so make sure the buy-in

23:46

Speaker 2: is there. A proof of concept is good, but it does need to prove something. And once you've proved that, move move on. And make sure you've got these dedicated results. As I say, we we do suffer a little bit from that one, I think. I'm building a temporary solution with quote unquote bad architecture. Again, not everyone agrees with my definition of bad here. when when I'm discussing this, but um make sure that you're you're planning in the next step to actually clean this up. Yes, we're doing this temporarily. Yes, it will be like that for six months, but then this is

24:31

Speaker 2: this is happening. We're getting rid of it and it's just it's an intermediate step to get to the end goal Do it sooner rather than later, otherwise it becomes part of the processes and it's a lot harder to remove. um and do establish areas of responsibility with other internal systems who who's responsible for uh for what for which APIs and where they're where they're being called and communicated. On the resource front, um, as I mentioned, we we didn't really have the Python jack. resources at the beginning. Find out who wants to learn a new tech and who doesn't. Some people were are very keen on doing that, some people are less keen on doing that.

25:21

Speaker 2: There's little point in having the less keen ones do it because they've got all the business knowledge, use them for the business knowledge in that case. So uh and try and make sure that you're a little bit in in line with Raphaela 's uh uh talk earlier, um try and sort of make sure people have got mental or someone to help and guide them through things as they're doing sort of foundational changes if they're still new to a lot of this stuff. Little bit further reading. Yeah, more off a time. If you've not seen these ones before, definitely uh

26:06

Speaker 2: worth doing both the Sam Newman books um and the cosmic python one there um have been very helpful There's a few domain-driven design ones as well, so which are which are quite good, but I think these are the main ones to think about. And I think that's it. Thank you

26:34

Speaker 1: Thank you. Um so you uh referring back to the early days uh when when when this decision was made, can you remember some of the reasons why Django was chosen for for this task?

26:51

Speaker 2: Um yeah, it was actually one uh one of the dev devs on our team was pushing very hard for uh for Django um because he'd started using it a little bit earlier. and and liked it. And um we got a proof of concept up and running and and yeah, they liked it. So the Python Django and Python was an improved technology already at the company which helped. So they've they've got a list of of languages and tools and all sorts of things which you are allowed to use and if it's not on the list you have to try and get it approved. Python was already there so that made made life a bit easier. So is

27:35

Speaker 1: the uh question? There there's there's a question there, oh and there and oh there. And there. Yes.

27:42

Speaker 3: Uh you Thanks a lot for the talk. And uh you mentioned that uh that you're still lacking a front end. What what is it that you're sort of missing from the Django admin that that cannot be built on top of the Django admin.

27:59

Speaker 2: There's nothing that can't be built on the Django admin, I don't think in that sense. But as we're trying to sort of refactor some of this stuff, the having everything in the Django admin becomes very hard to build tests for. Yeah, yeah. We are starting with tests now, yes. Uh Yeah, we have we have some tests. I've got the one I'm managing is is up to about 60% coverage now on unit tests, and there's a couple of the other the services I've got um they're do they're writing some end-to-end tests for them at the moment so so so that's positive but yeah the Django admin

28:44

Speaker 2: side that makes it very hard to to actually test. um properly. And also we are the because we've still got a lot of integration going on with the old system here and things, it's that uh trying to pull other data in into admin but also becomes a little bit complex. So we do have a little bit of front end at the moment. I I did the first front end stuff other than the login side last year so we we've got a few front end screens on there um and we're using we're using HTMX for because there isn't any there isn't really any JavaScript knowledge

29:31

Speaker 2: or not a lot of JavaScript knowledge in the team either. Um but HTMX seems to be working nicely with with some caveats. But so so we are we are moving on the on the front end stuff, uh I think. And yeah, a lot of it could be done in admin, but having recently refactored some stuff out of admin. the hoops you had to jump through to get some of this working , it's not worth it. Do it in the front end to start with if you've got that possibility. Yeah.

30:14

Speaker 1: Yeah, we had uh a lot more questions. Let's uh be good with the time. Yeah.

30:18

Speaker 4: Thank you so much for your talk. Uh my question you've partly answered. I was going to ask whether you've done any automated testing, but uh I think it's handled now.

30:28

Speaker 2: Yes, well it's Handled is a relative term, but yes, we are doing some now. They've also they've got Jenkins, um is their um CI tool of choice at works So one of our services is now hooked into Jenkins so it will run both linting and automated test on on the pull requests um so you can't merge domain unless unless they've passed that's own for the moment that's only on one one repository um yeah

31:09

Speaker 1: Thanks for the talk and that was a nice segue into my question, which is which is You didn't mention any challenges around

31:17

Speaker 5: uh version management, development process, release management. Has there been, you know, like a change there?

31:26

Speaker 2: Um what do you mean a change from the U.

31:29

Speaker 5: As in are there differences to how the dinosaur was built and developed?

31:33

Speaker 2: Yeah, absolutely, yes, most most definitely. release processes um which really don't align with anything that we're doing with this stuff now. The release process we've got at this point is fairly simple because um we we basically tag the main branch with a date-based release tag and then we pull we check out that tag in the production environment. when we're doing it. So essentially. There are some challenges there because the the repository definition for the legacy file system sometimes gets regenerated and that's sitting in the repositories at the moment which means

32:25

Speaker 2: Git technically looks out of sync. That's something that we need to change. So there are challenges with with the rebelease process. Um but at this point we're keeping it as simple as possible. I anticipate it changing at some point the first time we have to start doing something with AWS. So um and at that point I don't know w if if we'll be moving to you know Docker-based deployment or something else. They do have a Kubernetes cluster at work, so maybe we'll be using that. I don't know.

32:55

Speaker 1: Um we don't have so much more time, so I think we have to just take one question and then how do we pick that one question. There were I I saw two hands before. We take the first hand then.

33:14

Speaker 6: Yeah, my question is a really small one. Thank you for the great talk. And you mentioned that uh your team back in the days uh uh was uh wasn't really familiar uh deeply familiar with uh microservice uh concepts.

33:32

Speaker 2: Yeah.

33:32

Speaker 6: And uh my question is uh how much time has passed uh since you first decided to change your architecture until you developed your first proof of concept and uh how big the team was

33:46

Speaker 2: The proof of concept we we had before we we actually started writing anything the first time. Um I've worked at this place a few times and I I had a gap in the middle when I wasn't there which is where they started things and then I came. came back. But the the proof concept was there before we started building the main stuff. And that was very much just here's a service, we're deploying it. And we didn't have any of the architectural complexity in place then. We didn't have PubSub, we didn't have any of that. I did the rabbit stuff early last year. um for the for one of the services um for that. So that was when that was when rabbits and celery

34:32

Speaker 2: came in for things like that So the architectural complexity and the and the the the true microservices architecture has only really started coming into play in the last sort of 12 months or so Twelve, yeah. Twelve, fifteen months, something like that. If that answers your question, I'm not sure.

34:54

Speaker 1: Yeah. Really nice diagrams on the slides by the way. Uh I hope Danny is here. Yeah. Thank you so much.

Questions this talk answers

Why choose a distributed architecture instead of replacing a legacy monolith with a new monolith?

A distributed design lets smaller teams own more manageable areas, and makes it easier to replace individual systems later. It doesn’t eliminate complexity; it shifts some of it from code into architecture, and is primarily useful for managing organizational complexity.

Discussed at 8:13

What technology stack and service structure are they using to break up the monolith?

The proposed stack is Python, Django, PostgreSQL, Redis, RabbitMQ, and Celery. Rather than creating many tiny microservices, they chose roughly eight to twelve larger services aligned with the legacy system’s core functions.

Discussed at 9:46

How are the new services being connected to the legacy system during the transition?

They’re introducing RabbitMQ-based pub/sub, using an outbox table and worker to publish messages. The target design wraps the legacy system with a broker that can connect it to RabbitMQ and act somewhat like an API gateway, while eliminating shared database access and direct legacy reads.

Discussed at 19:53

What is the biggest risk when decomposing a legacy monolith?

The team could deliver only part of the new architecture, while the business considers the result good enough. That would leave the old monolith’s problems in place and make it more complicated by adding Python, Django, RabbitMQ, Celery, Redis, and other technologies.

Discussed at 21:25

What helps make a legacy-system migration more likely to succeed?

Get business buy-in and dedicated resources, and make sure a proof of concept actually proves something. If you use a temporary design, plan and schedule its cleanup, do that cleanup early, and clearly assign responsibility for integrations with other internal systems.

Discussed at 23:01

Why did the team choose Django for the replacement system?

A developer on the team advocated for Django, and a proof of concept helped demonstrate it. Python was already approved for use at the company, which made adoption easier.

Discussed at 26:51

Why move beyond Django Admin, and what are they using for the front end?

The speaker says Admin can do a lot, but it became difficult to test and awkward to use for integrations and refactored functionality. They’ve started building front-end screens with HTMX, which suits a team with limited JavaScript experience.

Discussed at 27:59

What testing and CI do they have in place for the new services?

They’ve begun adding unit and end-to-end tests; one service has about 60% unit-test coverage. Jenkins runs linting and automated tests on pull requests for one repository, and blocks merges if they fail.

Discussed at 30:28

What release-management challenges have come up during the migration?

The legacy release process doesn’t fit the new services very well, and regenerated legacy repository files can make Git appear out of sync. For now they keep releases simple with date-based tags; deployment may need to change when they move to AWS.

Discussed at 31:33

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Michael Nicholson

More videos from Django Day Copenhagen