Realizing Relé – Django Day Copenhagen 2020

This video features Andrew Graham Yooll at Django Day Copenhagen 2020 in Copenhagen, Denmark.

Realizing Relé  – Django Day Copenhagen 2020
0:26:41
Published September 27, 2020
96 views

At my workplace, our system architecture has evolved rapidly to fit our product needs. One of those needs was the to bring reliable delivery of user and system events throughout all our services. We adopted Google PubSub, and eventually Open Sourced our solution, RelƩ.

Django Day Copenhagen 2020

Summary

A product catalog serving several e-commerce services through daily REST imports created scaling costs and made updates slow and fragile; a late failure could force a long import to restart, delaying urgent changes such as a product recall. Andrew Graham Yooll explains how Google Cloud Pub/Sub lets the catalog publish events once while independent services subscribe, then describes the operational problems with Google’s Python client, including thread and memory use, database connection management, and weak failure handling. His team built Relay to make Pub/Sub practical for Django, Flask, and plain Python, with a simple publishing and subscription API, middleware hooks, logging, and Prometheus metrics; at the time of the talk, Relay 1.0 was released and handling millions of messages in production.

Key takeaways

  • Event-driven messaging decouples a service that publishes changes from the services that need to process them.
  • Replacing repeated full-catalog REST imports with published events can reduce API load and make important updates reach dependent services sooner.
  • The Google Pub/Sub Python client required extra work to manage threads, memory, database connections, and failure handling in production.
  • Relay provides a Django-first API, later extended to Flask and plain Python, with middleware hooks, logging, and Prometheus metrics.
  • Yooll’s team reported managing 136 topics and 317 subscriptions, publishing more than 20 million messages and consuming more than 125 million over 30 days.

Summarised automatically from the transcript.

Transcript

4,287 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:00

Speaker 1: But for now, let's give a big warm welcome, a big applause, big hand for Andrew.

0:11

Speaker 2: Nope, that's not my dog. Um that's not mine. Uh mine 's purple. There we go, okay, cool. Cool. Um man, I just want to say thank you very much for having me. And I'm actually super super bummed that I wasn't able to attend uh in person and see you all in person. Um I was really excited to come in April and and visit your beautiful or what I hear is a beautiful city. Um so yeah, I I hope your cake is good and I hope your coffee is good. I have uh my eighty-five percent chocolate here. Uh and in in the place of cake. So yeah, yeah. Well, anyways, I'm here to talk about

0:58

Speaker 2: um real li what I'm calling realizing relay. It's a story about Google PubSub, Python, and a bit of power. I think everybody enjoys power, so I thought that was cool. A bit about me, my name is Andrew Graham Mule. I'm a software engineer at Mercadona Tech Um Meradona is the largest supermarket chain in Spain and we are the tech portion of that, so we handle e-commerce and all the supply chain and logistics software and products that we that we build and maintain. Just on the side, since this is a Django conference uh conference, I'm also a Django import export maintainer. So if you have any questions about any of this, uh

1:44

Speaker 2: or Django import export, uh hit me up on Twitter, or I'm pretty active on GitHub as well. So, um yeah, I would like to talk to you guys about the problem that we had. the solution that we finally came up with and at the end if if we're lucky um a release so cross my fingers But before we get started, I think it's I think it's important to kind of define what um PubSub is. So this is something I took from the web. And basically, as defined, it's an asynchronous messaging service. that decouple services that produce events from services that process events. Now, I now that might sound crazy to you or like

2:30

Speaker 2: not You might not understand that fully, but that's okay. Um the kind of important things here are asynchronous, uh decoupling, and events. Um that's kind of just what we have to keep in mind when we go through this talk. And I hope by the end of this talk you'll actually understand what the PubSub is. So I'm just going to go through a problem that we had. This is the actual problem that we have and why we came to this the solution that we came up with. So since we're like an e-commerce platform , every e-commerce platform has a catalog of products. In our case, we have a we have a catalog. It was an API, it had 7,000 products. Um and every day it ranges between 7,000 to 8,000 depending on the season, um, what

3:19

Speaker 2: what's in stock, what's not in stock, etc. And then again, we have like different services. So we have a service-oriented architecture. In an e-commerce platform, we had a s we had an store API, and we also have a logistics API. Um and each of these services, they both need products. So they both need the 7,000 products every single day, and they need it very accurately. So we had we have um an HTTP REST uh REST A API which would go and fetch all the products every single day to each of these services. Now there are some problems with this. One of them is scalability, which we'll talk about in a second. Another is resiliency, and another is time to recovery.

4:04

Speaker 2: These were kind of like the three main points that we were we were feeling, um pain points. So scalability Now okay, we take the same catalog API with 7,000 products. Um we do this, we do these GET requests for these 7,000 products in two different services. And in the end, you're gonna have somewhere around 14,000 API calls. Now that that's not a lot, like we can handle that, no problem. But the problem becomes when you have like other services that pop up. So we have a supply service which does another request for 7,000 products. Um and that's a total of 21,000 API requ API calls in total. And then you add another service, let's say whatever service. multiply that by infinity

4:50

Speaker 2: basically and you have like the AP number of API calls get out of hand very quickly. So that is one one sticking point in scalability with with this type of architecture. We also have the problem of resiliency. So before with our catalog, we had a process that ran about 50 minutes And what this would do is that the 15-minute process would import all the products to the shop. Now what happens if it at 49 minutes you have like a timeout error and your entire process needs to be restarted? Well what happens is that you have to restart the entire 15-minute process and you've just wasted about an hour and a half of time. Now okay, you might say like this is alright for

5:36

Speaker 2: okay, you just push another button, you start the import process again, maybe you have to restart a cron job, whatever. But actually becomes like a really serious issue for us since we're in the since we're in the supermarket in d industry. Um say for instance if we had like I don't know, some recall on milk. Um and we unpublished in the catalog the recall of milk. Well, we want to see immediately in the shop That customers cannot add this milk anymore. And in order to wait 50 minutes to guarantee that they cannot see this milk, uh is kind of unacceptable. So we really wanted to figure out a solution to this problem. Um and as a group of engineers, we kind of sat down and and we we we brainstormed a bit and we we came up with the the solution that we wanted like an event-driven

6:24

Speaker 2: architecture, some sort of some sort of event-driven architecture in our system. And what that kind of looks like is that you have this catalog service which publishes to a broker, which is in the middle here. And then in other services, they can like subscribe to this broker. So in our classic case of the product, it would uh it would publish some sort of product updated event. Go to the broker and then from their services you can pull this information from your broker. And this can scale up to uh However many services you want. What is important to note about this is that you are decoupling the catalog from all the other services. So as a developer on the catalog API or catalog service, you only care about publishing to the broker.

7:12

Speaker 2: You don't care about how many services are actually using that data in their actual service. Um so for instance like um yeah like like if you were to use like a classic API and you were to have 28,000 API calls, you would have to worry about scaling at a certain time. um or my pods like scaling to you know X amount. Um you have to worry about that kind of stuff. With publishing, you just care about publishing and it's up to the services that are that are pulling that data from broker to they're they're the ones responsible for getting that data from the broker. So uh this broker, this this famous broker, um we kind of we kind of uh brainstormed a bit as well, like what technology should we use? We thought about Kafka, which is like the famous like

7:59

Speaker 2: Event platform. We decided against it. There was also Amazon has their own solution called SNS. There's also Redis PubSub, and there's also RabbitMQ. Um in the end actually though we decided on um Google PubSub. We decided on Google PubSub because it's Um super cheap, like super cheap and highly scalable. Um when I say cheap, I mean like you can publish like a million messages. uh for free every month and then after the million messages you pay something like forty bucks or something like that per month so it's really cheap and and scalable and reliant and and all these things that we were looking for

8:46

Speaker 2: Yeah. So in the end this is kind of what our our ecosystem looks like with that broker. Now before I continue, I just want to go over a couple concepts which I think I've mentioned but I didn't really define. One of these things is like what is a topic? Uh a topic is just basically a queue. Um it lives inside the broker and it houses all the messages The publisher is the per is the entity which pushes the data to the topic. So its responsibility is getting that data to the broker. And then the subscription is basically just a callback function, and its responsibility is pulling that data from the broker, sorry, yeah, from the broker that has been published from the publisher.

9:35

Speaker 2: These are quite important topics. Um okay, so we we got this Google PubSub system working in production or staging, whatever. And we were super excited about it. It was really simple to set up. You just literally just push a button on Google Cloud Console and you have it up and running. And we're like, this is awesome. So naturally, as software engineers, we went to the Google library and they have a Python library. We said, this is really cool. Like let's try to use this. Um the only problem is that like we ran into some serious issues when we tried it out and we were experimenting and doing spikes and seeing how it would perform in production. Um one of the biggest concerns that one of the biggest things that we saw was an issue of memory leaks

10:21

Speaker 2: and also thread management. The Google library does some like weird stuff with threads. Um they spawn like thread 10 threads for every process, for every subscription that you have, which cause like memory CPU to just skyrocket and our our pods were being killed and and our application was being killed and and yeah it just was pretty nerve wracking to see that happening. So we were like, oh like what's going on here? Um finally we s yeah There's also the issue of the uh database connections. So one thing we take for granted when using something like let's say celery or dramatic is that those libraries are actually creating and cutting off connections to your database. And so the Google Pi um the Google Python PubSub library

11:06

Speaker 2: query uh doesn't handle any of that for you. So you kind of have to like hack it away and and cut connections and and create connections with your database and manage that. So at one point in production, I think we saw uh all the connections, all the threads, or I should say all the connections to the database being used up to our Postgres database and I think it took down production. It to definitely took down staging at one point. So it was kind of nerve-wracking as well to see that happening. Luckily nothing really came of that, but we look we learned a big lesson from that. And then like one of the things we wanted, or one of the things that we were missing from the Google PubSub library, were um how do we handle failures properly? So they have like these like demo example code

11:51

Speaker 2: stuff on Google and it kind of shows you a little bit like how to handle failures but it really wasn't really wasn't like production ready to be honest. And and if we were Yeah, and like in a distributed architecture, you're always gonna have failure, so you always have to be able to handle those in a proper way And in the end this all just boils down to like a terrible developer experience. I mean we spent two or three weeks trying to figure out these problems and trying to fix them and and we we realized that okay like we have this pub sub like system sitting there. um ready to accept messages and and subscriptions. But it really was not production ready. And especially if we have like many services that are going to use this thing, we don't want to have to reinvent the wheel every single time we deploy it to production

12:38

Speaker 2: So that was like something that we really wanted to solve. It was very important to us. So in the end, um we created a project. The project is called Relay. This Relay project is basically just like a collection of all of our learnings with Google PubSub. And yeah, we put it there. It's our first open source project as a tech team, and we're really excited to share it with the world. Um so I I included this map. I don't expect you all to read this. Um it's pretty illegible to me. But basically what's interesting about this is that uh the time Top here, uh the top circ blue circle is basically the e-commerce site and you can see all the arms uh stemming from it, like an octopus

13:26

Speaker 2: And that's basically all the publications and subscriptions that are going to all the different services that we have in our ecosystem. Actually, this isn't all of them, it's just a subset of them. Um you can kind of get the point that like all these threads, all these publications and subscriptions going to each one of the services is doing something. So it's super it 's it's It's highly it's a it's a big project in our team and and lots of lots of systems are using it. And again, just this is a bit more like like zoomed in a bit. Basically you would have like a shop which when you click confirm on your order, it then publishes to a topic and Other services that are concerned with this order

14:11

Speaker 2: receive this order data. So like our last mile delivery team wants to know which orders to deliver. Our fraud team wants to know which orders, is this order a fraud or not, and mark it as fraud, if so. And then supplied because they need to know how much how much they're going to sell tonight. So they're they're concerned with that as well Um so Relay I think is has some cool features that I want to point out. Um One of them is, well, we're at a Django conference and and Django, we we are a Django team. All of our services basically run on Django or pure Python. And yeah, we developed Relay as a Django application first. Then later we added Flask support and then pure Python support.

14:59

Speaker 2: So yeah, Django was like a uh like the first first class season, if you will, of this project Another thing that we introduced in this project, um which I think is really powerful, is called middleware hooks. Um and basically what this allows you to do is give plug into the entire life cycle of Relay. Um we include Python logging out of the box and we also uh include Promethe Prometheus Prometheus metrics out of the box as well. And I think the Prometheus metrics are really powerful as well. It allows you to do cool things like show product like what's going on in your system at any given time. So I pulled this yesterday from our from one of our uh grafonographs, and it's basically like just like showing you uh all the

15:44

Speaker 2: all the order-arrived events that are that are being delivered at that moment. So it's kind of cool you can see like how how your system is performing and how events in real life are actually being used on your system. Uh yeah, so familiar API. So we came at this whole project from a developer experience point of view because we knew that the Python um pubsub library from Google uh really was not good or or production ready. So we really wanted to make the developer experience as good as possible. So we make publishing like super easy to do. You just tell it a topic and then you publish data in JSON JSON valid form.

16:30

Speaker 2: So like a regular Python dictionary. Uh and then to subscribe or to receive that data, you just decorate a s uh function, which is your callback basically, um and you just pass the the decorator uh your topic name and then you do whatever you want with that data. And that's it. It's super Super straightforward. And also we we have been using A Celery in our some of our services. So we wanted the API to be a bit like Celery. And we also took a lot of uh We also were influenced heavily by Dramatic, another another library. So yeah, if you're used to these two libraries, you'll be kind of used to to Relay. And one question I get a lot is like how production ready is relay? Because if this is going to be an integral part of your architecture, you really want to know if it's

17:17

Speaker 2: uh production ready or not. I think it is. And uh for the past year we've we've been managing 136 topics with Relay. Uh that's would be about like 300 well, we have 317 subscriptions being managed by Relay as well. Um I checked yesterday and over the past thirty days, we've published over twenty million messages successfully in our system. And we've consumed over 125 million uh messages successfully in the past 30 days. So I think I think it's quite production ready. Um and when I was when I was writing this talk, um I kind of came up with the name Realizing Relay, but then a couple days ago I was talking to a colleague and I was like, oh man, we should really call it releasing relay.

18:08

Speaker 2: Um and I think that's that's because like during this talk, we actually released Relay 1. 0 uh because we really believe that it's production ready. Um so yeah, I think I on PiPi right now, you should be able to see that 1. 0 is actually published. And that's Relay released. If you have any questions or comments or want to contribute somehow, you can always visit our repository at Mercadona Relay. And then I put Hectoberfest in here because October is a couple days away. And we hope to be tagging some issues on the repository with Hectoberfest. And that's so that the community can help contribute or if they have any questions, you can find us on on GitHub.

18:54

Speaker 2: And that's about it. Not sure how much time I used. There we go. Okay. Oh

19:40

Speaker 1: That's what

19:41

Speaker 2: we are live.

19:42

Speaker 1: Now you can hear me. Alright. So we'll just Gearing up here to see. Oh yes. Questions from the internet. They are not on my screen. They're coming on my screen now. Okay, great. So First question is from Joel, he asks, what is the Django ecosystem like in Spain?

20:11

Speaker 2: Ah, well we do have we have Python ES. Um we have Django uh I'm not sure if we actually have Django I'm in Valencia, so I should say that I'm Valencia. I don't I don't know much about the ecosystems in Madrid or Barcelona, the bigger cities. But the Django the Django ecosystem here in Valencia, I would say, is not probably not the biggest, but it's not that bad. Like I think we are actually the biggest Django shop in Valencia. In the greater Spain area, I know it's quite big. We have Django Girls meetups a couple times a year. I think this Friday actually we're doing a Django meetup online. um for Django, yeah, Django Girls, Django Girls ES. And yeah, I I think it's growing. I hope it's growing.

20:56

Speaker 2: And I know the Python community is growing in in Spain. And yeah, I hope I hope Django grows.

21:03

Speaker 1: These are exciting times uh in the sense that we uh connect our communities in in a new way than before.

21:11

Speaker 2: Yeah, exactly.

21:15

Speaker 1: I just need to uh switch my C There are no questions popping up here. Okay. You'll be uh with us on Zulib, I hope, uh if there's any more questions.

21:38

Speaker 2: Yeah, yeah, I'll be there, no problem.

21:41

Speaker 1: Oh, I'm being told there are questions. I just have to look

21:45

Speaker 3: have a look under the that one. Um maybe.

21:52

Speaker 1: All right. Cool. Yes, there are questions. So do you know if there's any other providers of uh subpub than Google Cloud and if there is a particular reason why you picked uh Google Cloud for this?

22:08

Speaker 2: Yeah, so um There are many like event streaming systems, like Kafka is the biggest one. Um I think I said Amazon has its own provider. I'm sure I'm sure there are a lot of cloud providers out there that provide their own managed solution. There's also open source, like Redis provides a pub subsystem. And RabbitMQ is like your classic pub subsystem. But yeah, we we decided on Google because um because A it was like the simplest to set up. Like we had we we run on Google Cloud, so it was literally just a click of a button. But then also again like I said it's super cheap uh to run. I have my own instance that I that I have personally and I

22:53

Speaker 2: I don't even come anywhere close to hitting the limits. And you can do really cool stuff with it. And it's really reliant. We haven't had any issues with it so far. Yeah, the service has been really good that we've had.

23:08

Speaker 1: Cool and um on on a related uh n note like when when is it uh good to start using relay? Uh is it uh already in in like the prototype hobby phase or is it when I have an enterprise system with uh thousands of items in the store?

23:27

Speaker 2: It depends. Like we had we had a full production system up and running with many, many services, and we introduced relay. only a year ago. Um and it's actually been very smooth to move from from what we had before to to what we have now, which is more of of a pub sub-based system. Of course there's gonna be like things that you're gonna have to watch out for, but but those things are just natural when you're body grading system from one to the other. I I would point out though that like you could probably start using it pretty early on in a project. I'd be hesitant to like say, yes, you should start out with this early on in a project because to use such like a big system in early project probably is not the most lean thing to do, but it actually is like

24:16

Speaker 2: we've actually replaced celery in our in some of our projects. with PubSub, just publishing and then subscribing to the same to the same topic in the same application. And it's actually worked out really well for us. So I would say if you have salary in your project, you can probably replace it with Google PubSub. Yeah.

24:34

Speaker 1: Which is probably a setup that people are familiar with. Uh

24:37

Speaker 2: yeah, exactly.

24:38

Speaker 1: Easy to comprehend where to use it.

24:41

Speaker 2: Exactly.

24:42

Speaker 1: Any other questions? Yes, there is one. Do you monitor performance and health of uh PopSup? Uh should you do it uh separately? Can it receive data as it And is data sent out? Uh so yeah, I guess the question is about mon monitoring the the performance and health of uh PopSub. Is that

25:06

Speaker 2: Yeah. So that's a good question. Um Yeah, we do monitor health and performance of PlobSub. Um Google provides metrics to monitor PubSub and how many, let's say, like how many messages you're publishing to to PubSub, how many messages you're you're acknowledging from PubSub , and it it publishes those metrics publicly. uh to us. And you can also monitor it not directly through Google, but you can monitor it through your application metrics as well But the only thing, yeah, is like it's a managed service by Google, so we can't see exactly, you know, how many CPUs how many CPUs our Google PubSub instance is running on or Or we can't scale it up or down really. We just kind of trust that Google uh does the right thing with our PubSub

25:53

Speaker 2: instance.

25:54

Speaker 1: All right. Thanks so much. Yeah, uh your sound was brilliant, your video was brilliant, and uh I can tell you uh that uh Perhaps we haven't told been I haven't been so good at sharing this uh all day, but there's uh people sitting around on uh tables here in uh Copenhagen. Not so many because we can't be a a a giant event. Um but uh it's been exciting to have you here with us and uh let's uh let's keep this up and connect in the future.

26:25

Speaker 2: Hey, thank you very much and I hope to make it there someday

26:30

Speaker 1: All right.

Questions this talk answers

Why use event-driven Pub/Sub instead of having every service fetch the product catalog?

Having each service fetch the full catalog creates more API calls as services are added, and long imports make updates and recovery slow. With Pub/Sub, the catalog publishes events once and subscribing services consume them independently.

Discussed at 6:24

What problems did the Google Pub/Sub Python library cause?

The team encountered high memory and CPU use from its thread management, unmanaged database connections that could exhaust PostgreSQL connections, and inadequate guidance for handling failures. These issues made it difficult to use in production and prompted them to build Relay.

Discussed at 10:21

What does Relay add to Google Cloud Pub/Sub for Python and Django apps?

Relay packages the team’s production learnings into a library with Django, Flask, and plain-Python support, middleware hooks, logging, and Prometheus metrics. Its API makes publishing JSON data and subscribing with a decorated callback straightforward.

Discussed at 12:38

Is Relay production-ready, and how much traffic has it handled?

The speaker considers Relay production-ready: it managed 136 topics and 317 subscriptions, and over the previous 30 days the system successfully published more than 20 million messages and consumed over 125 million.

Discussed at 17:17

Why did they choose Google Cloud Pub/Sub over Kafka, RabbitMQ, or Redis?

They were already running on Google Cloud, so Pub/Sub was simple to set up, and they found it inexpensive, scalable, and reliable. The speaker also mentions Kafka, Amazon’s managed option, Redis Pub/Sub, and RabbitMQ as alternatives.

Discussed at 22:08

When should I start using Relay in a project?

It can be used early, but the speaker cautions that adopting a large system too soon may not be lean. The team also replaced Celery with Pub/Sub in some existing projects, including by publishing and subscribing within the same application.

Discussed at 23:27

How can I monitor Google Cloud Pub/Sub health and performance?

Google provides metrics such as messages published and acknowledged, and application-level metrics can also be used. Because Pub/Sub is managed, users cannot directly inspect or scale its underlying CPUs and must trust Google to manage the service.

Discussed at 25:06

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos from Django Day Copenhagen