Pair Programming after the Pandemic and Beyond
Published July 11, 2024
This video features Tobias McNulty at DjangoCon US 2023 in Durham, North Carolina, USA.
By an order of magnitude, Celery remains one of the most popular Django-adjacent packages for Python. In this talk, I'll explore what continues to make Celery a go-to solution for background and scheduled jobs with Django, how to integrate Celery with a Django project, and some common patterns to use (and avoid!) when writing tasks with Celery.
This talk was presented at: https://2023.djangocon.us/talks/how-to-schedule-tasks-with-celery-and-django/
LINKS:
Follow Tobias McNulty 👇
On Twitter: https://twitter.com/tobiasmcnulty
Follow DjangCon US 👇
https://fosstodon.org/@djangocon
https://twitter.com/djangocon
Follow DEFNA 👇
https://www.defna.org/
Video production by the presenter and DjangoCon US 2023 volunteers.
Celery lets Django applications move work outside the request-response cycle, distribute it across workers, and schedule recurring tasks with Celery Beat. The presenters explain how to configure Celery, choose RabbitMQ or Redis as a broker, define tasks, use Django-backed result and schedule stores, and compose work with chains, groups, and chords. Their voter-roll example shows why tasks should be small enough to finish during graceful shutdown, while avoiding synchronous waits and excessive parallel database work; database-heavy steps may be better kept synchronous while CPU-heavy work is parallelized.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: Hi, y'all. My name is Kel Hanna. I'm the Chief of Curation at Cactus Consulting Group. As developers and tech professionals, I'm pretty sure you all are familiar with solving for unknown quantities, and as fate would have it. Tobias McNulty is out with a sore throat and unable to give his talk today. Cacti. Are known for withstanding some extreme and sometimes unpredictable conditions. And our team is a manifestation of that enduring quality. So we adapt and we support each other Tobias has asked a few trusted people to step in and deliver his presentation, and they graciously answered the call.
Speaker 1: Co-presenters Kenya Phelps, a software engineer at Cactus, one of my teammates, and Ahmad Faturi. Great friend to Cactus and a senior software engineer will develop deliver today's talk on behalf of Tobias. So, next, I'm going to tell you a little bit about our co-founder and the creator of this presentation, and then pass the mic to them to get to the good stuff. That's DeBys. He is one of the co-founders of Cactus and started using Django in 2007. Cactus is an employee-owned Django First consultancy based here in Durham.
Speaker 1: We support the full life cycle necessary to design, build, and maintain custom solutions for clients with a special affinity and talent. For Python, Django, and mobile app development. We are proud to own a building right around the corner from the convention center. And we'll be hosting sprints at that location starting tomorrow. So if you are still around, please do stop in and say hello. Um during Tobias' time, uh starting with Django, he 's a big advocate for the contrib. messages. He's also a member of the Django ops team. And he spends a lot of his time in our offices
Speaker 1: working, excuse me, most of his full time designing cloud-based and on-prem infrastructure for Cactus clients. And also he helps keep my crazy ideas grounded. In his spare time, he bakes amazing bread and grows vegetables. And if you're ever um lucky enough uh to try some of his bread, You won't regret it. At this point, I'm going to turn over the floor to Kenya Phelps to get to the good stuff talking about celery.
Speaker 2: All right. Um, this is my first tech conference, and this is my first talk. So please be gentle, okay? All right. So a couple of questions. Don't we all like polls? How many of you are familiar with celery? Good, good. How many of you are are actively using celery on a project? How many of you are using celery canvas? And last question, how many of you are here for more celery cooking ideas? Excellent. So I saw quite a few hands go up and I'm gonna s
Speaker 2: I'm a uh a software developer, two and a half years experience. And I just have a quick story about like you guys say I know you know what celery was it is. I didn't know what celery was when I first started. I had got to spend some time with a senior developer working on in some code base, a leg legacy code base. And we were working on a ticket, and I was just so excited that he had the um the time and the patience to help me through some problem solving. So, you know, as a junior developer, I come prepared and I'm writing my notes down furiously because he's talking and I'm like getting into it. Oh, I know what that means. I know what that means. Oh Don't know what that means. Let me write everything down.
Speaker 2: And then at uh sort of at the end, he he's like, yeah, and then you know, we're trying to I think the issue is a celery beat and whatever. I'm like, wait, did he say celery? And then he says something about a rabbit. And I'm like, hmm, okay, let me write these things down. They don't sound technical, but I'm gonna write it down. Okay, and so he gets done. And I'm like, okay, Joel, you said celery. Did you mean celery? Because I he was like, oh, I absolutely meant celery. And I'm like, well, I don't know what that is. And so We did a uh a a deep dive and no pun intended rat going down that rabbit hole on the rabbit too. We did that and I uh it it was actually amazing because I felt like this I understand It's easy to grasp this concept
Speaker 2: and I loved him for doing it. So I'm glad that you guys knew what salary was salary was, because I is, but I because I certainly didn't know. So I just wanted to share that little Little brawler with the thank you so much for indulging me. So let's hold on. Oh I can't see my yeah. My note thingies is not yeah, but I can okay, thank you. So our talk today, thank you, Ahmed. Our talk today will be broken up into two parts. First, getting started with celery.
Speaker 2: And in part one, we'll talk about what Celery is and how to add it to a Django project, how to choose a message broker, how to choose a task. scheduler and how to run tasks on a predefined schedule. The last thing we'll talk about is how to choose a result store and it won't be target. Okay. Okay. So then in part two, we're going to focus on a case study in which we can explore various ways of breaking down large tasks into smaller ones. That scale more easily, what we are calling good tasks. So we will also briefly summarize some of the celery patterns and anti-patterns that we learned. So, getting started with celery.
Speaker 2: What is celery, anyways? So I was supposed to insert a bad dad joke in there, but I don't have a bad dad joke for what is celery, but so insert your own. The Celery homepage says that it's a just distributed message processing system with a focus on real-time processing. That means in a nutshell, Celery is a means for executing tasks in the background, potentially on a large number of servers in parallel. Some of Celery's major features include support for running code that is triggered by but outside of the normal request response cycle in Django. It communicates with a message broker which is responsible for handling the work, and it can be configured to run tasks on a periodic schedule, so it can act as a cron
Speaker 2: job replacement. Because it can run code outside of the request response cycle, it is its own workers. It can help make sure requests are timely. And then it is reasonably good at distributed computing and high availability. So when it's configured with multiple workers on different servers. Celery remains immensely popular. We've written several blogs about it at Cactus over the years, and we're really pretty famous for our blogs, you know. real famous. This consistently remains one of the highest traffic pages on our site and sometimes by an order of Magnitude. Okay. There are other options for
Speaker 2: message processing and background tasks with Django, but we think that celery is a solid choice. And there's a good chance other Django developers like yourselves will have heard about or used it in other projects, which we already know since the poll, yeah, we know that you do. All right. So some specific use cases where celery is valuable. It's generating assets, for example, different image sizes. So after uploading a huge file, a large file. Notifying users, for example, when surveys open at a specific time during the day, updating the search index for a large content site from time to time. or running maintenance tasks, so um
Speaker 2: backups or cleaning up old files, database records that aren't deleted on their own. Adding celery to a Django project, pretty straightforward, relatively simple. How how's how How many times have we heard that, right? I just said that about my height. You need to install the package, okay? There is some boilerplate, so it should be, like I said, fairly simple. Or you can search celery first steps with Django and any of your favorite search engines. You need to import the app instance in your project's top level init pie file. Remember top level. I made the mistake of not see and we don't want to go down that road, right?
Speaker 2: So that's so that it's configured every time the Django runs. And before we can use Celery to run any tasks, we also need to configure it to use a message broker. The message broker is a task queue that allows work the workers to consume the task. These workers might all be running on the same machine or potentially on different servers. The two most common message brokers used with celery are RabbitMQ and Redis. Tobias recommends using RabbitMQ. I technically don't have a preference now and I haven't data, but Tobias does. Except for smaller projects where you already have Redis deployed and want to avoid adding a service dependency.
Speaker 2: RabbitMQ is purpose built to be a message broker and it excels at that task per tobias. It's highly scalable and it's very efficient. It comes with a built-in admin dashboard like Django itself and it provides visibility on different cues. So if you decide to use Redis as a broker, some of the caveats to be aware, and these are also documented in the Celery documentation. Tasks that are not acknowledged within the visibility timeout will be re-delivered to another worker. Example, executed again. The default visibility timeout for Redis is one hour. While you shouldn't have a task that is running this long
Speaker 2: If you do, you'll likely be quite surprised to see it getting executed again and again and again. If you forget to increase the visibility timeout, that's going to be the output. Now, while you might be tempted to simply increase the visibility timeout, its purpose to is to protect against task loss in the event a worker crashes. Or in case of a power loss. So increasing visibility timeout has the downside of prolonging the time before interrupted tasks will be rerun. So it's really not uh a a solution. You're just prolonging the terror because it's gonna be terrifying when you have this go on, run, run, run.
Speaker 2: So other more nuanced caveats include key eviction and group result ordering. If you would like to use Redis as a broker, we recommend familiarizing yourself closely. With these pages in the docs. Really pay close attention to those pages. So wealth of knowledge. In summary, we recommend using RabbitMQ whenever possible. possibly falling back to Redis if it's already deployed and the caveats are acceptable acceptable on that particular project. Once you've selected a broker, you can configure it into project settings. py file with the Celery Broker URL setting. As of Celery 5. 3, we also recommend setting the Celery Broker Connection Retry
Speaker 2: on Startup. Now that we have a celery configured, we can start writing a task. A Celery task and Python function that is wrapped with the task or share task decorators. Here we have a task defined called batch password reset emails, and it takes a list of user primary keys to send password reset emails for. You might cue this task, for example, when an admin user selects user accounts to activate in a web interface, but you don't want to wait to send all of the emails at once before returning a response back to the user. Here in our staff activate view, we might have a form that figures out which users the admin select
Speaker 2: to deliver invites to. Then we obtain those primary keys and pass them into the delay function on our salary on our salary. Task bat task batch password reset email. Say that. really fast. Then you can use the messages framework to save a status message for the user and redi redirect back to the staff list page. When delay is called, the web process will deliver this task with its arguments, the message broker and then return immediately so the status message can be queued and the HTTP response returned to the user. In the salary worker process, almost immediately if there's an open worker, the task will be picked up and the emails will be sent out.
Speaker 2: Voila. Here is a workflow diagram of the different processes and messages. In the Django laying in orange, the HTTP request is received. Then the Django view calls the task delay function and then returns the HTTP response. Rabbit MQ Intel receives the AMQP and then passes it off to a celery worker running in another process. The celery worker in blue then sends the necessary emails or whatever the workload might be and records the results if needed. Queuing tasks manually covers a large number of use cases, but sometimes it's helpful to be able to queue on schedule as well.
Speaker 2: Celery Beat is the name of a feature in celery that is responsible for scheduling tasks at predetermined times, predefined times. It does doesn't run tasks itself. Instead, it adds tasks to the queue so the workers can pick them up and execute those. It's worth noting that celery beet is optional. So don't go willy-nilly adding stuff into stuff. Don't go adding it if you don't need it. If you don't need it to run tasks at specific times of the day, you don't need celery beet. We don't want senior developers poo-pooing, you know, getting on you, okay? Celery beat is built in celery itself. But like a lot of packages it comes with c a couple of pain points. One of one of those is that it requires state
Speaker 2: By default, this is kept in a flat file, which can be annoying and to maintain in our modern container-based deployments. The second is that you should only ever have one celery beat process running at a time. Otherwise you might end up with two copies of the same task getting executed or at that scheduled. I'm sorry. There is a separate app called Django Celery Beat, recommended in the Celery Docs, that allows customizing the Celery Beat configuration specifically for Django projects. It includes an optional task scheduler that allows you to use the Django ORM to track state. Django ORM, yes. And it also comes with a Django admin interface, so you can monitor celery beat tasks with it via the admin.
Speaker 2: And we love the admin. So the installation steps for Django Celery Beat are similar and also straightforward. You install the package via pip, and that those are the installation instructions. While Django Celery Beat supports configuring tasks schedules in the database, Tobias still recommends storing those in your setting files. I also recommend it. I don't know if so he does too, I do too, whenever possible with the Celery Beat Schedule settings. Please note that some older versions of celery supported a setting without the first underscore in this name. So be on the lookout for that if you're copying, pasting code from the internet, please.
Speaker 2: To define a schedule task, you give it a name. You point the task or function to be called and define a schedule, either here in a period of seconds or in Quran syntax. Once the schedule has been defined and you have a celery worker and beat process running, you will start to see the output from this debug task in the Celery Worker logs. If you ever need to return values from the celery task, it's also worth configuring a result store. The results store is responsible for keeping track of the return values from task functions, so they can be retrieved async either by other task or rev request. For simple use cases, a better pattern is used to update the database and to end the task.
Speaker 2: But for more complex workflows, you might need a results store to keep track of the results as you go. There are many potential results stores you could use. Over 17 the last time Tobias check, but I checked too this morning and it was still 17. But since we're using Django, we recommend starting with the Django ORM. Redis works quite well as a results store. So while we don't recommend using Redis as a broker, it can absolutely be used to store the task results. So and and and just to P like like talk about that a little bit. That was confusing to me as um like as a beginner developer
Speaker 2: because it's recommended for this particular technology is recommended for one thing and not the other. So you might want to dive a little deeper in the docs for that because that that took a a couple of a couple of go-arounds for me to understand that. just to point that out. And not to say that it would for you, but for me it did. Like Django Celery Beat, there is a Django Celery Results reusable app you can add that includes a custom result. and cash back in for salary. It also uses the Django ORM and comes with Django admin interface for observability. So the installation instructions are the instru the installation instructions are similar and pretty straightforward.
Speaker 2: You need to install the package. add the app to the installed apps and configure the two settings in your settings file. We didn't quite okay so full disclosure, this is something that Tobias is saying, I love this about him. He is being completely um just you know vulnerable and full disclosure he's like we don't quite understand how these short strings work Django DB and Django Cash the first time we saw them but we believe that they are mapped behind the scenes So in the project setup dot file to entry points or classes within Django Celery Results Source Code. So don't come for us. We already said we're not we're not sure how everything is working under the hood, but we know that it works.
Speaker 2: Now that we have a results store configured, we can we'll be able to fetch those results of our tasks that are finished executing. And that that's my part. Now for the good stuff. Good task. I want to introduce Ahmed, and he's going to finish it off.
Speaker 3: My name is Ahmed Pitturi. I come all the way from Libya. I had to take four different flights just to be here. I'm really thankful for Cactus. They sent me an invitation. I was really happy. And uh thank them. I want to thank them for making this happen. So I'm here. This is my first Django Con and my first talk too. So we'll see how it goes. All right. So um picking the right size task One of the hardest parts about designing good tasks is fine-tuning the amount of work done in each task to best balance readability, scalability, and maintainability. Sometimes
Speaker 3: low-running tasks are easier to write. But it might be harder to debug when something goes wrong. On the flip side, really short tasks, for example, less than one second, might come with an undue amount of message overhead. If scheduled in the high volume. Ultimately, the goal is to design around units of work that are small but not too small and that can be executed simultaneously. Another good criter criterion to use for designing task size is how graceful shutdown will be supported by your application. So for example uh for Kubernetes, The default timeout of graceful shutdown is 30 seconds and for supervisor it's 10 seconds.
Speaker 3: These can certainly be increased, but I would not increase them indefinitely. In any event, if you if your task does not wrap up work, either by finishing what it's set out to do or by somehow saving its place and queuing another task to pick up where it's left off Later, it will be terminated and might leave its work in an inconsistent state. To help supporting To help supporting uh building the right size tasks, Silary supports a number of primitives of grouping and chaining tasks. These can be This can be combined to support arbitrary complex workflows. Before tasks can be chained or grouped together, however, we need a way to bundle tasks
Speaker 3: function. with its arguments so Celery knows how it's how to call it. So it calls it the signature. And it can be created with the S function. Which takes uh as its arguments, uh the arguments that you uh want to pass to it is uh the task when it's When it's eventually executed by a worker. Typically, this would consist of a JSON representation of the function that needs to be called along with its arguments. Once you have uh once you have a task signature, you can pass them into the chain function to execute them sequentially. The return value for previous tasks will be passed to the ta to the next task in the chain, and so forth
Speaker 3: Similarly, the group primitive can be used to queue tasks all at once and run them in parallel. In this case, all the return values from a group tasks will be collected together in a list A task group can be part of a larger chain of tasks. So you can, for example, chain the list of results from a group into another task to aggregate and report on the results. Finally, there is also a star map function, which can be used to pass a list of arguments into a task, which will be executed sequentially for each set of arguments. We'll take more about we'll talk more about this in the specific example shortly. There are other parameters you can explore in the Sili
Speaker 3: Docs linked here. As a general rule, try to avoid passing state from one task to another through task results. Instead, just use just as you would pass an uh an object primary key into a task instead of the model object itself. Try to minimize the size of importance of arguments passed to task in the chain. As Mr. Flavio went into more details on this in his salary talk yesterday, and Tobias recommends checking out the video if you didn't have a chance to see it yesterday. To help understand how these Celery primitives can work together ,
Speaker 3: it's helpful to explore an example together in code. As some brief context before we do that, Cactus has been working as a software development team for the High National Election Commission in Libya since 2013. Uh and and this is uh actually the same time as I have uh met uh Tobias too. It's like ten years now. Libyan citizens registered to vote by SMS and Cactus team built their voter registration system using Python, Django, and Sillary. One function of this application is the generation of water rolls. We think it's a good case study that we could use to explore how cellary primitives work.
Speaker 3: If you'd like to follow along in the companion repo, here is a link on the GitHub repo. It's uh cact. us forward slash Silery. While this code is based on the idea of voter rolls, I should mention uh it was written from the ground up. For the purpose of this talk, so it's not actually used for to run any elections. Uh if you would like to see uh running code for elections, there's um an open source Project that we have worked with Cactus for. It's called SmartSelect. Smart Elect.
Speaker 3: So the general goal of the code is to create lists of of all eligible voters for distribution on paper at the time. To all the 2,000 polling centers in Libya, 2000 plus. In a nutshell, the code assigns voters within each polling center to groups of 500 to 600 voters. Which is called a station. So this is the largest number of people that could reasonably be handled by a single desk or a table in a polling center. For example, uh there might be uh anywhere between two to thirty stations in a center, uh depending on the population density around that area. After splitting voters into stations,
Speaker 3: the code generates PDFs , one for each station, with the list of eligible voters in that station. So there are a lot of different ways we could write uh this code. Uh and Tobias wrote up several of them and the example repo. If you are looking for the code, it can be found in the voter roll forward slash roll rollgen. py. So let's talk let's talk through this each of uh these options now. In option one, all the code is executed in a single ciliary task. We it uh we do iteration through all the polling centers in the database. For each center, we split the voters in the center into reasonable size
Speaker 3: stations. Then we write a voter list to a PDF file. Along the way, we keep track of the total number of pages written And print summary at the end. This is a short and easy to understand, but it's not ideal for a couple of reasons. All the work done here is sequentially. The work can be distributed to multiple multiple workers in your config configurations. It takes longer than 30 seconds, even though we have given it a sample uh file of 10,000 fake fake users. Um a natural w I'm sorry, fake voters. A natural way to break up this uh
Speaker 3: work might be go um to go through each each center one by one and writing all the pdfs files for that center in one task. So this is option two here in sample code. Um what it does is splits the voters into stations only for that specific center. And then it writes all the PDFs files for each of that state stations , for each of the stations. And at the end, it returns the total number of pages for all the stations within that center. Uh we continue on on option two, and this is the second task. Uh we need a task or function to kick off all the individual centers levels tasks. We can do that with this
Speaker 3: primitive called group. So group task uh takes an uh terrible of of task signatures, what you get from the S function which can in in turn be called as a group with delay, the delay function, to add a task for each center to the queue all at once. The work will be processed by the workers up to the total number of workers in parallel. You might be tempted than to wait for the task to complete and get drop the results. Do not do this. This is not good practice. For one, it is not possible or even advisable to wait on a task within another task.
Speaker 3: And for number two, you will hold the process calling join, doing nothing. The join function here will be doing nothing. uh until all the other tasks are complete. So instead this is uh option number three. We can um Okay. Option number three, we can chain or pipe the results of the central level tasks into a final chord task That has the job of summing the page counts from the other tasks and reporting them. This task is virtually instant. It cues all the tasks with with group as before and tells Celery to kick off.
Speaker 3: The task sum pages task once the tasks in the group are complete. According to the Celery docs, the synchronization step is expensive. So this model should be used. sparingly, but sometimes it is unavoidable, so it's still better than the alternative of waiting synchronously for all the tasks to complete. Note you can uh you also you may also see the chord function here. It's uh it's doing the same thing as the pipe syntax, but we use the pipe syntax to be more explicit, just to be more explicit. And easier to understand. And this is the following function for uh uh task number three. Uh option number three.
Speaker 3: Here's um so it takes the list of the results from all the tasks in the previous groups, then calculates and prints the sum of those page counts. So option four. As a brief tangent, instead of splitting voters into stations one by one, we could do all the parallelized work up front. This turns out to be a bad idea. Why? Because work is relatively fast anyways. It even might be slower, highly parallelized, with all the ciliary message messaging overhead. And we don't want to pummel our database with all the uh with all too many queries at once. So this is a uh this is a database intensive category of work that is often best done synchronously.
Speaker 3: If the code is too slow, time is probably better spent on optimizing the underlying database queries rather than simply trying to do more things at once. In this case, trying to execute all the queries in parallel would almost certainly take more than more time than rather than less time. So in option five, we modify the um the task to accept uh two parameters, uh center ID and station ID. And the work it contains is limiting is limited to generating the list only for that single station in the polling center. In this demo, uh the running time for this task is uh between seven and eight seconds, which is an ideal for Sillary tasks because it simplifies graceful shutdown.
Speaker 3: Now uh you might be pursuing the Sillary Docs and see promising looking helper function called star map. If you're familiar with the Python multiprocessing uh library or the map or the other map produce type of workflows, you might think that Celery is a very important thing to do. Being the distributed systems that it is, would execute tasks in parallel when called with SARMAP, but this is actually not the case. Uh here, and noted in the docs, SARMAP queue is only a single task And runs each subtask sequentially within that task instead of in parallel like group, like the group function. So the performance of this uh version is is is really the same as what we did in option number one, which we did all in one task.
Speaker 3: But the code is harder to read. Uh Tobias wasn't sure what the use case is for star map. So if anyone happens to know, we would love to hear about it during the Q<unk>A. Um option 5B, uh our last and final option, we return to using group. To queue and one task for each center uh center ID and center uh station ID pair. Then we pipe the results into the task sum pages core task. This option gets the database work out of the way quickly. And then paralyzes the CPU intensive work. Again, anecdotally, this particular unit of work takes around seven to eight minutes,
Speaker 3: seven to eight seconds, sorry Which is a great length of time since it's less than the default graceful uh periods of Kubernetes and uh supervisor shutdowns. In other words, uh when a silly worker is told to shut down, for example, when new code is deployed, we can be confident that it will be that it will finish the task it's currently running. And then exit before consuming any any more tasks. If you'd like to try uh out the code, please check out the uh companion repo for this for this talk. And you can use the included management commands to generate some test data and run the various tasks. Watching the salary worker outputs
Speaker 3: will provide uh some insight into the into when and how tasks are run in parallel, when they're not, and how long the various options take to complete. Uh to summarize, some patterns we like to recommend for Sillary on Django projects are to first use RabbitMQ as a broker. Whenever possible, when RIM objects are needed in a task, pass the primary keys and fetch the objects in the task. Use Django Cillary Results and Django Auxiliary Beat database back backends for visibility and ease of use if you need those features. Divide the work into right-sized tasks, a good target is for all tasks to be complete within a graceful shutdown period.
Speaker 3: The don'ts here are just don't make a task that waits synchrously for another for another task and don't write a task that takes longer than your graceful shutdown period to complete. And don't massively paralyze database operations or other work that derives no benefits from running in parallel. If you'd like to learn more, Cactus has several popular posts on the Cactus blog about celery. As cactus. us forward slash blog dash celery. Last but certainly not least, yesterday at DjangoCom, Mr. Flavio gave a pre-recorded talk on mixing reliability with celery for delicious sync tasks.
Speaker 3: His talk focuses on some more advanced advanced use cases for Sillary and suggests additional options for tracking state in more complex workflows So we recommend watching when you watch video after the conference if you did not happen to catch it yesterday. Thank you for coming to our talk about Sillery. And we are happy to engage in conversations about this talk, and you are also welcome to reach out directly to Mr. Tobias via LinkedIn or faster down. Thank you.
Celery runs tasks in the background, outside Django’s normal request-response cycle, using workers and a message broker. It can also run tasks periodically and distribute work across multiple servers.
Discussed at 7:12Install Celery, create the required application boilerplate, and import the Celery app instance from the project’s top-level `__init__.py` so it is configured whenever Django runs. You must also configure a message broker before queuing tasks.
Discussed at 9:36The talk recommends RabbitMQ whenever possible because it is purpose-built, scalable, efficient, and provides useful queue visibility. Redis can be appropriate for smaller projects where it is already deployed, but its visibility timeout and other broker-specific caveats must be understood.
Discussed at 10:21Define the work as a Celery task and call its `.delay()` method from the view, passing the necessary arguments. Django sends the task to the broker and returns the HTTP response immediately, while a Celery worker performs the work separately.
Discussed at 13:27Celery Beat schedules tasks at predetermined times by placing them on the queue; it does not execute the tasks itself. Define a task name, target function, and schedule using an interval or crontab expression, then run both a worker and a Beat process.
Discussed at 15:45Django Celery Beat provides a Django-oriented scheduler that can store state through the ORM and expose scheduled tasks in the admin. Although it supports database-configured schedules, the presenters recommend keeping schedules in settings with `CELERY_BEAT_SCHEDULE` when possible, and running only one Beat process.
Discussed at 16:32Use a result store when later tasks or requests need to retrieve task return values, particularly in complex workflows. For simpler cases, the presenters recommend updating the database directly; Django ORM-backed results and Redis are both possible result-store options.
Discussed at 18:09Tasks should be small enough to run in parallel and finish within the application’s graceful-shutdown period, but not so small that message overhead dominates. Tasks that cannot finish or save their place before shutdown may be terminated with inconsistent state.
Discussed at 21:37A signature bundles a task with its arguments. Chains run tasks sequentially and pass each result forward, groups queue tasks in parallel and collect their results, and a chord lets a final callback process the completed group results; `s()` is used to create signatures.
Discussed at 23:57Prefer RabbitMQ when possible, pass model primary keys rather than model objects, use appropriately sized tasks, and use Django-backed Beat or result apps when visibility is needed. Avoid waiting synchronously for one task inside another, making tasks longer than the graceful-shutdown period, or massively parallelizing database-heavy work that gains no benefit from parallel execution.
Discussed at 37:20Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026