Why large Django projects need a data (prefetching) layer with Flávio Juvenal
Published November 3, 2022
This video features Flávio Juvenal at DjangoCon US 2023 in Durham, North Carolina, USA.
As the most popular and mature solution for asynchronous task queues in Python’s ecosystem, Celery is an essential tool for Django projects. But running Celery tasks with high reliability is a challenge. The settings are tricky, tasks can be lost in multiple ways, task code has opaque limitations, proper monitoring isn’t trivial, and more. In this talk, we’ll share what we learned to be necessary for running Celery reliably after years of running it in production.
This talk was presented at: https://2023.djangocon.us/talks/mixing-reliability-with-celery-for-delicious-async-tasks/
LINKS:
Follow Flávio Juvenal 👇
On Twitter: https://twitter.com/flaviojuvenal
Website: https://www.vinta.com.br
Follow DjangCon US 👇
https://fosstodon.org/@djangocon
https://twitter.com/djangocon
Follow DEFNA 👇
https://www.defna.org/
Video production by the presenter and DjangoCon US 2023 volunteers.
Celery is a distributed system with several ways to lose tasks: publishing can fail, large arguments can overwhelm workers or RabbitMQ, database changes can commit without their follow-up task being queued, and workers can disappear before acknowledging work. Use publisher confirms when needed, pass references to large data instead of the data itself, and use periodic database reconciliation to recover tasks such as verification emails. For workflows involving fan-out, long-running jobs, or data pipelines, Flávio recommends dedicated tools such as Temporal, Prefect, Airflow, Dagster, or Mage rather than Celery Canvas. He also recommends queuing tasks with Django’s `transaction.on_commit()`, enabling late acknowledgements and appropriate worker-loss rejection settings, making tasks idempotent and atomic where possible, handling exceptions and retries explicitly, and planning deployments around queued tasks, ETA tasks, broker redelivery, visibility timeouts, monitoring, and task expiration.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Hello folks, I'm super happy to be with you today on DjangoCon. And my name is Flavio Juvenal. I'm a partner at Vinta, and I'll talk with you about mixing reliability with celery for delicious async tasks. I work at Finta and we are a consultancy shop, web development, design and development shop from Brazil, and we've been building Django products for the last 10 years, and we've used a lot of salary And we've dealt with many salary incidents, salary issues, and I want to share with you what you can learn from our experience to really build reliable salary projects You can find the slides for this talk on this link and this QR code.
Be sure to check that because there are lots of links here and lots of references. I'm assuming you already use salary for this talk you and you need reliability on your salary uh project. If you are a salary beginner and you need to learn salary from scratch, you can check Tobia 's stalk. He'll do a deep dive into salary and we'll reintroduce all the concepts. I won't do this here, but you can check his talk. In any case, if you are a beginner, please stay. I'm sure you learned something from salary task loss risks and how the salary infrastructure works in general. Okay, so salary is our async tasks infrastructure and we want to avoid loss tasks.
We have to remember that salary is a distributed system. We have the web process, we have the broker, and we have the worker processes. And there are task loss risks when those things communicate with each other to really execute, to really delay queue and execute tasks. There can be task laws between the web process and the broker. For example, a test can never be queued due to a broken connection. There is a connection between the web process and the broker, and this connection can fail, and if it fails, the task will be lost. You have to remember that when you do task
delay, when you do my task delay , this operation can throw an operational error if the broker is down or unreachable. So the good news is that celery will automatically retry task delays. It will retry three times, waiting uh 0. 2 seconds between each try. But the task commit on the broker side can fail too. If the connection, even if the connection is fine, so even if everything is okay with the connection, the committing of the task on the broker side can fail You can configure that on RebTMQ. You can use the configuration confirm publish. So the delay will take a bit time more than the default. But then you'll be sure that after the task delay is done, the task is really on the broker side.
So this is a default that is not clear that you have to perhaps change on salary. Gigantic tasks can also bring down the worker. For example, if you pass large Python dicts as arguments We have a real case at Finta where we had 100 gigabytes of data inside rept MQ directories, data directories. And that happened because we were passing large amounts of XML data as task uh arguments, as task parameters. And instead, we should be passing data references, not data values. So we could use a blob storage like S3 and save the data there and then send the URLs as parameters, as arguments for the tasks.
So the tests can could download that and do the work. Or we could just use databases if we're dealing with relational data. But in any case. uh consider to use data references, not data values directly when passing data to to tasks It's worth remembering as well that lost tests break user flows. So there's the classic I didn't get the account verification email problem, right? The user just signed up and didn't get the account verification to finish. the account verification email to finish the sign up. So here we have a sign up view from Django and We are doing the form valid. We are doing the form validation. After the form validation, we already have the response and we already saved the user.
So on commit, I'll explain that better later. We are delaying a task to send the activation email. We are passing the user reference, the user PK, everything's alright here. But as we mentioned, the delay can fail, so this task can be lost. And there can happen a mismatch between the task state and the business logic. The database has committed the change, the user sign up is done, the user is saved on the database. But if the broker is down or if the connection with the broker fails, the verification email task is lost. So there's a failure of communication between the web process and the broker. Sometimes some applications just pass the button to the end
user, click here to resend the email, who never did that, and just click there. But is there a way for us to ensure an auto-recovery? So if the broker is down, eventually the user can get the email. Is there a way we can build something to fix that for us automatically? There is a robust solution that we can implement. We can use the database as the source of truth, and we can source tasks from the database. To implement something like this, uh we need pooling, we need uh some some way of checking the database from time to time. See if there are tasks that should run because the database state has changed. For example, a user is waiting for the account verification email.
And if there is, we can queue tasks on the broker. So if the task infrastructure fails, you have to manually do consolidation work anyway, right? You have to write a script, send emails, things like that. Instead, you could have periodic tasks with salary beat that queue other tasks that should run. So you can have this infrastructure where there is a task that will periodically check the database and queue other tasks. Here's the example. Here's the example of the verification email. We use salary beat and we have a task, ensure verification emails, that iterates over users. That have no verification email sent, verification email sent false. And for all those users, we are delaying the task sent
verification email task. We can keep the task, send activation email on the signup field, there's no problem with that. And we can rest assured that frequent periodic tasks, like every five minutes, will eventually queue the task. Send activation email task if anything broke on the sign up view. So this way we have the best of both words. We have we still have the fast behavior during sign up. Of sending the email if everything is right, but we also have consolidation work to fix if that breaks. Okay, uh, but imagine you want to build complex workflows, right? So you have a process day orders. workflow where you where you want to fan out uh many process order calls in parallel to process many orders
like even millions of them And then after all of those are done, you want to send them to logistics. You want to call a send-to-shipping task. So this task would have to wait all the orders to be done. You can also implement that with periodic tasks But it's difficult to orchestrate a complex and long workflows or data-heavy pipelines with salary and a database. Alternatively, you could use salary Canva workflows, but I wouldn't recommend that. We've used that a lot from the last 10 years at Finta, and we have many, many, many different issues. Here's a link That you can check later about uh that's a blog post from another company that shares uh their findings uh about salary
canvas workflows issues. You can check GitHub. Uh salary GitHub, and you'll find many issues related to salary canvas uh constructs like cards and stuff like that. So, what I would recommend is better to not use salary for workflow focused uh Problems. Instead, you should use workflow-focused tools. You have better monitoring, you have operability, you have error handling, you can run the tasks anywhere, like pods, containers, lambdas. It will be more integrated with the infrastructure. It won't be trivial to integrate with Django. It will be more like heavyweight. But I I think if you have complex workflows, that's what you really need.
The tools that I recommend for generic workflows are Prefect, Temporal, and Airflow. And for data-heavy pipelines, machine learning pipelines, data engineering pipelines, I would recommend you to check Daxter and Mage. Okay, so now let's talk about task loss between the broker and the worker. Remember the on commit that I've shown? Now let's talk about this. Why we need that. If you have atomic requests true in Django and many projects have that, uh Which means there is a transaction transaction wrapping the whole view. If you do something like this, which is the naive way of delaying a as Task send activation email, uh
you just like check the form valid, you get the user PK, and then you just delay the task. What can happen is that between the delaying of the task and the response. The test can execute in parallel to your view. And if that happens, you get this error. Does not exist, the user doesn't exist yet. Because the the view didn't finish. So the transaction wasn't committed, but the task ran. So the user wasn't on the database yet. So you can lose tasks uh because of that and you can have task errors uh in fact So in the correct fashion is to only run this task after the database commits. So you use transaction on
commit. Okay, so that's fine, that works, but you have also to consider that workers can shut down unexpectedly, and you have to decide a way to properly remove tasks from the broker. When it's safe, we have to figure out when it's safe to remove a task, to drop a task from the broker. Imagine you have the broker with the eat cake task. And the broker sends that task to the worker process. The worker process gets that task from the broker and says says that it can handle it. Okay, so the the this task is not anymore on the broker, it's only on the worker. But this task, the eat cake task, can be lost if the worker fails.
If there is not enough RAM on the worker, this task will just be lost. The task is not on the broker side anymore. And it's not on the worker side because the worker process has died due to out-of-memory errors. Workers can shut down unexpectedly. Sometimes it's so easy to forget that when we are using salary. Nothing is 100% reliable, even if you take care to not have like memory errors, stuff like that, even cloud containers will fail from time to time. Any kind of bad test code can also kill a worker, any time of bad arguments, uh memory errors, stuff like that. You can have errors that you don't have time, any time to capture any exceptions, because, for example, the the kernel might just kill
the worker process if it's uh wasting too much memory. And finally, deploys can interrupt tasks. Depending on the way you implement the deployment process, the deployment script, it can just like kill the workers and it it will interrupt the running tasks. Okay, so when the eat cake task should be dropped from the broker, if it's right after the worker gets it, we have the situation that we just described, and that's the default. That's the default behavior of salary. The broker loses, drops the task right after the worker gets it. And that that means task acts late false. That's the default. We can change that to only drop the task from the broker
after the worker finishes So if the worker fails, the broker will have to redeliver the task. The task will remain on the broker until the worker finishes, and if the worker fails, the task will be redelivered. That means task accelerate true. You can activate that behavior and you should activate that behavior. So you should activate that and that means that the broker will wait for the worker's acknowledgement to drop the task. In the case of the worker process not having an off-run and like losing the process, uh the process being killed by the operating system. The task will remain on the broker side and will it will be redelivered. Okay, but if tasks can be automatically redelivered, you need to ensure your task code is safely retryable.
They need to be indepotent. Your code need to be needs to be indepotent. It shouldn't matter that if the task is executed multiple times with the same arguments, there should be no side effects. That's what eDEPON test means. So you have to use uh RM tools, RM methods like get or create, update or create. You have to always be sure that you are Like not redoing work unnecessarily, not doing side effects unnecessarily if you want to have your tasks to be depotent. And also for making tasks depotent, you may need atomicity. So you perhaps should wrap tasks in transaction atomic, so in any case they are interrupted even due to a deploy or something like that. They will automatically roll back any db changes and eventually the
broker will re-deliver the tasks and it will be re-executed. But beware, you can have other side effects that are impossible to roll back, such as sending an email. So in this case you can be smart and try to couple. those uh impossible to rollback side effects with the db state change and rollback if uh fail. So the the database will tell you that you still have to send that email similar to what we showed before with the activation email. And finally, you can just keep tasks short and fast to ensure they are uh easy and reliably retryable. Okay, but how does red delivery work? If the task stays on the broker, how it knows it needs to re-deliver the task?
On RebTMQ, this is kind of automatic. You don't have to do anything about it. You just have to ensure that you have the basically the default RabTMQ configuration. That means we have a you have a TCP connection and heart beats between the broker and the worker. So if this connection is closed, uh RapTMEQD, your broker, will know that the worker has died and it needs to re-deliver the tasks that this worker had got. On the case of other brokers like SQS and Redis, you don't have this connection and you have something called visibility timeout. That means that if the task has not been acknowledged after n seconds, after a certain amount of seconds, it will be automatically re-delivered.
The default is 30 minutes for SQS and 1 hour for Redis. You can configure that with the visibility timeout on the broken transport options. There is a trade-off with the visibility timeout. After the visibility timeout, the task is automatically redelivered, even if the worker is still doing the tasks. Okay, so this is a problem because no task can take longer than the visibility timeout to finish, right? Otherwise it will just be redelivered and the work will be repeated. If you increase the visibility timeout to give more time to your tasks, it will also take more time for tasks to be redelivered if worker fails. So you have a trade-off here, right? You are trading off fast
re-delivery versus long task support. So visibility timeout is not great in general. What's the solution for that? You can just use ReptMQ as a broker. It doesn't need a visibility timeout as we discussed it. Or you should shouldn't use salary for workflows of long tasks, like we've discussed it before as well. And you could could also make task acceler false, which is the default. It also means no visibility timeout, but it's not recommended. uh because you you avoid task re -delivery but you you have task laws anyway for the reasons of test uh accelate false It's important to know that tasks are intentionally dropped on worker failure.
That's a very surprising default, but it is what it is on salary. By default, task reject worker loss setting is false. That means salary will intentionally drop from the broker tasks that cause segmentation faults, out-of-memory errors, or sick kills. So if your task is running and it receives a CQ, a lot of memory error, this task is automatically dropped from the broker. It won't be redelivered. If you change this to true, it will activate red delivery for all tasks, including the ones that can cause a worker process to die. So what Celery is trying to achieve here is to not have tasks that cause the worker to die to be redelivered, right? You can have a middle
ground if you have tasks that sometimes fail, but sometimes they work. You can on those tasks you can change the task rejectional worker log setting and you can keep the global default as it was before false. It's important to remember that handling exceptions is also necessary. Task accelerate true only solves unexpected shutdowns like deploys. If a task raises an exception, it's still acknowledged. So if you have inside a task something like that, a cake get object , cake objects get fantasy cake and fantasy cake doesn't exist, it throws cake does not exist exception. The task is gone. The task won't be re-delivered.
Because of a default on salary, that's actually a good default. That means a task acts on failure or timeout is true. You could change that, but you shouldn't. Ideally, you should handle all possible exceptions in tasks. You should treat uh task exceptions, you should handle task exceptions with the same diligence that you handle 500 errors on Django views. If you if you want, uh if you need, if you have like intermittent errors, if you depend on external APIs, things like that. You should use retries and there are many retries functionality. There is a great retry functionality in salary. Just be careful with uh long retries because there is a catch uh that I'll mention later. Let's talk about deploys versus queue
accelerary tasks. When there are tasks in Q, be careful with deployments. If you change the task signature, the function parameters, the new worker processes after deployment won't be able to deal with was already queued with the older parameters So the parameters have changed. So the task has changed, but the new workers won't be able to deal with the task. Imagine you had you had a task, task a cake with cake ID, but you changed the dessert ID and it's structured differently. the the new workers won't be able to deal with the old tasks that were queued and are being redelivered. So be careful with that. There we'll be very cryptic. And you need to ensure that the queues are empty uh if you are doing uh
if you are changing task signatures. Be careful as well when you are upgrading salary because the internal task payloads can change. The arrows will be very cryptic as well. And ETA tasks are also a problem here because they live on the worker memory. Yes, that's very surprising. I won't have much time to talk about that, but in general, you don't want ETA countdown tasks Because they live on worker memory waiting for the time to run. And they be equivalent of Q tasks and the like task signature can change. the payload can change and they are very complex to deal with. So in general you should avoid any ETA or countdown tasks that like are more than a few seconds.
About running salary tests when you are doing deploys, be careful with that. Because you have to decide when you what you are going to do about the old salary workers if you are like refreshing the seva. Are you doing a SIG term or are you doing a SIGQ with the with the worker processes? If you are doing a SIGQ, the task reject on worker love behavior that we talked about might bite you because salary will think that you want to interrupt the task and never retry and never re deliver it again. So check what happened what's happening with your old processes, double check if the rear delivery is working, double check if these processes are really gone from your machine after you do the deploys. There are other reliability concerns that I won't have much time to talk about, but please check those links, they are really important.
The ATA tasks that I've quickly mentioned. uh there they shouldn't be on the far future they should be only a few seconds after the same is valid for retries because retries are ETA tasks behind the scenes You can root priority tasks to dedicated queues and uh workers. That's very important, like to have different uh task latencies and stuff like that. You can prevent clogged queues with task time limits. If you have like badly behaved requests that take a long time, you should have timeouts and time limits for those sort of things. You can expire tasks that shouldn't run after a daytime. It's very frustrating for a use user to get a notification saying that something will happen in two minutes, but so that something happened like a day ago
So you should expire those types of tasks. You should set up alerts, monitoring, tracing, observability on salary. Please check those tools You can set up the letter cues as well if you want. Check those links. There's lots of information there. I've prepared a gist. with all the recommended salary settings for reliability that I've discussed on this talk and other uh other settings as well, please check it. Uh it it will surely be awesome for you. And It's all documented there. Please comment there as well if you have other settings that you are using for increased reliability on salary And that's it. Please feel free to reach me if you have any questions. Here's my link tree. Here's my Twitter. And if you need salary consultancy, please feel free to reach Vinta.
Please feel free to reach me directly or reach Vinta we are so we will be super help uh super happy to work with you uh with your salary project thank you very much
Celery retries a failed delay operation automatically, but RabbitMQ can also fail to commit a publish. Enabling RabbitMQ’s `confirm_publish` makes the producer wait for confirmation that the task is on the broker.
Discussed at 2:39Avoid putting large dictionaries, XML documents, or other data values directly in task arguments. Store the data in object storage such as S3 or in a database, and pass a reference such as a URL or primary key instead.
Discussed at 3:29Use the database as the source of truth and periodically scan it with Celery Beat for work that still needs to happen. For example, a periodic task can find users whose verification email has not been sent and enqueue the email task again.
Discussed at 5:49For complex, long-running, or data-heavy workflows, the speaker recommends dedicated workflow tools rather than Celery canvas workflows. He names Prefect, Temporal, and Airflow for generic workflows, and Dagster and Mage for data-heavy pipelines.
Discussed at 8:58Schedule the task with Django’s `transaction.on_commit()` instead of delaying it immediately. This prevents a worker from trying to read database changes that are still inside an uncommitted transaction.
Discussed at 10:32Enable late acknowledgements with `task_acks_late=True`, so the broker removes a task only after the worker finishes it. If the worker dies first, the task remains available for redelivery; worker-loss rejection settings may also need to be enabled for crashes such as out-of-memory kills.
Discussed at 13:37Make task code idempotent, so running it multiple times with the same arguments does not create unwanted side effects. Use atomic database transactions where appropriate, couple irreversible effects to recorded database state, and keep tasks short and fast.
Discussed at 14:29RabbitMQ can detect a dead worker through its connection and heartbeats and redeliver that worker’s unacknowledged tasks. Redis and SQS use a visibility timeout instead: an unacknowledged task is redelivered after the configured interval.
Discussed at 16:06A timeout that is too short can redeliver a task while it is still running, causing duplicate work. A longer timeout supports longer tasks but delays redelivery when a worker actually fails.
Discussed at 16:53With the usual defaults, a task that raises an exception is still acknowledged and is not automatically redelivered. Handle expected exceptions explicitly and use Celery’s retry mechanisms for intermittent failures such as external API errors.
Discussed at 19:09Be careful when changing task signatures or upgrading Celery, because queued tasks may contain payloads that new workers cannot understand. Keep signatures compatible or drain the queues, avoid long-lived ETA and countdown tasks, and verify that old workers shut down and redelivery behaves correctly during deployment.
Discussed at 20:41Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026