Lightning Talks Day 2
Published November 14, 2018
This video features Filipe Ximenes at DjangoCon US 2017 in Spokane, Washington, USA.
DjangoCon US 2017 - Tasks: you gotta know how to run 'em, you gotta know how to safe' em by Filipe Ximenes
Web developers often find themselves in situations where server processing takes longer than a user would accept. One very common situation is when sending emails. Although simple and relatively quick task, it requires the communication with an external service. In this situation, it’s not possible to foresee how long that service will take to answer. Not to mention the many unexpected situations that can arise, such as errors and bugs. The solution to this problem is to delegate long lasting tasks while responding quickly to the user. This is the point where we need async tasks. There are some tools available that can assist in this job. In this talk, you will learn about the concepts, caveats and best practices for when developing async tasks. For this, I will use Python’s most popular tool for the task: Celery.
Rundown:
Setting the context
The architecture:
Brokers
Workers
Use cases:
External calls
Long computations
Data caching
Tools available
Celery:
Callbacks
Canvas
Logging
Retrying
Monitoring
Tests and debugging
This talk was presented at: https://2017.djangocon.us/talks/tasks-you-gotta-know-how-to-run-em-you-gotta-know-how-to-safe-em/
LINKS:
Follow Filipe Ximenes 👇
On Twitter: https://twitter.com/xima
Official homepage: https://www.vinta.com.br/blog/author/filipeximenes/
Follow DjangCon US 👇
https://twitter.com/djangocon
Follow DEFNA 👇
https://twitter.com/defnado
https://www.defna.org/
Asynchronous tasks let a Django application respond quickly while workers handle slow CPU work, external API calls, cache preparation, bulk database operations, and scheduled jobs. Filipe Ximenes explains Celery’s producer–broker–worker architecture and results backends, and recommends passing simple identifiers rather than serialized model objects. He argues that reliable tasks should be idempotent and atomic, split into manageable units, and designed for safe retries with exponential backoff and jitter. He also covers time limits, automatic recovery, logging, error notifications, monitoring, testing with eager execution, and debugging tools.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: Thank you all for coming. I'll be talking to you about synchronous tasks, most specific specifically about salary. Shima is my nickname and you can find me on Twitter as Shima. I'm from Recife, that's a city in Brazil, northeast. And yeah, and I work at Vinta. So I'm one of the partners there and we are a service company from Brazil. Most of our clients are from the US. New York and San Francisco majorly. And we have a team of 11 experienced developers and work with Django and React.
Speaker 1: This is the link to our playbook. We compile there a lot of information about how we work and the things we do, so feel free to go there. We also do a lot of open source. So here are some projects uh you might be interested in. There's um our boyplate. Uh it's a Django React boilerprint. Um we have Django Rose Permission, Tapioc, and there is a few others uh on GitHub, so feel free to grab them. All right, so let's start with some context. In broad terms, the reason we use async tasks is because we want to answer quickly to our users. Here are some reasons um and some situations where we might want to use async
Speaker 1: tasks Tasks. The simplest case will be when we want to delegate it long-last CPU jobs. But the most popular reason people use async tasks is probably to execute external API API calls. Whenever you depend on external service, you no longer have control over how long things will take to be ready. It might also be the case that we will never it it will never be ready. The system you rely on might be down or broken. Another good reason to use a async task is to prepare for to prepare and cache uh values. And you can also use them to spread bulk database insertions over time. This can help you not to DDoS your own database.
Speaker 1: Chrome jobs are also another good example of things you can do with them Let's talk a little little about how they work. The problem of running a sync task can be easily mapped to the producer-consumer problem. Producers place jobs in a quill. Consumers then check for the head of in the head of the quill for awaiting jobs, pick the first one and execute. There are many tools available to manage async test in Python. Arquil seems to be getting a lot of attention lately, but Sari is the all-time champion so far. So let's talk about Telerary. First of all, let's introduce the correct name into our component
Speaker 1: From now on, producers will be web nodes, our quill will be called broker, and consumers will be workers. Since workers can also place new tests in the quill, they can also behave as producers. Now that we have um the basic components, we can dig a little deeper. The concept of a broker is very simple, but how do we implement this in a computer system? There are many ways to do this. One of the simplest would be to use a text file. Tasks can hold a sequence of job descriptions to be executed. Therefore, we do could use them as the broker of our system. The problem with text files is that they are not made for to handle real
Speaker 1: application problems such as network and concurrent access. Because of that, we need something more robust. SQL databases, on the other hand, are capable of running in a network and dealing with concurrent access. The problem with them is they are too slow. No SQL database on the other hand are quite fast, but many times they lack reliability. So when building clues We should use fast, reliable, concurrent enabled uh tools such as Rabbit MQL, Redis, and SQS. Study has full support for Habit MQL and Redis. Although SQL and Zookeeper are also available, they are offered with limited capabilities.
Speaker 1: Let's not talk about web and worker nodes. On the left, we have the code that should run asynchronously, that goes on the worker node On the right, we have the code that places the request for a job to run. This normally goes on web node. In this example, the web node places a net job and waits for the result to be available. When the res response is ready, the result sprint. Now I have a Django example. In it, the web node request uh the web node is requesting the number of attendees of the event to be uh updated asynchronously. Notice that we are passing the event objects to the test, to the to the task.
Speaker 1: Don't do this. Objects got serialized and stored in the broker. They are then deserialized before passing to the task. Passing complex objects such as model instance as the parameter comes with a few problems. First of all, in node versions of Celery, Pico was used to do as the default sterilization method As some of you know, PICO has security vulnerabilities. By allowing complex objects, you are increasing the chance of getting exposed. Later versions of this of salary address this by using JSON as the default serialization method Another issue is that database objects uh you pass might be
Speaker 1: might have changed between the time you place the the task and the time it gets executed. In that case, you'll be working with an outdated version of the object. What you want to do is to pass the ID of the object and fetch a fresh copy from the database We've been calling uh delay and get together all the time. But they are two separate things. Delay places the task to be executed by a worker and returns a promise that can be used to monitor the status of the status and get the result when it's ready. Calling get um in that promise will block the execution until the result is available. The ad task has
Speaker 1: has to store the result some uh somewhere and then uh and then um whenever it it finishes it will be accessible uh to the process that triggered the text the test. This means that we missed some piece of the of the architecture. Besides the web broker and worker nodes now components, there is also a results backend. The results backend will be used to store the task results. In practice, you can use the same instance you are using for the broker to also store results. There are other technologies besides the supportive broker options that you can use to um in the results backend. But there are some differences depending on what it use.
Speaker 1: In Postgres, for example, the get method will do polling to check the result uh when I when this is the the result is ready. In other situations such as as for Redis, this is done via PubSub. We've seen how to execute tasks and how salary works. But the real um hard problem about sync tasks is handling errors How to be prepared for bugs and failures. We are now going to go over two concepts that are essential for that matter. If you are to learn something from this presentation, learn them. The first one is uh.
Speaker 1: Idepotesy is the property of certain operations in mathematics and computer science. that can be applied to multiple times can be applied multiple times without changing the result beyond the initial application Multiplying by zero and multiplying by one are two examples of idepoting operations. Once you've multiplied by a number by zero, no matter how many times you do it again, the result will always be zero. It's the same for multiplying by one. No matter how many times you multiply the number by one, the result will never change Let's do the same for um HTP. Which of these methods are are the potent?
Speaker 1: Getting a resource should not produce changes in it. So get is a potent. Post is used to create resource. Well at least that use what we use to do before graphical so So every post request produces a change to the state of the application, therefore it's not a deportment. You cannot expect subsequent uh post requests to keep the application state the same Put 's a great example of a depotent operation. The first put request you make produces change to the state. But if you keep repeating the same request, no changes are expected. Delete is a bit delete is a bit of a gray area.
Speaker 1: From the application state perspective, it is a depotent From the resp response perspective, the first call might return 204, while the next one might return 404. So it depends on how you view it. The second comp concept is atomicity. An atomic operation is an indivisible and irreducible series of data-based operations such that either all occur or non-occur. Despite its being commonly associated with database operations, the concept of atomicity can also be applied in other contexts. In the first box, we have a test that sets the user status to updated, saves it, makes a request to Facebook, and only then updates the username.
Speaker 1: This is not atomic. If the request fails, we are going to have an inconsistent state in the database. The right way to do it is to first make the request, then update the user status and username at the same time. Another way to improve itomicity is to write short tasks. In the first example, we are iterating and sending emails to all users in a single task. If one of them fails, this task will stop and part of your users will receive the uh the email and the other part will not. In the second example, we iterate over the users again, but delegate the sending to another task. If one of the emails fail, it will affect it only
Speaker 1: it will affect only one user. The other advantage is that it's easy to rerun a task when it's failed. The order is an overhead on creating initialized tasks. Making them too fine-grained may harm performance. keep keeping in mind when you're designing. So I just told you about um how you can try a task that failed A fairly common situation is for tasks to fail while interacting with external systems that are not available. So it provides a retry method that can be called inside a task. This will make it try again to execute the task. In the example, we are fetching the user likes from Facebook. If it fails, we wait ten
Speaker 1: seconds and retry. If your tasks are indepotent and atomic, you should have no problems calling retry as many times as you need until the test succeeds. If they are not, retry may producing inconsistent states or can end up spamming your users There are some other caveats to trying. In the previous example, the request to Facebook might be failing because the it Facebook is under attack. So it might be a good idea to back off and give it some space to recover. You shouldn't kick a man on the ground. Another good idea is to make the interval between retries grow exponentially
Speaker 1: This will increase the chances of things getting back to normal before your next try. Also throw in another uh random factor. Imagine having a hundred tasks failing and retrying at the same time. The overflow might actually be the reason why the system we are interacting with is dumb in the first place. Alright. This is something I actually learned about the first day of the conference. I was chatting with uh David Baumgold. So if you use a salary for you can actually cut down a few lines of code by passing a out-retrive4 parameter. You can do the same thing I showed before, just using the out-retry four and passing the um the classes, the um error classes.
Speaker 1: I also learned that they've just got a pull request approved for a retry back of parameter. In the next salary release, we will be able to use retry backoff along with out retry4 or enter and get exponential backoff out of the box. So thank you, David. Another common issue are tasks that take too long to execute because they are there just there's a book there or uh we have a problem with network latency. To prevent this from harming the performance of your application, you should you can set a task time limit. Survey will input the task um Sorry we will interrupt the task if it takes longer than the time you set. In case you need to do some
Speaker 1: recovering before the task interrupts. also set test soft time limit. When that runs off, Sari will raise soft time limit exception and you can do some cooperation before the test killed. The XLATE setting is also something you should know about. By default, Survey first makes the call, uh uh marks the task as run and then executes it. This prevents a task from running twice in case of an unexpected shutdown. Having important anatomic tasks will give you the ability to turn on axolate cephine. By doing so, Sarah will automatically rerun the interrupt task once it re covers.
Speaker 1: We talked a lot uh a lot about how to actively handle errors But we we all know that bugs are inevitable inevitable. Bugs in a web application are generally easy to spot If something breaks, your user will get a 500 page and they will find a way to let you know that something went wrong for them For a sync task, this is not always the case. In many many situations, we'll be dealing with things that do not directly affect the user experience. This means that you should be extra careful with monitoring tasks. Monitoring will help you to capture errors early and not when they are already too late. Starting from the basics, logging.
Speaker 1: Make sure you log as much as possible. This will help you tracing what went wrong when bugs arise As usual, be careful not to expose sensitive information in the logs. This is a security threat. Logging is a general advice for any kind of application. But especially important for our tasks as they have no user interface and and the booking then is harder. Make sure people get notified when things fail. Tools such as Sentry and Opbit can be easily integrated to Django and Celery and will help you monitor in errors. You can also integrate them with Slack so you get notified every time something goes wrong.
Speaker 1: Make sure you fine-tune the what produce notification. Too many false positives and your team will stop play paying attention and let actual errors pass unnoticed. Have a well-defined process to deal with those errors. Make sure they are included in the backlog and prioritizing accordingly. Floor is a tool for live monitoring solid tasks. It allows you to inspect which tags are running and keep trace uh to keep track of the executive ones. It's a standalone application and definitely worth using in bigger projects. Django solid bit status is something that just came out. We didn't talk much about uh scheduled tasks, but they are not trivial to monitor and test.
Speaker 1: The Chrome tab API is sometimes confusing and might lead to mistakes. And this is also a project from Vinta. Django SolarBit status will add a page to the Django Admin interface. In it you'll be able to see all scheduled tasks along with the next ETA. This will give you a limit, uh this will give you a little more confidence uh that you got things right. Testing debugging tasks can be harder than w what we are using to in normal web applications. But there are a few things you can do to mitigate this. Task OSegger is a stuffing that comes very handy for testing and debugging. This code is from the salary
Speaker 1: source. If you had the task always eager eager something set to true, whenever you call delay or apply a sync, it will run the test synchronously instead of delegating it. This will simplify the bugging in local environments and facilitate automated testing. This last thing is actually more of a fun fact So it turns out there is a PDB for salary. RDB will allow you to telnet your breaking point and live debug it To be honest, I've never used it and I'm not really sure what situation we would use it, but I'm sure someone will find it handy. Alright, so time to recap. First of all, don't use complex objects in
Speaker 1: task parameters. This will help you avoid insecurity and inconsistent issues. It's inconsistency issues. Right at the pointed anatomic tasks, you want to be able to freely rerun your tasks. Back off when you try. Other systems may may need some space to recover. Make extensive use of monetary tools. Tasks that are harder to debug. Collect as much information as you can Get notified when something fails, but be careful to only notify when something important happens. And last, use Task Always Eager for testing and debugging. The last thing I want to show you is something that I've built in the process of making this test this talk.
Speaker 1: So task uh sorry task checklist has a compilation of the things we talked about You can actually click through the items as you verify what are missing in your tasks. And it's also open source, so you are very welcome to contribute and improve it. That's it. Thank you. I've sorry. I've got some um the links to the to these slides are there. Um so it's slash Vinta 2017. There is also the link to Sari test checklist, my contact information, and
Speaker 1: we have a newsletter if you want to sign up to this. It's the link's there. Thank you.
Speaker 2: Thank you so much, Philippe. If you have questions for him, definitely chat with him afterwards or during the They'll be happy to work together. All right, thank you. And we're about to have closing, so hang out. Closing happens here.
Async tasks let the web application respond quickly by moving long-running CPU work, external API calls, cache preparation, bulk database inserts, and scheduled jobs out of the request cycle.
Discussed at 1:00Web nodes act as producers that place jobs in a broker, while workers consume and execute them; workers can also enqueue additional tasks. Common broker choices include RabbitMQ, Redis, and SQS.
Discussed at 2:31Model instances must be serialized and may create security risks or become stale before the task runs. Pass the object’s ID instead and retrieve a fresh copy from the database inside the task.
Discussed at 5:34Calling `delay` queues a task and returns a promise-like result object, while calling `get` on that result waits until the task finishes. Celery uses a results backend to store and retrieve the outcome.
Discussed at 6:19An idempotent task can be run repeatedly without changing the result beyond the first run. An atomic task either completes all of its related operations or none of them, helping prevent inconsistent state when failures occur.
Discussed at 8:35Retrying is safest when tasks are idempotent and atomic. Use delays between attempts, preferably exponential backoff with randomness, so a failing external service has time to recover and many tasks do not retry simultaneously.
Discussed at 11:39Set a hard task time limit to interrupt tasks that run too long, and use a soft time limit when cleanup or recovery is needed before termination. For important idempotent tasks, enabling late acknowledgment allows interrupted work to be rerun after recovery.
Discussed at 13:56Log enough information to trace failures, notify the team through tools such as Sentry, Opsgenie, or Slack, and use Flower for live task monitoring. For testing and local debugging, enable `task_always_eager` so tasks run synchronously; Celery also provides an `rdb` debugger.
Discussed at 15:28Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026