Wagtail and Caching
Published June 27, 2024
This video features Jake Howard at DjangoCon Europe 2024 in Vigo, Spain.
Talk: Empowering Django with Background Workers by Jake Howard
https://pretalx.evolutio.pt/djangocon-europe-2024/talk/VDYCVB/
Background workers move slow, failure-prone, or hardware-specific work out of Django’s request/response cycle, improving responsiveness, throughput, and reliability. Jake Howard explains the trade-offs and demonstrates how email sending can be moved into a worker using Django RQ. He argues for Django Tasks, a first-party API contract that lets applications use Celery, RQ, or other backends without changing application code, and outlines its proposed ORM, immediate, and dummy backends, current status, and future development.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
There we go. Okay, I'm hoping you can hear me all at the back and the mics are gonna pick this up. Um so yes, as the slide says, I'm Jake Howard. I'm here to talk to you about background workers, what they are. how to use them and some exciting things that are hopefully coming to Django. So, who is this bloke who's standing on screen talking to you? The contrast there is somewhat readable, but I'm a senior systems engineer at TorchBox. I'm also on the security and performance teams for Wagtail. And as of last week, I'm also on the core team for that as well. I'm an avid self-hoster, and in my spare time, I um help students build robots. And I exist on the internet in various places. So, as hopefully
most of you know, Django is a web framework. It's a magic box which turns HTTP requests into HTTP responses. What you do inside that box is entirely up to you. For something like a blog, that's generally as far as it needs to go. Um but for a more complicated application, you're gonna need a little bit more than that. It's not just a case of taking information, putting it in the database, and at some point later retrieving it. You're gonna need things potentially like notifications. notification emails, talking to external services, transcoding video, complex reporting, and given it's 2024, lots and lots of machine learning, AI and everything else. Um and for many of those use cases you're gonna need code that runs outside of that magic box. You don't want the user to be waiting on those requests while they happen. If you had to wait while YouTube
transcoded all of your videos, you'd get pretty annoyed. And so to do that you need a background worker. But what are background workers? Well the idea is they let you offload complexity outside of the request response cycle to be run elsewhere, potentially at a later date. They keep your requests nice and fast. Usually the slow-moving bits can move somewhere else where the user doesn't have to wait. And that nicely improves throughput and latency. Now, how does that actually work? Well, generally your web process, usually Django, I would assume, submits a function to be stored in the QStore. At a later date, generally your um process runners will then grab that task, run it, and store the value back in the Q
tasks that Django can then review and retrieve later. Now, background workers are a very useful tool, sort of by design, but that doesn't mean they're useful for everything. Um and as with all great things, it depends on where they're useful. There are a lot of trade-offs between complexit complexity and functionality um when you're considering whether these make sense, there are a few sort of rules to consider. Things like is your action gonna take time or could it take time? If you don't want to make the user wait, throw it into the background, come back to it later. You don't want the user to have to be completely unable to close their tab. And then um not be able to continue off with the rest of their day. You want it to go off into the background and then retrieve the arction later, even if that's by polling
in the browser or something like service and events. Again, does control leave your infrastructure? The core components, things like the Django server itself, the database, the cache, etc. things like that, you can control and you can closely monitor. Um and you're in a good position hopefully to fix them if they go wrong or at least a different team in your company is. That's not true for external APIs. It's someone else's SRE team that is responsible for making that system work Um it's th their performance characteristics shouldn't affect your application. Maybe instead of about when the code runs, it's more about where it runs. For example, your Django application probably doesn't have a huge amount of GPU compute. To do any kind of machine learning training. Maybe the workload requires a bunch of RAM.
Maybe it needs specialized hardware to run that again your Django application servers aren't gonna have. Maybe it needs to run inside an isolated network, again, where Django isn't. Those are use cases that benefit from background workers being able to decouple that architecture. And background workers do have a world of usage. It's important to consider the scale when you're actually designing these features. It might work fine locally, but as your application grows, there's likely to be more data, and so that means that your processes are going to take longer. The user shouldn't have to wait for those things. It means they can get back on with their day and your web servers can get back on to processing the next set of requests. Now this list of examples is quite long, and it's even longer than the ones I've listed there. This is a very generic tool and a very generic process that can be applied to lots of different situations.
Um we're designing an application that scales with some of these features. Maybe consider moving them into the background with background workers. Now, moving back to Django, because as I've been told, this is actually DjangoCon. Um, in Python and Django there are lots of different frameworks to achieve Background workers and here is a very small list of them. Obviously right at the top we have Celery. It's probably the biggest and well known one. But it's not all that exists. There are so many different libraries out there, all with different strengths and weaknesses and different learning curves, or in the case of Celery, a learning cliff. Now, let's look at a slightly more concrete example, email. Um, it's very very common functionality for a Django application So let's consider a content management system for completely unbiased reasons.
When a page gets published, for example, we want to send an email to everyone that's subscribed. The reasonably common use case for an application like this. And here's the code we might use to write that. And it's sort of split up into a few key steps. We want to find the user, we want to construct the email content, and then we want to actually send the email. And this works perfectly fine and it scales relatively well, particularly locally. Um but it does have some issues Um if the if connecting to the email server takes a lot of time, usually it might only be a few milliseconds, but it could be a few seconds if there's latency issues. That's gonna slow down your users and slow down your requests. Uh if something goes wrong with one of the emails, the others won't get sent. You'll have an exception and they'll just be lost.
Uh what if your email gateway is down altogether? How do you handle that? You need to come back to it later potentially. If you're all living inside that one request, you don't have anywhere necessarily to persist that data How are you going to handle that? Um the web worker as well, it can't process things while it's trying to send an email. You don't want Gunnarcorn to be waiting on someone else's email infrastructure. You want to just do everything yourself. So let's look at how we might do this with background workers. In this case, I'm gonna use uh Django RQ. So this works in much the same way. You find the user that you want to email, you're gonna start a new task for each user. And inside that task you construct the email content and then you send it. And that function here happens
all inside the background. And most of this is exactly the same. If you knew nothing about RQ or Django RQ, you could probably still maintain this Moving it into the background just adds a little bit of extra stuff into the bottom and then the user can get back on with their day. You push a few things onto the queue and get back onto it. It doesn't matter if the email server is down or slow or has any other kinds of issues Emails get sent out by the runner, and if you've got multiple runners, you can send multiple emails at the same time, massively improving throughput and reducing latency on getting those emails And email sending is a very easy action to move into the background. It's a connection to an external API at the end of the day. It's going to have variable latency. It's running on infrastructure you probably don't control.
And all of those things can massively benefit from background workers. Now, there is a slight issue in what I just said. Um you may have noticed it when I said using RQ. And that's because it is very much that code is tied to RQ. And each worker library is going to have its own features, its own configuration, and its own caveats and implementation details. What what if we wanted to use Celerix instead? Well That's easy. We just change a few of the imports, change a few of the lines, and it works. But therein lies the problem. We've had to change some things. You've had to make some changes, sure. They're small, it's only a tiny amount of code. But what if you wanted to support both systems concurrently? That's a perfectly reasonable use case, but you'd have to then support
lots of things. And that leads to some competing situations. It's hard enough having multiple options, but how do you choose between them Maybe you've got experiences with certain libraries that you like, if you've got the time and patience to investigate the differences between celery and RQ and everything else. Great. But what if you don't? What if you don't want to? Maybe you've got a standard at work that you already need to work to. Maybe you need specific features that some libraries have, other libraries don't. If you're new to Django, do you really want to be spending your time weighing all of these things up? Probably not. You want to just get going and move on and build your actual application, not focus on the nitty-gritty details of your background worker implementation. Potentially if you pick the Love -Rong library as well, it might end up biting you as time goes on.
If you have scaling issues with a certain one, you need to move, you might need to end up rewriting a bunch of code. You really don't want to do that As well, this gets complicated for library maintainers. For example, like Wagtail, if you're trying to allow the user to use whatever library they like, whether it's Celery, RQ Or the plethora of other ones that I name. You really don't want to have to maintain separate integrations for every single library. And when a new hot thing comes out, you don't want to add more code to deal with that Do you just choose the big ones and go you either use celery or RQ or nothing? Or do you have to try and expose a hook so the user can tie the two together themselves? Or do you do anything like that? And it's just more
complexity. And to be honest, that's ridiculous. There should be one universal standard which combines them all. A single API to help developers use a library without tying their hands. Ideally, it should be first party. It should allow library developers to depend on it instead of having to expose a separate API. You want something that can scale easily. You want it to be there and started and easy for your smaller projects, but something that is feature-packed for larger deployments like Kraken, for example. Um and you want something as well that could be very easily tested. You want to change very, very few things and be able to test something very reliably without spinning up an entire salary cluster for your unit tests And that's what I'm here to talk about. Um I want to introduce Django.
tars. And I have introducing with an asterisk for reasons that we'll get onto later. This is a in-progress API spec for first-party background workers in Django. The idea is to build an API contract between worker library maintainers and application developers to sit as a compatibility layer between Django and their native APIs, hopefully fulfilling that promise of write once, run anywhere. Although potentially not with j with Java. Um the idea is this implementation should have a few built-in implementations or have something based on the RM to leverage Django's existing powerhouse that is the RM. You need something that is can immediately run cost, but in Local development or as part of testing where you don't want to run a background separate process, you just want to pretend everything gets run at the same time
As well, you might want a dummy backend for testing. You don't want to really run the task, you just want to prove, did I try and run the task? Because that's the bit in the unit test that you care about. This process is hopefully, fingers crossed, gonna try and land around the Django 5. 2 deadline. That depends on how much I have going on in my life Um and the idea is to have this process be entirely backwards compatible right back to Django four point two to allow anyone using Django at the moment to benefit from this work. Now, let's look at the same code example that we had before. In this case we'll look at the one that is tied to Celery. If I wanted to add support for RQ, I'd have to duplicate some parts of this. But instead, what I can do is I could rewrite this in a very small way with Django tasks.
It's still simple, it's still approachable, and it's still easy to use, if I say so myself. If we wanted in future to take this code and go, okay, well we're using something based on celery, they actually, as a user, I don't want to use celery, I want to use RQ. None of this code needs to change. It's separate stuff configured in settings. Zero lines of code change. If a new library comes out that I want to use instead of set of celery or RQ, zero lines of code change. If this is in a library, not in my own code, then I'm no longer constrained by their preferences. I can set my settings however I want, and as a library maintainer, I don't have to deal with the extra work and burden of supporting all of these background library processes. Again. Zero lines of code need to change.
And that's the goal. Now, in this case we can actually make things easier. Because email is such a common use case and is so easy to be extracted, we can hook into some of the things that Django already allows us to configure. This will jump back to the even simpler implementation that is just do a loop, do things, there's no background workers in site. But instead, all we can do is just change the email backend itself. Django then knows that when rather than sending an email immediately, kick it off into the background, do it later. And that means emails get magically sent with no additional work. No lines here changed. Your application works exactly the same, but with a single line change in your settings, now all of your emails are sent in the background
Now, I'm sure you're thinking, why something new? Celery already exists and it's got a borderline monopoly on the task queuing um ecosystem. Writing a production task queue is hard, writing distributed systems is hard. I'm not smart enough to do it. As I've been as I've been told many, many times, building these systems is hard. Um but why don't we just vendor something? If not celery, then why not one of the other ones that exists? Well That's not really what the point of this work is. The big value add comes from that shared API contract, being able to swap things out and interoperate as opposed to having to replace components. You swap out some configuration based on your scale as opposed to needing to rewrite a bunch of code. It's very much in configuration and it lets you use the ecosystem that you're already familiar with and comfortable with without having those decisions imposed on you.
Now, the Django backends, they will hopefully become great, but that must be done with careful planning and consideration. Django needs to remain that stable and reliable base that it always has been, and it's the reason we're here And of course we don't want to burn out the Django maintainers by having to rewrite their own version of Celery. Now, why do we want something built in? This is all something that could live in an external package, but where's the value in putting it in Django itself? Well the main reason is to reduce that barrier to entry. With it being integrated, there's no additional dependencies that you need to install. You have to learn one API one time. And that's it. And it will scale as much as you need. You learn it once and that's it. When a developer joins a project, if they already know the API, they're already familiar with how to kick code into the background without needing to learn your pro
what library you're using at that given point on that given project. A common API also helps library maintainers. Being able to maintain a library is hard work enough without needing to think about how to move code into the background. If Django can take some of that complexity off of you, great. You can use the tools that you like during development and advertise them. But as a developer, if I want to use your library with a tool that I like instead I can, and that's the benefit. And currently there isn't really the ability to do that with any of the tooling ecosystem that is available. The burden is just too great. Now, the RM at scale. For some scales, an ORM-based background worker might not be viable. The sanctuaries, the Instagram, the Krakens of the world.
It's probably not viable. Postgres scales incredibly well, but for some times it just isn't good enough running things on a database. And to be honest, that's okay. But the same is true for things like Elasticsearch versus Postgres full text search, and it's a debate I've had many, many times, and it's a debate that has been going on for a while. Elasticsearch is quite likely better for say ten percent of the users that need full text search. But that doesn't mean that the other ninety percent of people wouldn't be happy with Postgres. And they probably wouldn't benefit from that extra ten percent. from Elasticsearch anyway, but they would have to deal with the additional hosting and extra complexity that comes with running an Elasticsearch server. versus using the database they already have.
They could be perfectly happy with Postgres Full Text Search, save time, money and complexity, and that's what we're trying to do here. Let them start in the easiest way possible and in future if they want to move to Elasticsearch they can, but that doesn't mean they need to deal with all of that extra burden up front So, where are we now other than Vigo? Well, as of a few weeks ago, the Django enhancement proposal that I wrote was approved. In theory, this can now go into Django once it's written. But that all depends on being able to actually produce an implementation rather than just talking about the theory with a bit of hand waving. Because theory doesn't equal practice. But you can play with this right now. You can download it, you can play around with it.
There is a dummy backend, there is an immediate backend, and there is the ORM backend, which is where the actual magic happens The dummy backend is great for testing. The immediate backend can help you get started when you're not quite ready to move things into the background, but you still want to be building that functionality for the future. And you can install this now. You can pip install Django tasks. You will get a backend that works. The QR code is just a link to PyPI. Please do download it. Have a play around with it. Tell me where all of the bugs in my code are Do some testing so we can all work together to build something that is better. There are still features that we need to do around performance, um improvements, scalability, additional features, things like that. But they'll come with time. So, where are we going to be soon?
Well, there's a lot of testing, there's a lot of improvements that still need to be made. There's going to be the upstreaming itself, which is where the big effort of this comes from. Otherwise it really is just another competing standard. But once Django dash tasks is in a better state, it can become Django. tasks. Hopefully in time for the 5. 2 release, we'll see. And at that point we can then start the adoption process, or even we can start some of it now. The more people that know about this, the better it is for everyone. Developers can start working on the integrations now, knowing that they can trivially migrate once it's inside Django. Now I'm sure a question a lot of you are thinking is, is this the end for celery? No, not at all. Celery is s a great choice for those kinds of tools, and they really do have a massive head start.
This co this project started Four months ago and celery is a darn sight older than four months. This is much more about usability and flexibility than it is about trying to rewrite everything. If you need a certain feature that Celery has that I haven't talked about, keep using it. You haven't made the wrong decision. But now you have a Django native API that you can use that uses celery under the hood, and it means that you can swap things out as you need to in future. Now The world of background workers is huge. There are countless nice features. I've listed some here and you can probably name double that number if you've used background workers a lot before. None of not everything is gonna make it into the initial version. And that's okay.
External libraries have a head start, celery, RQ, they already exist. We're already running them in production But I hope that we can slowly catch up to them, bring the stability and longevity guarantees that come with Django into this ecosystem. That doesn't mean they'll never come, these system these um sort of features, but it will take some time to make sure that we can develop them properly and with the reliability guarantees that come with Django. And I hope that with your help, hypothetically, we can make these things happen But the future is bright. I see a time when more and more people can reach towards Django's task system and move towards background workers in general. Moving things into the background will make your Django application seem faster, even though you haven't necessarily changed much.
It can massively improve your throughput, it can massively reduce your latency, and it can improve your reliability. Gone are the days of needing to do the additional research and testing to find the tool that's right for you. You can use the ones that are built into Django, and as you scale and have more information to understand what your explicit needs are, then you can do the research. Knowing the information that you've gained then, rather than do a bunch of stuff preemptively, potentially crippling yourself as you scale. OGScale is easy to change without rewriting half of your application potentially with all the knowledge that we gained. And a lower barrier to entry helps everyone. That's the whole point. So, where next? Well it's time to turn my dream into a reality.
Um if you've realized from this talk that actually there are some use cases where background workers would be a benefit. Maybe give Django Tasks a try, test it out, report back your issues, give me suggestions and improvements. The issues and discussions are very much open and I welcome any and all contributions. If you want to get involved, please do. There is plenty of work to do and I cannot do it alone. If you maintain a background worker library, have interesting use cases in background workers, or have been burned by a background worker and want to make sure that I don't make the same mistake. Let's chat. Thank you very much
Background workers move slow or complex work out of Django’s request-response cycle so users do not have to wait. They can improve request latency and overall throughput by processing tasks elsewhere or later.
Discussed at 1:33Use one when an operation may take a while, depends on an external service, needs specialized hardware or an isolated network, or otherwise should not block the user or web process. The decision also depends on scale and the trade-off between added complexity and functionality.
Discussed at 2:18Queue a separate task for each recipient, and have that task construct and send the email. Multiple workers can process emails concurrently, while slow or unavailable email infrastructure no longer blocks the web request.
Discussed at 6:19Django Tasks is an in-progress API specification for first-party background workers in Django. It is intended to provide a compatibility layer between Django applications and worker libraries, allowing code to be written once and run with different backends.
Discussed at 11:01With the Django Tasks API, the worker implementation is selected through configuration, so switching from a Celery-based backend to RQ or another library requires no changes to the task code. This also lets library maintainers avoid supporting separate integrations for every worker system.
Discussed at 12:36Integrating it into Django lowers the barrier to entry: developers learn one API, avoid extra dependencies, and can use the same approach across projects. It also gives library maintainers a common contract and lets applications start simply before moving to a more specialized backend if needed.
Discussed at 15:01The project currently provides dummy, immediate, and ORM-backed implementations. It can be installed with `pip install django-tasks`; the dummy backend is useful for tests, while the immediate backend runs work without a separate background process.
Discussed at 18:12No. Celery remains a good choice, especially when its existing features are needed. Django Tasks is primarily a stable, portable API that can use Celery underneath while making it easier to change implementations later.
Discussed at 18:57Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025