Profiling Django & Python apps
Published June 27, 2021
This video features Jérôme Vieilledent at DjangoCon Europe 2021 in Online.
From Development to Production, Getting Actionable Insights to Optimize Django Code Performance
Jérôme Vieilledent
We’ll see how developers can see in real-time how end-user perceive the performance of an application, and the several levels through which Blackfire can drill down in order to find the root cause of issues.
We’ll see how Blackfire can be used within CI/CD or any testing pipeline, to prevent issues from being released to production.
And we’ll see how Blackfire can be used on a development machine to reproduce and analyze issues, as well as validate code iterations.
Sponsored by: Blackfire.io
Jérôme Vieilledent explains how Blackfire profiles and monitors Django applications from development through production. Its Python probe collects function timing, CPU and I/O time, memory use, database queries, and external requests on demand, while an agent aggregates and anonymizes the data; visualizations such as timelines, call graphs, metrics, and memory graphs expose slow templates, excessive queries, bottlenecks, and leaks. In production, monitoring shows real-user response times, errors, traffic, impactful transactions, deployment effects, and spans, while automatic and on-demand profiling plus synthetic YAML-defined scenarios turn those observations into actionable performance checks and recommendations.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Hi everyone. I hope you're uh enjoying your last uh conference day, I mean before the screen days and stuff. Um so today I want to um to show you um a tool, a specific tool um for Python developments that can help you um help you out debugging and uh developing your application uh from your development environment to um uh to uh staging and production and beyond. So this specific tool, well, uh is uh called uh a profiler uh specifically, and uh so I'm going to um to
talk about uh BlackFire uh itself which is a profiler and monitoring tool. So first, well, I'm so Jérôme Vieidon and I'm a developer advocate at BlackFire. io. So that will be, I will take on tour. on on what you can do with Blackfire and how it can be useful for you. The goal here is to make your development life easier. uh and uh so that you can focus on on what really matters for you for developing your applications and specifically to junk go So for this, I'm gonna first share my screen, of course.
Here we go. Right. So as I um explained just before, um Blackfire is uh first and foremost a profiling tool. Like many others in the Python ecosystem, I know that there are several profilers, open source profilers in the Python world. Here I will try to explain what can differentiate these tools from Blackfire. First, BlackFire is a tracing uh profile uh profiler it means that um it will take traces from every functions that are executed in your code. Um
how it is how it works, it um you have two main components. Um first Here we can see in the middle the probe. The probe is the components, is a PIP package for your Python application that will gather all the data coming from the Python engine. So we get uh when we uh talk about data, it's uh how how long functions took, uh how long Python I was waiting for external things to happen like file system operations or database queries and stuff like that. So that's what we call input-output uh how long um you had on CPU time, I mean
computate real computation, how much memory you took, etc. etc. And this for each and every function from your application. So that's the first component that needs to be installed like any other PIP package. It's written in C, so everything, every binding is in C so that it's the fastest possible. The only things in Python will would be the SDK so that you can interact directly with the probe to enable profiling. Very important thing to know about the probe in is that uh it is completely production safe.
You can and well you should actually install it in production because it's always an on-demand profiler. Uh it means that when unless you you don't ask for a profile, it does nothing. It stays completely inert. And when you do ask for a profile , of course there would be an overhead, because you know, measuring things from within from within the Python engine, even if it's very fast, it always comes with a cost. But this cost, this overhead, will won't uh uh
affect your end users in any way Because it's always on demand. When you ask a profile, the profile will only be for you. And the other users will don't know that there is a provider and that you're providing. So no overhead at all. The second component you can see here is called the agent. The agent is the demon that just sits there and waits for raw data from the probe. uh to because the the probe in order to limit to limit the uh the the overhead won't do the aggregation stuff uh aggregating the data and stuff like that uh they will it it the probe collects raw
data and then sends uh sends this raw data to the agent which lies into your infrastructure as well. It can be in a dedicated container uh you then usually um uh Docker container or within your Kubernetes pod or whatever fancy. They only need one agent and the agent We'll uh take this uh raw data, aggregates them because we take samples, uh 10 samples to have proper average. We'll see that in a minute. And then we 'll anonymize uh sensitive data such as uh um external HTTP requests with tokens and stuff like that and SQL queries uh uh so that Yeah, well, things are quite
uh uh straightforward and uh good to uh uh good to exploit for you. And the agent, when it's done, is uh uploading the aggregated profile to BlackRock servers. So that's the basics, let's say. So here, well, this is the same application. If you already followed the workshop, you would uh find this quite familiar. This is a very small application written in Django. Uh it's like a blog-like application for uh um so here it's all about uh hunting the Bigfoot. So instead of articles, we are writing sightings, reporting sightings of the Bigfoot, and members of the club can
comment this out, each sighting out. So here we're on the detail page and we're still in production. You can see I'm here locally. I'm running the development web server, the manage. py run server. um so pure um pure uh development environment and well before commenting and the and deploying i want to understand uh where I could have uh performance bottom next or uh well or maybe much simpler understand how my application is behaving uh exactly because you know um When you're running your code, your application is becoming
kind of a black box. You cannot really know unless you're profiling, you cannot really know what's happening within. Because when you're writing code, you're writing like the blueprints. But when everything is running through uh the the the ser the run server uh from the Django uh development server or with gunicorn uh well then you do not have any uh insights And this is exactly what a profiler is about. A profiler like Blackberry. It gives you like bionic glasses to know, to see exactly what's going on. So to trigger profile here, I'm I just have to uh click on this uh
very big red button to uh to to get started from my um my Firefox extension. It's also available for Chrome. And I just need to be authenticated. So let's click on it. And when I click uh to to trigger a profile, it's taking, like I said, 10 samples. So it can to 0 to 100%. and then averaging things. And now I have the summary of the profile for this page. Well I have like different dimensions there So wall time, which is the overall time execution for this page in your application, which is all split it into two flavors.
The first one is the I. O. time. So as As uh uh mentioned earlier, it's a time spent by Python waiting for other operations from the from the operating system. So it might be file system , it might be network network calls, HTTP requests, and well, stuff like that. Uh the second one is the CPU time, so the actual computation time when Python is really doing something. The memory, so uh very important aspect. Well, external HTTP request, well here we have none, and SQL queries And then we have two uh kind of visualizations available. So I will first click on the timeline.
So the timeline view is um a visualization of the profile uh on a time base. So this is quite unusual to uh most providers. Uh most providers are um only um showing interactions between functions or lines and stuff like that. And we if I with Blackfire, you will be able to have this as well. This is actually the main one. But as a developer, usually you have uh well uh this is my case at least i have a representation of the uh the the profile of how my my application is behaving on over time. I know that well I enter here from my whiskey um
from my whiskey uh scripts and then I'm going to my Django view etc etc Um and you get this uh visualization s thanks to the timeline view. If you're already familiar to um uh Flame graphs, this is kind of similar, except that here uh well it's like a reverse flame graph. I like to call them pit graphs. uh and well the the deeper the pits the uh the more complex your application uh is basically And each block here represents the function call. So we have all the middlewares from Django here getting to the WISD handler. And here well the our view
and um well uh goes going to templates and to uh uh um sql queries from Django models etc etc All the blocks are proportional to the time spent. So the longer the block, the longer the function. And with this representation, you can quite easily spot where from within your application you can have performance issues here. So for example, there I can see that these two template blocks are quite take a lot of time, right? It's really long.
Here it's citing show. html, which is taking well 564 milliseconds, which is almost all the time for all the requests, because we are here on 600 milliseconds. So there might be something to have a look at. I can click on it and to get the whole call stack and understand from where it's going. Well here it's very straightforward once I can see that it's coming from my uh my settings show uh Django view All right. So another very interesting uh part of this view is the blue graph in the background.
This blue graph represents the memory consumption growing, right? The peak memory actually, the peak memory, which is uh the envelope which is growing over time. And this is a very uh interesting graph because uh from this uh blue graph you could easily spot memory leaks. And thus uh issues, performance issues introduced by uh too much memory consumption consumption. So It's always good to have an eye on this graph. Also, well, timeline metrics. Metrics are combinations of uh aggregation of functions that are related to a domain, specific domain. uh
those uh are actually automatically gathered by black fire itself you don't have to do anything here to make them appear So we can see that uh well for the query sets uh uh we have 70 uh calls um bound to uh uh query sets uh and it's it took uh well 50 uh five more than 500 uh milliseconds so we definitely have an issue with uh with this metric and we can see that Thanks to the color code in pink we had a lot of calls there. So yeah, we definitely have an issue. And sample, well, Django templates and Django views, etc. etc. All right, so that was for
uh for for this one. Um Then well I have these representations uh of of the the profiles um from the different dimensions. And this is quite important because you can see that if I do not zoom, I can see what is my critical path. This is something very useful to find bottom line. When you find these kind of guys of nodes, well, this is where it's really intense. And well, actually it's been called more than almost 1,000 times. Then our goal is to understand well where it's coming from and get getting up
and see where it's coming from from my application, what triggers, what is the protofly uh effects uh where it came from. So well here clearly this is uh coming from this uh from this Django uh model function uh coming from uh specific uh uh um template tags right uh you can see um here that uh you you have a different representation uh of the call graph depending on how on on which dimension you focus on. And so with only one profile you will be able to see. have different aspects of the behavior of the application.
Okay. Same here with database queries. Have a list of the different data queries. All right, this uh you would do uh um quite often during your development cycle, right to understand what's going on, to find um memory leaks or performance bottlenecks. But then let's say you're done. It's time for deployments. And you have installed everything in everything needed in uh in production in your production service uh and well this is what i've done here i've deployed i deployed
Uh sorry, I had issues here. Okay. Thank thank you, Garrett of Tom. Sorry. Um So I've deployed on here on PlatformSH, actually. This is on the cloud platform with a container. And this is my final thing. But now, well, uh I I my supposition, my guess is that I fixed everything, right? Uh but um I cannot be so sure. And this is where monitoring comes into place because I'd like to know how my application is behaving.
I mean with real end user traffic and to understand if I did everything correctly. This is the point of monitoring. You probably know or probably uh already know different uh monitoring tools. uh with different services that's uh uh like well like insight uh data. com and URLic on these kind of tools. Which are very complex tools and very useful. Definitely. Well, if you use them, that's fine as well. But here I wanted to show you how Blackware monitoring can help you. So here, yeah, this is the same application, right? It's it's supposed to be uh super optimized. Well, now then
I want to show you this. So Here I'm on my BlackFire account. And the first thing as I subscribe to my uh to BlackFire monitoring. I want to know how my application how good is the health of my application. And so this is the first dashboard you're getting, the health report. And the health report will gather or will extract key information. uh from every request coming through uh your python uh application. So here it's pure Django. And
Everything is going through the same probe and agent goes to the there's nothing specific to do actually. If you already installed for profiling, you're good to go for monitoring And I can see here that from the last week, well, I have very few traffic, right? I have uh well in in average um 96 percentile is well 155 milliseconds and a major response time to 60 milliseconds and i can see the evolution of it uh the number of requests and also very just interesting the number of errors you might have. Well this one is pretty empty. So I'm going to show you
another one which is not in Python. I'm sorry about that. To get more data to show you, but you get the idea. I mean it's exactly the same. But uh and uh about monitoring here, you would see the uh most impactful transaction. And this this is key. Because when you have your your application live , well, usually you get analytics , and if you have already monitoring all logs. you might know where most of your users are going through. But here with the monitoring and the impactful transactions, you will get very precise information about
which transaction and uh which transaction which is actually an identified request. Um Which transaction is the most impactful for users? So here we only have two. Well, I don't have many pages here, but uh we can see that uh the the the different the response time distribution. uh on this uh on these uh uh transactions and where we have uh http status uh uh sorry uh distribution per um status and here for example the uh five hundredths the the number of errors um which then i can go to the monitoring uh itself
And going so here I can I get the different graphs of to to help me select the timeline itself and to understand what what went wrong at some point, maybe. uh and to be able to filter out uh here down to the transactions and these transactions then again click here uh to see the the same kind of um uh of information uh dedicated on the transaction themselves and also If we're jumping to some to to uh another one having more information, you can see that we also have spans. Uh the spans, well here we have a lot of spans, especially because we have a lot of
well quite so a few middlewares. It's this gives you like it's like a mini timeline. It gives you an overview of what kind of component software components are um interacting with each other just like you were seeing uh with the the timeline view but here it we're um uh more on the surface but you can still see overall what's going on Right, so now I'm gonna switch to a more extensive one. So this is basically BlackVerrier. io website, right? So I have a full um health reports. uh live which with real end user data and stuff.
So that's can help me show you well here the number of requests and the things I already showed you. Also the successful builds. which is another um another feature from Blackfire what is what we call the synthetic user monitoring uh based on uh scenarios and quickly show you this before getting back again uh to the health report you we you can define scenarios in a yaml file that the the root of your code base and uh in order to uh to test um on a regular basis here on using periodic builds uh like each hour
validates that my scenarios are ex uh meeting our my expectations. I won't get into details about it because I don't have time for this. I'm sorry. But you have like different constraints here, like about on the memory itself, on the number of SQL queries. um uh etc etc the number of tags there different things that you can define uh that you can test uh the and uh to for to to validate uh your performance profile okay um and this is important because this data is actually key for your health report We
uh because uh here in in this section in build and test, uh you will um you will see which well Blackware will extract which transactions are uh not tested by this kind of uh scenarios. And if you have quite uh high expectation, I mean um Um impactful transactions that are not tested, well, it means that you really should test them by this mean. Because uh these um mon uh synthetic in these bills uh they do profiles on each every every step. while monitoring it's uh are not profiling anything uh unless once a day and i will come to that
Right. So here we have quite uh quite a few um impactful transactions with time distribution. And also the top recommendations. This is something that is part of the profiler. Here it well it's all about PHP, but if I can have a look at this YAML parsing that should be cached in production. Yeah, uh well at least it depends on how it's it's taken into account by by the language, but uh here. We definitely have uh uh issues with that because we have to interpret the YAML file of here about SQL queries or RM entities, uh more too many uh RM entities that
have been created and i can view it within the the the last profile that triggered this recommendation and this is what we had here right Right. We had also this parts in our Python profile. Like here. Oops, I didn't uh enable the cache template loader because I still have the debug file. So this is actionable, directly actionable. You have this recommendation, you can definitely fix this. Even if it's in production, you can see it right away. Uh okay. Let's go back in. Okay. Then for the monitoring, as I said. Uh here you can uh visualize
um the graph, uh I mean the your application's constants. On the y-axis, you have two y-axis. On the left y-axis, you have uh time uh for each request. Um And so the basic roll time and the memory consumption. You can switch from one another. And on the right y-axis, you have the RPMs, which are the number of requests handled by minute. This or you can switch to the server load, like you know the one reported by top command, for instance, and the uh bandwidth output. So this can help you do a selection. So I already did a selection. You can also
see that these vertical bars, these vertical bars are actually events. That have been triggered. Like to uh when you deploy something, you uh you want to know if it had an imp a positive or negative impact on the overall monitoring. uh not the monitoring on your application so you so that you can have uh you know cue points yeah okay here i i did a deploy did it degrade the the performance yes or no And this is specifically what it's all about, all this virtual events. Then you can filter out. My HTTP status, by response time distribution as well. So we can see that we yeah, we have like uh more than 100
uh more than 1000 uh requests at our uh 1. 6 seconds long. So that's uh that's not good. So I can click to filter this out and then the response time visualization changes automatically To show well here the 96 percentile. It means that here the 96 percentile is 12 seconds. So it means that uh everything that is below 96% of the all the requests are below 12 seconds. And the major time is five seconds. So that's quite long actually. And then, well, I have the top transactions. So a transaction, as I said, is uh uh identified request. Uh we're not talking about request
anymore here, and it's sorted by impact. So here 80% in impact for the first one here, which is public keys actions. This is something uh uh an operation uh when you uh that is triggered when you want to trigger a profile it checks all signatures to uh to keep everything safe and I can see that I also had 44 errors so That's not good. Maybe we should have a look at it. And if I if I click here, I I I keep all my uh filters. uh to understand what's dedicated to this very transaction, right? And I even have SQL uh uh think
well queries here that are being extracted but this is uh an upcoming feature uh that i'm showing and the spans So the spans that's the mean timeline that are extracted from uh extended traces, that can that's usually one percent of all the traces, uh to to show you how your application is behaving but what is mostly important is here the automatic profile so Here from this part, you will get the possibility to treat your profiles yourself from real NGOs or traffic. So because sometimes you you know you don't you cannot really reproduce uh how um how things are going
And actually, Blackfire Monitoring for just do uh automatic profiling for you on the every day for the top 10. uh top top most sorry uh impactful requests transact most impactful transactions so every day uh money uh black air monitoring will trigger profiles And this is what monitoring uh monitoring automatic profiling is about. So you don't have to do anything. And I can then click on it and inspect what was going on to understand, well, uh yeah, what 's what where can be the problem. And well, you can also see say, well, actually I want to uh to profile the next request coming through this the uh um
uh this uh transaction Let's click on it. And then the next time a user comes through this transaction um a profile will be generated and you will be able still to understand what's how your application was behaving. Let's see if it worked here less than one minute ago. That's me So everything here is tightly coupled so that with a with a uh a logic within the things. Monitoring can lead to providing and profiling can lead then to monitoring from dev to production to give you real actionable impact. Uh and
and one of the best examples is the recommendations I already shown. Um and um You keeping the control on how your application is behaving from development to production to production. That was uh me speaking very fast because there is always a lot to show. Uh and so well I hope you enjoyed it. And um And I would be very happy to take any questions either from Slack or At our booth.
And oh yeah. Um about uh Blackberry monitoring for um Python. We're currently in a beta stage and we're currently looking for beta testers. So if you're interested to do the beta test with us. Please come by our booth and let's discuss. We have a form where I cannot share the link here, but I will drop it in the segones in the in Slack. um to if you if you're interested so that we can contact them contact you afterwards. Thank you for your attention, ladies and gentlemen.
Blackfire’s probe traces every function and records execution time, I/O wait, CPU time, memory use, and related data. An agent aggregates and anonymizes the raw data before uploading the profile for analysis.
Discussed at 1:42Yes. The probe remains inert until a profile is explicitly requested, so normal users incur no profiling overhead. When profiling is triggered, the profile applies only to the requesting user.
Discussed at 4:00Its timeline shows the call stack and the time spent in each function, making slow templates, views, queries, and other hotspots visible. The call graph and domain-specific metrics then help trace an expensive operation back to the code that triggered it.
Discussed at 10:08Yes. The timeline’s memory graph shows peak memory consumption over time, which can make memory leaks and other excessive-memory problems easier to spot.
Discussed at 13:12The health report uses real request traffic to show response-time percentiles, request volume, errors, and how those measurements change over time. It also identifies the transactions with the greatest impact on users and lets you inspect their traces and profiles.
Discussed at 18:32You define scenarios in a YAML file, and periodic builds execute them while checking constraints such as memory use and SQL-query counts. These synthetic profiles can reveal impactful transactions that your scenarios are not yet covering.
Discussed at 23:58Yes. Blackfire Monitoring automatically profiles the top ten most impactful transactions each day, and you can also request a profile of the next request reaching a specific transaction.
Discussed at 30:52Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025