Profiling Django & Python apps
Published June 27, 2021
This video features Sümer Cip at DjangoCon Europe 2021 in Online.
Introduction and Outline
Quick introduction of myself followed by an outline of what will be covered and what you will learn.
Why we profile?
A typical program spends almost all its time in a small subset of its code. Optimizing those hotspots is all that matters. This is what a profiler is for: it leads us straight to the functions where we should spend our effort.
What types tools are available and how they work?
-Deterministic Profilers
-Statistical Profilers
I will be walking over the different use cases, pros/cons for each type. Then I will dig in a bit deeper on how they work under the hood. Understanding the inner workings a bit might be helpful while analysing its output.
What kind of information do we get?
I'll describe what kind of output we get from different profilers. What kind of metrics are available(Wall time, CPU time, sub/cumulative time) and where are those metrics are most useful while hunting for performance problems especially
for Web applications. I will also walk over/explain different kind of visualisations that profilers generate: Flamegraphs, Callgraph, SpeedScope...etc.
Questions
Performance work should begin with measurement rather than guesswork or premature optimization. Sümer Cip explains how tracing and sampling profilers collect data, the differences between wall, CPU, I/O, and memory time, and how tables, call graphs, flame graphs, and timelines expose bottlenecks. He surveys Python tools including line-profiler, Scalene, Yappi, PySpy, tracemalloc, Django Debug Toolbar, and Blackfire, then recommends monitoring throughout development and production, examining architecture and data access before attempting micro-optimizations.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: Hello, welcome to DjangoConf 2021 everyone. My name is Shumanji Pen. I work for a company called Blackfire. io And we specialize in producing performance -related observability tools for various languages And as you might guess I work on the Python side of the things and today I will be talking about which and how uh how and which tools to use for uh debugging performance issues on Python code. Let me just say that this will not be a Python or Django performance tips and trick talk. While I will try to briefly touch on some general rules, most of the times I will try to focus on what kind of tools we have to identify and debug performance issues
Speaker 1: and how they work under the hood. So as can be seen in the outline, I will be talking about why we need these tools at the first place and then how they work under the hood And then what kind of data do they produce and some common formats, patterns that are used to visualize its output And at the end we will probably walk over some of these tools available. In my opinion, better understanding of these observability tools could aid more on diagnosing performance issues than to just strictly following some performance uh rules or tips and tricks.
Speaker 1: So I would start by this very famous quote from Professor Nutt. At least I heard this quote from time to time in my career multiple times, and it is so correct Doing pre-mature optimization not only did didn't solve your performance problem, but it also introduces unnecessary complexity to the code. And in my book, if there is one thing that all developers can agree on, it's this uh avoiding complexity is better. So I really forget the number of times I do various optimizations without doing any measurements and see at the end of the day it didn't matter that much. Maybe maybe one of maybe some of you are better than me at this, but I have seen this enough times to believe that we humans
Speaker 1: are very bad at estimating uh performance bottlenecks. So I really believe in this code from Professor Nut by heart. But the thing is this code here is not complete. It's probably the most seen part of it, but it's not complete. So here is the full code So as you can see, Professor Not clearly says to do not do premature optimization. But you should do it if you can measure the top three percent well enough. Like um I I think this changes and updates the perspective about the underlying message of this code, which is if you can find the top three person
Speaker 1: were top three person well enough Then you should try to optimize it. This is what I understand from the full code. So to chase this top 3%, I think it's it's key to understand what kind of tools are available for us. And how can we use them? And in order to better use these tools, I believe we need to first understand how they work at least at a very high level, and then we need to have a clear understanding of What kind of information can they provide to us? And this will hopefully be the main topic of my talk today. So before going any further, uh I'd like to uh rephrase something again. Which is you cannot improve something you cannot measure.
Speaker 1: So measuring is key. And the tools we use today to measure code are called profilers and um Okay, let's start from the beginning and uh let's go and search for some open source profilers to see what kind of options we have in the Python ecosystem If you do that, what we will see is yeah, we will see that the answer might not be that easy. This screenshot here is from GitHub by searching for keyword Python profiler. And as you can see, there are over thousands of profilers. available in the Python ecosystem and can you can clearly see that the EDs are not doing the same thing.
Speaker 1: Some of them are sampling profilers, some of them specialize only on memory usage, line profiling. uh etc etc and I will try my best to explain the differences between these. Um what are the basic types, what kind of data do they provide and in what cases they could be useful to us, especially for web application development. And by the way, this knowledge here is not only relevant to Python, for example, the Go language profilers also shares the same foundations that we will hopefully cover today. Okay, so the first thing I'd like to talk about is that there are three basic parts of any profiler that answers
Speaker 1: following questions And the answers of these questions are what that makes a profiler unique in its feature set So what is the data collected that uh by by a profiler? Um how is it collected? And how is it presented? For example, the answer to the second question here categorizes the profilers into two basic types, sampling and tracing profilers. Let's try to understand that a bit more. So I'd like to use um following analogy to explain uh the basic difference between these So imagine there is uh there are two coaches and an athlete.
Speaker 1: The athlete is constantly running laps and the coaches are watching from behind. and do doing some measurements. Like the first coach, let's call him Bob, is constantly monitoring the athlete and do doing measurements with his wrist roach. Whenever the athlete finishes a lap, he he he saves these information. The second coach, let's call him Smith, does the same thing. But he does the measurements uh whenever he likes. He constantly talks on the phone, do other stuff and then when he feels that it's the right time he looks at the athlete and save the current position to some notebook he has. So over time he makes more and more of these uh measurements, but at the end of the day he could be able to statistically calculate how fast
Speaker 1: his at was his athlete and uh what was the average lap time, right? So the first coach, Bob, was doing measurements constantly and thus is more accurate than the uh than the second coach Smith But Smith not only measures his athlete but also finds time to talk with his wife on the phone maybe. So coming to the reality we can say that Bob, our first coach, is a tracing profiler A tracing profiler does measurements every time when a certain event happens inside the language runtime, and these events are defined by the language runtime itself. This not only applies to Python but also PHP, Ruby or any other language that supports profiler hooks.
Speaker 1: And on the other hand, our second code Smith is Very similar to a sampling profiler because it samples the a sampling profiler is just samples the call stack of the application at specific intervals and saves this information somewhere and aggregates those information at the end. And as can be seen from this analogy, there is a trade-off between the two Bob, uh the first coach can provide more measurements and more accurate data, whereas Smith will work less, which means in competent terms will consume less CPU power So like I mentioned uh for for uh a tracing profiler there is some profiler events
Speaker 1: and for python Those events are for the time being Python supports hooks for every function, line or opcode executed. So this means you can do measurements whenever a function call exit happens or a line has been executed or even uh a single Python of code has has been executed. So the tools you might uh use uh you you the the the tools you use might change depending on uh what you would like to measure Like maybe you would like to optimize a single function that you know that it's heavily cold, or you would like to see which functions are most time consuming in your web application or your CLI script etc. And uh or more. Uh although I haven't seen any example, um
Speaker 1: I'm informed that Instagram has used some form of upcode profiling in their fleet, production fleet. And um Uh yeah, th there might be maybe times too that you'd like to see opcode um r work load as well. So uh yeah, the these are the hooks for the uh Python runtime. And uh coming to the sampling profilers, uh they they like I said they took another approach uh for to optimize for overhead and uh as you can see It directly looks at the costax of threads of the application and accumulate this call stack information at specific intervals, like the Coach Smith's note from the previous analogy It might seem inaccurate at first, but given enough time, statistically you will end up with an accurate view of what happened inside your application.
Speaker 1: I also would like to talk a little bit about the difference between internal and external sampling profilers here, because it's it's I believe it's an in in important distinction. The internal means that your your profiler runs in the context of your application and external means that it runs as a separate process. Uh this is this is for me important distinction because when you run the profiler externally um this means that it is safe to use um If it is safe to use in case your profiler itself has any bugs in it and it's impossible to somehow crush your application. And you can add it to your container system without changing a single line of code in in your application logic, which is nice
Speaker 1: since it can give you the benefit of separating the instrumentation library from the application logic in a maybe Docker file Yeah, so next time you use uh a sampling profiler, please give attention to this distinction as you might have different options there. Okay, um I hope that we now have a high level of understanding on how these profilers work under the hood. but and what what are the trade off uh trade-offs between the two but we haven't yet touched on um what these profilers actually measure So they obviously measure time, right? But the thing is
Speaker 1: there are different flavors of time when it comes to computing. The the first um yeah the the the first thing uh that uh comes to mind is uh the wall time actually uh which is like having a wrist roach. It's it's just uh just like the analogy we have with the coaches. The second one is the CPU time And it is it is this total amount of time your CPU actually spend on your code, including the kernel time. So if you have a CPU-intensive function, like as example, for example, like Cryptographic function uh that will consume much much more CPU than a simple function that reads a small text file
Speaker 1: uh for example yeah and cpu time is usually calculated by the support of the cpu itself it has some internal counters for for itself And there is this IO time. Let's suppose that you read a file in your code. In most simplest terms, your CPU will execute some instructions to read it from the disk. and then it goes to sleep. It waits for the disk to send the data it requested and when the data is available it will return to the application. I know it's it's it's kind of simple. Yeah, so at this scenario the whole time spent uh from requesting the read until the data is available is the whole time. And if you subtract the CPU time from the wall
Speaker 1: time, what you will get how much time your code actually waited on IO, which is IO time. And I think it's it's a very important metric for web applications because a typical web application uses usually is IO-bound and do lots uh lots source of I. O. And it would be very useful to see exactly where these I. O. bottlenecks are and try to optimize them, maybe by caching to memory or a faster disk medium, etc. And there are also some more metrics that can be useful as well. Like memory is one of them. Memory profilers aim to show you the memory consumption or peak.
Speaker 1: Uh and with using a memory profiler you can see how much memory does your application consume or peak and where is where it is allocated. Actually you can see the consumption and peak metrics via any CLI tool for any operating system. Right? Like you can open a process monitor in Windows and for macOS it's activity monitor. I don't for you can for Linux you can use top But those tools will not give you the information on where the allocation is coming from. If you have ever debugged, I mean I mean you cannot see uh the inside of your application your actually where it is allocated from from your code. If you have ever debugged a memory leak issue before, you will know how frustrating it could be to actually identify
Speaker 1: the exact source of the leak and these tools might really shine on those occasions. Like Trace Malloc comes in the standard library, for example, from if I remember correctly 3. 4 and up and it traces all malog free and realloc calls and saves a python trace pack along with the allocation And uh yeah, there is also this object ref library which uses uh the internal uh GC module to find out which objects are active on the heap, and you can uh we can you can easily see what kind of objects are currently active and after your request and maybe do a difference between them and see see what happened. Yeah there are there are very different um uh strategies to identify memory leaks
Speaker 1: And there are also application specific metrics as well. Like an example on this type is a famous Django debug toolbar which can profile Django specific metrics uh for example template render times and sql queries and and blackfire which is the observer titulate platform that I'm working on. uh that that can also provide lots of different application met specific metrics as well as like SQL queries, HTTP calls being made to different microservices and lots of other metrics as well So we have uh touched how profilers work under the hood and what kind of data we gather with them So now comes the question of how these profilers present their information.
Speaker 1: And in short, there is no common well-defined answer to this question. While there are different visualizations being used by different tools, I believe it's still an open problem. And I also believe that some visualizations might work better than the others for some problems. So it really depends on personal preference and the type of the problem itself. These are the common output formats that profilers usually use and this list is not only relevant to Python. Sometimes they might be called differently or there might be diff subtle differences between them, but the way they present the information doesn't change too much.
Speaker 1: So let's start with the simplest one actually. A table is just like its name implies. It just shows you a color callie pair in a row in a table For example, for a function tracing profiler, multiple calls get will get aggregated to a single row. Usually all profilers can output something similar to this Yappy C profile PySpy Trace Malloc BlackFire. All kinds of different profilers can actually support this kind of output. And There are tools like Key Cache Grind. If you would like to see and filter based on the output, most of the profilers there's some
Speaker 1: way to output as call grind. Maybe by using a separate tool. And here is a screenshot from KeyCache Grind. It's pretty easy to see the most cold or time-consuming functions with this visualization at hand. And it is probably the most uh use most used and the oldest probably. And here is a screenshot of a profile that uh For a Django application made by BlackFire , all these table view visualizations have some kind of exclusive and inclusive time concept, like you can see. Um exclusive time is is the time that's spent only in the function itself, whereas the inclusive time means the time spent
Speaker 1: in that function and all it is children. So in this view, for example, you can sort based on those metrics. The second visualization I'd like to talk about is a call graph, which is probably the second most famous one. A call graph shows you the color coli pairs in a graph. So you can see whole your whole application's called relationships in a big call graph and this visualization might help you to dig deeper at the problem and identify the root cause because you can navigate through the call relationships easily. And here is another example for the call graph from Blackfire UI Note that the
Speaker 1: call graph is not always very UI friendly because there are many many calls. There might be many many calls. And when there are lots of functions being involved it's hard to focus on the most interesting path. So I would suggest that using it to debug performance issues further rather than trying to spot the the issues itself We at Blackfire, for example, try to remove uninteresting calls, what we call noise, and aid the call graph with a table on the left view from the profiler data to highlight the critical path But even in that case, I must admit that there might be lots of noise and the thing is the cold graph is might not be the perfect visualization for every case. So another very common visualization is born out of these needs
Speaker 1: to reduce this noise that I mentioned and it is obviously the flame graph and when you enter it um you would see lots of tools are using it like a flame graph is called um uh flame graph because it uh it resembles flames obviously uh and there is an icicle graph uh which is the same data but upside time uh because it resembles ice uh And I have a friend that calls these icill graphs pit graphs because they seem like pits. So as you can see there is this pattern for disagreement here. So even though we did didn't agree on a proper name For this visualization the presentation and format is pretty consistent actually.
Speaker 1: If there is one format that's very close to being standard in this area recently, it's Flame Graph It is invented by um Netflix performance engineer Brandon Gregg and it's being used extensively since then. You could see it in action for any in any modern browser for for filing JavaScript code for example. Well there are different flavors for frame graph. It is usually generated by aggregating call stack data over and over again like we have seen in sampling profilers. We can even say that it is a natural output for sampling profilers The x-axis shows the times of functions seen on the call stack, and the y-axis shows
Speaker 1: caller-call-y relationships. With this visualization, it's very easy to see the interesting functions with a single look. Long empty lines, for example, indicate lots of exclusive time for this for the function. And the colors, uh yeah, the the the colors are just random. And uh there is also one more thing that I'd like to uh show you about the Flame graphs. In this previous flame graph example, what you see is a final image is probably an image that is generated from thousands of call stacks sampled and merged together But if you add the time dimension to the x-axis while retrieving that information, then what you will see or what or what what you will end up is that
Speaker 1: at what point your code is executing which function, which is similar to this one we see here. And this is heavily used in modern browser profiling toolset. At Blackfire we also use a flame graph with time dimension which we call timeline. As you can see it's pretty similar to the one seen in Chrome. And by the way, the blue background here shows the peak memory usage, which may help on situations where you're peaking memory too much. Okay, let's look over some examples on the ecosystem to understand what these tools provide for us. I'm not in any sense claiming these tools are the best.
Speaker 1: I just tried to show you most of the features available by these different types of profilers. So the first one I'd like to show you is that line profiler. Line profiler is a line tracing profiler. It measures per line volt time information and you can inspect its output from the command line. As far as I know, it's highly used in data science world, but you guys might know better than me. Scaline is a new sampling profiler that profiles per line information again , but it measures per line volt, CPU, and memory consumption simultaneously.
Speaker 1: And you can also inspect its output from command line X like seen here. Yapi is another tracing profiler. It's a function tracing profiler. It can measure per function wall and CPU time at the same time. And you can profile multi-threaded and async applications. uh and show uh show per thread uh trace information. It supports profiling async IO and G event applications currently. And PySpy is um is an external sampling profiler, uh which we have talked before, and uh that means that as um you can attach to any running Python application and dump
Speaker 1: traces uh on the fly with a top-like CLI command uh sor sorry command line interface and if your application is multi-threaded for example It can also show you how much time your application waited on Gill , which is a unique feature. If you have ever wondered multiple threads are uh suffering from GIL contention, you can use PySpy to detect this easily. It can output to flame graph or speed scope formats. By the way, we haven't uh we it's mentioned here, but we haven't talked about speedscope Uh it is an open source format that's very similar, or we can just say that it's it's the same as flamedraft with time. uh that we mentioned earlier and uh
Speaker 1: it has a JavaScript application which means that it works in a browser So you can do a profiler with a profiler that supports speedscope and then throw the output to the speedscope application you can see the traces in your web browser. Here is an example screenshot of a speed scope format. And here is a flame graph generated from PySpy, right and uh on the left side And the right side is actually the interactive viewer where you can directly see the code executing in the active thread When you attach to a running process, for example, you will see this gets updated pretty quickly with different functions during execution. If possible, PySpy
Speaker 1: will also try to retrieve this C extension call stack as well I said if possible because it depends whether the extensions compile with debug simple or any kind of stack trace information. And uh I also would like to mention about trace malog. Um tracing trace malog is a tracing, a memory profiler. You might have heard of it or maybe even used it because it's in the standard library for a while right now. It basically traces all low-level memory functions like malloc. uh realloc free calloc and that that that those C runtime functions and it saves uh a Python traceback object along with the allocation for later analysis.
Speaker 1: It can be a very valuable tool for analyzing which parts of your application is consuming or maybe leaking memory, because at the end of the day you will see how much bytes gets uh allocated per line, which is uh very uh valuable information. Uh however it needs to be enabled explicitly in your code and it will add some overhead to the application. as it basically hooks every memory operation. Like even you add uh some key to a simple dictionary, uh it will call uh pro it will probably call Maloc and uh Yeah, it will it will get uh some hooks inside the trace malloop. So if you want to use it on production, you are warned, you should be cautious about the overhead and calculate uh beforehand.
Speaker 1: And BlackFire. BlackFire is an on-demand tracing profiler that can measure ball I. O. CPU and memory uh metrics simultaneously, which is a bit different than a normal tracing profiler. For example, for HTTP profiling scenario, it is only enabled when a specific HTTP header is present, and when it is done, it unregisters the profile hooks. and you will have zero overhead for the um uh for for the other requests. So we can say that it is safe to use on production as only it only affects the profile request itself You can also run it on standalone CLI scripts,
Speaker 1: but it will be enabled during the whole execution in that case. It also is able to profile some application-specific metrics like SQL queries, cache calls to common cache frameworks, and HTTP requests being made to different microservices. some template and OIRAM related metrics, etc. We basically support most of the visualizations we have seen so far, so you can select and switch between different visualizations uh for the same profile to analyze the problem at hand from different angles. And even though I said I will not be giving any general tips on performance, I think I will make an exception for these ones So if there is one thing to take this uh
Speaker 1: take from this presentation, um it's this always measure If possible, always monitor for performance from development to production. Having an all-development lifecycle will also be more helpful to diagnose the issues earlier. And the second uh thing I would say is to focus first on architecture design and algorithm being used. When we first see a performance problem as humans, I identify the pattern among the people I have worked with, and myself included, is that we tend to we tend to hurry on fixing the issue right away without too much attention. However, uh uh maybe there is a fundamental design problem behind a seemingly subtle
Speaker 1: performance issue Uh I have seen this many times before, some wise uh refactoring this decisions being made after analyzing patterns for some performance problems. So at the end I think in my opinion these issues might help you design better systems if you give you give them enough attention or listen to them and to think more broadly. The third one is every application access has come some kind of data And an application uh access data is very uh application and how an application access its data is very specific to the application. For example, you should know that
Speaker 1: searching for an item in a sect, set or dict is very fast. It happens in constant time But if you are using lists internally and constantly removing and inserting into the list, then it might not scale very well at the end. For Django code, this concept might be applied to models, for example. Like you should avoid referencing a foreign key in a loop without a prefetch because it it might cause an N plus one problem. Or maybe you should try to avoid retrieving too many objects at a time from the DB, etc. You know it. So always uh know your data and how you access it
Speaker 1: And uh the the fourth thing is check if this standard library has a function. Yeah, again it's so important. I mean Uh Python's standard library is very, very powerful. It probably has a function um that does what you want. So if you have a function that seems to the to do lots of stuff that is not related with your business logic, then uh it's highly probable that uh there is a built-in function that does w what you want. Um so Just read Python docs some sometimes in your free time, maybe. Make sure you know and memorize most of them. And they will probably be highly optimized Uh not probably they are very highly optimized. And there are very gem
Speaker 1: modules like EaterTools, Collections, Inspect, OptParse, Globe, Struct, Re, Difflib, and yeah, many more. And I have said this before, but uh please just avoid micro-optimizations. A single dict access is pretty fast use usually and Python is a very high language high-level language and In most of the times you you can't you can't see any benefit from these kind of small micro optimizations. And the last thing is, if you really need performance, then there are tools. There are Cyton, Numba or other third-party uh libraries as well, or that you can directly write code and see. For example, with Cyton
Speaker 1: you can easily write C extensions with without dealing with C. Numba is also worth mentioning with a simple function decorator, you can optimize your CPU-intensive functions. But again, you need to measure in it JIT compiles your code. Not every function can benefit from this kind of optimizations. usually heavy numeric scientific calculations are perfect but yeah yeah for a typical web application that accesses database and cache or IO-bound uh applications cannot benefit from this kind of optimization. And the last one is you can always write a C extension. I always believe that C API of C Python is a bit underrated. It's a very well-taught out
Speaker 1: uh consistent and robust framework in itself And the documentation is awesome. You can always fall back to this. So that was all I want to tell actually. I hope that you enjoyed it and learned something from it. If you would like to discuss anything performance related, I would love to hear and discuss. I will be at the Blackfire booth. Thank you. Bye So
Speaker 2: maybe you can answer Michael.
Speaker 3: I'm sorry.
Speaker 2: Maybe you can answer Michael's question.
Speaker 3: Yes. Okay. I believe I understand what which which one do you you're saying?
Speaker 2: In in the chat box.
Speaker 3: Yes. Just a second.
Speaker 2: The question is uh is it possible to get full trace while in production? I mean it would affect your performance. And the first question was about what is the difference between monitoring and profiling interaction.
Speaker 3: So it's it's a very good question. So in in Blackfire terms it's it's pretty different. So profiling for us means that Whenever you start the profiler it uh instruments all your um code and it just um It just profiles all all the calls happening in inside a single HTTP request. Whereas when you enable the monitoring tool , The monitoring tool is just um monitors your HTT request. It's HTTP specific, like um It monitors how much time uh should be spent on a on an on an a uh on a single request and then aggregates this this data and shows in a
Speaker 3: you know in a single dashboard. Whereas in profiling you would have more more and more data in a single request up to any single function calls being being executed. So that's the that's the main difference between the uh monitoring and profiling.
Speaker 4: So so I sorry I interrupt. I in my mind I have that profiling is before deploying and monitoring is after Because so so for example uh for my projects I use Django debug toolbar uh because in order to use it uh because Django debug toolbar uh makes some uh some operations which are uh not Costless, I mean without cost. And when in production I use monitoring tools and logs analytics in order to find the bottlenecks And to to try to find solutions. But I haven't heard of profiling in production. That's my question
Speaker 3: Yes, that's actually uh that's a very good question and this is the tradition this is the uh basic the traditional way of doing things. But There is a current trend in observability tools that to do to do continuous profiling to include in the production as well. And in our case, we do the profiling but we only do it in for a specific specific HTP request. So uh in Blackfire you just um call click a profile button it will send a sp special http header to the web server and in that web server we have a special middleware automatically inserted And then we re-enable the probe for that specific HTTP request and nothing more.
Speaker 3: So that that will add no zero overhead to the application because It will only be profiling that thread, that specific request, not other other the other not the other concurrent ones. So this is a new approach and there are some other approaches as well like doing some using sampling profilers at some levels. But yeah, we will see most uh we will I I believe that we will see lots of stuff happening in this area like with continuous profiling in in production as well
Speaker 4: Okay, thank you very much.
Speaker 3: Yeah, I hope
Speaker 2: in addition with the Blackfire monitoring, um It it will take all um uh the the top t the top ten uh most impact impactful transactions that uh is being recorded in in the last day. And we'll send uh well do a profile with real end user traffic. So only for this 10 top top 10 transactions once a day. BlackFire monitoring will generate a real end user profile so that you can have a real insight of what's going on with these very transactions. So that will uh complement
Speaker 2: um what you're getting for um uh with the the the monitoring itself But it will definitely be not be on all uh transactions, all HTP requests. It will be it will cost way too much for this.
Speaker 4: Okay, okay. Thank you very much
Speaker 2: Welcome.
Speaker 3: Is there any other questions? So I believe that's it. Thank you very much for your attention. Thanks for the question. I hope you enjoyed the slides. I hope you learned something That's my uh so thank you very much again. Bye
Speaker 2: Krisuma. Bye.
Speaker 3: Thanks
Tracing profilers record measurements whenever runtime events such as function calls or executed lines occur, so they are detailed but add more overhead. Sampling profilers periodically capture the application’s call stack and use the accumulated samples to estimate where time was spent with less overhead.
Discussed at 7:17An internal profiler runs inside the application, while an external profiler runs as a separate process. External profiling can be safer if the profiler has a bug and can often be added without changing application code.
Discussed at 10:28Wall time is the total elapsed time, CPU time is the time the processor spends executing the code, and I/O time is essentially the time spent waiting for operations such as disk reads. I/O time can be especially important for web applications because they are often I/O-bound.
Discussed at 11:58System process monitors show overall memory consumption and peaks, but memory profilers can identify where allocations originate in the application. Tools such as tracemalloc can associate allocated bytes with Python tracebacks or lines of code, which helps investigate leaks.
Discussed at 14:14Common formats include tables, call graphs, flame graphs, and time-based flame or timeline views. Tables make it easy to sort inclusive and exclusive time, call graphs expose call relationships, and flame graphs reduce noise and highlight functions consuming significant time.
Discussed at 16:37Line_profiler measures wall time per line, Scalene reports line-level wall, CPU, and memory usage, and cProfile measures function-level wall and CPU time. Py-spy is an external sampler that can attach to a running process and produce flame graphs, while tracemalloc tracks memory allocations and their Python tracebacks.
Discussed at 23:40Monitoring aggregates request-level information, such as how long HTTP requests take, and presents it in a dashboard. Profiling instruments a request in much more detail, down to the functions being executed and the data associated with them.
Discussed at 35:34Yes, but the approach should limit overhead. Blackfire can profile only a specifically requested HTTP request, enabling its hooks for that request and leaving other requests unaffected; sampling profilers and carefully selected continuous-profiling approaches can also be used in production.
Discussed at 37:35Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025