Lightning talks
Published June 7, 2023
This video features Ahter Sönmez at DjangoCon Europe 2023 in Edinburgh, Scotland.
The Inevitable Tech Incident: The Lessons We Just Can't Seem to Learn
by Ahter Sönmez
Incidents and outages are an inevitable reality for software engineers. While there are always many lessons to be learned from them, there are certain lessons that are often overlooked.
Have you ever been in a tech incident? The kind that leaves your team scrambling, your boss furious, and your customers frustrated? They are inevitable. But do you ever wonder why we keep making the same mistakes?
I've seen my fair share of outages and incidents. And while we always walk away with practical takeaways (and that's great), there are certain lessons that just seem to slip through the cracks. The ones that we should have learned from before, but somehow, they keep happening. I've observed these patterns that make up the lessons (unfortunately) "not learned" from incidents and outages.
In this talk, I'm going to shine a spotlight on those lessons we just can't seem to learn. I'll delve into the psychology of incidents and explore some attitudes towards monitoring that need a serious overhaul. I'll share some practical practices that will help your team stay ahead of the game and avoid that all-too-familiar panic.
So buckle up and get ready for a self-therapy session like no other. It's time to face the truth and learn from our mistakes once and for all.
Tech incidents are inevitable, but teams often fail to turn them into lasting learning. Ahter Sönmez argues that asking better questions and looking beyond individual mistakes to team and organizational practices helps carry lessons forward. He recommends designing deployments so they can be rolled back or mitigated, defining incident ownership and boundaries in advance, building monitoring and alerting into production work, and writing reports for the people who need to understand them. A blame-free culture depends on leadership treating technology failures as opportunities to learn, while alerting should be calibrated to the risk of missing a real problem.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: Thank thank you. Thank you. So yes, we are going to talk about incidents and you might have heard Tim's wonderful uh talk yesterday about incidents. This is a different take. So um yeah, we are going to t talk about lessons not exactly learned from outages or incidents. And yeah, it's it's it's it's kind of inevitable that we'll have these but um sometimes We struggle to learn things. So we are going to talk about those things. I have a few disclaimers to make. First of all the word incident and outage they are
Speaker 1: kind of different, but we'll we'll use them interchangeably. So uh I think you'll you'll understand and second disclaimer is some of the things might be a little bit uncomfortable to digest uh but bear with me uh as They are uncomfortable but they'll be useful, hopeful in the future. So um before I start uh I'd like to see a show of hands. Who has been in an incident in a tech incident? Okay, uh that's that's most of you, so you know what I'm talking about. That's that's perfect. Uh the second thing I'd like to know is who has caused an incident
Speaker 1: Perfect. So you are you're on the right track. I I congratulate you. Um so there might be a few things to take away from this talk and I hope uh what What the questions I'll ask will enhance your understanding and we might avoid maybe avoid few things in the future. We'll never probably avoid them all. So About me. We are great. I I encourage you to apply if we are recruiting. So we have a boot. And we have lovely recruiters, please go and talk to them if you're interested. So here's the story. Um how
Speaker 1: this talk was born. So about two years ago I've I've been into lots of incidents really uh before Kraken and During Kraken , lots of incidents. And one day, maybe it was a rainy day. Maybe I wasn't feeling my best. But again an incident happened. And this time I was so frustrated. I was so sad. I I I I I I was just I said I I can't I I need to speak with this some someone. I mean I I I need to lay down uh like I need a therapy I need to talk about this. And uh the only thing at that time what I could do was I I started to write a blog post, but you know We all know blog posts take a bit too long.
Speaker 1: So after two years, uh the blog post hasn't finished, but I am giving this talk. So Um hope you uh enjoy this. So before I I jump onto the actual lessons and some of the findings or my observations, I'd like to talk about two fundamental things. And These are fundamental because it will help you understand where I'm coming from and what I am trying to say. The first one is about Level of analysis and level of analysis is is my lens about how I look at things And there are set different levels of analysis. I'll talk about them. This is a concept from social sciences. So and this kind of changes how we see things.
Speaker 1: And second thing is about asking questions and it's much simpler. So I'll start with the second one. I'd like you to meet my mentor Robin. Uh he's my son. Uh my son Robin, he is three years old, and he doesn't have the cognitive biases we most of us have yet And this means he has this non -restrictive per perspectives about anything. Anything is possible, and um the approaches to everyday phenomenon is just uh limitless. So that's what I needed to to understand this. And what I mean is he has this endless in interrogation of reality. And what does this mean? So
Speaker 1: I'll give you an example. So he observes something in in his standard or in his life and his trip, okay? And then he starts this interrogation game and interrogating me usually. And like let's say he he saw the car pass at red light and he says like Daddy, Daddy, the driver didn't stop at red light. And she wants me to explain because he says like red light stop Greenlight go. And I say perhaps the driver was in a hurry. And he says um Why was the driver in a hurry? And then I said perhaps he left the house a little bit late
Speaker 1: And then he asks, why did he did he leave the house a little bit late? And I say, m perhaps he didn't realize how half time passed. And he says Why didn't he realize the time has passed? And you you you see the thing. This goes about maybe 30 minutes, 40 minutes. And eventually we really We really get to the fundamentals of things. It's it's fundamentals of of the questions. And one of the things I've learned through this inter interrogation is what questions you ask matter a lot. Because It is the question that determines the answer. If you don't ask the right question, you don't get the right answer. And in other words, what you ask changes the answer.
Speaker 1: And In my opinion, uh eventually it boils down to human psychology, molecular or evolutionary biology, or um maybe uh atom science. But we'll only talk about the psychology bit here. This was the first fundamental. The second was level of analysis. How do we look at things? What's our perspective? And We have three levels of analysis. So the first one is the developer, the individual. I see things as an individual and my learnings are my individual learnings. And the second level of analysis, one step above, is the team. We we have it under we have the understanding in a team.
Speaker 1: And finally the organization. And while these three things seem related to you, actually how we transfer knowledge from one to uh to the other uh is not that simple and many many people are trying to answer this question so um Before I jump into actual practical lessons, I I'd like to say that how do we generalize the learning from the developer to the organization? Because in literal sense, teams or organizations or buildings, you know, they're they're abstract things. They don't learn. We are we learn. But we might not always be there when the next incident incident happens. So we must have this team or organizational level uh
Speaker 1: understanding so that we could kind of transform our understanding or psychology or or d utilize things we we we see every day into a strategy or a plan actually we we have a plan when the outage happens and that's what I'm trying to say. So Uh the real root cause is maybe being human and yes, uh no need to say more about this really. And this is the end of my presentation. Thank you very much No no j I'm just joking that I'm just getting started. So uh please uh so we have an agenda here Uh I have defined these into three categories and the numbers represent
Speaker 1: the number of actual findings. Uh these will be much briefer. Um you you might group them differently if you want, but I'll start with the psychology uh thing. So I I'd like to ask you a question, and this is an analogy I I'd like to make. So Suppose you're you're a diver and you dive and you maybe you are seeing this undercave underwater caves. What do you optimize for? What do you hope for? Um and Maybe I hope that I I can get out uh when I'm entering a maze or a cave underwater. So Maybe this is not uh
Speaker 1: the uh the simplest thing, but I'd like to say I'd like to see deployments that way. And maybe Maybe if you think about how you can roll back or revert or mitigate the situation , maybe the solutions will reveal themselves to you. Yeah, this this this this is the first analogy I'd like to start with. They will get more and more practical as I go towards the other categories, but bear with me here. Oh sorry, this is a glitch.
Speaker 1: But not common. And we'd like to I'd like to ask why don't we share about this the details of incidents so others could learn We do write reports and we do maybe share the reports on Slack or something, but do we really talk about the details and Um d do we talk them openly and um I I I don't know. I I have a whole section about this uh uh about reports and things so I'll I'll skip those things. And this is the most cliche thing ever, but uh Are we really talking about how we can avoid our code getting more complex?
Speaker 1: And this is a really psychological thing because you know if we all go through these PRs and code reviews and coding conventions, of course But we don't really take the effort to make things simple sometimes, and this is still a psychological thing because I think you can Overlook things once, twice in a PR or so. But if you don't remind yourself every day Maybe you end up with a complex and complicated system and that's not what you want really. So the next thing is the incident definition Well not everything is an incident. A page might be slow, it might not be an incident, or it may be an incident, but we
Speaker 1: should really have um clear boundaries around what is incident or not. And the the reason this this is important is if if we don't have these clear definitions maybe it results in unnecessary discussions conflict or just we lose time and not escalate things So as I said it will get practical. So this is one of the practical things. This is about psychology of acting. So when the incident is over, we say, okay, over, back to normal. But Is it back to normal really? Maybe we need to act fast and prioritize things.
Speaker 1: Because we have to do something to maybe avoid those those uh short shortfalls in the first place. And I think this is the last point about the psychology thing. So um There's a blame culture and I think this is embedded in our DNA and it's who we are because as a person you This is like an evolutionary thing, but um and we we tend to hide incidents or we we are uncomfortable about them and I I think that's totally fine but when the the the the The culture when you talk about a blame culture culture is is a is a concept related with organizations or teams
Speaker 1: and this should not be part of the culture. It doesn't help, it has never helped And we should probably embrace embrace the uh the fact that we make mistakes and we kind of learn from it uh and hoping we won't repeat them again. So yeah, there there we uh you know making mistakes we don't suppose we'll make mistakes but uh there's always Um we should probably uh expect unexpected sometimes. We are not AI and probably AI couldn't have written this the song I had in my mind while I was r writing these slides. So um so yeah I I think this this was the
Speaker 1: uh this was last thing. So let's jump into more practical things. Uh I'm aware of time so So we do write instance reports. Maybe you do write them too, and this is really encouraged to to share information. But About audience. So are we really writing those reports with audience in in mind? And I have lots of questions here for you. So So why would a developer read an incident report? And yes, there are reasons, of course, but um is reading incentivized for developers? And maybe you have
Speaker 1: one outage a week or one a month, but how about five a day And are are the reports b just being written because they have to be written? And Are we writing them to the right audience? Maybe we have to write them for developers too, not just for stakeholders or higher upper management or other things. You know, those are different audiences. We can't expect to have the same things in those different reports. I just curious who writes incident reports here? And things wrong go wrong. I encourage most of you to do it because even if it's for yourself, for your future self, you'll be quite thankful for it.
Speaker 1: Or if your team grows, they'll they'll be thankful for it. Um and the other thing is understanding. So suppose you you have the reports, do you understand them Is understanding in sanctificed and do we maybe give the time give the time and resurs to people to understand them and internalise them and digest them. Well, I'll tell you what. Vi we really fail uh with this in my company so. And it's it just becomes uh one of the 200 things that they'll probably have to do in their daily lives, unfortunately. And We don't really have someone who follows this after
Speaker 1: process and make sure people have understood it, has have embraced it. And we can't even measure the understanding. I'm I'm not sure I'm not suggesting practical tools here, but you know, we we just um don't don't know. We don't even know who to convince that we should have this as part of our organization so that we should understand the reports. And the last thing in this category is incident ownership. Yeah, who who is the ownership? Is it the is it the person who is responsible from the outage or the team or Who just a random person who happened to be online? Or or is it someone who has expertise there? And I think you should have very clear and strict rules about who is the
Speaker 1: owner of this incident and we it those that person should have clear responsibilities like We should get this down before the incident so that you know we we don't uh waste time during the incident. So um this is the last part I'm aware of the time, so uh I'm kind kind of rushing, but I I'd like to take your questions. So this is not this is misleading, this is not system design, but this is like more design related questions. So uh not in a technical sense perhaps So again, monitoring. Do we have the right monitoring and alerting tools? And I say please, please, please do create the tools, set them up, and don't wait for the outage. Be proactive
Speaker 1: Even if you if you can afford it to create teams for this or create dedicated days for for for this kind of thing. This this even starts when you are actually writing the production code. Of course you do write in a test or integration test, blah blah blah. But also when writing the code, think about how it should be monitored, how it will be monitored. And this decide, okay, we have the monitoring. What's critical about this monitoring and how are we going to get alerted? On which cases? These are very difficult questions by the way. So there is no definitive answer here. But look forward to the outage really because it will happen. Um
Speaker 1: and it's inevitable. We we'll all have o outages and maybe it's a good strategy to start uh the mitigation from De minus an, meaning you know do whatever you could do in advance And then, you know, while it's happening, just do the things you have to do, w the egg the the minimum you could do probably during the outage and uh that will make many people's lives easier. Maybe do nothing really. It's outage. Anyway, um I think this these are one of the few final things I have. So s sometimes Okay, five minutes. Perfect. Uh sometimes we uh
Speaker 1: we see that unfortunately we don't really have The right tasks assigned to the right people. And um yeah Uh this this is this is maybe for the team leaves or uh other product managers, but sometimes uh this happens and If if there's no good match , then it things c get a bit ugly. I think we need real honesty from team leads and the devs. uh about uh their technical skills and their understanding of of the whole thing. And of course I'm not saying we should always have the skilled people to do things. We
Speaker 1: all all learn from doing things. But there should be a good balanced learning developing ratio. Meaning of course you are going to learn but you should probably know more about the thing you you are doing and then learn in smaller steps. You just don't learn the full ho whole terraform thing when you are trying to do your first deployment. Yeah, well the numbers can change but yeah the the bigger the learning becomes I would say the more riskier it usually becomes And I think this is the last one. Sometimes we fail to draw boundaries. And I think the best example that we all know, the major one Not the only one, but is the distinction between the infrastructure, back
Speaker 1: end and front end. And those abstractions are kind of there so that we have the skill required to do them properly. But sometimes we those boundaries disappear and that's totally fine. But when when when when we are designing things this way, it's easy to couple things together so that let's say one part of your system goes down and other part of your system is taken down too. And um Yeah, uh it's it's just kind of frustrating because you don't know anything about that area but th that outage has stopped your work or has caused outage in your system which is not ideal really
Speaker 1: And perhaps I could just skip this. Um the the main takeaways are I think Uh the main takeaways are please do ask the ask the right questions. I I'm aware I just asked you questions the whole talk. But this this was the talk itself. And if you ask the right questions There is a higher chance that you'll find the right answers. And the second takeaway is look look at things from this level of analysis perspective. The lens is important because Sometimes we are too focused looking at things from a personal individual level, but we also need to think about the team and organization because that's at the end of the day that's what that's what we need. And if you could fill in the gaps between the understanding of an individual and the team or the organization, then I
Speaker 1: hope uh you you take away something And uh enhance your understanding about these things and yes this is the real thank you that was so much content. to hammer you hammer your brains at this hour but
Speaker 2: yeah I think we can take uh questions we have all five minutes okay if you can uh some come here or yeah
Speaker 1: Who goes first? Hello.
Speaker 3: Hello. That's a great talk, Otto.
Speaker 1: Thank you.
Speaker 3: Thought inspiring. How do you prevent or how do you propose preventing a blame culture if somebody is responsible for outages?
Speaker 1: Yeah, the blame culture I think comes from uh from the leadership If the leader is open to learning and open to embracing the mistakes, I think this is very important. You can't have mid-level majors do this. I know leadership very well because it it was my expertise decades ago. The leadership decides on the blame culture based on how they see the organization. For example, let's take octopus energy. Does the CEO of Octopusc think Octopus NG is an energy supplier or a tech company? So what's the what's the main thing here? And if I was told
Speaker 1: directly by him when I was joining him that he thinks it's a tech company. And when it's a tech company, then such things like tech mistakes are quite okay because that's how you learn and he is aware of this. But if he thought this was a business for energy suppliers mainly, of course it's a energy supplier, we all know that. But um if he If he placed the technology to the side as a support thing, uh not in integral part of the company but as a as a other thing, then I don't know maybe there would be blame culture or other things, but uh in its current form we we really embrace uh failure
Speaker 1: and um He literally comes and thanks us when we have an outage because we now have an understanding and we won't make this same mistakes again, hopefully. Oh, sorry.
Speaker 4: This is a very quick question. I'm one of these weird people that really enjoys reading outage reports because you learn so much from them. If you know of any good resources, newsletters or websites or whatever where these things are collected that you know of, could you please put those in the um the the chat for this this session afterwards? Thank you. Um
Speaker 1: I do know few uh companies publish them. We do sometimes publish them. Sometimes you can't, but yeah, I I'll I'll I'll let you know.
Speaker 5: Okay, thanks. Um often when you're setting up your monitoring and your alerts, you're gonna have an alert that uh has a lot of false positives and then People get alert fatigue, they stop looking at these alerts, um and then something a real issue goes undetected and you know the cycle kind of continues. Do you have any thoughts or ideas of how to break that cycle?
Speaker 1: Yes, happens to us every day. Uh and there's no easy question. But Uh personally I'd like to see things in two ways. For example, if it's something risky, I'd rather have the false positives and look at them And then tick off okay this is a false positive we can go back to work. So i uh but if it's something less important uh I'd rather have r some real failures and not be aware of them And then you know just get alerted hundred percent when things are going quite bad. But you know, it's not the end of the world really. So that's that's how I kind of see it.
Speaker 1: And There are some systems that kind of learn based on things like they have this detection algorithms and they kind of change the thresholds of alerts and other things. It might be a good idea to have a look at them and um Uh see how those alerts or thresholds or monitoring systems change with the different traffic or different requirements. Perfect. Any questions?
Speaker 2: That's I think the end of our QA session, unless we have any other questions Let's give uh tell another set of applause
Optimize deployments for getting back out: consider rollback, revert, or mitigation options before you go in, rather than only focusing on the deployment itself.
Discussed at 8:18Write reports for their intended audience, including developers, rather than using the same report for everyone. Give people time and resources to read, understand, and absorb the reports instead of treating them as a checkbox.
Discussed at 13:56Set clear rules in advance for who owns an incident, and define that person’s responsibilities before an outage happens. This avoids wasting time figuring out ownership during the incident.
Discussed at 16:15Set up monitoring proactively, starting while production code is being written, and decide what is critical and when it should alert. Don’t wait for an outage to build these tools.
Discussed at 17:04Leadership sets the tone: leaders need to be open to learning from mistakes and treat technology failures as opportunities to improve, rather than placing blame. The speaker’s example is a CEO who frames the company as a tech company and embraces learning from outages.
Discussed at 23:19There’s no single easy fix: weigh the risk of the alert. For high-risk issues, tolerate false positives and check them; for less important issues, it may be preferable to miss some real failures and alert only when things are seriously wrong. Adaptive detection systems that adjust thresholds may also help.
Discussed at 26:00Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025