Closing session
Published June 13, 2025
This video features Zelma Gist at DjangoCon Europe 2019 in Copenhagen, Denmark.
Zelma Gist presents a lightweight way to catch unintended CSS changes by taking screenshots of the same Django page in two environments and comparing their pixels. The approach uses headless Selenium/PhantomJS, masks dynamic content such as dates and feeds before capture, and draws a grid over differing regions so failures are easy to inspect. She runs selected visual checks separately from functional Selenium tests, cleans up screenshots after successful runs, and connects the process to CircleCI through a browser-testing service. Approved visual changes become the new baseline; she says production-based comparisons are a future step for more complete regression testing.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: All right. Hi everyone. So today I'm going to talk about simple visual regression testing. The emphasis here really is on simple, simple specifically because This project, this talk is really generated from an issue I was having at work. I'm a full stack engineer, and so I was having issues with knowing that my CSS changes were rendered correctly and I needed a simple way to test. So a little bit about me, I work for a company called Wealthy. It's a New York-based company. And what I've realized here now that I've been traveling in Europe for about a month or two now is We solve a problem that maybe is a bit unique to the United States. So we work with families who are caregivers for people who are with chronic and aging loved ones and we provide them care coordinators to really help them manage their care
Speaker 1: Unique to the US healthcare system, we have a whole massive insurance market and negotiating that can be challenging, especially if you're trying to manage a full-time job. So we work with families for that. So, what is a full stack engineer? So I decided to go ahead and start with this little comic here. I think it does a good little summary of what is full stack. There's different types of full-stack engineers. There's those who are more DevOps and they sometimes do some other things on the side. There's back-end engineers who sometimes do some CSS occasionally. There's those who are front end and are occasionally forced to go into the back end. And the little comic goes on to continue. It's from Commit Strip. It's a good little comic. I recommend it. It's super nerdy, super fun. And so looking at this, I would say that I'm probably the top right.
Speaker 1: I'm a full stack developer, feel extremely comfortable in the back end, feel comfortable with Django. I can work my way around HTML. JavaScript's an old friend. My first job was in Java, so these are old friends. But CSS is a bit of a mythical art in that respect All of you front-end engineers who are handling this well, I respect you, I value you. We're hiring a front-end engineer, just gonna throw that out there. At which time, maybe I'll go back into the back end. And so considering now that I'm a front engineer, I still have to work with CSS. It's work needs to get done. And so the problem I was having was I would get some trying to answer this question. How do you know your CSS changes? don't have any unexpected side effects, being that
Speaker 1: if I move change this class, how do I know that there's not unexpected changes on another page that shares the same class? Perhaps one way is, you know, don't reuse CSS if it's all unique, but then you end up with a massive code base that's unmanageable, nobody wants to use, and people like me will just avoid CSS even more. Alternatively, you send it over to QA, QA magically, they do whatever they do, and potentially they come back to you with the request that this is off by a pixel. But sometimes, even for a human eye, it's challenging to catch these differences that you see in two web pages, especially as they click through So for example, if you have something like this, who can tell me what's the difference between these two images? So for this demo
Speaker 1: I went ahead and just spun up a sample site. I was being art through the other day, so I took a picture at a park of some flowers and threw it on here. And so it's subtle, it's visible, you can see. Any guesses? Yes. Exactly. Yeah. So he's correct. It's the padding here on the images are off. And here there's not it's not high contrast, but you can still see it. It's a border, it's there. And so he's right, the padding's off a little bit. This is running just off my local host. I spun up two instances One's in a subdomain, calling that one staging, and then I have my dev environment, which is to mimic dev. So when I have this problem, I sometimes will
Speaker 1: catch it, but to be honest, sometimes I won't And so rather than relying on QA or the person doing my code reviews to come up with these changes, I decided to come up with a solution to really solve this problem for me. And it should be super simple. The code's pretty straightforward. It definitely could be architected more into something larger. I was working on making part of this open source, but did not get quite finished before this talk, so coming soon. But it's just being able to integrate this into your Python environment. So to do this, the requirements I needed, I needed something That would compare the same page in two environments. So really whether it's your dev environment versus staging, staging versus prod, I really just wanted to look at two images programmatically
Speaker 1: and let me know if it's different. Second, I really wanted to know exactly what's different. I don't need an error message that says your CSS is off. Well, what does that mean? This color's wrong, the spacing's wrong. An image that should be there isn't there anymore because you broke everything. There's like a fire. Like what does off meet So I wanted a way to know that things were different, and that's why I decided to go something visual. That I could see if there's an error, something that would return so clear what I've done wrong, that way I can easily fix it For me, the third requirement was I wanted to be able to run in my existing Django environment. So that means being able to run the test locally before I commit being able to run the tests in our continuous
Speaker 1: integration tool. I personally didn't feel it was necessary to run it between stage and prod, but that was a me choice. It definitely can be done. And so these were the requirements that I had when I started with this project. And so I really started with something super simple. It's just a basic setup function, and I'm using Web Phantom. js Phantom, I think they recently went to like they're not developing it anymore, so we'll see what that does long term. But the real rationale behind using that is I wanted something headless. I wanted something that would be fast to execute. And as you can see, this is super simple. In the real world I have these broken up into functions, but I just put all the lines there because it's so simple. It's first we stop with I called it pick viewer because Why not?
Speaker 1: And so the dev one is running on the dev, so this is based in my local host right now. And so I have to get the fix width when I get the window. Then I go into the page and I save a screenshot. Really, when you're starting up Selenium, depending on the page, these pages have no validations. But before you would really access these pages, if you need validations, you enter username or whatever you need to do to actually access these pages And so I'm doing the exact same thing, which is why I pasted it here for both, for both staging and dev. So in the end, I should have Two images. So to make my life easier, I made clearly different rather than the last screenshot where it was pixels off. And so I have just a basic page, one's in dev, one's in staging.
Speaker 1: And it's an image viewer. Same image, different borders. So one's the one running in dev. I made a CSS change clearly to just change the color. So it should be blue because that's what it should be Because I said so I don't know, and the other one is red. But the real challenge here that you're gonna have with this implementation that we ran into fairly quickly is the date. So I just arbitrarily stuck 2019 on the top. what happens in 2020 if this is a changing if it's a daily thing like happy Monday or a fixed date that's a field that you know is going to change and so if you're running if it's running seconds if it's running minutes When you're trying to process these images, they will be different always. And so you don't want a failure based on a time
Speaker 1: because you know the time should change at least in a regular fashion. There's also the same issue with RSS feeds, if you have a feed that's going into your site or things that update fairly regularly but you know can be different and having different content isn't wrong. So to address that, I simply wrote just a script that's executed. Once the DOM is rendered before I take the screenshot, I literally just go in, I have this class on there just for the demo as date. And within our platform, we do have fairly consistent CSS classes. And so it's generally fairly predictable of what elements should change. and I block it out. So what I'll end up with is two screenshots like this. So now we can see that the date has been just
Speaker 1: completely removed from it. It's really more of a practical functional decision. I think that there are better ways to handle dates. If you do want to validate the content in the date, perhaps you actually parse the jum and you check the field when you're running the Selenium test. But for me, I don't care if it's 2019, I don't care if it's 1999 for this case. I really just care about the CSS And did the changes get rendered? Because I have confidence perhaps I'm rendering the date from a view in their test already. So I'm not so much worried about that field. And so what I've done here is I got rid of the date. So really what's happening from a functional perspective of how we're doing this is I just use Python image library and we're comparing pixels.
Speaker 1: I kind of broke it out for this talk just so it's a little bit more visual as to what's happening. But basically I take the two images and break it up into a grid and I just compare the brightness of the grid. So we can the size of each image that we're comparing from the dev image and the staging image, that can like vary, yes. But here this is just a simple going through and comparing and we create a grid. And so just for visual purposes This is kind of what that grid ends up looking like. So the top right, they all could have compared. In generally when I'm running this in dev, I don't print out this image because I find that there's no need for it Really the next step here is comparing the two images. Or so it's top corner, top corner.
Speaker 1: And Generally by comparing the brightness, the images theoretically should be the same. I'm not brave enough to run live code on stage, so I have it all set up and it all works But that's just really how it's working. It's super simple, super straightforward in terms of comparing the two images. And so really just building out the same add grid function. We have the images and generally in the tests I know the URLs and I've already had the images, so I just throw it to the function. And this function, the rows and columns, they can be smaller or larger. Yet again, this is more of a functional tool. And I'm just trying to build something that'll work, that'll tell me what's different. And so we just really go through, we make columns, and then we get the dimensions of the page.
Speaker 1: or of the each image, which has been fixed when we originally take the screenshot, and I really just go in and I compare the image. And so for the case where the images are different, that's the real case when I go ahead and I draw the grid. The real purpose of the grid is for me to visually see what's different. So the image is the same and the same basic header is the same and all the white space around is the same. But for me this allows me to see what's different. And so what happens when I'm running my test cases is I run my Selenium test generally locally first, and things run, and then the cleanup, which I don't think I put. In the teardown there should be just a function. I don't think I included it here, but
Speaker 1: there should be a function to get rid of these images that have just been generated. The reason behind that is if the test succeeds, I don't need all these spare images just sitting on my local database or in some like hard, like um storage that I'm using. So I'll generally go ahead and delete that. But for the test, the case where the tests fail, aka it generates this result. For me it's really simple to see. exactly what's gone wrong. And that's really what I need, especially when I'm dealing with CSS. Because my big challenge is Maybe I'm not observant enough, maybe I don't catch these things, but it's a way to still do my job and execute it well without having to spend time clicking around on a site So once we have that, the last
Speaker 1: most challenging part of this entire implementation is the continuous integration. So all these things work so nicely on your local environment. I'm running Selenium, it's headless, life is good, it's giving you this output document. But what do you do when it's time to go to continuous integration? And someone else needs to validate that your code actually works and that the CSS changes you've made are fine. So for me this was probably the biggest challenge of maybe this entire little project. Originally I was like, oh, we're using Circle CI, we'll just throw it up there and maybe we'll have a Docker instance that's reading that URL. But the real challenge here is running this instance separately without running it in the Docker
Speaker 1: instance. So for us, we're using cross-browser testing, and cross-browser testing allows us to take the screenshots of the different images. and process them. And what that really means is I'm just running my Selenium test in cross-browser testing and they have a built-in integration to CircleCI, which is really what I'm using. All right, I'm done and I think I spoke too quickly. Bit nervous, sorry. Yeah.
Speaker 2: No, that was great. Uh so now we uh have some time for questions. Um so if you would like to ask questions either online using uh hashtag Django QA or in the IRC channel and in person.
Speaker 3: Hi, that was brilliant. Thanks. You said you were going to open source that, so
Speaker 1: Yes.
Speaker 3: Are you going to?
Speaker 1: Yes, I am. I am. Now, yesterday we learned even if it's not perfect, we should put it out there. I wasn't bold enough last night, but yes, I will definitely.
Speaker 4: Hi, really really cool. I was wondering how you handle stuff like very iterative pages. So you're working on something that hasn't been finished yet, but you just keep updating. Um how do you go about did you just leave it out or?
Speaker 1: That's actually a great question and I realize now that I omitted that. So there are chance there are times when CSS changes are valid, right? I made an update and I do want every page to change. And for those cases we have the process of we approve this change. So it's do we acknowledge that so it's a warning is what I have running in Circle CI and it's you acknowledge that warning and if that warning is sufficient by the code reviewer, we go ahead and squash and merge. And that way the next time it's run, it will always, it'll take a new image. And so that's why in Circle CI it's we try to merge integration or merge the branches first and then run the test runner. That way any new updates, especially if you're in between like it's in staging but it's not in prod, we'll
Speaker 1: be able to get those changes as well.
Speaker 4: Okay, cool. Thanks.
Speaker 5: Hi, great talk. Uh thank you. Uh just a very quick question. Uh what wasn't very clear to me. Uh so you normally have lots of uh uh Selenium tests where you like test the behaviors of your uh pages and then all of them become uh screenshots uh automatically and then you compare that
Speaker 1: So we have them running separately. So the image comparison one is run separately on really certain pages. We don't have them running on every page because we're There's parts of our site we don't use quite as often, but we have broken them out as separate because the Selenium tests for functionality I think of as a little bit different because you're testing functionality. And there is some correlation as if if I type this, does this render? If I type this, does this render? There's a correlation there, but generally we choose to run them separately. Perhaps over time this is going to become time consuming, but for now that's what we're doing.
Speaker 5: Nice. Thanks.
Speaker 6: Hello. Thank you so much for this talk. Um is there anything you've learned from having something automated like this? I'm guessing a lot of changes you maybe didn't catch before are now caught by the machine Uh is are there things about CSS and those differences between integration and production or like your dev environment and your integration environment that you've learned through having all this feedback?
Speaker 1: I think we've really learned it's We have a standard now. So if I work at a startup and we're very lean and we're trying our best to build an amazing product, but we have now a CSS standard that's really helped me personally, but it also helps as a team, what class should we be using? And we've been following that, and that's really been the most valuable thing in to hear these changes. And so that we're not just saying, okay, today it's afloat. All right, now it's like block and like whatever it is. It's we're using the same con classes consistently. And so it's really cut back on all the issues.
Speaker 6: Okay. Thank you.
Speaker 7: Uh hi. I was wondering to turn that into um or to use it for regression testing. Uh you need to have some like base against which you're comparing your change. How do you are because like in your demonstration you were comparing death and staging but I guess in like for regression testing you would compare the current state with like the suggested change. So do you track the Current state in the Git repository or how do we do that?
Speaker 1: The current state is really tracked based on prod. Um our releases in prod are a little bit more stable and it comes as a chunk of work. And I mentioned that we're not currently doing the stage versus prod regression test, but I think that's where it should be done. But it's the issue that we were having was the continuous integration as we're constantly deploying into integration. And so the integration of prod I think is really the next step of what we need to do to actually have proper regression testing. But that's a work in progress.
Speaker 7: Thanks.
Use headless Selenium/PhantomJS to render the page in environments such as development and staging, fix the browser width, and save a screenshot from each environment. Compare those screenshots programmatically instead of relying only on manual QA or code review.
Discussed at 5:25Run a script after the DOM is rendered but before taking the screenshot to hide or remove known dynamic elements, such as a date field or RSS content. Validate the dynamic data separately if its content matters.
Discussed at 7:44It uses Python Imaging Library to compare the screenshots pixel by pixel, dividing them into a grid and comparing the brightness of corresponding regions. When the images differ, it draws a grid over the result so the changed area is easy to locate.
Discussed at 8:17Run the Selenium screenshot tests through a browser-testing service rather than trying to launch the application separately inside the CI container. In the example, CrossBrowserTesting provides the screenshots and integrates with CircleCI.
Discussed at 12:23Treat an expected difference as a warning that a reviewer must acknowledge. After the change is approved and merged, the next run takes a new baseline image; the workflow merges the integration branches before running the test so approved updates are included.
Discussed at 14:38They can be related, but the speaker runs them separately: functional Selenium tests check behavior, while image-comparison tests target selected pages and visual changes. The comparison is not run on every page.
Discussed at 15:52The current production state is treated as the stable reference because production releases are more stable and arrive as a chunk of work. The speaker says comparing staging or integration with production is the next step, but that workflow was still in progress.
Discussed at 18:01Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025