Testing Modern Web Apps Like a Champion with Andrew Knight
Published November 22, 2023
This video features Andrew Knight at DjangoCon US 2021 in Online.
Good test data can be a nightmare to manage! It can make-or-break testing efforts. Should we preload our databases? Should we use dynamically generated dummy data? What about collisions? Let's cover practical strategies for handling data both in our products and in our test cases.
This talk was presented at: https://2021.djangocon.us/talks/managing-the-test-data-nightmare/
LINKS:
Follow Andrew Knight 👇
On Twitter: https://twitter.com/AutomationPanda
Website: https://automationpanda.com/
Follow DjangCon US 👇
https://twitter.com/djangocon
Follow DEFNA 👇
https://twitter.com/defnado
https://www.defna.org/
Video production by the speaker and DjangoCon US 2021 Volunteers.
Test data has two related but distinct forms: product data stored in the system, and test case data used to drive tests, configure their execution, or check their results. Product data can be prepared statically or dynamically: static data is useful for complex or slow-to-create records but needs maintenance, while dynamic data improves test isolation at the cost of execution time and cleanup. The speaker compares manual setup, automation, database cloning, and mocks, recommending a combination based on data size, freshness, skills, constraints, and cost. For test case data, browser choices and environment details should be externalised into input files, environment variables, configuration files, or secret-management services rather than hard-coded. Behavioural values may be literals, outputs retrieved from the system, or input references to existing product data; data discovery can make those references more resilient than fixed names. Finally, parallel tests need isolated environments, immutable shared data, and as much per-test dynamic data as possible to prevent collisions and brittle failures.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Hello everyone, my name is Pandy Knight and I'm the Automation Panda. I'm the lead software engineer and test for Precision Lender at Q2. And today. I'm going to talk to you about test data. Test data can be a nightmare. Getting the right data in your system as well as into your test cases can be really difficult if you don't have the right strategies or know-how. And that applies whether you are doing a Django project, some other kind of Python project, or really any kind of software project. So in this talk I'm going to give today for DjangoCon 2021, so excited for it, I'm going to show you how to handle your test data. I'm going to provide you strategies for handling product data
as well as test case data. So by the end of this, you'll be able to conquer your own nightmare without fear. So if you're ready, let's learn! Let's say we have an application for a bank to provide loans. The bank could configure this application for many different types of loans. Such as a home mortgage, a car purchase, or student loans. All of the information the bank needs to provide loans must be stored as data in the system Each loan product is different. It comes with its own rate, maturity, and payment schedule. The bank must also store information about borrowers. Funding curves, and profitability targets for loan opportunities.
It's fairly complicated. Testing requires that all data already be present in the system as a prerequisite. We could write a simple test case to exercise the basic application behavior. Scenario. Create a new loan application. Given the Chrome browser is open, and the page mylooneapp. com is loaded, when the user creates a new loan application for home mortgage And the user enters all their personal information, and the user submits the application Then the page displays a success message with a reference number, and the loan application is sent to the bank.
Now, bear in mind, a real loan application would probably need several pages of information. But let's keep our example simple. This test creates and submits a new home mortgage loan application for the user. There are many test data points in this short scenario. Most apparently, all the user's personal information. The type of loan. There's also the record of the loan application sent to the bank. And the reference number shown to the user. Furthermore, the URL is configuration info, and the browser type is a test input. Test data is everywhere in this short, simple scenario.
The data is inextricable from the test. Without specific data, this test would be meaningless. Unfortunately, the term test data is ambiguous. We've applied it to both the product data in the loan web app, as well as to the various pieces of test case data that make even the most basic test work. Product data refers to real data living in the software system. For the loan web app, product data includes all the bank's product configurations and lending information. Test case data refers to data used to define test cases. It may include values to enter into the product under test, inputs to control how testing is performed, or records to retrieve from product.
In the latter case, test case data is a reflection of product data. Its values refer to entities existing in the product data The two types of test data are separate but connected. Distinguishing these two types of test data is important to avoid confusion. The dependency of test case data on product data can be brittle. For example, consider our test case step to create a new loan application for a home mortgage. This step works as long as the bank's web app is configured for home mortgages. However, the product data can be changed at any time, just like product code. What if the specifics of a home mortgage application change?
What if the loan is no longer called a home mortgage, but a personal residential loan? That would make the test case break. Compounding breakages cause nightmares for test management. So, how should we manage test data? For feature testing, test data is just as important as test cases and test code. How do we handle both product data and test case data? Are there strategies we can use to avoid brutal dependencies In this talk, we will explore multiple ways to handle both product data and test case data. Unfortunately, there are no universal or perfect solutions to the test data problem. But you can avoid nightmares by picking strategies that work well for your needs.
Let's start with product data. As stated previously, product data is any live data in the product or system under test. In simplest terms, it is everything in the database. That could include user accounts, administration settings, product customizations, records or files created by users, and more. For our example loan app, product data would include user accounts, loan product settings, loan applications, and behind the scenes bank data. Data must be present in the product as a prerequisite for most testing. There are two primary ways to get that data into the system.
On one hand, you could set up the data before running tests. This would be static data creation. For example, the loan web app could be set up with a set of pre-registered users and a collection of loan types. Test cases, whether they are manual or automated, can presume that this static data is already in the system and simply reference it. Static data prep is a good strategy for comp for complicated data or data that is slow to create dynamically. For example, user accounts may need email verification, so it might be easier for automated tests to simply use a set of pre-registered users. Tests will run faster if they can simply reference existing data
instead of creating new data each time. However, static data must be maintained. Any changes to static data could impact tests too Static data may also become stale over time, as data formats are updated or if data is time sensitive. On the other hand, you can set up data during test execution. This would be dynamic data creation. In the example loan test case, the loan application document is dynamically created. The test does not reference an existing loan application, it creates a new one. Dynamically created records avoid the brittleness of hard references to static data.
They can also be used exclusively by the current test case, protecting them from interruptions from other test cases. The main downside of dynamic data prep is the execution time. It slows down tests. Dynamically created data is essentially disposable too, so it should be cleaned up eventually. Which strategy is best? Typically, testing requires both strategies together. Data that is slow to set up or considered immutable should use static preparation While data that is quick and easy to set up should use dynamic preparation. When I develop test solutions, I prefer to create as much data as possible dynamically per test case
to preserve test case independence. When a test creates the data it needs dynamically, it will be the only test using that data, and there is a much lower risk of test collisions. These two data prep strategies are a bit complicated when implementing them. Dynamic prep depends directly upon the test case using it. But static prep has a few general cases that are independent of test cases. The simplest data prep strategy is manual configuration. That's exactly what it sounds like. You log into the system and manually create whatever records need to be there. That could be those users, config settings, or any records.
The nice thing about manual configuration is that it's low tech. Anyone can do it. It doesn't need fancy or complicated tools. However, manual configuration is slow. It does not scale well for large systems or large test environments. Furthermore, manually configured systems can easily fall into disrepair without any automated mechanisms for maintenance. A better strategy might be automated configuration. Rather than manually setting everything up, automated tools can create the desired data. This could be accomplished in many ways, reusing UI interactions from tests, calling REST APIs, or even possibly using tools like Puppet and Chef.
Automation could generate data deterministically or randomly. The main benefit of automation is the ability to create fresh data at any time. Automation can also clean data, like scrubbing private fields, or updating time-sensitive records Unfortunately, automated configuration is not a free lunch. It requires extra skills and code must be maintained. If you want a shortcut, you could try to clone the databases. Cloning databases is easier than ever with cloud management tools. You can maintain one database in a golden state and then create a copy before running tests. Once testing is complete, the copy could be deleted.
Granular cleanup would not be necessary. Database clones make it easy to copy all data at once without worrying about any damage that rogue testing could cause. However, databases can have a lot of data. I'm talking b -gigabytes. So cloning large ones may not be practical. Clones may also need extra refinement to scrub special fields and hook them up properly. Finally, if managing real data is too much of a hassle, then you could always mock endpoints. This would completely remove dependencies on databases and even services. All data returned by the mocks would be deterministic too, yielding consistent results.
But mocks are not always a good solution. They often require a lot of effort to set up, and mocked data can make tests overlook unpredicted real-world variations. Mocks also mean that tests will not truly be end-to-end in coverage, for what that's worth. These strategies can also work together. For example, you can use automated scripts to configure product data in a golden database, and then you can make clones of that database. In another example, in a large testing environment, you could choose to mock some endpoints while using real data for others. There are multiple factors that should be considered when deciding the best test strategy for static data prep.
How big is the data? If it's small, you can do it manually. If it's large, you probably need automation. How fresh does the data need to be? And how frequently will it be updated? Again, automation can help for frequent updates and time-sensitive scrubbing. How difficult will it be to try advanced tricks like mocking or cloning databases? This might be difficult in old legacy systems. Is there any bureaucracy in the way of an automated solution? Hey, it happens. Companies will be companies. Bureaucracy can stonewall advanced solutions that might need extra support.
Do folks have the skills required for automation, database administration, or mocks? Skill level may be a barrier at first, but training and learning can help any team overcome any limitations here. And finally, the beginning. What about cost? Each strategy has a cost. Teams should do a cost-benefit analysis when deciding. So, that's how to handle product data. What about test case data? Let's look at that next. Test case data is inherently part of test cases. Let's revisit our example test case from earlier.
As we saw before, there are multiple bits of test data throughout the steps of this short scenario. They represent different types of test case data. First, let's look at the first step. Given the Chrome browser is open. The Chrome browser is test data because it specifies the type of web browser in which to load the web app. This is what we call a test control input. It directs how tests will be run rather than specifying feature behavior. Theoretically, this test should run the same on any browser type, but the steps dictate that this test should run on Chrome. As a best practice, test control
inputs should not be hard-coded in test automation code. Don't do that. Instead, they should be passed into automation as inputs That way, tests can easily be retargeted. There's a few ways to do this. The simplest way would be to create a flat file with input values. I recommend using a format like JSON or GAML, because they are easy to write, easy to read, and easy for programming languages to parse. For example, in Python, to read a JSON file into a dictionary takes only one or two lines with Python's JSON module. Test automation code can read the file before any tests start, and it can eject input values as appropriate.
For example, using this JSON file, Automation could read the browser type and construct Selenium WebDriver objects for Chrome for each test. The path for the input file would need to be hard-coded into the automation, but it could be as simple as a standard file name in the current directory. Another way to handle inputs is using environment variables. Testers could set variables from a system shell profile, and automation could read those variables by name. This can be useful for integrations with continuous integration servers or Docker containers. However, it can be a little more dangerous because anyone could change variable values. Again, automation can read these variables before tests run and handle them appropriately.
Let's remove that hard-coded step for browser type from the test scenario. That can be handled as an automation concern. Next, let's look at the second type of test case data. Notice how the browser URL is hard-coded. This is also not good practice. Typically, development teams host multiple instances of products under development, like a developer environment and a staging environment. Hard coding configuration information like this limits where tests can run. Any information about a product's configuration is called configuration metadata. This can include things like URLs, usernames, passwords, and possibly other descriptors.
There are a few ways to handle config metadata. You can use flat files or environment variables, like for test control inputs. However, I recommend using flat files, and I also recommend separating test control inputs from configuration metadata. Create an input to refer to the target configuration and store multiple configurations in the config metadata files. That way, testers can change just a few simple inputs to target any configuration, and they won't need to change multiple configuration files regularly. If you want to be fancy, you could create a web service to provide config metadata. For example, in Azure, they have something called Azure Key Vault.
This would be especially helpful for keeping secrets like passwords safe. However, creating such an endpoint may be overkill for your needs. Either way, the test step can be rewritten to refer more generically to the web app. Automation can select the target environment using the inputs and config metadata. The remaining pieces of test case data all fall into a category called test case values These values pertain directly to the behavior exercised by the test, not to any configuration factor. Even in this classification, there are subtypes.
The first kind of test case value is a literal value. These are values that are hard-coded in the test. In this example test, the table of personal info contains literal values. Literals are simple to use, and they provide specification by example. Literals should also be independent of any statically created product data. They should be values that can be safely originated by the test case. The literals in this info table will be entered as input values into the web app. Theoretically, they could be any values. The second kind of test case value is an output reference. These are values that are retrieved from the product under test.
Typically, they are the outputs generated by exercising a behavior. In this example test, the reference number can be scraped from the success page and verified for correct format. The loan application can be retrieved from the web app's backend to verify that it was correctly submitted. These values cannot be literals because they originate from the product. Tests must refer to them by reference and retrieve their values from the product. Also, quick side note here. This loan application created by the test is an example of dynamic product data prep. The third and final kind of test case value is the trickiest, the input reference.
At first, these may look like literals. However, input references are values that directly refer to product data. While personal info like name and address are created dynamically by the test case, the name of the loan type refers to the loan configuration in the web app. Thus, this test has an input dependency. It must specify the type of loan, and the loan type must already exist in the product data. The simplest way to write this test is to simply hard code the reference. That's what's done here. The name Home Mortgage refers to the name of the loan type in the web app. Automation can use that name when selecting the loan type from, say, a button or a drop-down. Hard-coded references make it easy to write tests, but they require statically prepped data to exist in the system.
References also become hard to maintain when the product data changes, or when the same test might run against different configurations with different names. One way to avoid the pain of static data is to dynamically create these records or configurations. If the test calls the back end to create a new loan product, named Home Mortgage, for each test run, then static pre-prep isn't needed. However, we already know the pain points of dynamic prep. In this case, let's say dynamically creating a new loan type is just too slow. A more robust solution could be data discovery. Let's say the target web app is already configured with multiple acceptable loan types. Instead of hard coding the name of the desired loan type, the test could describe the loan type
and then use automation to search the web apps config to find a loan type matching the desired criteria For example, if different regions of a bank have different names for this type of loan, the discovery mechanism could look into the config for a satisfactory home mortgage loan and return the specific name for the current bank region. Discovery enables tests to search existing product data for required records instead of hard-coding records. Discovery makes tests more resilient to changes in product data. It's great when testing multiple environments with ever-so slightly different configurations. But it does require extra coding, and it may be overkill for small test projects.
That's a lot of info about test case data. Let's summarize Test control inputs direct how tests will be run, not what behavior is covered. They should be supplied via flat files or environment variables. Configuration metadata describe product configuration for the target environment. They should be supplied via config files or service API calls. Test case values direct the behavior covered by the test. They may be literals, output references, or input references. And those input references may be hard-coded or discovered. At this point, you're probably thinking, wow, that's a ton of information.
That's it, right? We're done? Time for QA. Move on to the next talk. Well, frightfully, the nightmare isn't over quite yet. There's one more problem to address. Collisions. Collisions can happen whenever multiple actors operate on shared resources. For example, they could happen whenever multiple testers simultaneously access the system, or when automated tests run in parallel Additional considerations must be taken to avoid collisions. First and foremost, isolate the test environments Prevent external actors from interrupting tests. If you have a shared test environment, block other folks from accessing the system when running tests.
You might need to schedule tests to run during off hours. You might also want to set up multiple test environments. If the product on a test is containerized, or if databases are clonable, then you could easily dynamically create fresh environments for each test launch. That would guarantee perfect isolation Second, treat any shared data as immutable. When I say immutable, I mean unchanging or constant. Sometimes shared product data is unavoidable. For example, if tests run in parallel against one test environment, then they may use the same product data. Or, if an application has multiple components, certain components may be difficult to isolate for testing.
Whenever data must be shared, treat that data as constant or immutable. Any changes to shared data could break tests. For example, one test might require a car loan, but another test might delete the car loan type from the bank's configuration. Any tests that must alter shared data should be run serially instead of in parallel, and they should always undo any changes they make. Do no harm, leave no trace. Third and finally, use dynamic data preparation as much as possible. Tests cannot collide on data they don't share. Keep statically prepped data to a minimum.
Statically created product data is more likely to become shared data, and shared data is more likely to cause collisions. Oh, we covered a lot of material in about half an hour. So let's recap what we've learned. There are two primary types of test data, product data and test case data. Product data can be prepared statically or dynamically. Test case data either control how tests run or they reflect product data. You should handle references and shared data carefully And overall, choose the best strategies to defeat your nightmares.
Every product is different, every team is different You can use these strategies on your Django app that you're working on. You can use them for any Python project you're working on. You can use them for any type of software product, whether it's implemented in Python or not. Take the strategies I shared in this talk as suggestions to help you. So, thank you very, very much for attending my DjangoCon talk. I really appreciate it. Again, my name is Andy Knight, or Pandy for short, and I'm the Automation Panda. Be sure to check out my blog and follow me on Twitter, and I hope you enjoy the rest of DjangoCon 2021.
Product data is the live data in the system under test, such as users, configurations, and records. Test case data defines how a test runs, what it enters, or which product records it retrieves; the two types are separate but connected.
Discussed at 3:43Static data makes tests faster and is useful for complicated setup, but it must be maintained and can become stale or break tests when changed. Dynamic data avoids hard references and test collisions, but it increases execution time and must eventually be cleaned up.
Discussed at 6:58Use static preparation for data that is slow to create or effectively immutable, and dynamic preparation for data that is quick and easy to create. In practice, combining both approaches works best, with as much per-test dynamic data as practical to preserve test independence.
Discussed at 8:30Do not hard-code test control inputs such as the browser type in automation code. Put them in a readable flat file such as JSON or YAML, or use environment variables so the same tests can be retargeted easily.
Discussed at 15:30Keep configuration metadata separate from test control inputs, using configuration files or a service API. Select the target configuration through a simple input, and consider a secrets service such as Azure Key Vault for sensitive values.
Discussed at 17:55Literals are values entered directly by the test, output references are values retrieved from the product after exercising behavior, and input references point to product data that must already exist. Input references are the most brittle because they create dependencies on product configuration.
Discussed at 19:21A test can create the required product data dynamically, although that may be too slow. Alternatively, data discovery can search the existing configuration for a record matching the desired criteria, making tests more resilient across changing or different environments.
Discussed at 21:56Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026