factory_boy: testing like a pro
Published October 14, 2022
This video features Camila Maia at DjangoCon US 2022 in San Diego, California, USA.
How to test complex objects using the library factory_boy. The lessons I’ve learned using the tool in a Django monolith containing 230+ tables and 75k+ relevant lines of code for over 3 years.
This talk was presented at: https://2022.djangocon.us/talks/factory-boy-testing-like-a-pro/
LINKS:
Follow Camila Maia 👇
On Twitter: https://twitter.com/cmaiacd
Website: https://cmaiacd.com
Follow DjangCon US 👇
https://twitter.com/djangocon
Follow DEFNA 👇
https://twitter.com/defnado
https://www.defna.org/
Camila Maia explains how to use Factory Boy effectively when Django applications have complex models and relationships. She argues that poorly designed factories spread implicit assumptions through test suites, causing duplication, fragile tests, and difficult refactors, and presents seven practices to keep factories aligned with the database and tests explicit: represent model relationships accurately, avoid relying on factory defaults, include only required data, prefer `build` over `create` when the database is unnecessary, choose `SubFactory` or `RelatedFactory` based on foreign-key direction, avoid sharing factories and fixtures across files, and use local fixtures to wrap factories when that prevents duplication. Examples from a poll application show how these choices improve correctness, maintainability, and test performance.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Hello everybody, I'm Camila Maya. I'm super glad to be here presenting at DjangoCone US and today I'm going to talk about Factory Boy and some best practices when using it This presentation you can find at speakerdeck. com slash cmycd. So let's get started. Who am I? So I'm a back-end developer at SoundCloud. I'm a Brazilian living in Berlin, Germany. I have a Bachelor of Computer Information System. I'm coding since 2010, it's a long time ago, more than 10 years ago. experience. I have more most experience with Python and Ruby programming languages. I love to work with open source contributing with it, using it, and also I love the communities like tech communities in general
I love to help in attending to conference. I have helped, for example, to organize some of this conference, PyJamas, Europytal 2020. Python Brazil 2020 and other and others. So yeah, hi. Also I'm creator of Scan API that's an open source library. that helps you to create integration tests like automated integration tests and also auto-generate documentation for REST API. So it's a pretty nice project that I am helping uh to maintain since 2019 and uh because of these contributions and other contributions contributions Um my GitHub profile was the first one to be accepted
for the GitHub Sponsors program in Brazil and uh the first sp first sponsorship that I received was from mini rally a pretty nice um sponsorship for as I started so got pretty uh proud about this this achievement. But anyways, let's stop bragging and let's uh start with uh the presentation I'm going to talk about Factory Boy. This is not an introductory this is like not an intro um talk. I will talk about more about best practices, but I will just give you an overview of what is it if you don't have uh if we didn't work with it before but um uh as I said I I will probably assume that a lot of context concepts here I already know
So basically Factory Boy it's a fixture replacement, so we're we stop using mostly stop using fixture using fixture and start using factors instead It's based on Factory Bot, that is a Ruby uh library. First, before uh been called factory bot it was called factory girl and then maybe that's why we have like a factory boy in title and the first time that I used was the factory boy uh the factory bot And uh I thought it was pretty nice when I saw that there was like a corresponding library, the same uh the same idea in library uh for Python, so quite nice Uh the first version for a Factory Boy uh was only working with Django, but nowadays
it's framework independent, so it works With Flask, Django, and wherever, uh, and also work with unit tests, PyTest, and so it's quite flexible. Um, basically, here we are going to talk about the advantages of using um factory boy for complex objects because like if you don't have like a really complex uh database you don't have complex models with complex relationships The fixture probably fixture will work fine and uh probably you won't need all this best practice that I'm going to talk about uh in this in this presentation. Uh this presentation is more focused when you have like huge uh uh data models, relation, complex relationships, complex objects. So in this scenario, white
fixtures are not that good. Because they are static, they are hard to maintain. while factories they are super easy to use, they are customizable and you can use customizable and you can use the specific fields that you want for the specific test you want to create So fixtures, you need to create an exhaustive test setup with every possible combination of coronary cases. So you have just to handle all the corner cases and the fixtures like separately While with factories you can customize your objects for the current test so you can declare only the specific test attributes that you want. So you can use specific things that you want Another really good thing about using Factory Bi for me is like
it comes with a lot of different tools that help you in the creation of the test. For example, you have sequence Let's say that you want to you have a model, user model, and this user has email, but then if you hard-coded email like saying user at gmail. com The next time that you are going to create a new user and if it is like a unique field , you're going to have a problem because it's saying that the two users have the same use the same email So a sequence would be like a pretty good feature that you could use. So saying that programmatically you will increment our number um by the end of the email. So let's say that the user one will have will have the email user one
at at gmail. com, then user two at gmail. com. So it will incrementally change this this sequence so it will uh change the email. This is like one example where we have a lot of different tools like faker to create uh fake address, fake questions, fake names. You can use a lot of uh a lot of fake stuff here mainly to avoid like using hard coded and you can like check more different edge cases for your tests. You have fuzzy attributes, lazy functions have a lot of different tools I'm just listening here so you can check around, but it's really really powerful this set of tools that comes with Factory Boy.
Okay, so to start talking about the best practice, first we need to go uh one step uh back and talk about this demo API app, this demo application. This is just like a super simple application that I'm going to show you here. And based on this application, we are going to show some examples, bad and good examples. of how to do it following this best practice. Of course, I was talking before that um this best practices in the presentation focus on complex objects And complex relationships. And the ones that I'm going to show here in the demo application, they are not complex, they are super simple, but they need to be simple so I can explain better here in the presentation. But the idea is like to extend.
the same principles, the same practices for complex objects and complex relationships. So let's let's see this demo application Basically it's a web app and the first main page, the index page, actually the slash post page, list uh a sequence of polls. Basically polls are like questions that will have like options for a person that that can vote And we have some uh some specific things here in this uh in this page. First thing is that every pool that is new, this means that uh it was created in the last 24 hours we will we'll present this label here
new also all the premium premium questions like premium uh uh uh premium polls we have this star here in the front besides uh a question or poll um can have an outer. So uh can have or or not. So if there is an author it will be appended by the by the to the end of the the the poll. So here for example it's saying coke or peps is the question and this question was created by Camilla. And is if you notice here we have different languages. So if there is a so this is like the the the final thing here
that we need to take notice of is that if the question is in English, then we are going to append by to the name of the author. So we are going to say Cocker Pepsi by Camila. But if the question is in Portuguese, we're going to translate this by preposition to the Portuguese corresponding one that is por So we have the question in Portuguese and then it's saying por João Silva. So this is like the name of the author. Basically, this is all the business logic that we have. And this is all the logic that we have for this application. Let's see more in details. Also, if you if you click on one of these polls it redirects you from to the vote section to the vote page that's the detail
of the poll so you can choose here it's if it's coke or peps that you want to vote if you click vote you are redirected to the results page that says the results of this current poll. And you can vote again or you can see the results and so forth. So here is uh the relationship the models and the relationship between them. First we have the model poll which contains the date published date, that's a date time field. We have the premium Boolean, so it's saying that a poll can be premium or not, true or false. Also we have the author that is a charf char field, so it's an open text. Uh to put the name of the author And the poll has a relationship one-to-one
to the questions. Question model, a question has a text that is a sharp field, an open sharp field. and a language that is a chart field but it's it contains like choices that in this case could be English or Portuguese And the last model that we have here is choice. And the choice is related to questions. So one choice. can have one or any questions and um a choice has mainly the the fields choice tags that is the option of the choice And also the number of the volts. So an integer that sums the number of the volts in that choice. So this is like the relationship between the models Nice.
Um also how can we add and change a poll as a question add questions? So this is is done by uh via Django Admin. So you have Django Admin where you can add a poll and then you can add the information about it, the published date, the question, and also you can add directly the choice for that um for that question with the votes so you can start with all zero or you can like put start with any different votes you want Nice. This is like an overview of the admin where you can see basically the question text, the language, how it's formatted, the name of the poll, and the published date.
Now talking about code, let's see the models the code of the models. So the first one is poll. As I already present, we have one-to-one relationship with questions, we have the published date. premium and outer fields. And we also have a the uh string method, uh dunderscore, yeah the string, that says uh that basically is responsible to format the text of the poll. So it grabs the question the question text and then pulls the star or change to buy and pour if there is an author. and also append the name of the author. So basically the string, the under string method is responsible to to format this. There is also uh one property of the model
poll, uh the poll model that is uh was published recently that basically check if the published date was uh uh was uh in the last 24 hours and not in the last day. Awesome. The question model is uh as I already described, it has the question text has the language and also has uh two uh two properties. One is to check if it's in English, this returns a Boolean, or it's in Portuguese question. And the string just returns the question text And we have the choice model here that is basically that basically contains a Ferengi key to the question, also has the question text and the votes. And this is like a simple a simple model.
Awesome, so let's talk let's start talking about the best practices. Why I started uh to why we started to work with this best practice and what happened that we came that you start thinking about you came with this this list of practices. So the scenario was that I I I worked with a huge, huge Django monolith. more than three three years with it. And in this Django Monolith we were using Factory Boy and um and uh in this Django monolith Besides of its size that was huge, it also has like really complex uh relationships So
to get a sense of how big this this monolith was, it had uh more than two uh 230 tables More than 2200 relevant files and more than 75k relevant lines of code. So it was huge. And not only huge but also super complex. The business logic was super complex and uh to work with it was pretty hard So basically we starting uh work uh working with this monolith and we're starting to notice that bad factories are like ver viruses So basically, when you start working with factors and they are not really well designed, you can get super tired
factories that they are not flexible and you cannot use in in the way that you would like to. So they are not customizable as you would like to. So also if you have like bad factories And you don't understand how the relationship with the factories are if they are not like really mapping from the database uh what can happen is that a lot of implicit errors are starting occurring. So you're testing, you take a look at your test and everything seems fine. But then you realize that the problem is like in the factory. So when you were creating the factory, when the declaration of the factory, the problem is like there. So you can start facing a lot of implicit errors. Also If you you create a factor that's not well designed
and um you this this actually happened a lot like we had one factor that was not well designed it and then we couldn't change it because if you try to change it it would break a lot of stuff. So what we did was we started uh using the factory but then changing everything that we need every time in a test. So we we fix it the this factory every time in the in the test suite. So what happens is that everyone is starting to copy and paste the solution from one place to another, handling like starting having a lot of really bad code on it. So a poor design factor might affect many tests. And the test like w this is like the thing
that we call it as factory-oriented testing. Like we started to think of how we are going to make this test passing in a way of the business article, but we started to think how we are going to change the test to make uh it passed using this factor that's not well designed. So it started to change everything in the test to accommodate that's bad design at factory. So this was like factory-oriented testing and not the way around it should be. So What happens is that the developer is starting to bump into the same issue again and again and again and copy and pasting, copying and pasting without being able to change the factory And also using it and trying to fix in every single place.
So it was hack um hack uh after hack after hack after hack So what we started to do is like, okay, we have to refactor this, we have to improve this. So let's get this really big factory here, this really important factory here. Let's start to like recreated it. Let's create it in the way that we think it should. Like afterwards we have like now some some experience and now we know how how is the best way to create it. So let's um put everything on trash this and start like recreating one only one single uh factory so we refactoring uh everything from that factory. What happened is that from a test suite that we have
more than thousand tests , more than twelve thousand tests When we changed that factory to be the way that we wanted, more than 1,500 tests were broken. So basically the problem is that they're really like virals. Once you create one not well designed factory and starting to use everywhere, you it's really hard to maintain it and to fix it And it generates a lot of duplication. So, okay, based on that, we start to figure out, we understand that we need best practices, and we start like creating this list of best practices And uh here the idea is to pass through uh everything that we learn
and I hope it helps. So the first one is that factors should represent their models This means that everything that we expect that it has in the database should be reflexed, shouldn't be mapped to the factory. It's the same way. So let's see an example. Here we have the poll factory and we have the question factory. If we see here in this factory we are creating a related factory pole Inside of the question factory. So we are saying that the relationship here is being meted inside of the question factory If we see here in the models , in the question model, there is no mention anywhere of poll.
So here in the question factor we are mentioning and we are creating the relationship with poll, but in the other hand in the question in the database itself doesn't have any mention. If indeed if we see the relationship is in the question, is in the poll. So poll points to question So the best way to do this like that we believe is that changing. So now the relationship reflects what is in the database. So basically the pole factory now points to the question and not the the other way around. This is also avoid having uh implicit errors because if you if you are already used with the database You already know the relationships, you already know how things work
in the database. When you see a factory, for example, you are going to use a factory and you are not seeing like how it's written, like how it's like initializated. You you only see like all factory, you already would imagine that there is a relationship with question. You don't need to go and to see the class itself Because you already imagine it. So the idea would be like you would expect that there is this relationship. And then if there isn't, or if the relationship is in another way around. Probably this would generate implicit errors, like errors that you wouldn't expect. The second best practice is do not rely on the false root factory. So basically if the default if the default value
changed, so if you are like uh relying on the default value and this these values change All the tests that depend on it will break. So basically the idea here is that you need to set up your tests to have everything to pass and not depending, not relying on the factory creation itself. So this mainly says that we want to explicitly say what we are going to use in our test and not uh get the attributes defined by the foe implicitly. So we want you to be like implicit here. So implicit is better than implicit. What would be in the terms of an example?
Here we have the poll factory, and as you can see here, we are setting in the question factory the default text question the the default value for the question text that is whatsapp so when we are testing the poll now we are getting We are certain that the poll text that comes from the question is WhatsApp. But if someone comes and changes this default value to something else, this test will start breaking. So we are not defining everything that we need in our tests. What would be the best way to do this? So we have the pull factory and in the question factory, instead of hard code the value like WhatsApp We could use faker here, for example.
So it would generate random sentences for you, and then you you cannot rely on this. And then on the test, you would explicitly say everything that you want from your factory. So in this case, we want a question, but it's not um is uh we want a factor in here for example but uh we want a pole here for example but we want a specific one right So we want one that is not premium with known author and with a question that's that 's explicitly saying. So here we are saying explicitly like it's a Pepsi or Coke And then we are going to assert it, we assert with this value. So if someone here changes or like the factory, like the factory
faker change, it will not break our test because we are explicitly saying everything that we want in this test. Okay, cool. Let's go to the third best practice. Factors should contain only the required data. So basically the idea here is if the field is nullable, where this means that in the declaration of uh in the database it can be null, so nu equals true The attribute should be under a trait and not as a default value. So let's see an example here. We have pull factory here and question factory again. And here we are seeing that for the pull factory the author is Joan. If we come here, we see that the author
in the model can be no. So we are here saying that explicitly for every pole factor that we create, if you don't say anything In CLIST we are going to have an outer Zoom already. Okay, and also here the question factory. So if you see the question factory here We have language, but the default is English. And here we are setting the default as Portuguese. So this means that everyone that Creates uh every time that we create a question factory, it will come implicitly. The language as Portuguese Brazilian, a Brazilian Portuguese and not the English one that we would expect reading the model, like
knowing the model. So what would be the best way to do this? The best way to do this is like first of all putting the alter under a trade. So since we can have polls without alter Let's make it if you want this outer we you need to be explicit. So if you say nothing, you not receive an outer. But if you want to have an outer, we are going to you need to pass with author equals true and then it will come an author with a random generated fake your name. So in this way You are being explicit and you are uh following the same uh the same structure of the database And here for the question factory, we just don't say anything about uh
about the language because by default it will come already with English. Okay, so the idea is that for the outer thing, now you can use pole factor with alter equal true. And this is especially important because like if in the database we can have uh there is the possibility of having an auto-requal known And you explicitly say that auto will be always have will always have a value. When are you going to remember to test auto equals none? It's hard. So we should not assume there is another uh there is an author when in the database actually allows to not have it. So the idea here is to be explicit If you want an alter, you have to pass through. If you don't want an alter, then you don't say anything
because this is already what the database would do. Nice. Let's go to the fourth best practice. This one is more related with performance. So basically it's to use build over create Let's first see the difference between them. So if you run myfactory. build, it will create a local object but in memory. If you run myfactory. create, it will create a local object, but it will also store it in the database. So when you use create, it really hits the database. And because of that, of course, build will be way faster comparing with create. Actually, oh okay, let's see
an example here. So Let's say an example of using create that is not like the best uh scenario for performance. So if you use create notice that you have to use the notation Django Tb because there because then you need to explicitly say that you need to use the database. Otherwise this test will fail if you do not if you don't use the notation So to fix it like to make it uh faster and also to be more like unit test because unit test wouldn't hit the database as at least we would expect that. So to fix it we use uh we would use build instead of create and then we don't need the notation and actually
When we noticed that uh because we as you can imagine we have this monolith with more than uh 10,000 tests and 12,000 tests. And it was like super super slow to run the test. And then once you figure out we figure out that the build was faster than create it was really game changing just to to show you an example um we did one uh one test with uh one file. So in this file we we only use it uh the create strategy so 14 tests passed in 3. 26 seconds. It took more than 3 seconds to run the 14 tests.
And using the same test , the same test file, but just changing the strategy. So starting using build instead of create, the test started passing in Less than two seconds. So the thing is, uh it was really, really way faster. And this is like only for 14 tests. Imagine for more than uh 10,000 tests. how fast it should be, right? So of course there are some times that you need to like really use the the create, you really need to hit the database and so on. So it's this this is just like a um Another device that's saying if you want things to be faster like you really want to be like doing unit tests and other integration tests and so on, uh use build instead.
So now we are in the fifth uh best practice, and this helps a lot because uh it took us a while to understand this this relation. But once we figure it out, it was super nice because we could map the database exactly as we want in the factories. So basically the rule says that if the forengic is in the table, you are going to express this relationship using a subfactory. If the Ferrangi key is in the other table, it is not there. So we are going to express this relationship with related factor plus trait. Why this happens? Basically because of the difference between superfactory and related factory. Super factory, it builds and creates the subfactory during the process of creating the main
factory. So the main factory and the subfactory they are created at the same time. Why on the other hand related factory creates the related factory after creating the main factory. So first you create the main factory and once it's already finished it creates the related factory. So let's see how we should do that So we have here choice factory in the question factory. Okay, so if you see here the Choice factory. We do have the relationship explicitly saying here in the model. So choice has a foreign key to a question. So when we have this relationship declared, like explicitly saying here, then we are going to use a subfactory.
So if you see here, we have the subfactory being declared here. Cool. And on the other hand, if you go here and see the question, question does not um mention anything about um about choice. So in this way we are going to use related factory with a trade. Why? First because choice in this in this first scenario Choice depends on question. So we need to have the question and then have the choice at the same point. But a question doesn't depend on choice. We can have actually questions without choice. So we cannot explicitly say that we always have choice. So that's for this reason we are going to use
trade because then we can create questions without choice And then if we want, we use the trait with choice, and then using the straits, because we are using a related factory What will happen is that factory boy will create first the question factory and then the related factory that is the choices So basically this is the way that you would create this relationship and in this way you can map exactly what is in the database The sixth uh best practice is to use fixtures. So I said that Factory Boy is like a Fixture replacement it is, but in this case we can still like in the there's some cases that we can still use fixture like if you have like really really simple uh things you to wrap
Or in this scenario specific, we can use fixture to wrap factors to avoid duplication. So this is a really nice use of fixture combinated with factors. So let's see this example here. We have two tests in the same file, but where they declared uh they they create this this factor this pull factor so we have the same pull uh being declared two times twice And uh if you see they are they are using the same attributes, so it's the same thing, duplicated code that we could avoid using uh fix a fixture here. So we create a fixture, give the name, explicit name of English, no premium with outer pole. So here it's like
this fixture returning this factory. And then we can reuse But this brings us to the next uh best practice that we can use fixtures we can use factors as we were doing so far but we should not and we should try to avoid sharing these factors of our fixtures both among different files Why? Because if you start sharing these factors and features, many tests depending on them. So if you want to change one factor or one fixture, you need to fix a lot of different tests in different files that makes even harder. Also, what
happens when you start sharing is that it tends to inflate the factory of the fixture. Because in one scenario we have uh these criteria and then you add the attributes that you want, but then for another file, another test, you need more different uh different fields, you need more traits, you need more a lot of more uh things and then you start putting everything in the same fixture or the same factory because they are shared So it's starting getting really big and it's hard to maintain, it's hard to to read. problem uh in the monolith where we have like a fixture file and uh one really really huge fixture with a lot of edge case
a lot of different things and we use it a lot of places so to change it was pretty hard And uh yeah, so if you change one, you break a lot of stuff. And this also uh brings the idea of you when the things are too too tight, you start and just changing everything on your test, fixing the factory, fixing the fixture in a way to make it pass. And you are not more thinking about your test, you are more uh you are now thinking way more in the fact in the factory in the fixture you are just trying to pass it so uh it it's not a good scenario so that's why we should try to avoid sharing fixture and factories in general So basically
all this uh this code, this fact, this best practice and uh are in the GitHub repository under my profile so it's Camila Maya slash factory boy best practices you can see all the best practices and also uh the demo app And the demo app also has like folders with good examples and bad examples so you can check and see uh the code running. Um and just as a reviewer, we kept The best practices are factories should represent their model. I think this is like the most important because if you really follow this one you almost uh have all the others The second is do not rely on the false or the false
factories because they can change and then they will break a lot of different tests. The third one is factory should contain only the required data to avoid implicit errors. The fourth one is performance, so build over create build is way faster. And also it's uh it it's it doesn't hit the database, so you're testing on like only the unit you want Fifth one is about uh the relationship. So if the Ferengia key is in the table, use a superfactory. If the Farrangia key is not there, is in the other table, other table, if the relationship is not there then user-related factory under your trait. The sixth one is avoiding sharing factory or fixtures among different files
uh because this makes uh makes harder to maintain, makes harder to to to to to use the fixtures and it's it's hard to read also so and the seven one is use fixture to wrap factories to avoid duplication so you can use fixture But uh it's it's a good way uh especially to avoid duplications or to avoid factors. So these are the seven uh best practices that we try to map And I'm trying to uh pass here. I hope this helps. I hope this helps you in your daily day-to-day work. And thank you very much. The links are here. So
this is the documentation from Factory Boy. You can check it, the official one. They also have the common recipes that basically are some tips of how to use the factories. There are some best practices also there. Reading this can also help you solving some problems. And if you want to check the code itself of the library, here is the repository Thank you very much. It's a pleasure to be here at DjangoCon US. If you want, and please let's let's keep in touch. My website is cmycd. com. I hand on Twitter. There is mostly where I'm more online. So if you want to get in touch. Send me a DM there.
And also my GitHub is Camila Maya. Thank you very much. It was a pleasure. See you. Bye.
Factory Boy is a flexible replacement for fixtures that lets tests create customized objects instead of maintaining static, exhaustive test data. It is especially useful for complex models and relationships, and provides tools such as sequences and fake data generators.
Discussed at 3:34Factories should mirror the database models: relationships belong on the factory corresponding to the model that actually defines the relationship. This makes the factories predictable and avoids implicit errors.
Discussed at 18:55Use Faker or other nonessential defaults in the factory, and explicitly set the values required by each test. That way, changing a factory default does not unexpectedly break tests.
Discussed at 21:10Nullable fields should not be populated by default unless the model requires it. Put optional data behind an explicit trait, such as an author trait, so tests opt in when they need the field.
Discussed at 23:27Use build when the test only needs an in-memory object, because it is faster and does not hit the database. Use create only when the test genuinely needs a persisted database record.
Discussed at 26:36A local fixture can wrap a factory configured with the same attributes and let multiple tests in one file reuse that setup. This avoids repeating factory declarations without requiring a large shared factory.
Discussed at 32:49Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026