Django for AI: Deploying Machine Learning Models with Django with Will Vincent
Published October 23, 2025
This video features Will Vincent at DjangoCon Europe 2025 in Dublin, Ireland.
Keynote: Django for Data Science: Deploying Machine Learning Models with Django by William Vincent
https://pretalx.evolutio.pt/djangocon-europe-2025/talk/STZPLT/
Will Vincent argues that Django is a practical, approachable way for data scientists to turn machine-learning models into usable web applications. He demonstrates the full path from an Iris dataset in a Jupyter notebook: train and evaluate a scikit-learn support-vector classifier, save it with Joblib, load it in Django, collect predictions through a form, and store them in the database for later iteration. He explains that data science usually involves large, messy datasets and spans statistics, machine learning, and data analysis, but the deployment needs of many projects are ordinary web needs such as forms, CRUD, authentication, storage, and administration. He also stresses that production deployment, large models, retraining, performance, version compatibility, and database-backed training require more investigation than the simple example covers.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: Good morning, thank you all. Thank you all for coming. I want to talk today about Django and data science. So I started a new job this year as a developer advocate at JetBrains. You can see here with the swag, uh working on the PyCharm IDE. And this means I get to focus on web tooling, which is fun, but also around data science, which I'm not gonna say is less fun, but it's certainly less familiar to me. And like many of us, I've been head down in the web for years now. And there's plenty to keep us occupied as we've seen with the talks today. Lots of things going on with Django, double-digit PRs every day, new advancements. But it's clear that in the wider Python world, data science has taken over and is where the momentum
Speaker 1: is So it's been very eye-opening for me and really good to be surrounded by colleagues who focus more on data science and to not just have an include an audience like this where we all agree with each other that Django is what you know what we should focus on So JetBrain started with a Java ID IntelliJ, and that's still our biggest product. PyCharm is one of the top ones, but it's definitely second fiddle. And so I come to this talk, again, having spent time without a web focus and without Python as even the main focus, so hopefully that's a bit of an outside perspective. So the TLDR version of this keynote is that it's surprisingly easy and quite fun to train our own machine learning model. I'll walk you through that today in a Jupyter notebook. And then while deploying a production-level machine learning model like ChatGPT takes a lot of resources and engineers, we can do it in Django
Speaker 1: very easily, and I'll show you how to do that. And I have a GitHub repo So you don't have to take notes or anything, but I'll do it all in one talk. Um just to prove I'm not doing the typical hand-waving thing. So I will do a bunch of talking, but there's also real code involved as well. And then I want to talk about what data science even means these days, because it can often seem like everything but web. You know, is it statistics, is it AI, is it machine learning, is it data analysis, is it scientific computing? We'll get into that. Um at the end of the day, it's all Django, uh excuse me, all Python under the hood, so it's not that foreign to us. So very briefly, this is me and the slide. Uh three books in their fifth edition. Um if you go on libgen, the pirated book database used to train all the LLMs, there's double-digit versions of them.
Speaker 1: So I take that as some validation that you know. There's more than more than three, more like 15 out there. Um the last six years I've co-hosted Django Chat Podcast alongside Carlton Gibson, as well as writ co-written Django News newsletter with Jeff Triplett, who's on the board of Django Um I spent much of last year building out Learn Django. com, which is an online site for all my resources, and I want to do a future talk on building a subscrip uh payments site in Django from scratch. And then since January I've been a developer advocate at PyCharm focusing on the IDE. We just launched a ton of AI features last week, so I'm happy to talk to any of you outside this talk about that. But this morning the focus is on Django and data science. So let's try to define these terms a little bit. All right, I want to do a quick show of hands.
Speaker 1: Who considers themselves here a web developer? If you could raise your hand. Okay, almost everyone. All right, how about data science? Data scientist. Okay. Three. How about so that's pretty standard. Uh the for the three of you, do you consider yourselves equally data scientists and web developers? Okay, I want to talk to you after the talk. I think that's pretty rare. Most of the time we have web developers, we have data scientists, you throw information back and forth over a wall, but there's very little actual overlap. You know, data science can seem scary and we'll talk about it because there's lots of data, lots of maths, but they're terrified of the web in Django. I mean, I gave a version of this talk in Boston before to mainly data scientists and I think I sometimes and we sometimes forget how
Speaker 1: intimidating and how much knowledge there is actually in building websites and using Django. So we'll talk about that. Um JetBrains has run this annual Python survey for many years now. And if you see the top two results, you may not be able to read that. The top one, 44% says data analysis for what do Python people do. followed by web development at 42%. And these trends have continued over five or six years. So it's pretty clear even back in 2017 when they started this survey, this is what Python people are doing. They're doing data science or they're doing web development. But again, what even is data science? This is an old Twitter account I followed back in the day that because it's easy to feel like data science is everything but the web, but in some sense It is kind of just statistics on a Mac.
Speaker 1: And it is true. I don't know. Nobody seems to use Windows these days. Again, just one more, but what do you actually do? Again, 80% of time prepare data, 20% of time complain about preparing data. Like this is kind of the real-world reality. If you study in university, you have these beautiful algorithms, these nice clean data sets. Then you go in the real world and you spend all your time cleaning the data and it doesn't fit, it doesn't match. Um and so I saw even back 10 years ago, people with PhDs from Pick a School go in the real world and it's fairly frustrating that they wanted to use you know all their academic mind and they're just cleaning data the whole time. Um but the basic point is we have Lots of math, we have big data which requires cleaning, and in a hand-wavy sense, that's kind of what data science is. But the interesting thing is, again, the
Speaker 1: amount of data is just Hard to conceptualize. So we're just going to use LLMs as an example here, which is just one form of data science in AI. But this shows how much data you have uh to be trained how much public text is available in the world. Um so GPT three which came out a couple years ago used 10 to the 11th number of tokens. A token is just Simple explanation is characters in a word, so dog would be D-O-G, three tokens. Um so we're now at 10 to the 15th tokens for all human-generated public text ever. Let me say that again, that's one quadrillion tokens, which we can say is roughly equivalent to a character and a word. You know, Google estimated in 2010, 15 years ago, there were 129 million books published.
Speaker 1: That's probably at least double now. There's over 1 billion websites, tens of trillions of index pages. Add in social media content, emails, forums, newspapers, it's easily one to two quadrillion tokens available to train. And the thing is, that's just text. We're not even talking about audio, video, real-world information such as a self-driving car. So hard for us to conceptualize when we talk about millions of rows in a database for Django. But ultimately, this is kind of what you're doing in data science, right? You're taking unimaginable amounts of data and you're trying to focus it and extract insights that you can use using statistics, machine learning, and computer science. And so just common examples we see in everyday life, spam filters in email, mapping technologies, right, with Google Maps, Apple Maps, recommendation systems.
Speaker 1: Healthcare early detection, LLMs like we talked about, finance fraud detection, weather prediction for farming, crop yields, on and on and on. And I do want to make the point that today Python is the dominant, almost the dominant programming language, and certainly very dominant in the web space and the data science case. But this was not the case. It was really the rise, like this has risen in the last ten, fifteen years. If we go back all the way to twenty ten Um Django was only five years old. Flask had just been released on April 1st as an April Fool's joke about how small can you make a framework. Django Rest framework wasn't released until 2011. Starlit, the lightweight ASGI , also from Tom Christie, was not until 2018.
Speaker 1: Fast API, which uses Starlit, didn't come out until 2019. Um and Django start starting then also rolled out its own asynchronous support. And then Django Ninja, which is quite popular for APIs, only came out in 2020. So It seems like Python's everywhere now, but certainly when I started programming, Python and Django were not dominant choices. And the same thing is true in data science, actually. So again, back in your time machine to 2010, R and MATLAB were much more dominant. Um pandas, which is now a default for data manipulation data manipulation, excuse me, only hit point one release in 2010. Uh NumPy for numerical computing, so large arrays and matrices. Uh became mainstream only really in recent years. CBORN, which you use for visualization, which we will I'll show in the demo, 2013.
Speaker 1: Jupyter Notebooks only spun off from iPython in 2014. And then the machine learning tools that again we're going to use that are common now, such as Scikit Learn, only first came out in 2010, TensorFlow in 2015, PyTorch 2016. So this talk assumes we're, you know, everyone uses Python for everything, but it's important to sh say even in my not super long career, that has not been the case. All right, so let's let's train a model. Um source code is available on GitHub, so you see it all at the end. Don't take notes or anything. I want to talk to you about how you do this if you've never done this before. And this is an intentionally simple example, but the general process applies to training any machine learning model. And again, machine learning means we give the inputs and the outputs in the computer through algorithms.
Speaker 1: comes up with a reasoning on its own. That's what machine learning means. All right, so if you start with machine learning, you learn there's basically two big data sets for classification problems. There's Titanic dataset, which is it's a bit morbid, but it's literally who lived and died. And then iris dataset for what type of species of iris flower. So Titanic is often considered the hello world of machine learning because it's clean enough not to overwhelm beginners, but it's messy enough that you have to do a lot of core machine learning concepts, regressions, decision trees, random forests, these types of things. For Titanic, there's a total of 1,309 passengers, 891 survival outcomes, and then you get information like passenger class, first, second, third, sex, uh, male, female, age, number of siblings.
Speaker 1: etc. And then you can make predictions and parse out the data from there. But the iris data set, which is the one we're going to use, is even simpler. So in this case There is only 150 rows, so 50 sets of measurements for each of three species. Um there's no missing data, so we don't have to do any pre-processing. So we can focus just on doing the model. So in short, if you're starting with things, I would recommend starting with Iris and then move to Titanic. So we're skipping all of the cleaning data, which is a big part of machine learning, just to focus on doing the model itself. Okay, first step, Jupyter Notebook. Right? There are multiple ways to do this. You can do it on the web with uh jupyter. org. You could use Anaconda, um, or you could use a text editor. Uh
Speaker 1: PyCharm, VS Code have their own versions as well. We're just gonna use PyCharm here, but it doesn't really matter how you use your Jupyter. So I know you can't see this, but this is if you created a new Jupyter notebook in PyCharm. It comes with data and models folders. There's a requirements. txt file, a README file, and a sample. ipynb file. That's where the notebook is. It's not uncommon to train and retrain multiple models. That's why you have a models folder. But we're just going to focus on one here for simplicity. So this is what we're working with. This is what the iris flower looks like. There's three different species. There's cytosis, versicolor, virgenica. And they have different pedal and sepal uh measurements for width and length.
Speaker 1: So it's a balanced data set, which makes things a lot simpler for us. And the goal is we want to train a model so that we we will add in our own petal and sepal information and it will predict for us what flower that is. Again, this is just another look at the CSV file. On the far left, you have the ID, and then you have sepal length and centimeters, sepal width, pedal length, pedal width, and then species. I would mention if you do this yourself, there's actually different versions of this online. So you get slightly different CSV files, which blew up a couple days of my life. So just be aware. This one comes from Kegel, but there's they're all slightly different for some reason. But this is so common that it's included by default in machine learning libraries like ScikitLearn.
Speaker 1: um our MATLAB because again it's sort of where you start if you're getting your uh your feet wet. All right, let's do a little code. So how what do we do? Install two packages. So right now we're just doing pandas. So it's a library for data manipulation and analysis using data frames, that kind of thing. And then Scikit Learn. Excuse me. This is a machine learning library. uh for data mining, pre-processing, model building. Um so we'll use both here. And again, this is making it as simple as I can. Okay, I'm gonna try not to dryly talk through code, but there's a little bit of code. So this is just one Jupyter notebook. Um again, source codes on GitHub, but just to talk you through how do you train a model. So we import the libraries at the start. Panda is to load and manipulate the dataset.
Speaker 1: Train test split from ScikitLearn, so that's to split the data set into testing and training sets. something you do in machine learning, an accuracy score to evaluate our model, see how well it performs. Um SVC, so that's support vector classification to train an SVM, which is a support vector machine classifier. I'll talk about that in a sec. And then joblib um to load and save our model. Joblib is a binary file that um stores a serialized Python object It's included with Scikit. The end result of what we're going to do is we're going to have a job lib file that we can move over to Django and show to people who view the webpage. So an SVM classifier, so supervised means trained on labeled data. So that's what we have here where it's actually labeled. In the real world, often you have non-labeled data, so that would be unsupervised.
Speaker 1: Um you load the data set. Creating a variable DF here to read the iris. csv file into a pandas data frame. We're going to extract the features as columns and rows. So there's four feature columns, three species labels And then we're going to split it into training and testing data. So 80% training, 20% testing. 80-20 is a common split. You could use a different split and in the real world. You would depending on your needs. So 70-30 is more conservative. If you really want to be accurate, 90-10 would be for large data sets where you don't need as many tests. All right. Train the SVM model. This is kind of where all the action happens. So we're creating an SVM classifier, setting gamma to auto.
Speaker 1: So that's the kernel coefficient. So this is how tight our model is around the data uh around the data. A lower kernel smoother might underfit. Um higher kernel be wiggly might overfit. As a default, we can just use auto here to not worry about that. And then model. fit trains the SVM model on our training data. Then we are going to make the predictions to test the model and evaluate its accuracy. And then we save the model using job and reload it. So we trained it once, and then we have it. We don't have to retrain it every time we want to use it. All right, this is the last little bit. So we get the user input for predictions, and I'll show you what this looks like in just a sec. So there's prompts to enter in the four values, so sepal length and width, pedal length and width.
Speaker 1: And then try to predict the species based on user input. All right, let me just show you. Whoops. Uh-oh. I have a live demo. Where'd it go? Well, that's too bad. Huh, that's two hours of my life I don't get back. Well You just have to trust me on this. You can load it yourself that if you hit run in the Jupyter notebook, you can enter in the inputs and it will show a prediction. That is deeply unsatisfying. You know, I when I gave an earlier version of this talk, I was flipping between screens at the live Jupyter notebook, but from Tim's example and others, that never worked, so I thought loading a video would work, but Okay. Trust.
Speaker 1: Oh, oh it's oh it's not playing on mine. Wild. Yeah, so you can see this is in the Jupyter notebook. We're entering in our four things and scroll down to the bottom. In this case it says Versa color 97% accuracy. Thank you, Adam. Okay, so it does work, but it doesn't show here. That's very odd. Okay. But now we can do something cool. Now we want to visualize our data in our model. So we can install cborne and matplotlib to help us do that. And then we add a new cell to our Jupyter notebook, imports both, runs a basic pair plot. There could be a whole talk on what a pair plot is, but it's basically an easy enough default to get some visualizations, which I'll show you in a second.
Speaker 1: Um but it's it creates a grid of scatter plots and histograms. So this is a very basic visualization, but you can see the clumping of data, and this is actually important. So in the top left, is the ID, so those are the same, but then you have visualizations for each four. So you can ignore that that top row. Those are nicely separated. But you see uh sepal length and CPAL width here and you can see that Orange and green are clumped together a bit, whereas blue is separate. That's actually good because that creates a challenge essentially for our model. It's not if they were all separate, there wouldn't be much for our machine learning model to do. So it makes our classifications look uh work a little bit harder. And then this is the bottom, the bottom two for petal length and width. And again, you can see clumpings for orange and green, which is Verstacolor, and Virgenica
Speaker 1: where citosis. Is on its own. The important thing is that we have a trained model here. It exists as an iris. joblib file, and now we can deploy it to Django as a web app. So now get to the comfort space. This is our game plan. So we're going to create a new Django project, load the job lib file. Add forms so the user can make predictions, share user uh store user info in the database, and then maybe I'll show you deployment if there's time. So an earlier version of this talk, I waited until the end, but if you actually pull out your phone or laptop right now, you can see what we're going to build on Django for datascience. com. And that'll maybe help help get you through talking through the code that's coming.
Speaker 1: So I recommend you check that out. Do I have? Is it gonna work? This is what this is what you will see. So you web page, you can enter in your information, make a prediction. And it varies depending on the type of flower. So again, this is just basically our job lib file thrown into a Django website, and we're storing the information too. Again, a different version of this talk. I would show you the live admin as you're typing things in to prove that this all works, but hopefully you trust me that it actually works. So that's what we're gonna build and this process will apply to any model that you make with a relatively basic Django website.
Speaker 1: Get my thing over here. Let me see how I here we go. Okay. So I'm I'm gonna walk through the process for a new Django project. I'm gonna go a little bit fast, but I do wanna show all the steps because I know it's easy. Many people in this audience are very familiar with Django and this part maybe is not as interesting as more technical deep dives, but anyone watching or new to Django I remember being so frustrated when someone waved their hands and skipped a step. So I'm gonna go through the steps. I might go a little fast, but all the code is in the repo, and I want to show how you do it because it's not that many steps and it's the same thing again and again and again. So we're going to create a new pro Django project from scratch. So
Speaker 1: you could create a new if you did it in terminal, you could create a new directory on your computer, new Python virtual environment, uh Django admin start project. Start app to create a predict app, um, update installed apps and settings. py, create a GitHub, uh Git repo If you're in PyCharm Pro, you can just do it from this screen. But again, it doesn't really matter how you do it. All right, so this is the layout of our new Django project. And this is one I would recommend in general for for projects. So we have Django Project, that's our project folder. People can call this anything. I like to just call it all Django Project, but that's a whole separate forum discussion. Predict this is our app, this is where we're gonna put our focus
Speaker 1: um templates, right, for our template file. Um and again, I think it I sometimes forget, but all of this structure, this exists for us as Django developers. Django doesn't care, the computer doesn't care. You can have one app, no apps, like people in this room can and do do very different things with this. But if you're new to Django or you're just doing kind of a vanilla version, I think this is as safe as it goes. But again, this is just one approach. Django doesn't care how we structure things, but I like to take advantage of the project and app separation. And we're going to do it that way. Right, so you you're going along, you do your project manage. py runserver, Jenkha welcome page, which we all know and love. And then again, we just want the iris. joblib file.
Speaker 1: So we can copy it over, you see it in the middle there, to the project level in your in this um in our application. That's all we have to do. Now again, if we had multiple models, we'd have a models directory, and those are more complicated real-world things, but for demonstration purposes, we just pull over the one model that we want to work with. You probably want to run python managed. py to migrate and get rid of all the um warnings about unapplied migrations. And then as ever we need URLs, views, and templates. So the order really doesn't matter, and that really trips up beginners. But I like to start off with URLs, so we're going to do that. So we just have project level Um file here we have got uh uh open empty string, excuse me, um because we're just gonna have it at the homepage.
Speaker 1: Um And we're including the predict URLs, which we'll do in the sec uh do down below Again, I gave this talk previously to a data science crowd, so I focused a lot on the Django piece, but I'm going to go a bit faster because we're a Django crowd here. Um then we add a view. Um again, function-based view. whole talks you could do on function versus class-based views, but we'll just do a function-based view call it predict, render a template that's called predict. html And then we just create the simple template. So to step through this iteratively, we don't do anything other than just have hello in it for right now. Run server, right? Bob's your uncle, you're good. Okay, so now we get to interesting stuff. So this is the views
Speaker 1: file. This is Not that scary. This is really where things are happening. Um so at the top we're gonna install job lib in numpy. Um and then we're gonna load the model from our base directory. We have post requests for the form, four inputs to match the four options Um make a prediction using a numpy array, and then return that as a variable prediction that we send uh to our template. Um I should note that we also need to install ScikitLearn if we want to load the job lib file to create it. So separately and then your requirements. txt, you're going to need ScikitLearn. Again, that code is in the repo you can see. Okay, update our template file. This is a little bit hacky, but some basic CSS.
Speaker 1: And then here's the form that has the information where the user can put in um their guesses or their measurements, um and then we would get this. Right? So this is our basic form. Um intern predictions. And then here's the results if you did one, two, three, four. So you would get iris forgenica. Now one, two, three, four are terrible, like that's not actually what widths and lengths are for some of these things. That's why in the live site I put some boundaries, but just for demonstration. All right, so let's keep going in the view. Now we're going to add a dictionary called form data to store user inputs because it's nice to store what people had. Um this is often the case where you make a machine learning model, you test it on users, you s you see how it works, and then you iterate on it.
Speaker 1: So for example, if you had a recommendation engine, you would Build it, test it on users, store that information, retrain the model, and you do these feedback cycles. This is why adding uh storage in the database is important. And I thought it was I thought it was kind of cool how easy it is to do. I'll show you in a second. So dictionaries populate by values from the form and the request. We're storing them because Django clears form fields by default. And then we pass this form data to the template context at the bottom of the file so it can be rendered on the page. All right, and then finally we're adding the inputs and the form data. Um nope, that's not correct. Uh
Speaker 1: Why am I showing this again? We'll just skip that. All right. So very ugly. Here's where we are. This is cool. All right. I feel safer now. Let's talk about the models. So obviously, if we want to store data in a database, we need a models. py file. Um we're going to create one here called iris prediction. Just use float fields for the four inputs. And also store the prediction and just for the heck of it, we'll add a created at um date. And then the string method to show the prediction date and time when it was made. Um if you were doing this step by step, you'd you know you'd make migrations here, run migrate. And then this is almost the last uh slide with code, I promise. Um we update the view to save the prediction. So we import the model at the top. And then we do the
Speaker 1: iris prediction. objects dot create to save the prediction to our database. And this is the last line of code. So we would update the admin so we could view it. Again, if you were doing this from scratch, create a super user account, log into the admin. You know, it looks something like this, very vanilla, but very functional, and you can customize it as we want. Um I had a version of this where I was going to show you how to do deployment, but then I realized that's probably a 40-minute talk. But I just want to give you the short version, which is this to me is the deployment checklist you can and should use. So for example, I was able to take the live site and in 15 minutes Put up the version you have now on a custom domain because I've done this a bajillion times.
Speaker 1: So I'll quickly talk through it. You can do it differently. I think this is pretty much the bare basics to have a not wildly insecure site. So configure your static files, environment variables. A lot of people like Django environs. I'm partial to environs. It really doesn't matter as long as you have environment variables. Create a. m file, update your dot git ignore file, so to ignore the. m file, otherwise what's the point? Um update your settings, right? So debug allowed hosts, um secret key, CSRF trusted origins. Um update databases to run Postgres in production, install Psycho PG. So if you install environs, there's a j uh Django. configuration where it will automatically instore um DG DG database URL, um
Speaker 1: extra goodies that do that for you. So production whiskey server. Gunicorn proc file because uh because Hiroku, but that would vary depending on what uh hosting provider you'd use. Umdate requirements. txt file. And then just create a quick Heroku project, push the code, start a dyno process. Again, it seems like a lot. I can never remember any of this, but that's why we have checklists. And I would strongly, I feel pretty strongly recommending this for a basic Not wildly insecure setup. Of course, you could do a million more things. All right, this is the last slide, I promise. So here's basically the takeaways. So Django is great for deploying machine learning models. I think most data scientists, they just want what I showed you. They want store in a database forms. They want all the basic features that Django gives you out of the box.
Speaker 1: There's often this sense that Django is hard to use and so they'll use Flask Or maybe Fast API, just because they think Django is difficult. Um and no disrespect to those frameworks. They have their uses, and if you know them, use them. But Django is built for this use case. Just take a model. forms like we give everything you need out of the box and you can more or less follow the code here and apply it to almost any basic machine learning model. Again, Iris is a great data set. Titanic, it's really fun to train machine learning models. Like it's you don't have to know all the maths to do it. In fact You don't really I mean to use it you don't need to know any of the maths to understand it you do but you can go a long way just playing around and following tutorials um and then deploy it in the real world, right? If you have your machine learning model, there's no sense having a Jupyter notebook
Speaker 1: thing locally, like you can easily share it with friends, others, colleagues. And in a real-world setting, this is what data scientists want to do. They want to take their model, they want to expose it to users and do that iterative loop of retraining it. Okay. Thank you for your patience. I'm happy to take any questions.
Speaker 2: Can I just pick up at that point at the the end? Do you think we're mar when failing to market Django to the data science?
Speaker 1: Thousand percent
Speaker 2: Dope.
Speaker 1: Yeah. Yeah. I mean I don't think that Nets I don't know that other web frameworks are doing a better job, but Again, it's I think for us in this room, Django doesn't seem so scary and difficult, but I'm telling you, people with PhDs and machine learning Are scared of web development and Django. Um and I think just it's just, you know, it's batteries included, it it could be presented better. The whole point of this talk was to show you It's really not that much much code to train a model or to do the Django bit, and the process is the same. So hopefully this helps market Django a bit better for that. You know, and you don't have to install a third-party forums library like Django comes with every most of what you need in you know a basic setting to deploy your your model.
Speaker 3: Thanks, Will. In the okay example here you you trained the model and then provided that to users. Are there any gotchas that you can think of for your users to be able to train their own models to then show or provide to other users. Is there anything different that you would want to
Speaker 1: probably, but I Can't speak to it off the top of my head.
Speaker 4: Hi Will. Um you went through all the steps to set up the you know, showing off your model in Django and showing how accessible that actually is. And you you didn't skip through anything.
Speaker 1: No.
Speaker 4: Um But then in the deployment checklist.
Speaker 1: Yeah.
Speaker 4: As you say, you know, that's that's a quick process for you. You've done it a lot, but I think to someone new to web development, that process has actually got A lot of warts and a lot of hairs. And is there anything to make that more accessible for new people?
Speaker 1: I mean there are books. I've written a few. In-depth step-by-step guide. I think the thing is it depends on the project you have. So the steps, you know, because you have so much flexibility with Django, you have to know what the project is before the steps that you say apply. And if you know one thing is off to a newcomer, they're gonna get totally frazzled. Um so for example, my Chango for Beginners book I show you how to do a bunch of projects and I show you how to do all the steps. Um but yeah, I do think about this. I think you know Django has a great deployment checklist. Um I would like to make that more accessible, but I have seen that beginners their prod their projects are different enough that they get tripped up.
Speaker 1: So unfortunately it's difficult to just say this is Exactly how you do it unless I know exactly what your project is. But I I completely agree. Um you know, read a book and you got it.
Speaker 5: Thank you, Great. Just a wild idea slash suggestion. Perhaps this can be a great official Django tutorial to leave alongside the other one that you can send to data science folks and say, hey It's not that hard and this will help push Django to more people. That's all I mean.
Speaker 1: And we could have a Hello World tutorial too that's simpler than polls while we're at it.
Speaker 5: As well.
Speaker 1: Yeah, I mean I'm not you know I'm not in the board or important anymore, so I'm happy to give it to Django if they want it.
Speaker 6: I realize you glossed over it in the talk. But you spoke about briefly this Django underscore project thing. Yeah. Where should I look to learn more about that particular convention or that idea? Because I've seen it in a few places and it looks like it solves one of the continual problems I come across myself
Speaker 1: Tutorial or book I've written on learnjjango. com has that pattern. I mean I I I feel less comfortable saying everyone should use that. I use that because I see that people name their Django project different things. And to me, I just want to know um what the project and what the apps are. You know, in a real world repo, the the structure is a bit different, but you know if you have six, ten apps. I just like to say the name and project. So that's that's more of a personal thing. I used to call it um config because Uh my friend Jeff Triplet likes that approach, but other things are config too in projects offer. So Yeah, I'm partial to it. I don't feel confident saying like that's the one way to do it, but for me, just having having it called something
Speaker 1: project is helpful
Speaker 6: Thank you.
Speaker 7: Thanks, Will. Great talk. This was a small model and you committed it in the repo. Yeah. Most models are big. And we don't want to commit them and make our Git repo giant and we won't have push updates. Like what would be the next steps you take for larger models?
Speaker 1: Yeah. So I have the same question. I mean I would love to know Uh where is that limit with the Jupyter notebook? Because I think it's actually a lot bigger than we think of so in the broader world, um data scientists think of themselves as not great programmers, like below web developers, right? Which I think we're relatively low on the spectrum. We're not like you know nuclear submarine programmers. Um I don't know exactly. I would love to know where at what point can you like what is the limit of a Jupyter notebook and then when do you write Python scripts and how do you Do all those other things. Um so yeah, if I did another version of the talk, or if somebody knows, please tell me. But that's a very good question. I have the same one.
Speaker 7: Okay, thanks. Uh
Speaker 8: we actually have an online question, so I'll read it out. If Django were to be marketed better to data scientists, can you imagine it becoming one framework to rule the to rule them all for data science and web development.
Speaker 1: I'm old enough to say no. I don't think there's ever going to be one to rule them all, but I do think What data scientists want is crud with auth, with guardrails that just works. And so if you're not, I would say if you're not comfortable with web development, Django is great for you because it just It gives you batteries, it gives you things to do. It doesn't require you to be an expert. It doesn't ask you to make some of the decisions that other web frameworks do. So I think it could be I think it should be like the the top default. I think it often is not. Um and that's probably related to just not having tutorials or ways to do it. I think also, again, there is this perception that Django is really, you know, it's batteries included, it's really hard to learn. Whereas Flask is considered simpler.
Speaker 1: And you know, the first part of Flask is simpler. And if you need more advanced stuff, maybe you want all the flexibility of Flask. But if you're right in the middle and you want CRUD and auth and forms and stuff that just works. I think Django, yeah, sh should should be more prominent. But of course I'm biased, so.
Speaker 3: Hi, thanks for the talk. Can you comment a bit about uh the real-world uh let's say examples of using model? So if the resulting file size is big. uh about the performance of the whole thing. So how much time would it take to actually train it and how much time would it take to uh query the thing and get the results back. Thank you.
Speaker 1: Yeah, that's a great question. I don't have a great answer because I'm not a data scientist, but I'm spending this year learning a lot about data science. So hopefully maybe next year I'll have a better answer for that. Yeah, that's a very good question. Sorry I don't have the answer.
Speaker 4: No, no, no, thank you.
Speaker 3: Going to a bit more in-depth for long-term project maintenance Um you use JobLib here. Joblib uses pickle under the hood or replacement or something like that. Yeah. How do you ensure that that model that you trained once will run on whatever future Python scikit learn whatsoever version.
Speaker 1: So I didn't mention, so you could use Pickle or JobLeb um and JobLeb is preferred for larger data sets. I don't know the answer to that question. That's a really good one. I mean, there's still so much of data science that's mysterious to me, to be honest. I mean, most of what I know is up here. So, but yeah, I want to find out. I mean, I don't I haven't found resources of people talking about deploying models outside of, you know, at massive, massive scale. Um because I think people just don't do it that much. Um but yeah, that's a great question. I'll research it, but I don't know.
Speaker 3: Thank you.
Speaker 6: Thank you for the talk. In in this demo you you use the uh CSV file for the source of data for training the model.
Speaker 1: Yeah.
Speaker 6: Uh in a Django project usually you have a lot of data in your database. How how easy it is to take like a query set instead of a CSV file as a source of data?
Speaker 1: I don't know exactly. I could make predictions, but I haven't done it myself. So um Yeah, that's another great question. Again, I mean even for me presenting this, I proposed this talk with the idea of like how hard is it? How hard can it be? I mean to data scientists, the idea that you can just move the job lib file over was sort of mind blowing because they're used to big production scale. So this that's a great question I have. I want to push it further and see like where is that limit? How much can we put just within standard Django structure I know I would love to do a demo to retrain the model because that was feedback I got from an earlier version was hey had you know that's what we do in the real world. We have the model users take the data from the database, retrain the model Maybe next year I'll have a demo showing that.
Speaker 6: You started by saying these data sets are the hello world of um uh of data science. Um can you tell us a little bit more about
Speaker 3: what
Speaker 6: people in data science see as Not just the Hello World data set, but the the Hello World problem for us.
Speaker 5: We know what it is, it's to get a web page up that's displaying something we want from it from a database. How should we be thinking about data and data we could use or data we could try and construct for purposes like this.
Speaker 1: Um so Kaggle is like a big data one of the big places with Tons and tons of data sets that you can use. I think it depends what you're trying to do. I mean if you're a data scientist, again, iris is good here just because it's easy. You don't have to pre-process it or clean it. So much of what you do is around that. So Uh I hope I'm answering your question correctly. Like a lot of what you're doing is cleaning data and trying to get the accuracy prediction and you know how is the data clumped, which classifier do you use? That's a lot of what data scientists Do I feel like I'm not answering your question directly though. Um that's you know for me to go further. I mean there's a number of books. I mean the thing is you can be overwhelming if you read an entire book on You know, pandas, an entire book on psychic learn. Like I would recommend people go to Kaggle and
Speaker 1: follow tutorials and learn kind of the basic Cleaning and training kind of steps.
Speaker 6: I just realized that actually Hello World is the wrong metaphor. The right metaphor is the first thing we do in a tutorial, which is for example to make a to-do app. Yeah. Yes, exactly.
Speaker 1: I mean it's basically what I showed. Uh that's and again I I played around with a few. That's that's as simple as it is. We actually do something and then the fact that you have two of the three are clumped together means your model has to do some work. Because of course if the data was just completely separate, you wouldn't need A model for it. So yeah, I I iris is as simple as it gets. And then most people they focus on Titanic because it's busy enough and big enough that you can get a taste of the problems you encounter as a data scientist without it being overwhelming. Again, but it's it's so morbid, like really?
Speaker 3: Uh thank you for the talk. It was a great talk. Uh I'm going to put you a little bit on the spot and I blame Corton who uh said we should do that today. Um you showed that the results for the PyTrum survey for uh the past few years. Do you know when we're going to get the one the results for twenty twenty-four? uh
Speaker 1: soon I've seen it and reviewed it months ago. Um I've JetBrains does a lot of work and a lot of teams to uh Put it out, but soon I hope, and I hope that cycle improves. Um but it's it's there, it's just inching its way through the the process. But yeah, that's a good question.
Speaker 3: Thank you.
Speaker 1: I I will say there's no there wasn't anything like Crazy that changed. Like if there was, maybe we'd put out a code red or you know in the responses. But you know, for next year, um, I think we'll probably ask about UV for package management. Um And maybe actually that's I need to do maybe on the forum, I want to get some more community involvement around the questions. I mean the board runs it now, but I'm trying to help um Yeah, make sure we ask the right questions because it does matter. It does help. You know, we had we saw that Redis had a lot of um support, so there was work done to make that official for caching. Um it is kind of our only public way to get feedback. Uh
Speaker 6: thanks again for the talk. So this presentation was using a bunch of pre-made data and of obviously it's it's data science, but do you see value in there being something like some sample data sets for Django itself. So people can use maybe the books with examples of real data to play with and things like that to demonstrate parts of the ORM a bit easier.
Speaker 1: Yes. I mean I think the challenge is uh I mean you could just make up data. It's nice if it's actually real world, but Yes, like especially if we had a tutorial or had sort of a get your feet wet, something beyond just iris and Titanic would be great. Yeah, I don't know why yeah but yes we should do that.
Speaker 8: Okay. There's no more questions Thank you, Will.
Speaker 1: Thank you, everyone.
It involves working with large, often messy datasets and using statistics, machine learning, and computer science to extract useful insights. In practice, much of the work is preparing and cleaning data.
Discussed at 4:02Use a dataset such as Iris, load it with pandas, split it into training and test sets, train a scikit-learn classifier such as an SVM, evaluate its accuracy, and save it with Joblib for later use.
Discussed at 9:31Copy the saved Joblib model into the Django project, create a view and template with a form for the model inputs, load the model in the view, make predictions from the submitted values, and render the result.
Discussed at 19:37Define a Django model with fields for the input measurements, prediction, and creation time, then save each submitted prediction with `objects.create`. This also enables feedback loops in which user results can be reviewed and used to improve or retrain a model.
Discussed at 25:01Configure static files and environment variables, protect secrets with `.env` and `.gitignore`, set production hosts and CSRF origins, use PostgreSQL, configure a production WSGI server such as Gunicorn, update dependencies, and deploy through the chosen hosting provider.
Discussed at 27:20Django already provides the forms, database support, authentication, CRUD features, and other building blocks that data scientists typically need to expose a model to users. The speaker argues that it can be used for many basic models without much code or deep web-development expertise.
Discussed at 28:53Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025