Recap
Published September 16, 2025
This video features Will Vincent at DjangoCon US 2025 in Chicago, Illinois, USA.
This talk was presented at: https://2025.djangocon.us/talks/django-for-ai-deploying-machine-learning-models-with-django/
LINKS:
Follow Will Vincent 👇
On Mastodon: https://fosstodon.org/@wsvincent
Website: https://learndjango.com/
Follow DjangoCon US 👇
https://fosstodon.org/@djangocon
https://x.com/djangocon
Follow DEFNA 👇
https://www.defna.org/
Video production by the presenter and DjangoCon US 2025 volunteers.
Will Vincent argues that Django remains useful in an AI-focused web ecosystem, even though FastAPI is increasingly popular among AI engineers. He distinguishes classic machine-learning deployment from large-language-model inference: a trained scikit-learn model can be saved and served directly from a conventional Django application, with forms, the ORM, admin, authentication, feedback storage, and familiar deployment practices. LLMs require much more computation and specialized inference engines, but Django can still provide the surrounding application, including a streaming chatbot interface using server-sent events, Django’s streaming responses, HTMX, and local models run through Ollama. He concludes that the Django community should make these capabilities clearer and more approachable to data scientists and newer Python developers; the recording cuts off as he is asked how to choose between local models and hosted AI APIs.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: AI is all around us. My entire trip here was powered by it, from the mapping software that guided my taxi to the search browser that found my hotel, the songs recommended on my playlist, all of that is AI and more specifically machine learning. A branch of AI where we provide the data, add algorithms, and the computers come up with the connections themselves. No explicit human code. Today, an even newer flavor of AI, large language models, LLMs, is transforming our industry. You don't need me to tell you that. And like ChatGPT, they can write essays and increasingly code. So right now the hype is as high as it's ever been. And I want to talk about this new AI feature, uh future, excuse me, and where Django fits into it. Let me see if this remote works. Boom. So really quickly, who am I? I'm Will Vincent.
Speaker 1: I've been involved in Django for a long time. I have some books. I do the Django chat podcast with my friend Carlton Gibson the last six years. Also write the Django News newsletter with Jeff Triplett going on six years. And I have run a site, learnjango. com, that has a lot of free tutorials and courses. And this year I've started working with JetBrain's PyCharm, who makes an IDE. So it's been really interesting seeing AI being integrated into that. And if you have any complaints or compliments, I would love to take them after the talk. So this is from the most recent Django survey, which will come out in probably two weeks. But there's those of you thinking, yeah, I don't use AI, but you're in the minority. Most people, uh
Speaker 1: 17% of developers said they don't use AI, but over 70% use ChatGPT, a third for co-pilot Claude, JetBrain's AI assistant. And again, this was taken almost a year ago. When we do the survey again, I expect these numbers to all be higher. So maybe you're able to not use AI, but you're in the minority. More interestingly to me as a content creator, how are people learning Django? This is again from the same survey results. Fortunately, it's still 79% are using the official docs. Followed by Stock Overflow, AI Tools, YouTube, way down to 22% is books, which breaks my heart a little bit. But since the LLMs are trained, for free on my books and those of others. Maybe there's some good things in there. But it does bring up some existential questions around why write a book when it's just going to be, you know, stolen.
Speaker 1: And by the way, just a quick shout out: DjangoBook. com is a website that I run again with Jeff Triplett, featuring all in-print Django books. It's not just mine. Whoops, excuse me. This was the original domain of the first Django book written by Jacob Kaplan Moss, if he's here, and uh Adrian Halavati back in the day So that's a good resource. Again, just there are still great books being put out in Django. I'm happy to recommend. I think all those ones on there I've read. I'm still a fan of books. What I want to talk to you about is I've been to a few conferences this year through my job, so EuroPython, PyCon US. And guess what everyone's discussing? It's not Django, it's AI, but when they talk about AI and or the web, excuse me, what they really mean is fast API. Now, this is from the most recent Python survey, which the Python Software Foundation runs with JetBrains.
Speaker 1: And this is a data point, but we can see that there's more usage of Fast API than Django. Now we can quibble about what does usage mean, but I don't think we can ignore these trend lines for what's happening in the industry. Again, GitHub stars. I can quibble all day long with whether this is a good metric, but guess which one is in red, right? That's Fast API. It now exceeds Django as well as Flask. And so I think for a long time we've in the community looked at these metrics and seen, you know, Flask is just ahead or behind and said, oh, how does that match up with our real-world experience where we know so many companies using Django as opposed to Flask being a slight add-on? But Fast API is really having a moment, and I want to talk about why that is and what we can learn about that. So what I don't want to do is stick our heads in the sand.
Speaker 1: So I think Previously, before I started going around to conferences, I was very Django all the time. And so it's easy to say, you know, Django's where it's at. And it still is where it's at. But there are changes to the traditional web framework and hosting that I want to talk about today that I think we should all be aware about. So this is what I hear when I go to conferences and I'm sitting in a booth, right, for eight hours a day. Python developers, right? And I say, oh, the web, like why not Django? This is the top three I get again and again and again. So Slow. Django is perceived as slow. Why is it slow? Fast API has fast in the name and it has async and everything needs to be async, right? No, but people think that. Django's old, right? We're rightly proud of 20 years
Speaker 1: history. But a 20-year-old says, I don't want to use a web browser older, a web framework older than I am. To Carson's point earlier, we need, we, the people with gray hair, need to explain the benefits of season technology a bit better. And then number three, and this is the one that's shocking to me, it's perceived as bloated and hard to learn. Now it's always had that reputation because Django's battery is included and has so many features. But compared to building an LLM from scratch, I don't think it's so hard. But people still feel that way. So again, I think for me and other content creators and people in the community, we need to make Django as approachable as we can. Because we're at a risk of this issue, where this new generation of, again, Python developers thinks that the web is just API endpoints.
Speaker 1: And part of that is Fast API is built into a lot of these new AI tools. It's just like, oh, web, boom, fast API, endpoint. Done. Now why do they think this way? Part of it is because when they read websites on you know hacker news and other things, all they see is people talking about Fast API and these things. Because it's people from OpenAI or Anthropic and these big companies. And this is a version of what used to happen with React and Angular. We'd have people at large companies, Facebook, Google, talking about the benefits of those approach. Those approaches and they have benefits, but not everyone is at that scale or has those problems. By the same token, I think data scientists and AI engineers need to know that like they can do web development too. I'm going to give some demos later. Like it's not that hard.
Speaker 1: There's a lot in there. That's why we all have careers doing it. But it's not an unapproachable thing, especially if you have the expertise in Python to begin with. And this is my second big point, which is that the web itself has actually never been more important. So these fancy models are useless on their own. They need to connect with users, hopefully, paying users. And they do that via inference, which is a radically different way of connecting with information than the database-driven approach that we're used to with Django. So this is Jan Lakun, who's one of the godfathers of AI, pointing out the real cost is not the training, but actually the inference. Again, that's serving the models. This is this post was from January. Remember when Deep Seat came out and for $6 million or whatever it is rivaled or exceeded the best ChatGPT models at the time?
Speaker 1: You know, wiped a trillion dollars from the stock market. And so there's a big I think people need to understand the difference between training and inference. And that's something I'm going to go into a little bit So every company is not going to be an AI chatbot, but they're all going to be integrating with them one way or another. And so again, Django has a role to play in this new web world. So my key point is the web has never been more important than it is now, but it's changing in some perhaps subtle ways. All right, so the game plan. I'm going to talk about classic machine learning models and then I'm going to talk about LLMs and again how Django works with both of them. So we're going to train a model from scratch and deploy it. We've never really, we, Django has never really owned this niche, but we really should, because Django is perfect for what most data scientists, small medium companies who train a model, and then they say,
Speaker 1: How do I how do I get users? How do I get feedback? How do I retrain it? They feel completely stuck. They have to have a web developer to do it for them Django's here. Like most of I'll show you in the demo. Most of what they want can be done. They just clone the repo and vibecode it. So I'll walk through that. And then the second part, I want to go into a little bit of L LLMs and how they work and again why it's different than the database-driven approach of Django and talk about the role Django can play despite that. Okay. So AI is a is a problematic term. From the beginning, it was always a bit of hype in marketing. There's a lot of areas within it. Machine learning is just one. Um the most high profile, and we'll fill this out in a moment when we get to LLMs.
Speaker 1: But again, machine learning means we as programmers don't explicitly program the model. We get a ton of data, we apply the right algorithms, we clean the data. And then we get results and we evaluate whether that works. So for example, image recognition. If we loaded a million pictures of dogs and cats, added a machine learning algorithm that computed for a while. The computer would figure out ways of identifying them that would be different than we as humans would. And then we can evaluate how accurate are they based on, again, having humans look at the results. So the classic workflow model is you have a training set, you make a model, you expose it to the real world with humans for feedback, then retrain. Okay? And Django again can work well with this. So forgive me if you've done some data science. Actually, how many people in this room have done data science before?
Speaker 1: Please raise your hands. Okay, maybe a quarter. All right. So the rest of you will learn something here. So when you do data science, there's the two Hello World data sets are Titanic, which is 1309 um people on the Titanic and whether they lived or died, sex, gender, all these other things, and then iris, which is about iris flowers. Titanic is the one most people use because it's just messy enough that you do a lot of the data cleaning and algorithm stuff that you do in the real world. But Iris is a great starting point, which is we're going to use now. Because clean. This is what the iris flowers look like, by the way. There's three different types: Setosa, Versicolar, Versicolor, and Virgenica. And they have different uh petal and sepal
Speaker 1: lengths and widths. That's all the measurements we care about. All right, so let's train a model. You need a Jupyter notebook. You can do this three different ways. You can go to web version jupyter. org. You can use Anaconda. Or you can use text editor like PyCharm, VS Code, that has built-in support and a lot of features. PyCharm has a lot more features than VS Code, but I'm not going to go too deep into that. So here's what the data set looks like if you download it, and you can get it anywhere. In fact, it's built in to most of the machine learning libraries, including ScikitLearn, R MATLEB. So there's 30 different species, 50 measurements, so 50 for each of the three, 150 total, super clean. We don't have to do any cleaning or processing as you would in the real world. This is another way you can look at the data.
Speaker 1: So this is if you pair plot in your Jupyter notebook, you can, and this is using the CBORN library. This is like two lines of code. The blue dots are Setosa. The orange and green are Versicolor and Verenica. So what what I want you to see is that blue is distinct, but orange and green, depending on how you look at it, have some overlap. And that's important because it means our model has to do a little bit of work. Right? If they were completely separate, we wouldn't really need to train much of anything. A few more visualizations. You know, you can slice and dice it how you want. All right, so this is how we do it. I gave a version of this talk at DjangoCon Europe and I walked through all the code, and I'm not gonna do that to you because I don't think you need to hear me do that and it's in the GitHub repo. But this is the process for training a model broadly.
Speaker 1: So you it do your imports. We're gonna use pandas and scikit learn, two of the staples. We're gonna load the iris. csv file. split into training and testing data. So 80-20 is a good default that a lot of people use. So we'll take 80% of the data and that's what we'll train our model on. And the other 20% we can use for comparisons. We can check the accuracy after we have the train model. We're gonna use an SVM model. I'll talk about that in a second. And then make predictions and evaluate its accuracy. And then finally, this is the important part. We save the model as a job lib file. That's a basically a Python object that we can uh Serialized Python object for saving loading trained machine learning data. One more point, SVM, supervised learning model classifier. So that's A type of algorithm that works well in this specific use case.
Speaker 1: I'm not a data scientist. There's a lot of different models, but this is a good one to use and to start with, and it's relatively simple. So it's about 20 lines of code in the notebook, again, which is online and I'll link to the repo. This is the important part, like training the model at the top there, it's two lines. Like that's it. That's the training part. And then it runs. It's almost nothing with the modern tools that we're using. And again, the data is clean, you know, but a lot of what data scientists do is actually cleaning and getting the data ready to be run. It's not most people are not, you know, handwriting these algorithms and doing all the maths for that. And again, I know I'm going fast, but I don't want to get bogged down the code too much. But around 20 lines of code, we can train a model.
Speaker 1: And then I should have a live demo here. Oh no, it's not showing up. Oh, is it showing up? Yeah. So this is a live demo. So if you ran this Jupyter notebook, you would see you can enter in different uh inputs for Cl and Pedal length, and down at the bottom. You can see it says accuracy 97% and identified it as um versicolor. It's kind of cool. Like, you know, again Maybe not the most awe-inspiring demo ever, but this is the process you would use for data science. What's super cool, let me go to the next slide. Oops. Here we go. All right, so this train model is pretty lonely on its own, right? It needs a web interface to connect to the outside world.
Speaker 1: Django. So what do we need our web app to do? We need forms, we need to interact with the model and maybe store the model predictions so we can retrain the model later. That's the kind of data processing loop that often happens. This is how you would do it. Again, I have a repo with all the code. I'm not gonna walk through all of it. Create a project, load in the job lib file. It's not that big. Even in larger data sets, it's not the size of LLM models, which I'll get to. We just need forms, models. py to store the user, the information in the database, and then deployment. Again, Django, 20 years we've been perfecting how to do this. We have an approach, it works. This would be the model uh the project setup, excuse me. So in this case I'm calling the project Django project at the top just because I prefer that, but you
Speaker 1: call it whatever you want. Notice that the iris. joblib file is there in the root directory. There's an app called Predict. And the other thing I would mention is in the real world, often you'd have multiple models you'd want to play around with and try. So you could create a directory for models. And you could actually do this with different models and different web pages. Same process. This is the key bit. This is a views. py file. It's a little long. The key point is at the top we load in the job lib file. Like was it, line four or five there? And then we have one function, it's a function-based view in this case called predict, accepts form values from the user and then makes a prediction using the model using numpy. Again, not gonna dwell too much on it, but
Speaker 1: code is there. This is the process. You do it a couple times and you're like, yep, yep, I got it. We can also add a model to store the predictions. This is super cool. Like FastAPI doesn't have an RRM with it. You know, Django's ORM is arguably a top three, if not number one, technical feature of the framework. This couldn't be easier to do. But you know what? You can try it for yourself. If you pull out your phones or lap or laptops, you can go to Django for datascience. com. And you can see this actually working. So I'm not just like, you know, saying I know how to do something I don't. I might be, but how can you tell? So while you're doing that, this if it works, this is what you're gonna see. This is the live website. You can see it's at Django for data science.
Speaker 1: com You can enter in your predictions. I added a little bit of you know graphic there for the flower. So you can again try it out. And I think it will go and show you the admin in a sec. Yep, okay. Are people getting to the website? Is it working? Yeah? Okay, right. And then this is the Django admin. If I wasn't scared of presentation mode, I would hop into the live site and show you know in real time. Trust me that this is from DjangoCon Europe. There'll be results from this room where we can see all the information and we can use it to retrain our model as we want. Okay, deployment. I could do, I probably should do an entire talk on deployment
Speaker 1: because it is complicated and it's not my favorite part of web development. But it's largely a solved problem. It really is. Like I have two books that cover this in great depth, Django for Beginners and Django for Professionals. The key point is we don't have we have this relatively small machine learning file that we can just serve with Django. We don't have to do anything fancy the way we will with an LLM. This, by the way, I would say is the refined deployment checklist if you're using a platform as a service. Again, I'll put links to these slides after. I was able to go through in 15 minutes, take the code, custom domain, deployed, ready to go. Again, because I've done it a ton of times.
Speaker 1: So This is why we have checklists, by the way. Like I forget these steps all the time, even though I've done it a hundred times. So don't feel like you need to memorize this. Just have a checklist Blast through it. I like platforms as a service. You can use whatever you want. But it's my point is that this is one of the benefits of Django, is this is a solved problem. There's ways to do this. You don't have to reinvent the wheel. And we can do, we can take relatively small, medium-sized models and serve it to thousands, millions of people, no problem. That's a use case that works for a lot of data scientists. They should all be using Django. Where can we go from here? Of course, we could add authentication, security, we could have some API endpoints. Again, you're using Django Ninja or Django REST framework. Solve problems in Django land. But we have to tell AI engineers and younger engineers
Speaker 1: like this is what we've been doing. Like we have answers. You just are asking the wrong questions. You're actually not even asking questions. Okay, so now LLMs under the hood. Let's talk about why they're different. That was pretty simple. I'm sorry to say simple. That was Easier than LLMs will be because we had just a job lib file that we could just put into our existing Django project and do all our standard web stuff. So at a high level, large language models are next token prediction machines developed for language models that turn out to have other use cases that are relatively unexpected. This would be the hierarchy of AI. So we have machine learning as one of the branches within machine learning, neural nets, deep learning is another one, and then within that are large language models.
Speaker 1: If you yeah, this is what a neural network looks like. So note that you have an input layer, hidden layers, talk about that in a sec, and then output layers. This is not a new idea. This was first proposed in the 1940s, actually, as a way of You know, how does the human brain work? Let's make a computer version of that. Now the irony is we d we still don't really know how the human brain works, um, but somehow this Actually seems to work relatively well. If you stack a lot of neural nets together, frontier models like ChatGPT are doing like over 100 at this point. You get a deep learning network. So deep learning just means multiple neural networks stacked together. Oh, one last point. Um
Speaker 1: so why didn't we have this in the 1940s? Uh for a long time there were problems getting neural networks to work. They were actually very much out of fashion. All the gods of AI right now, like Jeff Jeffrey Hinton, Jan Lakun, and I'm forgetting the French other person. They've been doing this since I think like the 80s. And then backpropagation was introduced, but it wasn't really used. It was actually AlexNet, which is that image um system that won a competition in 2012 by a bunch of grad students that showed, hey, there's something here. That kickstarted a revolution around uh revisiting an older um concept that had been around, but the problem was When you is that people would hit a certain level of um success with their models, but they couldn't go anywhere else because they didn't know what they don't really know what the hidden layers are doing.
Speaker 1: Like Okay, like I have my algorithm, I have my data, how do I eke more out of it? For a long time they didn't know how to do that. And I want to make a point. So this slide is taken from Simon Willison, Django co-creator. He has the analogy of large language models being alien technology. And he said, I'm going to quote, one way to think about it is that three years ago, he wrote this two years ago, aliens landed on Earth, they handed over a USB stick and then disappeared. Since then we've been poking the thing they gave us with stick trying to figure out what it does and how it works. So we're two years after he said that. We have a little bit better sense, but it's still We still don't truly understand how it works because it's a machine learning model. Like by definition, it can't explain to us how it does the things that it does
Speaker 1: This was a big paper in 2017. Attention is all you need. Google researchers that introduced the transformer architecture. That's the T in GPT, by the way. generated pre-trained transformers. Previous architectures were sequential, so one token at a time. That's part of why you couldn't scale them. With Transformers architecture, you could look at all the tokens at once. This meant you could throw big data sets at it, especially using GPUs to do all the linear algebra, matrix multiplications across, at this point, billions of parameters. I'll talk more about that in a second. So even if we had all the NVIDIA chips, like this paper in Transformer Architecture is what enabled those chips to really run, you know, because you could also use CPUs for stuff.
Speaker 1: Oh, and just final point. Um the AlexNet was the first paper that showed how you could use GPUs with a neural network. He used two in his bedroom. um that ran for a week to get these incredible results. So people didn't think that GPUs would have this use case for a long time. You know, NVIDIA was a video game company. They're doing just video games and it turns out that's what AI wanted, or this type of AI wanted. And then the second part was in 2020 there was this paper. So scaling laws for neural language models that came out of OpenAI, which basically showed that scale matters. It showed that the bigger the data set, the more compute you use, the better results. And again, I want to emphasize this was unexpected because neural networks had been around for a long time. People thought they'd max them out. They just didn't have
Speaker 1: didn't think of they didn't have transformers, they didn't think of using GPUs, and they didn't spend, you know, a hundred million dollars to do it. But if you're a big company, like this is great, right? So money is all you need. So who wants that, right? All the companies we can think of who also have all our data. I'll make a quick note that these are starting to decrease a little bit, which is kind of interesting, but from 2020 to 2022, it was clear that just smash them with as much information as you have and as much compute and you would get better results. So a quick note on different models. So w one of the things with AI is that it was an academic open area for most of its history. You know, open AI. There's some interesting books on this, you know, whether or not they really intended to be nonprofits and stuff, or whether that was just marketing to get Google engineers to come work for less.
Speaker 1: But it used to be open. Since 2018, it's all closed. It's all skunkworks RD in these big companies. But one thing you want to think about is the size of models. So this is from OLAMA, which is an app you can download. It's a great way to consume local models. In this case, I have a couple from Gemma from Google. And I have uh oh yeah, these are just Gemma here, excuse me. Um so the B stands for billion of parameters. Now those are like fine-tuning the models Generally speaking, the more parameters in a model, the better the model is going to be, also the better this the uh bigger the size. It's not a free lunch. You're not just gonna have you know quadrillion parameters and have it work For technical reasons, I won't pretend to get into. But generally speaking, in a published model, more parameters
Speaker 1: means it's going to perform better. So this is just on my computer. You can just again using OLAM, which I highly recommend you all play around with, you can see the relative sizes of these models. So at the top is the new one from OpenAI that's 20 billion parameters, that's 13 gigabytes. Uh and then Gemma has a 27 billion that's 17 gigabytes, you get the sense. Um I'm just using my my MacBook and I can run all these no problem. There are bigger open source models I can't run. But I would highly recommend giving this a play. It's click to download Olama, click it'll in the background download any of these open source models. That solves all your privacy concerns. That solves cost concerns. And personally, I think that is the future of where things are going rather than these frontier models.
Speaker 1: It'll be fine-tuning and having these, but that's a separate talk. All right, so how do you build an LLM? So fans of Luigi's Mansion, Jeff, I know you are, we've talked about this. Um use a vacuum to suck up the world, right? This is a video game, you just go around and vacuum everything. That's effectively what these companies do, right? So there's two stages. There's training and inference. You don't even have to suck it all up yourself. There's open um uh All the LM companies have a private version of this, and there's a lot of cleaning and other things they have to do. But there's public options like Common Crawl, which is a nonprofit that crawls the web and has um Tens of billions of pages, compresses them down. It's 45 terabytes. Last I checked in size. You can just, if you have a big enough computer, go pull that if you want to go crazy training models.
Speaker 1: So there's more to the stage than just copying. There's a lot of data cleaning, duplicates, offensive content, but effectively you just suck in the web. It's important to note that not all data is treated equally. So this um just like in search, things are weighted differently. Um so what has the highest rated? This is a graph from SEMrush for perplexity in ChatGPT. It's Reddit Which to the point of data cleaning, you know, I mean I use Reddit, but I wouldn't trust all of it. And then Wikipedia, right? Reddit is ahead of Wikipedia, which is kind of crazy in terms of what kind of responses you're going to get back. YouTube transcriptions go on down. But you can see you have to there's a lot of cleaning that's involved with this data. One more, this is for Google AI. So instead of giving you search results, Google just wants to give you its own LLM.
Speaker 1: And the top one for them is Quora, which is kind of interesting, followed again by Reddit and then LinkedIn. So download the whole web, but then the companies have to decide the emphasis they put on the data sources. Okay, so this um tokenization. I'm not gonna go through the whole process, but I want to mention tokenization. That's a key part where basically you transform words or subwords into numerical IDs that the model can compute. This is a website called Tokenizer from OpenAI. Definitely a director just Well you can search for it and find it. Um you can enter any text and see how it's tokenized. Tokenizers differ by model and by company, but they kind of work the same way. Note that so I said hello DjangoCon US.
Speaker 1: Note that the punctuation receives its own tokens as do longer words. In general, you can in English about three quarters of a word equals a token. If you were in Chinese or Japanese, each character would be a token. But so this is what's happening. It's turning text into numbers and then eventually spitting back text to us. But the computers to the extent that they think are just looking at numbers. And tokens, how many tokens do we have, right? So it's something like 10 to the 15th, which is one quadrillion. And yeah, I had to look that up, what that number was. That's a huge number, but it's not big enough. All the companies now are using AI to create more data, synthetic data, for training. So they're it's a little bit like an auriboris of a snake eating itself.
Speaker 1: It's still not enough data because they're trying they can can't just slam um compute at these problems. They want, they think, well, what if I had twice as much data? But that like it's crazy to think this is like all humans ever on the web in print form. That's how many tokens. Like that's the entirety of human writing. It's kind of wild to me. This is then compute times. So from Epic AI over the last 15 years up into the right, increasing 4x a year. The y-axis is logarithmic, by the way. Uh that's a key point. Um it's just showing flops, so that's floating point operation, basically one arithmetic step, so two plus three, done on floating point numbers. Um we have gone from in 2010 10 to the 14th.
Speaker 1: To now closer to 10 to the 26th, that's a trillion times more processing that's happening And the end result is that the frontier models, so think of ChatGPT, are terabytes in size. Now why does that matter for us? We can't just throw that into a Django project and serve it the way we could with the JobLib file, you know, amongst other things. Let's get to the other things. So this is why also Django would not be used directly to serve the models because there's something called inference. So inference is the process for which inputs are fed into the LLM, computed, and then stream out, right? When you use ChatGPT, you type something in and it spits out streaming text, which are actually tokens, that's what's happening. The key point is this is not like querying a database.
Speaker 1: You go to the database, it does something, or it doesn't. Maybe we've cached it, maybe we have indexes, maybe we have a CDN. Like every single time you ask something of one of these LLM models, it has to spin around, which is why you know they're building, restarting nuclear power reactors and spending billions and billions of dollars. This is a slightly more detailed guide of what's going on. So you send a prompt into the GPU box, because again it's GPUs, they're doing the processing here. Um input text pass through all the layers of the model, multiplied and transformed by the weights. Parameters and weights are kind of the two big things of the models. And then there's a probability distribution distribution over what the next token will be. Interesting point, I think, is if you look at the output prompt, it doesn't just give you the whole thing. It'll say, in this case, if I said, why is Django the best web framework?
Speaker 1: It'll calculate away, it'll say Django, and then it'll take that output token, feed it back into itself, and use that to generate the next one. So it's Django is, put in Django is uh Right? It's just a lot more computationally expensive than querying a database. If we go down one more level, our text is tokenized. There's a lot that goes on in the GPU. Like I've watched a lot of Andre's Karpathi has great videos. There's so many resources. It's super cool. I don't fully understand it enough to talk to you about it up here. But this is effectively what what it's doing. So take your text, tokenize it, throw it through the GPU, detokenize it, output prompt that it again is fed back in. So just
Speaker 1: Computationally expensive. And this is, if anything, this is the payoff. This is maybe the most important slide of this talk. for all the LLM internals you had to sit through. It's just that it's radically different. These LLMs are radically different than what we do in web worlds. Again, you can't use CDNs, you can't use um caching, indexing. It's a totally different paradigm. And when we talk about internet scale for companies like Google or Facebook or Instagram, built with Django, um They have crazy scaling challenges, but LLMs are just a whole nother order in terms of the computation required. Part of that is training, but again, most of it is inference actually. So I think that's kind of cool. And then last thing, um, what is the web piece for um LLM companies?
Speaker 1: You know, why is Fast API so popular? They want something that's fast, lightweight, and can stream tokens asynchronously as they are produced. FastAPI is built on top of Starlit, the ASCII framework. that provides request response um cycle and routing. Fast API adds request parsing and validation with um via Pydantic. And a couple extra features for API development such as OpenAI and Swagger. But it's pretty lightweight. So for web handling, do you have your LLMs run through an inference engine like VLLM? There's other ones that's Seems to be what has the mind share these days, which has Fast API built in. And then you just toss a fast API endpoint on top of it and say, boom, we're done. Now this requires a JavaScript front
Speaker 1: end and a lot of other things like authentication, security, and ORM that you know we know you need for a web app, but AI engineers just think, I'm done. And so where does Django fit into this new world? We're not going to be on the hot path of inference. Like that's not what we're designed for. But only the LLM providers themselves really need that. Most people are consuming these LLMs as APIs. That's not too dissimilar to what we do now with websites. So I have a demo coming up. I want to show you Django can do a lot more than we think it can do. So like let's say we wanted to re-reproduce an LLM chatbot. We can do it just like right now in Django. We don't even need WebSockets, right? Two-way communications, even though Django has that.
Speaker 1: It has channels, Daphne, channels, Redis, um, that are all maintained by Carlton Gibson now. Andrew Godwin, I saw him earlier. Did a lot of that work initially. Um, we don't even need WebSockets for this. We can use service sent events, the boring old web. So you send one HTTP request and you can receive receive streaming responses. Django has a streaming HTTP response class. Can anyone guess when this was added? Year or Django version? I'll like buy you a beer if anyone can get this right. Right? So like what is that 12 years ago? You know, this is not new. It's just Maybe needed a good use case. Um
Speaker 1: if you wanted to get really fancy, you could do HTML streaming as well. I did that at one point. HTML streaming is like newer and cooler and I'd love to do a whole talk on uh server sent events versus HTML streaming, but I don't have time for that. But let me show you a demo. Okay. So this is running locally. Um I've wired up a Django project to an OLAMA model. Uh it's Gemma Uh 3. 4B in this case, you could use any model you want, streams tokens to an API endpoint. So again, so the local uh OLAMA is just streaming tokens to us. There's no JavaScript here. I'm using all HTMX and the templates. When I saw Carson was giving the talk, I thought, ooh, I gotta do something cool. Oops. Well, come on, show it again. Um using a Python generator, so just yield to uh send the tokens one at a time and render them.
Speaker 1: This is a synchronous view. If you look, it spits it all out and then it formats them in Markdown. Like I could do a lot of things to make this better. And to be honest, I vibecoded this whole thing, which wasn't a one-shot deal. It took me a couple hours, but I was also playing around with vibecoding. Oops. And um yeah, I have uh I have the repo up. I think this is kind of cool. This is not new technology. This is something Django Could have done some version of, you know, 13 years ago, 12 years ago, whatever it is. Just we didn't have these L local LMs to do it. So there's a lot we can do just out of the box with Django as is. To solve the problems that AI engineers are going to have, we just have to sort of tell them about it.
Speaker 1: And so that's it actually. I've covered a lot of ground here today. I hope you have a better sense sense of the AI landscape, how the web is changing, how Django fits in. Personally, I'm biased, but I think Django is perfectly suited to this new world. We as a community just need to do a better job of telling the new generation of Python developers. that were here and when they're ready to share their models, whether it's classic ML like iris or fancy LLMs, Django has the answers for a lot of questions they don't know how to ask. And I'll conclude with Django has been here for the past 20 years and hopefully it will for the next 20 as well. So thank you for your time. And I think there's time for questions, yeah? If someone close wants to go, I can repeat the question.
Speaker 1: Yeah.
Speaker 2: So in your demo you were running uh LLM locally?
Speaker 1: Yeah, in the demo I was running uh using OLAMA to run uh the model locally. And I was using um Gemma, but you could use any local model in that case. The process is the same. Yeah. Okay. Okay, got a mic.
Speaker 3: Hi. Is there any concerted effort um to To make Django more um I don't know popular in this space or anything like that?
Speaker 1: Well there's a working group on AI uh that is being spun up. Um But we could do more, right? I mean I think part of it is I wasn't fully aware of this until I spent time outside of Django. Um so we're open to ideas that people have. I mean I think it's People need to know what questions to ask before they see that Django has the answer.
Speaker 4: As these local models get better, uh how do you make the decision to outsource, you know, if you're incorporating AI into an application, outsourcing it to API calls to like OpenAI or any of the other large models versus using a local model. Like how good are the local models? How do you decide when to use which?
Speaker 1: I'm not gonna say it depends. Um but I think so this is just me, my personal opinion. I think you're still gonna have there's only a handful of companies that can afford to do these frontier models, and it's really to be seen how the economics of that works. What happens a lot of times if you have a request you'll put most of it in different local models and then only a little bit of the high-level reasoning to the frontier models. You can also, as long as those frontier models are putting out things you can fine-tune them. So you can take a local model that has cost whatever amount to do and then fine-tune it on Django, you know, using Hugging Face and other things. So you can have a much smaller focal local models. So I personally think that focal local models are the future because of privacy, because of cost, both in terms of training and serving.
Speaker 1: So that's where I would put my money on them. There is a question of how big do they take to how much um how expensive is it for them to run? But you know, again, you don't need the whole internet if you just want to use Django code. Right? So I think compilation of local models is where I would place my bets. And fine-tuning is super, super cool. I wish I could do a talk on fine tuning. You should look into that. One of my colleagues from um JetBrains has a post on how to do PyCharm hugging face tuning. You should check out.
Speaker 5: Thanks for the awesome session. I saw that um there are some things that you want to do, especially as developers, for example, explainability, and maybe you want to leverage Django to do maybe some of the models, but Maybe what you're building has explainability of the model front and center. Wha how how would you do that with this infusing AI into Django or would you rather outsource that to the model itself and you know maybe call some explainability um functionalities outside of the framework?
Speaker 1: Yeah, I think that's the right question. I don't have an answer. But that's exactly it. Yeah, where where does the logic live in Django? When does it go to the model? How do you stack the models? Um, that's the right question. I will say sorry anecdotally, um where I work at a co-working space, there's a company that's a healthcare startup that's using AI. So they get all the data from your healthcare provider, and then they're using AI to um do diagn figure out diagnostics for you. They got funding and they prototyped it using uh LLM. So they had AI in the title and then they prototyped it with the LLM. They've now rewritten 90% of it in boring old Python, because it turns out they didn't need it.
Speaker 1: And it's and they want something deterministic for like lab results. So, you know, separate from the funding, it's interesting to think you could just prototype with it and then rewrite in just old Python. So I think we'll see some of that too.
Speaker 6: Yeah. Hey, so uh as you were like I guess jumping through the vibe c vibe codes, right? Um When, if or if at all, um did you find yourself having to use things like JSON fields or JSON blob storage to record like I guess the the set of text messages and the chat history and things like that
Speaker 1: That's a really good question. I was super fast and loose with this prototype is the short answer. So I didn't du you know how long did it take me? I think it took me like Two or three hours to get it total. A lot and you know 90% of that was just waiting. Um for this particular one I was using Claude within PyCharm. There's a beta plugin. Um we also have uh Juni. Um but yeah, I want to play around with it more and do that. And you know, I'll say as a plug for AI, one of the things you can do is I can take the existing data set and just ask in the chat bot exactly those questions. But I don't have a good answer for you, but that's that's where I'd like to go next.
Speaker 7: Yeah, you say that you used uh sync view, but in real world in production, um have you organized streaming in Django, like chat response back to user and all that stuff. So it like if you are going to use sync view, it's very easy to reach a denial of the service, right? So you used WebSocket or something else?
Speaker 1: So in this case it's just um It's just it's we're asking the local LM model and it's streaming to us, but there's no back and forth. It's just one input and it's giving us the full If we had a two-way, then we could use WebSockets. But I thought this was sort of interesting because again, it's just we send the input and then it streams the responses, but we're not streaming anything back to it. So it's not two-way. So it was a f it was a lot simpler than maybe I thought it. I actually wanted to use an async view before I thought about it and was like, I don't need to. I think that's part of the issue with WebSockets is how many things are truly two-way. A lot of things aren't, just like a lot of things aren't really async. So, you know, synchronous views are pretty easy to reason about.
Speaker 8: Thank you.
Train and evaluate the model separately, save it as a serialized Joblib file, then load it in a Django view that accepts form inputs and makes predictions. Django can also store those predictions in the database so they can later be used for feedback and retraining.
Discussed at 14:19Django already provides the forms, ORM, admin, authentication, security, APIs, and deployment patterns needed to turn a trained model into a usable web application. Relatively small or medium-sized models can be served directly by Django without the specialized infrastructure required by large language models.
Discussed at 17:26Training creates a model from data, while inference is the repeated process of feeding inputs into the trained model to generate results for users. For large language models, inference is especially expensive because each request requires substantial computation rather than a simple database lookup.
Discussed at 29:48Large language models are often terabytes in size and generate output token by token through intensive GPU computation, so they cannot simply be placed in a Django project like a small serialized model. They normally run behind a specialized inference engine, with the web framework handling application-level requests around it.
Discussed at 29:58FastAPI is lightweight, handles request parsing and validation, and is well suited to asynchronously streaming tokens. Inference engines such as vLLM commonly expose a FastAPI endpoint, while a fuller framework like Django can provide the authentication, security, database, and other web-application features around it.
Discussed at 33:04Django can connect to a local or remote language model and stream its generated tokens using server-sent events and Django’s streaming HTTP response, without requiring WebSockets. The demo uses a synchronous Django view, a Python generator, HTMX, and templates to display the response as it arrives.
Discussed at 33:56Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026