Evolving Django: What We Learned by Integrating MongoDB
Published June 13, 2025
This video features Anaiya Raisinghani .
Presented by: Anaiya Raisinghani
At MongoDB, we have several Django enthusiasts who have jumped at the idea of backing a long-term solution to combining MongoDB and Django. Having historically provided support for SQL-based open source frameworks like Entity Framework in .NET/C#, Doctrine in PHP, and many more, we are familiar with the territory. We're happy to say we've successfully created a MongoDB Backend Library for Django and want to share all that we've learned kitting out a new NoSQL backend for a traditionally SQL framework including how we believe we have influenced -- and will continue to influence -- changes in the core Django library.
Anaiya Raisinghani explains why MongoDB built an officially supported Django database backend after earlier integrations left developers relying on sidecar libraries, losing access to Django features such as the ORM and admin. The backend maps Django models and querysets to MongoDB documents and operations while supporting forms, validation, authentication, migrations, custom fields, aggregation, Atlas Search, and vector-search workflows. She demonstrates how the package enables both conventional text search and fuzzy, relevance-ranked Atlas Search, then shows how she built a Dublin pub finder by combining Django, MongoDB Atlas, Voyage AI embeddings, LangChain, and vector search. The project is still in public preview, with planned support for further encryption, transactions, index management, GridFS, change streams, schema validation, and other features; she asks Django developers to try it, report issues, and contribute feedback.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: Right, hello, hello everyone. My name is Anayare Sangani. I am a developer advocate with MongoDB. I'm super excited to be here today to share what we learned with integrating MongoDB and Django in our new Python package. And please bear with me, my voice is a little bit rusty today. Don't know why, but we're gonna do our best So a couple weeks ago, I actually had the opportunity of speaking at DjangoCon Europe in Dublin, Ireland. And for it, I built this really incredible, well, unbiased biased opinion, really incredible Dublin City Center Pub Finder. And I did it because I had never been to Dublin, nor had I ever had a Guinness. And so I went ahead and I built this little application using our very new Django
Speaker 1: MongoDB backend library. And it's a pretty simple app. It actually allows you to specify any query you like. And then it'll show you three places that you can go to in Dublin depending on the query that you picked. So for example, my query here was actually Guinness with outdoor seating. And then it shows me three great options in Dublin where I can have a Guinness outside. So here is a screenshot of the Pub Finder in use. And you can find the full-fledged tutorial on my Dev2 account where I will walk you through the code and help you build it out yourself. But also near the end of this talk, I'm going to go over a very high-level understanding of the tutorial to show you how great our new Django backend
Speaker 1: package is and also some of the cool stuff that you can do with it. So now that you have like a high-level view of applications that you can build, I want to quickly run through what Django and MongoDB are from a very high level as well, just to ensure that we're all on the same page. So, Django is an open source web framework for building web applications. It lets you build database -driven websites. with very clean, reusable code without having to reinvent the wheel every single time. And by default, Django is designed to work with relational databases because Django's ORM is built for relational models. So how many of you in here have used Django before?
Speaker 1: Okay, great. Awesome. And then now for MongoDB, same question. How many of y'all have used MongoDB before? Okay, hopefully by the end of this talk, many of you will want to check us out. So MongoDB is a NoSQL document-based database. So instead of rigid tables and rows like relational databases, it stores your data as flexible JSON-like documents Complete with nested fields and arrays, making it very, very easy to model complex, constantly evolving data structures, and also allows for you to scale out horizontally without any costly join operations. So MongoDB's flexible schema means that developers are able to iterate very quickly by adding or removing fields without needing any downtime.
Speaker 1: and our very powerful query language and aggregation capabilities allow developers to filter, transform, and analyze data directly inside of our database. We also have official drivers in every single language and have tools like MongoDB Atlas, our Managed Cloud, Compass, our GUI, and MongoDB Charts. meaning that we have a really great ecosystem of tools to help teams move very quickly and keep related data together for efficient reads and deploy globally resilient applications very simply. Okay, so now that we have a very foundational understanding of the platforms that make up our new library, I want to jump into our agenda for today. But I do know that I went over all of this a little bit quickly
Speaker 1: So if you're looking for more resources or information or you just want to chat about NoSQL databases, come find me at the booth either today or tomorrow. Or Saturday too, I believe So I am going to be taking you all through a little bit of a timeline to help just give a very holistic view of how we got to where we are today and what our view for the future is. And then I'll aim to leave a little bit of time for some questions at the end. So, first of all, let's go over our inspiration and our motivation for how we ended up today with our new Django MongoDB backend library, which I'm going to give y'all all the full details on in a bit. But in order to fully understand our motivation and our goals with this package, we will have to dive into a little bit of history first.
Speaker 1: before we can chat about what exactly the package is and where we are hoping to end up in the future. So, the idea of backing a long-term solution to combining MongoDB and Django really is nothing new. We have several enthusiasts who have been pushing for this integration within MongoDB itself. And we have historically provided support for SQL-based open source frameworks like the Entity Framework in. NET and C sharp. doctrine in PHP and many more. So it was actually always top of mind to support Django, and we even had some previous attempts with MongoDB, which I will go over in the next slide.
Speaker 1: But that kind of begs the question, right? So why did it take us so long in order to be successful? Well, we needed a very strong narrative from the Django community. To truly justify the organizational cost in order to develop and also maintain this project. Luckily for us, this narrative actually came last year We saw a very growing presence of MongoDB usage within the Django community. So a couple of statistics, because I know we're all very data-driven. In the 2021 Django Developers Survey, MongoDB wasn't even listed as a back-end database used by Django developers. And this all changed in 2022, which is crazy how much can change in a year, right? So in 2022, Django
Speaker 1: Developer Survey, MongoDB was actually cited as the most used in the 6% of the other databases category. And then again in 2023, that number rose up to 8%. So in the grand scheme of things, 8% is not that big of a number, but it equates to about 100,000 developers. So this is what gave us that push that we needed. And we learned from developers and existing third-party platforms that a common pattern of integrating MongoDB and Django was just by connecting to MongoDB using the APIs of PyMongo, which is our official Python driver library, or through open source ORMs like Mongo Engine. So this sidecar
Speaker 1: was of course different from the standard way of connecting Django to a database by specifying the back end in the database's setting. So what was the consequence from this? It meant that a lot of really great aspects of Django and MongoDB weren't able to be fully utilized in a simple way, such as things like the admin panel And it was awkward because developers weren't able to use the Django ORM even if there were some very, very valid use cases. So this kind of made us realize that projects that went ahead and tried to integrate Django with MongoDB really needed the scalability and also the flexibility that we offered. So we decided to build a more comprehensive compatibility layer.
Speaker 1: And so, which is fully developed and also fully supported by MongoDB. So that we could free these projects from any extra development burden and let them build fully fledged Django applications backed solely by our document database. And that kind of also begs the question, why did we go ahead and build an entirely new library? Raise of hands, how many of y'all have used one of the libraries on the screen if you've ever tried to use Django and MongoDB together? Okay, one person. Yes. Um so one person in the audience has tried one of these. Um maybe some of y'all have heard about it, but there have been some very impressive attempts at integrating MongoDB with Jango over the years.
Speaker 1: One of these attempts even came from us, and unfortunately, as I'm sure you're aware, they fell really short. These libraries had developers dealing with maintenance issues. technical complexities, poor performance, and just general overall unideal ways of working with Django and MongoDB. But they were very helpful as well because they showed us exactly what had been built and how we needed to pivot going forward. So we created a new library that was actually started with code from Django NonRel since it took the exact approach that we were looking for. A database backend that allows for developers to use Django models and query sets.
Speaker 1: However, it was very outdated So we used it just as a starting point and then our updated back-end work actually became the initial code commit to our new MongoDB Django MongoDB backend repository. I'm going to be saying Django and MongoDB so much throughout this. Don't get sick of me. So why did the other libraries fall short, right? While some of the other libraries simply worked by just translating SQL queries generated by Django's ORM into MongoDB queries, which had many, many drawbacks, as you can imagine. Other libraries are very innovative using, for lack of a better word, very hacky solutions to force different parts of your code to work without actually adhering to latest versions of MongoDB.
Speaker 1: And then others were started and then deprioritized due to a lack of Django specialist knowledge to make properly calibrated decisions around the platform itself. So many developers ask, especially Django developers ask, and it's a great question, why would a Django developer even want to use a NoSQL database, right? Because you guys and Relational are like this. But you'll get your answer soon. I have an entire slide on that. And you'll want to stick around to find out. But for now, I want us to chat a little bit about the learnings that took place while successfully creating the Django MongoDB backend. So, first of all, our work with Django and third-party libraries like Django Filter
Speaker 1: involved fixing numerous tests and boosting overall code quality. stability and provided a library with a very low barrier to entry where developers were actually capable of using it to create super complex projects And out of the five libraries that we were targeting, Django Filter was the library that we collaborated most closely with. So employees at MongoDB were actually boots on the ground. fixing tests in the Django filter package and connecting directly with the author of the library. Second one was that one of the biggest positives was being able to directly engage with the incredible Django community. The Django Software Foundation has taught us that building this library required so much more than just code. It demanded a really deep understanding of the Django community's needs.
Speaker 1: and required us to really fill in any NoSQL knowledge gaps. We also wanted to create another option for Django developers to use and this was our biggest attraction with developing the package the way that we did. And another huge learning we faced was discovering that transitioning from SQL to NoSQL isn't a one-size-fits-all solution. And this really reinforced the importance of flexible migration guides. And our support of the Django ORM goes very, very deep. We took on the challenge of making this work despite the fact that MongoDB is a NoSQL database. So these are some of these are sorry these are a few of the very crucial learnings that we discovered holistically.
Speaker 1: while working on the public preview for our package, but without getting deeper into our learnings while actually programming it. These are the three aspects that we learn the most about while actually going ahead and creating the library, testing, generating queries, and documents. So testing was actually one of the more challenging parts of building the MongoDB backend for Django because traditional Django tests expect a relational database environment, right? Despite these difficulties, the effort to develop comprehensive tests really paid off by building much-needed confidence in our system's reliability Not only did these tests help catch very nuanced bugs early on, they also served as a safety net for future development changes.
Speaker 1: I made sure that the integration of MongoDB Schema with Django's ORM maintained its overall integrity Tim Graham, who if you know Django, you may know him, was super beneficial in this process because he actually started testing right away against the Django test suite and he made some very, very crucial modifications in order to make it work. Because of his changes, it's actually much easier to use our package in your developments than other previous versions. Secondly, transforming Django ORM calls into MongoDB's query language was another very complex task that we faced. The process of query generation required rethinking how high-level queries are interpreted and executed in a document-based database.
Speaker 1: Insights from our database experience team was super instrumental because they helped clarify and streamline the logic needed. To support a seamless and efficient translation between Django's expectations and MongoDB's capabilities. And this very collaborative approach approach ultimately led to a much more robust query handling that really preserved performance in the long run. And third of all, documents. So of course in MongoDB, the natural way to store data is by using documents, which pose a unique challenge to Django's model system that is built around relational field types. And to address this, the development team actually created custom fields specifically designed to support Django's ORM
Speaker 1: This custom solution not only maintained the flexibility and power of MongoDB schema, but it also allowed for developers to work within Django's familiar framework. which meant making sure that data structures could be managed as effectively as traditional relational fields. And all of that was to say that this brings us here to present day. In February earlier this year We actually announced that our official Django MongoDB backend library was available for public preview. So what exactly does this mean and why should it get all y'all excited? Let's begin by answering exactly what this package is. It's a third-party database backend
Speaker 1: that integrates seamlessly with Django. When a developer defines the database backend engine when creating their application, they can go ahead and specify our custom backend. And as long as a library is installed in your Python environment, Django is going to be able to connect without any issues It also aligns with all of the familiar steps that Django developers know and love. This process is very smooth, especially if folks use the templates. that were made to set everything up properly, particularly because MongoDB is determined to support as many of Django's core features as possible. There is a little bit of overhead with our package at the moment, but MongoDB is very committed to seeing everything through and ensuring that the process is not only great, but requires a very low lift
Speaker 1: from all the developers themselves and the community. We're very well aware that a successful library requires so much more than just pure technical expertise. It needs to be a perfect appreciation of Django in its entirety, which means its ecosystem, convention, and most importantly, the needs of the developer community. And we are fully committed to making sure that our back-end package is up to par, that not only meets the technical requirements of developers, but that it's also very painless and a very intuitive process. acting as a very natural complement to the base Django framework. In this public preview release, we are offering developers these various capabilities
Speaker 1: First and foremost, we want to ensure the community can use Django models with the utmost confidence. Developers can go ahead and use Django models to support MongoDB documents with support for Django forms, validations, and authentication. We also wanted to keep Django admin support intact, so this library allows developers to use the Django admin page as they normally would, with full support for migrations and database schema history. Third of all, configuration is as simple as pip install Django MongoDB backend followed by Django Admin Start Project. And as I said earlier, when developers follow the templates, the settings are actually configured for you. It can also be as simple as using your MongoDB connection URI string.
Speaker 1: And I'm going to be showing some code snippets in a bit just to fully describe how intuitive this connection process is. We also wanted to enforce MongoDB-specific querying optimizations. So field lookups have actually been replaced with aggregation calls or aggregation stages and operators. Join operations are represented through $LOOKUP, and it's actually possible to build indexes right from Python code. We are also allowing for advanced functionality. While the package is still in development, there will be a ton of really cool features to come, which I will go over in a handful of slides as well. And there is already support for time series and projections.
Speaker 1: We also encourage utilizing this package to create AI applications like the one that I showed at the beginning of this presentation, since it is compatible with various AI frameworks. such as Langchain and Llama Index, and you can use it to do advanced search such as vector search. This package has aggregation support as well. So you can use the power of our aggregation pipeline in conjunction with your Django framework. Raw query actually allows for aggregation pipeline operators. And since aggregation is a superset of what traditional MongoDB query API methods provide, it gives developers more flexibility and functionality. And you all might feel familiar with this with Django 's RAW query.
Speaker 1: So I know that I went over a ton of information, but the best part about all of this is that it really is just the start. We have more functionality and features. such as Bison data type support and embedded document support in arrays on its way. So definitely stay tuned for our general availability that is coming out later this year. So at this point, we have gone over why we decided to create this library and some of the functionalities that are available for developers. But now let's go ahead and turn our focus onto the benefits of utilizing this package inside of your projects. So our integration really is the glue between the best of Django and MongoDB. With the integration, you're able to truly highlight the best aspects of both.
Speaker 1: As a Django developer, you are going to feel right at home with MongoDB's document model because it will map directly to your Django models, lets you store related data together, which means that you can ditch complex joints. and handles hierarchical or semi-structured data naturally. This means faster development and cleaner code, which is exactly in line with Django 's philosophy, as we all know. Beyond modeling, MongoDB Atlas brings features that boost productivity. We have Atlas Search, which offers built-in full text and vector search, while the aggregation pipeline Which is essentially an assembly line where you can isolate documents through individual transformations, lets you run very complex analytics
Speaker 1: and transformations in database, which means that no extra services are required. And if you need a local test environment, it's very easy to pull up a Docker image and spin up a single node replica set with your Atlas Connection string or convert your existing Atlas setup into Docker Compose. Our security is fantastic as well. Data is actually encrypted in transit, at rest, and even in use with client-side, field level, and queryable encryption. And when it comes to AI and semantic search, MongoDB integrate seamlessly with Langchain, Lama Index and haystack, which means that building chatbots, recommendation engines, or hybrid vector tech
Speaker 1: search over text, images, and even audio is easier than ever before And for real-time applications, MongoDB delivers low latency reads and writes at scale, which means that you're keeping your Django apps responsive under a very heavy load Plus, whether you're kicking off a side project on our free community edition or architecting enterprise solutions on Atlas, you have a ton of deployment options that fit not only your needs, but also your budget. And then finally, MongoDB's ecosystem, our VS Code and PyCharm extensions, the MongoDB Shell, and Atlas CLI. Mongo Import, Mongo Export, and our GUI based compass means that you can work however you like
Speaker 1: and however your team seems fit. However your team sees fit. So Just how simple is it to get up and running with this library? It truly is one command. Pip install Django MongoDB backend to install our Django integration. And then we have our easy to use starter template that works with the Django admin command start project making it simpler than ever to see what typical MongoDB migrations look like in Django. And you know it's not necessary to use these templates. But it is highly encouraged because they are the same as the default templates with a couple very important changes It includes MongoDB specific migrations
Speaker 1: and our settings. py file has actually been modified to ensure that Django uses an object ID value for each model's primary field, primary key. It also includes MongoDB specific app configurations for Django apps that have default auto field already set. So we recommend that you start a new project or app using a new template. Otherwise, you can check out our GitHub, which I have a little QR code for at the end, for the other steps to set it up however you see fit for your project. With all this in mind, I want us to jump into a little bit of a demo to see all of this in action.
Speaker 1: Okay, so this is going to be walking us through our OmniSearch tutorial here, and this is actually going to look very similar to how one would do it. First party backends that are provided by the Django MongoDB library. So I've just gone ahead here and I've created a very basic search experience And to show you what that looks like, I'm just going to go to this route that I've already made, which is search. And you know what? I know the front end is incredible, but hopefully it can communicate our search properties very well. And as you can see, if I go to the view that I have defined here, it's quite simple as I get the query as a query parameter. Defined in this title bar.
Speaker 1: So let's say that I wanted to use the word transformer. I type in Transformers, which will give me all of the results for the Transformers movies, right? And how it works is that it does a simple query where I get the query set. It will then take the term that we passed in as a keyword argument. And then it checks to see if that term is contained in the title. So that'll show you the movie. We also went ahead and just took the liberty of adding in some extra metadata pieces to make it, you know fit easier and have this demonstration flow. So this is great, although there are still some limitations within this traditional system, right? Let's say that you don't know the title of your movie, but you do know what the plot has.
Speaker 1: In this case, it is harder to search because the search can only do or contain a search about the title. To make that work, I could simply go over to the objects filter and include information about the plot, but there are still more elements that we want to see that we may not know as the user ourselves what is best to rank it by. And this is actually where the power of MongoDB really comes to light. In MongoDB, we have the very powerful Atlas search operator that works seamlessly to take a search term that you provide. Search it against all of your specified documents within the database, and then give a database ranked value. Exactly on how relevant the document returned is
Speaker 1: to the actual term being searched So in a handful of seconds here, I am going to go ahead and I'm going to scroll down on this page. And here's our example of Atlas Search. I've gone ahead and I've actually inherited from the basic search experience just to stress how easy and simple it is to create this new experience without leveraging any additional things other than the Django MongoDB backend provider. If you look through this query, we define what is known as the dollar search operator. And this dollar sign search operator actually allows for us to use the Atlas search feature Within that, I define another operator known as text, and what this does is it looks over tokenized words or phrases provided on each field inside of our document.
Speaker 1: So let's say that you have the word transformer. That in of itself is actually a token. And with these tokens, you can use them and then search against them. Then within this text operation, we can go ahead and we can specify the path because I already know that I want to search along every path defined. in my model, I can just specify the field's name and then say, hey, I want every field's name to be searchable under this specific query. And then you can define what the actual query is, which once again is just going to be our term. And then like the next part is what we know as fuzzy. So what is fuzzy? Fuzzy is actually the ability for us as users to say when I search something. I may have made a spelling error
Speaker 1: or I may have made a mistake or two. Fuzzy allows us to include words that are off by maybe one or two characters to represent that actual slight misspelling. And then finally Because I am going ahead and leveraging the power of our raw aggregate and I'm no longer able to use the built-in limit feature. I just went ahead and I specified a limit. But as you can see, that operation is fairly straightforward and easy to understand So again, using our Atlas search, we can actually head over to our URL. py file where we define this, and I can go ahead and I can switch out our basic search for our Atlas search. And then let me go over here and let me refresh my page.
Speaker 1: And then I can type in Transformers. And we can see that Transformers, thankfully, is still searching. Um, but what if we didn't know the exact Transformers movies, right? That's the entire point of this. Let's say that we want to try searching against the Decepticons, which if you know Transformers, you'll know that they're the natural enemies of the Transformers and as you can see I spelled it wrong that was on purpose. And even though Decepticons is misspelled, we are still able to get all of the results that we were looking for. That was a little demo. And I have another demo for you that isn't, I'm not going to be playing the video version of it, but I will be walking you through the code. But hopefully that was a good way of seeing Atlas Search and how you can do that with
Speaker 1: our new backend. So you might be wondering what's next. And if this is it, and I'm really, really happy to say that this is just the beginning of our journey together. We have huge plans to pull out all the stops and ensure that we have a very incredible general availability release later this year. We actually have someone working full-time on third-party library integration to make all of this possible. Alex Clark, if any of you know him. along with the continual beta releases until our general availability is out later in the year. So that kind of begs the question, right? What can you expect from our GA? You can prepare for programmatic management of a vector search, atlas
Speaker 1: search, and geospatial indexes by using the Django API. Vector search, atlas search, and geospatial queries through the Django API, queryable encryption, and client-side field level encryption. database transactions, storage of cached data in the database, and even more great features and functionalities to look forward to. We even already have plans for post-GA releases, such as Grid FS for large file storage, change streams for data monitoring, and schema validation. So we have a ton to look forward to and hopefully As our developer community, y'all can let us know what you think with our continual releases and help guide us on features that you would like to see integrated in your specific
Speaker 1: projects And a big part of looking into the future is our acquisition of Voyage AI. So this is groundbreaking for developing complex AI applications. As Voyage AI has a state-of-the-art embedding model, which is solving the hallucination and fragmentation problem, having all of your AI needs in the database layer removes a ton of friction. And Voyage AI services are going to become fully embedded in Atlas with auto embedding for vector search, native re-ranking, and enhanced multimodal retrieval. So what does this mean for the Django community, right? It means that Django developers are going to be able to model and store embeddings as native vector fields in your documents.
Speaker 1: perform semantic searches directly from Django views, build multimodal experiences using the MongoDB query language that hopefully you're familiar with, but if not, I have a lot of great resources to get you there and a ton more. One example of this that I've now shown so many times is my Dublin City Center PubFinder that I built using the Django library, Voyage AI for the embeddings, and Link Chain for semantic search. that I chatted about in the beginning of the session. But for now, I want us to go through and take a closer look at the platforms used to build this demo and then take a look at some of the code. So we've been chatting about the Django MongoDB backend package, of course, and we went over what Voyage
Speaker 1: AI is, but let's quickly talk about Langchain and MongoDB. Langchain is an open source framework used to build applications that are powered by large language models or LLMs. There are a variety of use cases for Langchain, such as document analysis. chatbot, retrieval augmented generation, otherwise known as RAC, and so much more. It actually works by chaining together various components or links to create a very comprehensive workflow where each link performs various tasks in the process, such as accessing your data, calling the language model used, processing data, and more. Because these links are very, very malleable, Langchain is actually known for its flexibility. And Langchain and MongoDB offer super
Speaker 1: complementary technologies for rapidly developing production ready AI applications. Their integration streamlined the development of LLM-powered RAG and agent-based applications by combining Langchain's LLM integration framework. and Langgraph agent orchestration with MongoDB's scalable database for operational and vector data. And they have like the Langchain MongoDB package, then they also have a Langchain Voyage AI package. So it's it's nice because all the platforms are very compatible with one another. Now that we have a pretty good foundational understanding of what Langchain is and what our Langchain MongoDB package can accomplish for developers, let's go
Speaker 1: over on a very high level how my Dublin PubBinder was built. So to provide some background on how I collected my data, while poking around on Google Maps, because I'm a Google Maps fanatic. I don't know how many of y'all love Google Maps. I love Google Maps. I saw that Temple Bar was one of the more popular locations in Dublin and the new Google Maps API allows for a maximum of 20 locations per API call. So I did two calls in this location, one for the tag of Pub , then for the other with the tag wine bar. This way I had 40 data points for my file, which meant more variety on where to grab a drink. Then I also made sure that I went ahead and I concatenated five Google reviews into a single string with each call, since the reviews are what I wanted to embed and do my semantic search on.
Speaker 1: Then from there, I saved all of the places located from my Google Places API calls into a JSON document that I was able to upload into my project folder. So I had 40 places just highlighted like that in my JSON file. From there, I used the commands that I showed a couple slides before. I went ahead and I installed the Django MongoDB backend library, created my new Django project, configured my database in the settings. py file to connect with my MongoDB Atlas cluster URI. And then I created my Django application. And just as stated before in the talk, I use the templates that are available and highly encouraged to use. So once I had all of that done, I went ahead and I created my models instead of my models.
Speaker 1: py file. And Django models are especially useful because they define the structure of our data based on our JSON format. So based on my JSON file, I wanted to ensure each document recommends it a place with a list of types or my array of strings. My formatted address, which was just a string, a nested display name, which became my embedded document And then my reviews field, which is the very, very long text field that we will generate our embeddings from, and an embedding field. And this is where we are going to store the embeddings that we generate using Voyage AI. Once that was done, I just wrote a little script to embed the reviews field of our JSON file using Voyage AI's Voyage 3
Speaker 1: Lite embedding model, creating a new file to hold the embeddings. that had the same name when we clarified in our models. py file. And then I went ahead and I created a new JSON file that had exactly the same as before, but now with the embeddings as a new field. And then I saved this new JSON file so that I could upload it into our MongoDB cluster. Then, once I had my places and their embeddings saved inside of my MongoDB cluster, it was time to actually integrate MongoDB Atlas Vector Search with Langchain to really get the most out of our embedded reviews. And then before I could write my script though, and I really wanted to highlight this
Speaker 1: part because I feel like a lot of people don't know it or forget about it, you need to create an Atlas vector search index if you are going to be doing vector search in any of your projects. Which is very simple to do if you're working with Atlas because it's just clicking a handful of buttons. And you can go into your cluster, click on the Atlas Search tab, and then on Create Search Index. And you want to make sure that you're keeping the default name of the index as vector index, and then choose the database and collection that our data is stored in. Then you want to choose the path that contains the array of vector embeddings, and I like to just keep mine as embedding to keep it very easy. And then the number of dimensions is going to be 512, and that is because we used Voyage AI's Voyage 3
Speaker 1: light model. And depending on what embedding model you use, you might have different number of dimensions you might want to use a different similarity method but mine was just cosine and 512. And the cool thing too about MongoDB when you upload if you upload your embeddings into Atlas it'll show you how many dimensions are there. So it keeps things a little bit, it makes it a little bit easier to keep track of everything as you're building these kind of complex AI applications. And then you'll know that your vector search index is ready for use when you see a status change inside of your Atlas cluster. Once my vector search index was ready, I wrote a script using some skeleton code from both the Voyage AI documentation and MongoDB's documentation.
Speaker 1: And I just kind of changed it up depending on my use case. And this is what part of the finished product looks like. I wanted to make sure that I had my proper embeddings object with the model I was using and my Voyage API key. And then I wanted to ensure that I created a vector store for my documents where I was connecting from my MongoDB connection string while specifying that all of my reviews were located inside of my text key. Once I got this part swerk, because what I did was I did all this lang chain part in a separate file, tested it, and then incorporated it into my Django project. Once I did that, I just uploaded it and changed it so I wasn't hard coding in my query. I was getting it from a user. I put that in the views.
Speaker 1: py file. I created all of the templates for my front end and then I just edited my URLs. py file so we actually had some place to go when we ran our server. Then once that was done, my Dublin City Pub Finder was up and running, and I got to make this funny little view for it. So it was partly a side project, partly an excuse to experience Dublin, but it was also 100% built with tools that I am super excited about. And I hope that this tutorial can show how powerful and fun the intersection between Django, MongoDB, and AI is. And it shows like some of the cool projects you can make. that can be highlighted with the Django MongoDB backend package. And of course if you would like the full tutorial
Speaker 1: with even more details, images, taking you through how to build something like this, similar. completely different to if you just want to use it as a place some starting code. You can scan this QR code and if you do check it out please leave a reaction or a comment or if you have any questions I can help answer those as well And in this article, which is linked on my Dev2 account, you will see that the repo is actually in something called our MongoDB Gen AI Showcase. And this is where all of our incredible AI demos live. Most of them are in notebooks. So you can go through and you can run the code yourself It's a really great starting place if you're interested in building AI applications or learning more about MongoDB and all of our AI use cases, but are a little bit unsure
Speaker 1: where to get started or just need some inspiration. And then I just wanted to include this picture because I had so much fun at DjangoCon Europe a couple weeks ago. And I just want to say that I was really appreciative for being able to be there and meet the entire community. I wanted to take a second to highlight how impressive the J community is. So it's definitely a community worth celebrating. I'm really excited that MongoDB gets to be a part of it now. So we are, this is the QR code with the repo for our Django Long B backend, if y'all would like to try it out, because we are incredibly passionate about this project. Not just because we have seen the data points telling us that there is a market fit, but more importantly because we like Django a lot. And the joy of developing is truly what has kept us going
Speaker 1: and rigorously hashing out our MVP details pouring through lines of code to make every necessary sequel to MongoDB conversion, and grappling with individual tests in Django's test suite to ruthlessly document what we can and cannot achieve. So we are very committed to working on this backend for Django. We have made significant strides to ensure that we are going to ensure, to, to ensure that we are going to provide a very first class experience to the open source community. That all being said, this is a continuous learning experience. We would love to have similarly passionate developers try out the library, give us feedback, file issue tickets, or submit pull requests. Tell us what you like, tell us what you hate.
Speaker 1: We are here as your sounding board. We want to make sure that is something that the community uses. So please check out our repo and then let us know your thoughts inside of our MongoDB developer forum. And then, of course, stop by our booth later today so we can invite you to our VIP community lounge that is happening tomorrow. We also have a happy hour that we would love for all of y'all to attend. I know Veronica was giving out some cards earlier, but you can also join us here by scanning this QR code. And thank you all so much for coming and for letting me talk to you for the last 45 minutes. I hope that you learn something new. If you have any questions, you can ask me them here.
Speaker 1: I'm going to be at the booth, I believe, from 5 to 7 p. m. today, and then I'll be at the VIP lounge tomorrow, and I'll also be around on Saturday. So thank you all so much. Do we have any questions?
Speaker 2: Yeah, you mentioned these other tools in the beginning that were solving this integration, but they had quite hacky solutions like translating SQL code and all that.
Speaker 1: Yeah
Speaker 2: Can you uh talk a bit more about the solution that you guys came to by looking at the problems in the existing libraries?
Speaker 1: Yeah. Well it was great because a lot of these libraries had a lot of foundational issues where, you know, you can't take something like SQL queries and just translate them. You have we the great thing about our new package is that we kept it as a Django developer would want to use and we kept the ORM as stable as it can be. Meaning that you don't have to force your code to fit. Like I I have some resources that I can show you if you wanna meet after, but it's exactly the same way that nothing has been translated, which I feel like is the best thing when you're coming from a relational database and you want to go to a non-relational database.
Speaker 1: But I'll show you, I have a quick start that will take you through it. And then I also went ahead and I translated some of Django's like original quick starts and show like showing the difference between them where if you use our templates and if you use um just like our initial commit, like everything stays the same for the developer experience. So it's it's pretty cool. It was a little bit complicating because like I don't come from a SQL background and some of the commands I was like wait like Will this work? And the engineering team was like, yeah, it will work. Like they wanted to make sure that everything stayed as simple and as digestible for a Django developer.
Speaker 2: Thank you.
Speaker 1: You're welcome. Hi, yeah.
Speaker 3: So for our use case, we have to register schemas and to get them changed is like a big process that we have to go through.
Speaker 1: Uh-huh. And
Speaker 3: MongoDB being like flexible and stuff. Is there a way that you can limit the changing of schemas by like as you're using it or is it like you we need to figure out a way to restrict that?
Speaker 1: That's an interesting question. Um, I think I'm gonna have to see your use case to like have a proper answer for that. But we can talk if you want to come find me at the booth or even after. Yeah.
Speaker 3: Sounds good. Thank you.
Speaker 4: Yeah. Um his this kind of goes that that is uh calling uh like uh migr uh make migration Yeah. It's calling make migration even a thing?
Speaker 1: Yeah, it is.
Speaker 4: Okay.
Speaker 1: Yeah.
Speaker 4: I wouldn't think that that would be the case. case.
Speaker 1: No, you wouldn't think that, but that is what like our engineering team took so long to make sure that everything stays as consistent. So make migrations is that was the one part that I was confused about because I was like, how would that work for you know a NoSQL database? But it does work. And I can show you. It would obviate the need for that. Yeah. Okay, great. Well, thank you all so much. If there are any other questions, um come find me later today. But I hope that I hope y'all enjoyed this. I enjoyed talking to you. So thank you again.
Existing libraries had maintenance, compatibility, performance, and design problems, often translating relational SQL queries or using workarounds. MongoDB used Django NonRel as a starting point to create a supported backend that works with Django models and querysets more naturally.
Discussed at 8:33It is a third-party Django database backend that lets an application select MongoDB as its database engine through Django’s normal configuration. The package aims to preserve familiar Django workflows while supporting MongoDB documents and features.
Discussed at 15:44The public preview supports Django models, forms, validation, authentication, the admin site, migrations, and database schema history. It also provides MongoDB-specific querying, aggregation support, projections, time series support, and index creation from Python.
Discussed at 17:18MongoDB’s document model maps naturally to Django models, keeps related data together, avoids complex joins, and handles hierarchical or semi-structured data well. It also provides capabilities such as full-text and vector search, aggregation, and scalable low-latency reads and writes.
Discussed at 20:27Install it with `pip install django-mongodb-backend`, then create a project with the provided starter template or configure the backend and MongoDB connection URI in the project settings. The templates preconfigure MongoDB-specific settings, migrations, and ObjectId primary keys.
Discussed at 22:47Use the backend’s raw aggregation support to add Atlas Search stages, such as the `$search` operator with a text query and searchable path. Atlas Search can search across specified document fields, rank results by relevance, and use fuzzy matching to tolerate small spelling errors.
Discussed at 26:40The planned GA features include Django API support for vector, Atlas, and geospatial indexes and queries, queryable and client-side field-level encryption, database transactions, and cached data storage. Later plans include GridFS, change streams, and schema validation.
Discussed at 30:39The project stored Dublin pub data and concatenated reviews in MongoDB, generated embeddings for the reviews with Voyage AI, and used an Atlas vector search index with LangChain to retrieve semantically relevant places. The Django app then accepted the user’s query and displayed matching pubs.
Discussed at 34:35Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.