Elasticsearch DSL

This video features Honza Král at DjangoCon US 2014 in Portland, Oregon, USA.

Elasticsearch DSL
0:42:12
Published September 17, 2014
9,180 views
110 likes

By, Honza Král
Elasticsearch DSL is a new library for integrating Django apps with Elasticsearch, enabling users to utilize the full power of Elasticsearch.

Help us caption & translate this video!

http://amara.org/v/FOP7/

Summary

Honza Král explains Elasticsearch as a distributed JSON document store and search/analytics engine, covering dynamic mappings, nested and parent-child documents, full-text queries, filters, caching, scoring, and aggregations. He presents Elasticsearch DSL as a Python query builder that sits on top of the low-level elasticsearch-py client, replacing deeply nested dictionaries with composable Python objects while preserving Elasticsearch’s own semantics rather than imitating SQL or Django’s ORM. Through a Stack Overflow dataset demo, he shows migration from raw query dictionaries, automatic composition of queries and filters, aggregation handling, convenient response objects, custom scoring, and basic Django indexing and signal integration; planned work includes mappings, persistence, and deeper Django integration.

Key takeaways

  • Elasticsearch is a distributed JSON document store that supports search, analytics, dynamic schemas, nested documents, and parent-child relationships.
  • Queries determine matching documents and relevance scores, while filters only narrow results and can be efficiently cached as bit sets.
  • The low-level elasticsearch-py client handles transport, serialization, node failure, and load balancing but leaves query construction to the developer.
  • Elasticsearch DSL composes queries, filters, and aggregations automatically and provides attribute-based response objects instead of deeply nested dictionaries.
  • The library deliberately exposes Elasticsearch’s own concepts rather than pretending to be SQL or a Django ORM.
  • Existing raw query dictionaries can be wrapped in a search object, modified with the DSL, and serialized back to support gradual migration.

Summarised automatically from the transcript.

Transcript

6,441 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:21

Thank you. Um so I'd like to talk about Elasticsearch DSL, which is a new library for interacting with Elasticsearch that I've been working for. working on. But first let's let's take it a little slower. So let's talk a little bit about what Elasticsearch is. I did a little talk yesterday about what search engines are in general and how they work, so I will try to be brief on the on this part. So what Elasticsearch is, it's an open source distributed search and analytics engine. That's quite a mouthful for uh essentially a distributed data store that can store your documents, search through them and analyze them. And by analyze them I mean run different sorts of aggregations.

1:08

By distributed, I mean just that. If you have one instance, it will work. If you start two instances, they'll find each other, form a cluster, and automatically share your data and spread the load. So that's where the elastic uh part of the title comes in. So as I mentioned mentioned is a document store. So it's JSON BASO. Anything that you can express as JSON, you can index and search through using Elasticsearch It's not exactly schema-free, but it has a dynamic schema. What that means is you don't need to tell us what your documents look like. We'll look at them and we'll infer the schema information from the from that data. Only in some cases where you have some knowledge that we don't, for example, you know that

1:56

this uh this number will never get above 256, you can tell us and we'll index it more efficiently for you Or in some cases, you actually need to inform us uh what the data type is because there is no way to know from the JSON. For example, if you index a geopoint or a geo shape, there is no way from us to automatically distinguish it from just a list of two numbers. So typically you want to tell us the schema, but if you don't, if you just want to uh play around with it, just start indexing documents and you should be good to go. We also have support for some relationships, so you can actually have nested documents, which is essentially a sub-document

2:42

as part of a bigger document that can be queried independently We'll see an example in just a bit. And we also have parent-child, which is essentially a one-to-many relationship that you can use to query across. So you can query the parents while asking conditions on the children and vice versa. So to sort of give you an example and don't worry, I don't expect you to be able to read this. This is a sample document from indexing data from Stack Overflow. I will later be doing a demo, so this is the data that I'll be using. You can see that I have several interesting fields that I have highlighted. One is a title and body. Those are just text fields that do exactly as as you would expect. And we have a we have a date time

3:28

as a creation date. We have comments, which is a list of nested documents, because each question in Stack Overflow can have comments, also each answer. What we don't have here, this is a question. We also index the answers and we use the parent-child relationship in Elasticsearch to map the relationship between the question and an answer in OnStack Overflow. And finally we have um uh also highlighted the field rating, which is an integer field, which is the rating from the Stack Overflow, the quality of the question or the answer. And That's important because you can actually take that into account when uh when sorting or when outputting the results. You can either sort by it or you can just take the score

4:13

uh that the search engine gives you and combine it with this number to produce the optimal sorting. So it's not one or the other, it's a combination of both. Unfortunately that that will have to be left as an exercise for the user. Not enough time. So I've I've talked about queries and so what do they look like? Well how do we query Elasticsearch? So Elasticsearch is HTTP and JSON. So everything we do is HTTP and JSON. So if you want a query, you send JSON over HTTP, surprisingly. And the j uh the JSON that contains the query is essentially an abstract syntax tree It's essentially a serialized version of an of an expression

5:01

of an expression tree that contains, amongst other things, but the most important ones, are queries and filters. There's important distinction between them, but what you need to know is there are they are fully indistinguishable from the outside. If you can use one query, you can use any other It's uh really uh the queries in Austicsearch can be overwhelming for for beginners when you look at a query and it's a it's a full page of JSON. It's It's really distracting. But if you start to think about it as as as a tree, as an expression tree, which has a very easy grammar. You have a query and each query each query type has a different grammar.

5:46

For example, a filtered query can contain a query and a filter. So it's simple, simple grammar that can be recursive, whatever. So Once you understand these concepts, it's fairly easy. So queries represent the unstructured part of Elasticsearch. It is the full text. It is the part that not only tells you which document matches your query, but also how well does it match. Is this a good match or so -so? So uh that's why uh we have several different types. We have match queries which do what you would expect. We have fuzzy queries that are able to take into account typos in form of

6:31

matching across uh Lavenstein distance, so across different uh mi uh mispronunciations or mistypes of of the word We also have uh queries like regex or wildcard which allows you to do partial matches. You can you also have compound queries, so if you have multiple of those core queries, you can put them together Typically you do that using a bool query, which is a short for booling, which essentially just takes a bunch of other queries together and says, you must have all of these, some of these, and none of these. Again, we'll see example. Just to uh there are queries and there are compound queries. The queries rely heavily on analysis, which I talked a lot about uh yesterday, and they produce score.

7:21

And because the score is uh r uh dependent on the actual form of the query and on the state of the index, these queries are not cached. So I wouldn't say that they're slow, but filters are faster. Because filters Do the same thing as queries in that they limit the result set, but they don't have to bother with the score with the relevancy because they only narrow it down. And because of that, they're much more suitable for caching. What we actually do inside is we represent a result of each filter as just a bit set So we literally have one bit per document that that uh match that filter.

8:06

So you can imagine that's a very efficient storage and also something that can be very easily cached. And also, once you have these caches for each of your individual filters, so you have a term filter, so you you're looking for an exact match, or you have a range filter where you're looking for a range of a numeric or date value So if you have multiple caches which are bit sets and you want to combine them using a bool query, uh bool filter in this case, sorry. That's very efficient because you have multiple bit sets and you want to see a document that are in both of them. Well, that's an AND a binary bitwise AND uh pretty much one of the most effective uh operations you can do on any CPU.

8:53

So it it gets very fast and it allows us to cache the individual core filters. So you get a lot of reuse from the caches. This is all transparent for the users. It's just important to keep in mind that there is a difference between filters and caches. uh and queries and you should always use filters if you can if you don't care about the relevancy. So again with with filters you have the core filters and the compound filters. Uh pretty much the only compound filter that you need to care about is the bool filter that allows you to exactly compound the individual filters very effectively. It by default uses the caches and the bit sets inside. So this is actually one of the smaller

9:39

typical queries that you would ask Elasticsearch. So this is a query. It it is a filtered query inside, and a filtered query has two components, a filter and a query. In this case as a filter we use a range filter. So we're looking for uh questions on SecOverflow that were 20,000 and newer. So that's the filter part. For the query part, we have a bool query. And then we have three parts in there. We saying that the fee uh title or body must have PHP in them and There must be an answer to this question, so a has

10:25

child , which has Python in the body And we're also saying in the must-not branch of the bool query that title and body must not contain Python. So effectively what we're looking for is some poor SAP on on Stack Overflow asking a question about PHP and some smartass replying, yeah, you should use Python. We've all been there, we've all done it, and this is this is how we can identify ourselves. So that's sort of what this what this uh what this uh query is. Uh You can see that I was right that it can be confusing to people. It's a lot of text, a lot of weird characters, and that's one of the reasons why

11:13

why I'll uh created the DSL. And just the bottom half of the of the text is actually aggregations. So I'm actually looking to see the distribution per tags and for each tag I want to see the average uh average comment count So in the result set I will have that with the tag design patterns or something, I had I had 24 documents and on average they had three comments each. Again, something that we'll see in more details. This is just to give you an overview how it would look if we had to write everything by hand. So that's pure Elasticsearch. Now uh we're at a Django conference, so that means Python. So how do you interact with uh with Elasticsearch using Python?

12:00

Unfortunately, many people immediately jump to jump to this question, like how hard can it be? It's just HTTP and JSON, right? Yeah. So the problem is Elasticsearch can be a little difficult Not that it's unpredictable or anything like that, it's just it there's a lot of things going on. For example, it's distributed, so which node do you talk to? You run the risk that if you only talk to one you're gonna overload that node and the rest of them rest of the nodes in the cluster will just be there blazing around. Even though they will share share all the load, some of the work will always go through that one node. Not ideal. And what happens if that node goes down? The cluster is still fully operational, but your application cannot reach it anymore

12:47

So that's one aspect, the distributed aspect. Then there are dis different environments. So many people will would deploy Elasticsearch behind a load balancer. Or they would try and use alternate transports. For example, you can use Thrift as as a plug-in to Elasticsearch doc thrift , because some people prefer binary protocols for some reason. And then there is also the fact that we have almost a hundred API endpoints with almost uh 700 parameters each. So if you if you want to use the raw HTTP, that's knowledge that you have to carry around in your head And trust me, it's not pleasant. It's just a huge amount of information that's essentially useless and you just want something to do it for you.

13:34

So that's why we uh last year we released a bunch of official clients that are very low level. It's for all those people who would prefer to use the HTTP, but we think that they shouldn't. So they should use Elasticsearch PY instead. Elasticsearch PY is what you get when you do pip install Elasticsearch. It's a very low-level client. It's essentially just one-to-one mapping to the REST API. There is nothing added, there is no opinions. Because we really wanted uh we wanted that nobody would have an excuse not to use this client. So it's very it's very extendable, it's modular, you can override any any different parts of it. And it supports all the API and all the parameters.

14:20

and actually have documentation for that. It's tested as part of the release cycle for Elasticsearch itself. So if you're using Python, if you're using Elasticsearch, there should be no reason not to use this client. But as I said, it it's very raw, it's very low level. So the only thing it will give you on top of using raw HTTP Is the different methods for different API endpoints and it will do the serialization for you properly. So it will uh take your Python dictionary, serialize it into JSON, and send it over the wire It will then do some smart things. For example, if it cannot reach a node, it will it will put it on a timeout and talk to a different node instead.

15:06

Or It can even ask the cluster, hey, what is what is the uh what are the current nodes that are part of the cluster so it can do the load balancing properly. But aside from that, it's it's fairly dumb. You still have to write the queries yourself. S Python dictionaries, which is much better than than JSON because you can actually use trailing commas. Yay. But It's still it's still fairly painful. So instead I I said I don't want this. I want there should be a simpler way how to do this because for example imagine that you have a query like this And you want to add a filter. So first you need to determine like is it already a filtered query?

15:52

Can I just add a filter? Or do is it just a raw query and I need to convert it to a filtered query to add a filter? Then inside the filter, is it already a bull filter that I need to just add something into or do I need to convert it to a bull filter and add the filter to the filter that's existing already there? So that's painful. It's certainly doable. It's just Python dictionaries and and it's nothing complicated, but it's it's just painful and it should be easier. So enter LSICSearch DSL. This is, for now, it's essentially just a query builder for LSICSearch. It relies on the on the Elasticsearch PY, on the raw client for transport and everything network related and communications related. So what it essentially only does is

16:37

it will build a query, serialize it into a Python dictionary, and send it over, get the results back, and present it to you in a nice wrapper. So again, you don't have to get a a dictionary uh that contains a dictionary, it contains a list of dictionaries which contains a dictionary with your actual data So this is this is how it looks. You basically define a search, you associate it with the low-level client, so it will know how to communicate with the with a cluster. And you start querying. We'll look we'll look into into it uh deeper how it looks. Uh just suffice to say that you can just issue individual queries or filters and anything against a search object. And we'll figure out beneath uh beneath the hood how to combine them into

17:25

the compound queries and filters and we'll do the same for aggregations. And then if you want to get result back, we'll give you a nice class that you can actually access attributes into and you don't need to use brackets everywhere and have a lot of lot of work with that. So this is sort of the high level overview. So how what was what was the design decisions that we that we made? Well first one is first one was I was just sick and tired of typing brackets. Square, curly, it I I felt like a list programmer and not in a good way. So that's the first thing that I really wanted to to get rid of. It's

18:11

dictionaries are easy. It's a great data structure and and it's very fast. and easy to work with but it's not really fun to write. So that's one part. We we want it not to have any uh any more brackets than we absolutely need it. The second part was we wanted to do the automatic composition. So you don't need to know how to combine two queries and what's the logic between combining a bool query with the match query. How does it work? We have simple rules that will actually do that for you. All you need to do is say, yeah, add this query to the mix. Add another condition essentially. And we'll we'll figure it out underneath. All

18:56

this while still allowing you to do it yourself if you absolutely need to, if you if you know what you're doing. and hopefully without any additional pain in that case. Also one of our in very important points is we don't want to pretend what something that we're not. We are not SQL. The query DSL is completely different from SQL. It has different capabilities, different semantics, different syntax for sure. And we don't want to shoehorn something like a Django RRM onto Elasticsearch. That would make no sense because it wouldn't allow you to access the 90%

19:42

of Elasticsearch features while still uh not supporting all that the ORM can do. So it would be sort of the the least common denominator which is in this case very small So we just we just own up to the fact that we are not SQL, we are not anything else, we are still Elasticsearch, and you should be familiar with the with the queries and filters that you can run against Elasticsearch. We'll try to take the pain away, but not the actual work. Sorry. So if you actually look at the example again, you can see that I'm actually manually specifying that, yeah, this is a match query. And essentially

20:27

what I'm passing in, the title equals, is the same absolutely same thing that I would create a dictionary for in the raw DSL. So it's essentially just a just a syntax sugar for uh in this case creating a dictionary with one key match and which would uh have as a value a dictionary with the key title and the value Python. So it maps very s very easy. So you don't need to learn another tool. You don't need to uh to learn another DSL. You just you know LSI search or you should if you don't And then you can just you can just start using this immediately.

21:14

And You can see we do the same for filters. So there is a range filter with the creation date equals. And here in the DSL you would have a nested dictionary. So here you have the same. Because like I I didn't want to try and invent some syntax or borrow overleave with the underscore underscore because that would get really hairy. So again, explicit is better than implicit. Go for it So it's very it's very close. You can however see one thing, uh that in the in the query in the second one, I have something with a capital Q that's a It's a name that I borrowed from Django and it's essentially a shortcut. If you want to create the query manually, like outside of the search object, if you need to manipulate it.

22:02

For example, if you need to negate it or if you need to uh combine it with another query using an using an OR operator instead of an AND. So uh uh we have those uh those shortcuts for all the important uh important objects in the DSL. So queries, filters, aggregations, and some others that that we'll keep secret for now. So you can see how you can create it. Underneath it will actually do what it will do, it will look up the class that corresponds to that given query type or filter type and just instantiate it. So it's a it's really literally just a shortcut. It can you can even just pass it uh the raw dictionary

22:48

that you would otherwise use as the query. So we'll see later how that can be used to actually facilitate the migration process if you want to switch over to this new library. And So you can do and or you can do negations and it will actually do the right thing. It actually tries even to be a little smarter. So if you, for example, do uh double negation, you will end up with the same uh with the same filter or query. just so that you don't have a ridiculously big ridiculously big queries uh once you once you work with them a little bit. So this is how you can how you can construct them outside of the search object

23:37

and how you can work with them. Then if you if you construct a query this way this way you can just pass it into into the search and everything will work as expected. So uh the way you pass it into the search is is by using the dot query or dot filter uh methods. And those long as everything else will just uh actually return you a modified copy of the search object. Here we we uh we borrowed from from Django's design where the query set is essentially immutable and every time you do something on it you will get back a copy. So you shouldn't be afraid to pass it over to someone else or

24:23

uh or anything like that so you can actually f uh fork it and have two different uh two different versions. The only exception to this is aggregations because there we needed to the chaining behavior to be a little different So for queries you can do search. query dot query dot query dot query and add multiple queries on the same line. With aggregations , You want to do uh something a little bit differently, at least that's what we we came to uh expect. So it is Uh in Elasticsearch, when you when you define aggregations, you define a s uh a bucket and then metrics inside. Because essentially any form of aggregations, be it SQL or NoSQL

25:10

or anything else. is essentially dividing your data into buckets and then calculating a metric or a computation inside each of these buckets. So if you have a Group by, you say GRUBY this this column in SQL, so you'll have a bucket for each column, and then you want to see account or a sum over this value. That's the calculation that you run inside the bucket. Elasticsearch is very explicit, so we actually yeah we call it bucket. So here in the f the first line we are creating a bucket per tag. And inside we're looking for an average over something. This is just a this is just a shorthand. I'm omitting all the parameters so it can actually fit on the slide.

25:58

And then we're adding another bucket uh another metric. So we have one bucket with two metrics. On the other line, however, we have two buckets that are nested. So we have one bucket and then a subbucket and then inside of that we have a metric So you can see that the chaining behavior is a little different because bucket will actually return itself, so you can call an aggregation on it, a metric. Whereas a metric will return its bucket so that you can add another metric next to it. So just something to just something to keep in mind once you once you uh once you start start using this that the behavior there is slightly different.

26:45

So the last thing we have is is a response object. I uh mentioned it several times that you get back a fancy response object instead of just uh just a huge dictionary. That contains nested data structures. So you have a response object which has a success method, which will tell you like, did I actually reach all the data that I needed? Because Elasticsearch will happily keep uh keep uh serving you search requests even if half the cluster is down. It will tell you that half the cluster is down and that it couldn't reach half of the data, but it will still try and return to you something So you can ask, hey, was this a success? Did I reach everything? Yes. And then you can just iterate over it and get individual hits.

27:31

With the raw response, you get the metadata, and as part of the metadata, you have the source. That isn't really that practical for normal use case, so in this case we reverted it. So you get the the object back. So you can see that I'm doing h. title. And if you want to access any of the metadata, you just do h dot underscore meta underscore ID, document type, or index, or any of the metadata that are typically associated with the with the document in Elasticsearch. Also the score. So and so you can see you can use attribute access. You don't need to use square brackets and strings to to access the data. And the same goes for the overall response.

28:19

So you can just do response dot aggregations dot per tag dot buckets uh then you access the first element and you can do dot value and stuff like that So it's much more convenient to work with. We even add it , we'll see that, we'll see that in the demo, hooks for uh for introspection, so IPython will correctly auto-complete and everything. So this is essentially all that all that we have done. So now what do you do if you want to start using it? If you have a if you have a fresh project, congratulations, I envy you from all of my heart. If you don't, hopefully you're you're already using using uh the low-level client. In that case, you already have the dictionaries with your queries lying around.

29:06

So what you can actually do is you can just create a search object from the dictionary, manipulate it however you wish, and then either execute it directly or you can again serialize it back to the dictionary and plug it back into your existing code. So for example, if you have a query somewhere and and you you wished that it was simpler to add a filter to it, just create a search object from it, add a filter to it. So you realize it again and nobody needed to know that you actually cheated and used a different library instead of doing the work yourself. So now let's see if everything works.

29:53

So uh Can you read this? Somewhere in the back. Can you read this? Thank you. So The first thing that we can do is I'll I'll show you how the how the migration actually works. So let's assume that we have we have a dictionary like this that actually contains a one a typical uh A typical query to Elasticsearch. It's a pleasure to read. So what we can do is we can create a search object from it. we can actually already see how it would look otherwise if we were

30:40

wrote it using uh using the DSL, using the Q notation. that's that's representation. We can associate that with the low-level client ES is just an instance of the of the Elasticsearch client. And now we can finally execute this to get a response. So you can see response, it has hits total. So totally we have hit 48 documents out of the approximately 500,000 that I have currently loaded You can you can get the first one, and you can see that it has title, it has

31:25

comments, It's a plural. And it even has something like an owner, which is actually a nested uh nested document, so we can continue, so we can have nested. owner. display name So the first question that we found was actually asked by Joel Fan. That's unsurprising. given the data set. So this is this is sort of the basics. This is if you already have a query and you just want to you just want to plug it in. You just create a search object from it and uh and start querying. If instead you actually are starting fresh, so what you do is you just create a search object yourself. Now, this search object will actually, if I do a count, it will match absolutely everything.

32:15

So we have, okay, so just 200,000. We can also limit it to just a certain doc type. So we are we're only looking to search for questions. So we can see how this has changed. And we have no questions. So it actually should have been questioned. And this is what happens because I'm not copy pasting things as I should be. Because the corresponding question, then if I just query for doc type, it will actually just add the doc type to there. So if I do a count now, it correctly uh it correctly returns.

33:04

So now let's say that I want to actually do us do a query. So I want to have a match. And again for the title just use Python And you can immediately see that it's exactly the same as as the as the di as the dictionary would look like. So if I just now eat add some more qu queries and filters and aggregations, which I will not type in, you can see that it it gets more complicated. So through gradual steps, you don't need to you and know that you should have used a filtered query with a

33:51

with a bool filter with all this all this stuff, but just adding a filter and then adding another filter, it will first be converted from a query to a filtered query and then the filter in the filter filtered query will be converted to a bool query. You can see also for the aggregations that I defined. In this step, I defined a bucket per tag, which is a terms aggregation over the field tags And inside I actually am asking for a metric. I want it to be returned under the name of max score. And I am saying it's a max aggregation over the field rating. So when I execute this

34:39

I get the aggregations back, so I can see the per tag, I can see the buckets. So these are all the tags that were that were in the in the result set. So I can get the first one and just get the key. So the first one was obviously Python. So we just learned that when you query a Stack Overflow for Python, the documents with the most tags will be actually tagged with Python. What a surprise. And we can actually also ask for the max score that is that is in Python. So for m uh for Python the max score of a question was 134. So this is sort of the

35:24

analytics part uh of of Elasticsearch where you could easily Just take all these values and visualize it very nicely using using JavaScript or or plotlib or anything else. So uh last part of the demo, I'll just show you how to uh construct the queries yourself. So let's start with creating a query. So we are looking for title Python and not body Ruby. then we'll we'll create a filters the same way. So we're looking for tags Python or range which is

36:10

uh smaller than now. So we're looking either for documents that are older than in the future. I know it makes no sense, but bear with me. Or or that are tagged with uh with Python. Yes, this filter will match everything, but it's it's really hard to come up with demos that actually do something. So then what we can do is we can actually manually wrap this query. So we have this construct in OSICS called a function query, and that allows you to uh take a query and uh provide uh elastic search with the formula how to actually calculate the score if you know better and we do So we have we have a field in our document that's called

36:56

writing that is that is human contributed. Some humans actually said that this is a good question. It would be shame for us to ignore that information So by this line, I'm saying query is now a function score query, which is wrapping the original one, so I'm saying query equals q. And the function that that I want to that I want to run on it is a script score function, and the script is I just multiply the score by 10. Here instead of 10 I would typically just use uh just use access the field and everything, but that wouldn't fit on a slide. So that part is left to your imagination And now we just create

37:42

a search object from it, which hopefully we'll be able to execute. Yes. And also you can see this is what we created by the by the three steps that I showed. This is the query that that we that we put together. It would not be impossible to write this by hand, but I certainly would not want to. So that's sort of the uh the grow of this of this library. to allow you to run this uh run this query create this query easily. So that was uh that was the DSL and now let's let's see how it how you can actually how you can actually plug it in, how you can use it. So if you want to use Elasticsearch from your Django

38:28

application, this is all the code you need, more or less. So this is this is a code how to actually index all your data into into Elasticsearch. The first part is you just do a bulk load and you iterate over all your models, over all your uh Yeah, models. The model is called model for some reason. You iterate over all of them, you call a method to dict on them, and then just index that. And then the second example is a simple function that you can register as a signal handler for PostSafe, and it will update Elasticsearch after any change of in the document. This is literally all you need.

39:16

You might you uh probably want to get a little fancier by by specifying the schema that's the line with the put mapping But this is all you need. And I do want to make this more automatic in the future. But for now, this is this is what you need. And then you query as usual, as as I showed in the demo, that's all you need Just construct the query, run it, you get data back. That's all you need. So that was that was a Django integration, the helicopter overview. So what what is what is next for what is next for the this library? The first part is I want to extend it to be not only for queries, but also for mapping. because that's also something that that people struggle with.

40:02

How do I define the mapping? The mapping is the schema. And that also has a fairly complicated syntax and semantics, which is very powerful but sometimes a little overwhelming. So that's the first part. Once we have the mappings, we also have the information about the types that are stored in in OSIC search. So at that point we can implement a persistence layer. So essentially model, something that has a. save method. And we can do that because now we know how to serialize and deserialize even things like nesty documents so we can wrap them in their respective document classes or we know how to deserialize a date time because currently we return date time just as as a string

40:48

Because JSON has no support for daytime, so the only way how to do it is by matching a regex against every single field. That's not very good and uh by far it's not uh performant enough. And once we have the persistence layer, it's only a short step to uh do a proper Django integration to actually be able to to correlate the documents with the models. So that's all for me. I would love to uh thank Rob Hudson and uh and William from Mozilla, they helped me a lot when designing this library and they they tested this library. Mozilla has been brave enough that they already run this in production. So kudos to them. I still haven't gotten any any complaints, so I'm guessing I haven't uh interfered with with their operations

41:38

by creating this library. So that's good news for me. And now if you have any questions, I'll be more than happy to answer them.

Questions this talk answers

What is Elasticsearch and what can it do?

Elasticsearch is an open-source, distributed search and analytics engine that stores JSON documents, searches them, and runs aggregations. It can form a cluster automatically across multiple instances and dynamically infer much of a document’s schema.

Discussed at 0:21

What is the difference between Elasticsearch queries and filters, and when should I use each?

Queries handle unstructured or full-text matching and calculate relevance scores, while filters only narrow the result set and can be cached efficiently. Use filters when relevance does not matter; use queries when you need scored, relevance-aware results.

Discussed at 7:21

Why use Elasticsearch DSL instead of writing Elasticsearch queries as Python dictionaries?

Elasticsearch DSL builds and combines queries, filters, and aggregations for you, reducing the need to manage deeply nested dictionaries and brackets. It also returns convenient response objects with attribute access while still exposing Elasticsearch’s native query concepts rather than pretending to be an SQL or Django ORM layer.

Discussed at 16:37

How can I migrate existing Elasticsearch query dictionaries to Elasticsearch DSL?

Create a search object from the existing dictionary, modify it with DSL queries or filters, and either execute it directly or serialize it back into a dictionary for existing code. This lets you introduce the library incrementally without rewriting every query.

Discussed at 29:06

How do I use Elasticsearch with a Django application?

Bulk-index the Django models by converting them to dictionaries and sending them to Elasticsearch, then register a post-save signal handler to update the indexed document when a model changes. After indexing, construct and run searches with the DSL as usual; defining an explicit mapping is recommended for more advanced use.

Discussed at 38:28

Presenters

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Honza Král

More videos from DjangoCon US