Observe!
Published October 14, 2022
This video features Honza Král at DjangoCon US 2015 in Austin, Texas, USA.
Beyond the basics with Elasticsearch
Elasticsearch has many use cases, some of them fairly obvious and widely used, like plain searching through documents or analytics. In this talk I would like to go through some of the more advanced scenarios we have seen in the wild. Some examples of what we will cover:
Trend detection - how you can use the aggregation framework to go beyond simple "counting" and make use of the full-text properties of Elasticsearch.
Percolator - percolator is reversed search and many people use it as such to drive alerts or "stored search" functionality for their website, let's look at how we can use it to detect languages, geo locations or drive live search.
If we end up with some time to spare we can explore some other ideas about how we can utilize the features of a search engine to drive non-trivial data analysis.
Elasticsearch’s inverted index supports more than keyword lookup: sorted posting lists, positional offsets, and document statistics enable efficient Boolean and phrase searches, while TF-IDF-based relevance can be combined with popularity, distance, business rules, or controlled randomness. Honza Král shows how the percolator reverses the usual workflow by indexing queries and matching incoming documents, supporting alerts, live updates, geolocation and language classification, and precomputed metadata. He also explains aggregations as tools for exploration, recommendations, meaningful graph connections, and trend detection, with significant terms and sampler aggregations helping distinguish what is distinctive from what is merely common. The main argument is that Elasticsearch’s understanding of relevance and data distributions makes it useful for ranking, classification, and pattern discovery—not just filtering and counting.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: So hello and welcome. Thank you for having me here. And the talk is beyond the basics with Ausic Search, and it is essentially about all the use cases where you can use Elasticsearch, but might not be immediately obvious that that is something that you can do. So we'll walk through several of those and see how and why it is it is suited for uh for that particular scenario. But before we go beyond the basics, we need to talk about what are we going beyond? So what is what is the base functionality of Elasticsearch? And how does how comes that we can do all these other things? Like where is it coming from? Well, it's all coming from search. Search, especially full text search, is the primary function of Elasticsearch.
Speaker 1: And search is not a new problem. It's been around for a while and it hasn't actually changed much. The first essentially index over a book, over some text. has been created in 1230 and we still use the same data structures to this same day. Of course there's been a plenty of improvements, but the underlying infrastructure, the the Inverted index remains the same. It is the index that you're familiar with if you've ever read uh any book, which I hope that you have, sincerely. And this is how it looks You have the list of interesting words. And then for each of these words, you have a list. In the case of a book, you would have a list of pages.
Speaker 1: For us, you would have a list of documents that actually contain this word. And notice several things. First of all, the words are sorted. Of course it makes sense because you need to be able to find the word that you're looking for so you can go to the page that actually contains it. And also the pages or the documents are sorted as well. And this is not uh this is not accidental, this is very important for us. And we'll s we'll see how And also, when we're talking about uh search using a computer, there are other things uh involved in this data structure. Notably some statistics. For example, how many times is this word contained in this
Speaker 1: in this document? or uh how what is the length of the list, etc. Things that will be very important later on. So when we have a data structure like this, how does the search work? Well it's super simple. If we're looking for a document that mentions both Python and Django, we locate these two words and we get back the lists. And now we just walk the lists. And we merge them together. So whenever we find a document that is uh present in both lists, that is our result. If we wanted to do something like a phrase search, that we're looking only where uh uh Python is immediately followed by the word uh word web.
Speaker 1: All we have to do is add another information into the inverted index. We just need add offsets. What is the position of this word in the document? And then when we are going through the merging process, we just say we care not only that the document is in both lists But the offset must be uh immediately following each other. So Python would be on the position n and web would be on the position n plus one So you can see that actually doing a phrase search is not any more expensive than doing a regular search. You're just adding one more comparison, a numerical comparison at that. So it is fairly efficient What else you can
Speaker 1: you can sort of imagine here is I can get the list of documents from anywhere. It doesn't have to come from the same index. So I can have multiple indices, I can have index on every single field in my document, and I can use them all. If I have one condition on the title, one condition on the category, and one on the body. I will just uh query those three inverted indices to get these posting lists, is what they're called, and merge merge them together. So we don't have the limitation of many other data stores that we limit the number of indices you can use per query, per uh collection. And that will uh that is also something that we benefit greatly from, is it's this
Speaker 1: this data structure. And finally , Uh the last thing that you do when you when you do this merging, when you find your match, you quantify how good a match it is. That is the primary difference between a search engine and a database. We not only tell you which documents matches uh your query, But also how well does it match? Is it a good match or is it just meh? And we know that because we we have the information about uh the statistics. So this is called relevancy. We tell you how relevant the document is to your query.
Speaker 1: So How is relevancy relevancy computed? At the base of it, there are two uh there are two numbers. Numbers that we call TF and IDF. Tf is term frequency. It is just the number of occurrences of that word in the given document. or in the given field. So if I'm looking for Django uh in a document, so how many times does this document contain the word Django? Is it there only once? Is there three times? And obviously, the higher the number, the better the relevancy. IDF is inverse document frequency. Which is just a fancy word of
Speaker 1: saying how common or rare this word is in your entire dataset. Is this a word that is contained in every single document that you have? Or is this a word that is only present in 1% of your documents? And we can get this information right away from the inverted index because that essentially uh the length of the list attached to the field compared to the number of documents that we have overall. It's fairly easy to calculate and there's actually the exact formula if you if you're so inclined. And this number has has the opposite effect. The more common the word is, the less relevant uh this document is
Speaker 1: for uh for the result. Because if we find a word uh that is in every single document Yeah, who cares? It's in every single document. Of course we're gonna find it. That doesn't mean anything. So this is sort of the base formula for anything that has to do with relevancy And it works very well for text. Now Lucene, the library that does the indexing and the heavy lifting for Elasticsearch, adds some some more stuff on top of it You can see the exact formula there, and uh you can see in the middle that that's the TFIDF, the big big sum. What it adds on top of it is it takes into account, for example, the length of the field.
Speaker 1: Because if we find the word Django in the title versus in the body, that also gives us different information, right? If we have a short field and we still find it there, it's more relevant than it if we have a full text of the book and we find it there as well. Those are different different types of information So it improves on the basic TF IDF formula, but it's still it's still only relying on the statistics that it learned about your dataset. And sometimes you want to go a little further. Sometimes you have other information about your data So let's say that imagine that you you have a QA
Speaker 1: QA website where people ask questions and and give at give answers. Let's call it like buffer overflow. I don't know. And uh you have the users rate the questions and the answers. This is a good question, this is a good answer. And that is an information that is that you want to take into account But you don't want to sort by it, because that would completely destroy your relevancy So if you had one very high quality question that has many, many words, it will always be on top no matter what the what the people would search for. That's not what you want. You want to take the relevancy and you want to take the popularity
Speaker 1: and combine those those numbers together to get something that you can then sort by So another use case would be I'm looking for a hotel or I'm looking for a conference for an event and I want it to be in this location I'm not strictly limited to that location, but I would prefer it to be there or around this time. So again, we can take the numerical indicator, the distance from our ideal. and use that to and feed that into into the relevancies mechanism. So this is the theory behind it and this is this is the practice. This is how it actually looks This is a this is a fairly simple query in Elasticsearch when we're looking for a hotel.
Speaker 1: We're looking for a hotel that's called the Grand Hotel, and we have several other criteria. We would prefer to have a balcony We are not limiting ourselves to just hotels with balconies, but if it does, like we we want to add two to the score. So we want to bump it up. Also, we want it to be uh in central London And we want it to be within one kilometer of the centre of London. Well, not centre, Greenwich. You can identify it by the zero in the coordinate. And we want it to be there. We don't want to limit ourselves to uh to hotels that might be so good that they would get to the top even without fulfilling this criteria.
Speaker 1: So we say that we we want to use the sort of the uh the Gauss function to calculate the score. It's one of the one of the eh sorry it's one of the shapes on on the image on the right. That actually determines how fast the score drops once you get outside of your ideal zone. And we also want to take into account the popularity of the hotel. And then we want to add some random numbers. So random numbers are always good. They make everything so much better. Now in this case, the random numbers are there sort of to to shuffle the results a little bit around. so that uh people have chance to discover new things. Because uh we actually took this from from a from a customer from a real example how they're how they're doing it
Speaker 1: And they do have this random score there because otherwise they would have some some hotels or some uh some results that would never be hit because they would always be just just behind uh just behind the fold And also people would perceive the results as stale. But if you if you sh uh if you shuffle it uh around a little bit, they will always find something new and they'll always be excited and hopefully come back to your website So, uh, this is one of the ways how we use the statistics, the the TF-IDF, and all the things that we know about your data. This is the most straightforward way. We use it to calculate relevancy and we allow you to hook into that process yourself.
Speaker 1: If you are so inclined you can just uh remove all of this and just say hey instead of all these different criteria I just want to use a script and give it an expression in your favorite programming language whatever that is even even Python So and uh do do all these calculations yourself. These are essentially just pre-built scripts that we have we have built So that you can you can use it and you don't have to expose the scripting functionality because obviously that can have some uh that can have some issues. So this this was this was the first use case. Sort of how to get more out of the out of the relevancy that we already have.
Speaker 1: Another interesting use case we have revolves around reverse search, or how we call it the percolator. And it is exactly what what it sounds like: it is reverse search In normal world, you index your documents and then you run your queries. With percolator, you index your queries and then you run your documents. So uh what is what is uh what this is useful for is for uh example alerting. If you have something like stored search functionality on your website You have classifieds or something like that and you allow the users to search and then nothing shows up But the user wants to say, hey, I I I'm interested in this search.
Speaker 1: Like save it. And whenever there is a new new item, new document that matches that search, just send me an email and I'll I'll come back Normally that's a fit that's a fairly hard problem. With percolator, it comes out of the box. You index the query that they're that they're running Including all the bells and whistles that uh that Elasticsearch allows you to do. And then when a new document comes in, you just ask for it to be percolated And you will get back all the different queries that people have registered that they want to be alerted on. Some people even use it to power uh something like live search. If you've ever been on any website, you're searching, and suddenly there is a pop-up, hey, in the time that you've been looking at these results, there are 10 new ones.
Speaker 1: That again can be powered by the percolator. When you do a search, and at the same time you register the percolation, and every single new document that comes in gets percolated as well. And you will know which browser you need to push this document into. Who is actually looking at the results right now So those are the fairly obvious use cases. I'll talk about my favorite one for percolate, and that is classification. Because there is a bunch of stuff that's super easy to to search for, but not that easy to do it the other way around. For example, geolocation. It is fairly easy to construct a query that will look for
Speaker 1: events in Austin You just have to have the shape of Austin somewhere. You either pass it into the query, but that's not optimal. So typically you have it indexed somewhere in Elasticsearch. In my case, I have an index called shapes, where I have cities, and I also have Austin there. So I say, hey, I am interested in anything that falls within the city of Austin. And that is a very simple query to run. But what if I want to do the opposite? I have a geo point, I have a set of coordinates, and I want to know where they are, what city, what's the address. That's not that trivial unless you have something like this, where you have a bunch of queries indexed in your Elasticsearch and then you just show it a document
Speaker 1: with uh with a geo point. And it will tell you, yeah, oh yeah, these these queries matched. The one representing North America, the one representing uh United States, the one representing Texas, the one representing Austin. Maybe even the one representing a city block or something, so you can really pinpoint down the exact address. At that point it's just a matter of how good data you have and how much CPU you want to burn on this. But it gives you sort of the non-obvious reconstruction of the data. Another interesting one is language classification. This operates on an assumption that every language, and I've chosen Polish because English would be way too uh way too weird.
Speaker 1: Uh every language has a few words that don't exist in any other language. So the assumption is I can run a query that looks for these words. In this case, I'm I give it a list of words and I'm saying I want If at least four of those are in the document, like I consider it a match. So in this case, I'm looking for documents that are written in Polish. Because nobody else, no other language in the world would ever write a a uh a word like this. So it's a fairly fairly good indicator. So again, uh very easy to write the query and once we have the query we can reverse the process and actually uh ask for uh ask for which queries
Speaker 1: matched. So if I then have a document for an event, so I have a DjangoCon US with some Polish description because otherwise my demo would fall apart. Uh this is what I get sorry about that. This is what I get back. I get back the identifiers of all the queries that actually matched. So I know that this is an event that is uh in the city of Austin. It is in Polish. And it is it it deals with the topic of Python. Please don't try to look where the coordinates are.
Speaker 1: They're nowhere near often, but So this is how you can use uh percolator for classification. You typically do this before you index your document. You have your document, you're about to index it, so you run the percolation, you get all the dynamic classifiers, you get all the topics and all the uh the language and the geolocation, you put you add it to the document and then you index it So then when you're looking for something in the city of Austin, it's a super simple exact lookup, which will obviously be much faster than running the Geo lookup every single time. And you can take the percolation a little bit further. You can attach metadata to the percolation queries. For example, who requested this
Speaker 1: percolation? Is it a user that paid me or is it a user who can wait? And and uh other other criteria like that. Uh you can also uh not run all the percolations every single time. You can then use this metadata to filter those, etc. etc. You can also use it to uh to highlight. So if you want to highlight some passages, but it that's that can be a fairly hard problem, but it's a problem that search engines are really good at So you can ask uh ask Elasticsearch to actually highlight some some passages for you and then index index the already highlighted text. So this has been uh this has been
Speaker 1: our journey into the depth of the percolator. And then there is one last big thing that I want to talk about. And that's aggregations. Now many different many different uh databases have aggregations. How come a search engine has one too? Well it all started with something like this This is an interface you might be familiar with, and it's called a faceted search or faceted navigation. You type something in, you get the results that you get the ten blue links. But you also get on the left side the overview, the overview of what actually matched your query So in this case I'm looking for Django, so I immediately can see
Speaker 1: that Django is mostly connected with Python. And I can see how many repositories and how many users have actually matched my query. And that is the huge difference between facets and search. Search is great when you know what you're looking for. If you know how to spell Django or how how it sounds or something like that. The facets are great for exploration because you don't need to know. You look and you see. It is one thing to do with code, the other thing is to do it with hotels or books or if have you if you've ever shopped on a website like Amazon like you can see the categories you can see the brands you can see the price distribution
Speaker 1: And you can see it. You don't have to read all the results to get that information. So we have taken it one step further with Elasticsearch and we we power some analytics based on based on this stuff. And we visualize it because humans are essentially pattern-recognizing machines. You're very good at recognizing patterns You can probably spot several weird things about this picture, like the gap in the uh in the timeline, or the fact that the two last pie charts are completely different. And you can see that immediately. If you wanted a computer to see that, you would have to tell it what to look for or have something very, very sophisticated.
Speaker 1: But any human can spot this immediately. So that is why why facets and aggregations become so important and why we continue to develop them But this is boring stuff. I don't want to talk about this. This is just counting stuff. Any database can do that. We can do better than that. We are a search engine. We understand your data and we can use it. So let's see how we would actually use Elasticsearch and aggregations to do something like recommendations. Let's say that I have a music website and I have different users and then they like different artists So I have a document per user and there is a list in that document in that
Speaker 1: JSON that has all the all the artists that the users like. So one sort of naive way how to do recommendations is just ask for the aggregation. Just give me a look for people who like the same thing I do. And then give me the top ten popular artists in that group that I I've not been exposed to yet That's easy to write, easy to run, known as useful, because popular doesn't mean relevant. Just because everybody listens to one direction doesn't mean it's relevant to my uh to my group. So what can we do instead? We need to find
Speaker 1: we need to identify the artists that are more relevant. to my group, to the to the group of people that like the same things that I do compared to the background. So the code looks looks remarkably different. We replace the word terms with the word significant terms. And that's all all there is. Because now we are essentially telling Elasticsearch, hey, we have we have this group. We have defined it based on based on the results of the search. And now Give me the stuff that is more relevant for this group compared to all the others. And we can do this.
Speaker 1: We are the search guys. We understand the data. We have all the statistics. We have all the numbers. You can even see the graphical representation of what's happening on the right side. Normally, when you select sort of a random uh uh random selection you would expect that there would be the same distribution of people who like something in that group compared to the general populace. So you would expect all the data to be laid out on on the diagonal. What this did, what using the significant terms did, it selected all the old values that are pretty much on the vertical, which means they are much more liked in my group
Speaker 1: Than they are in the general populace. Which is exactly what I was asking for. I was asking for a recommendation based on the people who are similar to me. give me what I'm more likely to like. Using using the data, using still using just the dump statistics about uh the distribution of the individual values throughout the data set. So no learning was involved. This is actually a fairly simple aggregation that you can that you can run. You can see the code is not that expensive, not that expensive to write. It's not that involved. But this still has some problems.
Speaker 1: This is an aggregation we've had for a while, but we've noticed that there are some cases where it doesn't work as well as we would like it to. And that uh the one uh problem is the the terms that everybody likes, the term that everybody has. So One Direction is my go-to example of this because you know everybody likes them, especially around Django Cons. And so what do I what do I do there? How do I make sure that I don't suffer from the bias that every single document I have actually likes this? Well, a lot of it is already filtered out by the significant terms, but I can actually do much better.
Speaker 1: I can also ask for a sample. Of the documents. So I will not do this analysis on all the users that have something in something in common with me, but only those that are the most relevant. So I I've I've included relevancy right now at least twice in uh in my query that I'm running. First of all, I'm looking for users that are most similar to me. And most similar means they have the highest relevancy to my query. Because let's repeat, the relevancy is based on TF, IDF, norms, etc. So what does this translate to?
Speaker 1: TF, term frequency. It immediately translates to people who likes more things, who have more things in common with me. IDF, inverse document frequency. It means that people who prefer the the rare choices that I have, I prefer the people who have like the same thing with me. That are rare. I will ignore the one directions because that doesn't bring me anything. But I will actually hang on to uh to the weird groups that nobody else in the world knows about. And then uh norms, the stuff that Lucine adds on top of it when it takes into account the length of the field.
Speaker 1: I'll prefer people with shorter lists. That means that people who like pretty much exactly what I like. So not people who like everything in the world because that will not be relevant. So I can use directly the tools that we've built in the beginning of the talk for text matching. And I can use the same numbers, the same formulas. to get the people who are most similar to me. And then I say, and take like five top five hundred of those, like on every shard because Elasticsearch is distributed, so everything happens on a shard level. And and then run the significant terms.
Speaker 1: Tell me what is specific for that group. So we have refined the selection, not anybody who has anything to do with me, but the most relevant people, the most similar people, and then give me what is specific for that group So this is the sampler aggregation. It is currently currently in the newest uh releases of uh of Elasticsearch being ready to release. And these two together allow me to try to do the recommendations and have it be more relevant and also have it be faster because we are actually looking at the subset of the data. We have just used all the information that we have about the data to identify the most relevant part so we don't lose any precision, quite on the contrary.
Speaker 1: So sort of to to uh generalize what we have done here is in a connected graph When we have uh people and artists and they like each other, we have identified uh the connections that are meaningful. Not the ones that are most popular or most common or we've managed to hopefully circumvent sort of the super nodes, the superconnected nodes that are that are the hubs of uh of any of the any of the graphs. And we can use this and and go further with it. We can actually use this uh in a graph algorithm. So imagine that you have a an algorithm that's uh used to calculate the shortest path.
Speaker 1: So you want to go from point A to point B, or in our case, let's model it on Wikipedia. You want to go from a page A to page B, and you want to see what are the connections. How do you get from one point to another and still only take into account the relevant connections? Because if you just uh use a naive uh graph algorithm, you will you will fall victim to to the supernotes. Like you will s you will see that Hey, Beatles have concerted in the in the USA and so has pretty much every other band, so there's an immediate connection there. And that is not relevant at all. And I don't mean to insult your country, but that's just not relevant connection.
Speaker 1: So when you when you use this approach to identify only those connections that are relevant, that are not just an accident based on statistics that everybody has that connection then you can uh get much more interesting information out of this. So you can actually use this to uh to find meaningful connections. After all, aggregation and relevancy. That is how we look in the world. I said earlier that we are a pattern recognizing machines. We we look at we look at things and we immediately make assumptions like
Speaker 1: hey this room is not as full as it was in the last talk. Yeah that makes sense like Elasticsearch is not as relevant as Postgres. I I get it I I do. And I can see that immediately. And because I have the context, because I can because I can see I don't have to count all the chairs to to know that. So that's that is the aggregation part. And the relevancy part is if I ask you what is the most popular website, the website that you visit most often? Many people, when I I actually asked this at conferences, they'll tell me something like GitHub or Stack Overflow or something. And that is actually not true. They probably spend more time on Facebook or Google or something like that. And it's not that they are ashamed, even though they might be
Speaker 1: too It's they immediately recognize that that is not a relevant answer. That's not interesting. Everybody goes to Google. Like I'm not interested in that information whatsoever. I know that, you know that. Let's skip to the interesting part. GitHub. That is more relevant to our group than to the general public. If I ask on the street what GitHub is. Like hopefully I won't get punched, but I don't know that. But de I won't I'll definitely not get the correct answer. So we do the exact same thing. That is what what is special about humans compared to compared to computers. So I have a I have a question for you. If I do an aggregation per
Speaker 1: time period and then I ask for significant terms on the tags or anything, what do what do I get back? How do we call that Anybody's willing to guess? Okay. We get to the trending information. What's trending? Not what is the most popular for that given time period, but what is more specific for the time period compared to any others. So if we only filter the last five minutes and we ask for significant terms, we get what is currently trending, what are people talking more often right now than in general, than any any other time. So something that might appear as
Speaker 1: sophisticated algorithm, we can replicate it with two lines with a single query to LSIC search. So this will this will this is a nice nice sort of shortcut if you just want to want to throw something out there and you don't want to spend uh spend billions trying to come up with your own algorithm, like this is exactly what it is. So I uh so there is only one other pitfall that I'll I'll warn you about. If you try to play with this. The way uh aggregations work in Elasticsearch is we take all the possible combinations that might come up and we create a a bucket for them, a placeholder.
Speaker 1: And that can blow up very fast. For example, if you if you're look if you have a data set from IMDB and you're looking for author uh for actors who acted together most often If you just run this naively, it will it it will work and the query is super simple. But the query will also blow up your memory like crazy. Because it will essentially do uh do uh Cartesian product of actors versus actors. So it will be a huge essentially table, a huge data structure that we need to then fill out What we can do, however, is we can limit one of the dimensions before we get into it. So
Speaker 1: uh by default we will try to do everything in one pass over the data because it's the most effective way and because we uh we are a distributed distributed database so we need that sort of for us to function But in this case, if you really insist, if you know that this will blow up your memory or because you've tried or because uh uh because you you can count. Uh what you can do is you can say do it do it another way. Do a breath-first search. So first identify the top ten actors and then only find the coactors of those ten. So just simply identify
Speaker 1: like what is the what is the uh dimension that you can limit most effectively And do that, and then then you'll be fine. Then you can actually ask for all the information in the world. And we will give you a tiny, tiny, tiny little sliver of that. So remember, information is power. We have the information. We have information about your data and we can we can use them. That is the one leg up that we as as a search engine. uh have over the the more traditional data source we're limited to to the boring kind of stuff filtering counting stuff It can be useful, but it's not exciting. If you want exciting, use Elasticsearch
Speaker 1: Thank you very much. If you have any questions, so
Speaker 2: we have five minutes for questions if anyone has any questions. Uh
Speaker 3: yeah, great presentation. Uh question about language detection. Have you tried to use language detection within Elasticsearch? uh in in a production environment?
Speaker 1: Uh no. So what what I typically recommend when people deal with multiple languages is just use everything at the same time. The problem with language detection typically is that even if you can detect the language of the source, you have enough information to go on there and you can do it. You can identify the language of the document. It's very hard to identify the language of the query. Because if you if somebody just types in two words, it's very hard to say what language they are. So what I typically recommend people is if you know that you're going to be dealing with these five languages, analyze everything five ways. Analyze it as English, as as Czech, as German, as Japanese, and then do the same for the query. When a query comes in, query all of these fields at once.
Speaker 1: And Elasticsearch has tools to allow you that. You can specify that I have this one field, but I want to analyze it multiple ways. And then I have this query and I want to run it against all these differently analyzed fields. So it's what I call the shotgun approach. Just throw everything in there and see what sticks. Because of how the relevancy works and how different uh relevancies from different queries are combined. Without trying so hard to actually think about the problem, you will actually get the most relevant results. Does it make sense?
Speaker 4: I may have missed it, but is the vector space model still a a common way of uh combining information from different query terms, or are there more sophisticated
Speaker 1: TFIDF and everything, but that's only to give weight to the individual parts of the vector. But overall it's still it's still essentially we're talking about uh vector metrics and vector differ uh uh distance by default essentially what what I showed the the formula that's actually a cosine uh uh uh metric pretty much. Like there are some modifications and stuff, but yes, it's still based on based on that. Okay, thank you very much for having me.
You can combine the normal text relevance score with external signals such as ratings, popularity, distance from a preferred location, or proximity to a desired date. Elasticsearch’s scoring functions can boost or reduce results without discarding otherwise relevant matches, and custom scripts can replace the built-in scoring when needed.
Discussed at 8:47The percolator reverses the usual search flow: you index queries and test incoming documents against them. This supports saved-search alerts, notifications for new matching items, live search updates, and other applications where documents need to be classified by registered queries.
Discussed at 13:26You can store classification queries—for example, geographic shapes or distinctive words from a language—and percolate each incoming document against them. The matching query IDs identify attributes such as the document’s city, country, language, or topic, which can then be stored for fast filtering.
Discussed at 15:47Use significant terms rather than simply selecting the most popular values: first find users or documents similar to the current one, then identify terms that are unusually common in that group compared with the broader dataset. A sampler aggregation can limit the analysis to the most relevant matches, reducing bias from universally popular items and improving performance.
Discussed at 24:18Filter the data to a recent time period and run a significant-terms aggregation on the relevant tags or fields. This returns terms that are more specific to that period than to the dataset overall, rather than merely the most popular terms.
Discussed at 34:24Large combinations, such as actor pairs, can create a Cartesian-product-sized set of buckets. Use breadth-first execution to identify the top values in one dimension first, then calculate combinations only for those values.
Discussed at 35:54Instead of relying on language detection—especially difficult for short queries—analyze the content separately for each language you support and query all of those analyzed fields. Elasticsearch can combine the resulting relevance scores, which the speaker calls a “shotgun approach.”
Discussed at 38:46Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 15, 2026
Published July 14, 2026