Pseu, Pseu, Pseudio. Pseudonymization in Django.

This video features Frank Valcarcel at DjangoCon US 2018 in San Diego, California, USA.

Pseu, Pseu, Pseudio. Pseudonymization in Django.
0:29:40
Published November 10, 2018
409 views

DjangoCon US 2018 - Pseu, Pseu, Pseudio. Pseudonymization in Django. by Frank Valcarcel

The General Data Protection Regulation, better known as GDPR, is a regulation on data protection and privacy for all individuals within the European Union. GDPR went into effect on May 25, 2018 and was the cause of the “Great Privacy Policy Update” that occurred in the weeks prior.

This talk will cover what GDPR is and why you should care about it, but we won’t stop there. This is not going to be another talk on data protection policy. No.

In this talk, we’re going to jump right into discussing HOW to implement data patterns that comply with regulations like GDPR by examining a pattern known as pseudonymization.

Pseudonymization is a data de-identification procedure where fields of personally identifiable information (PII) within a data record are replaced by one or more artificial identifiers. These artificial identifiers are also called pseudonyms. Pseudonyms make a data record less identifiable without sacrificing data analysis and processing. GDPR requires that PII undergo either pseudonymization or complete data anonymization.

For the hands-on portion of this talk, we’ll construct a Django User Model where we apply pseudonyms to the data attributes which qualify as PII. We’ll explore a couple strategies for implementing a compliant pseudonymization pattern, examining their individual approaches and performance, and we’ll discuss limitations of pseudonymizing certain attributes and how to achieve compliance through consent.

GDPR sets a precedent for responsible data management. Whether your application serves citizens of the EU or not, the regulations serve as an encouragement for protecting your user’s identities. This talk is great for everyone from beginners to expert Django developers… and fans of Phil Collins :)

This talk was presented at: https://2018.djangocon.us/talk/pseu-pseu-pseudio-pseudonymization-in/

LINKS:
Follow Frank Valcarcel 👇
On Twitter: https://twitter.com/fmdfrank
Official homepage: https://www.cuttlesoft.com

Follow DjangCon US 👇
https://twitter.com/djangocon

Follow DEFNA 👇
https://twitter.com/defnado
https://www.defna.org/

Summary

Pseudonymization replaces personally identifiable data with reversible artificial values so it can still be used for analysis while reducing exposure. Frank Valcarcel distinguishes it from anonymization, approximation, encryption, and tokenization, then shows two Django implementations: property-based masking for legacy models and a custom Django field that masks values before database operations and unmasks them when read. He also covers query filtering, deferred fields, Django admin forms, validation, audit logging, and the need to choose a masking method based on the threat model rather than treating the simple demo algorithm as production-ready.

Key takeaways

  • Pseudonymization reduces identifiability while preserving data utility, but unlike anonymization it remains reversible with the right information or process.
  • Personal data should be identified broadly: information that can identify someone alone or alongside another data point may be subject to privacy requirements.
  • For legacy Django code, properties, custom QuerySet methods, deferred fields, and customized admin forms can provide end-to-end masking and unmasking.
  • A custom Django field can centralize masking by transforming values before database interaction and restoring them when values are loaded.
  • The simple character-shifting example is for teaching only; production systems need a stronger masking strategy and an explicit threat model.
  • Access to re-identified data should be controlled and logged, including who accessed it, when, and for what purpose.

Summarised automatically from the transcript.

Transcript

5,019 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:05

Yeah, I think that's a good thing.

0:16

Speaker 1: Uh, this is on. Okay. Hey everybody. Thanks for joining me for what is probably going to be my uh my silliest talk of the year. Let's just get one thing out of the way. How many people who know who Phil Collins is by show of hands? Alright, we're gonna have a lot of fun. For those of you who do not know who Phil is, I've got plenty of background information on him. And um He's uh yeah, we'll we'll get to that part. So this talk is called SueSudio, um and it will cover pseudonymization techniques in Django. Hi, I'm Frank. I'm FMD Frank on Twitter, but I am also

1:02

Speaker 1: on quite the extended Twitter sabbatical. You're welcome to go look at my greatest hits. They are there for you to peruse. But I may not go back. I don't know if I will ever return to the Twitter sphere I work at a company called Cuddle Soft. We have offices in Denver, Atlanta, and Tallahassee, Florida. And I'm an avid Pythonista. I've been using Python as my primary programming language for the better part of eight years. This is my first time at DjangoCon. And I'm very excited to be here. I'm also the co-founder and chair of Pi Colorado. We'll be having our inaugural conference next year in August. I'm happy to talk to anybody more about that if you're interested.

1:47

Speaker 1: Please come visit me in Denver. It's beautiful. And then I also run BoulderPython in Colorado. So yeah, thanks for having me. So my speaker in spirit is Philip. Oh my god, I forgot his name. Philip, sorry, it's right here. Philip David Charles Collins. He's an English musician and he's a drummer, singer, songwriter, multi-instrumentalist, record producer, and an actor. He was the drummer and singer of a rock band known as Genesis. And during the 80s, Collins had more US top 40 singles than any other artist, which if If you're old enough like me to remember the 80s, that's actually quite impressive. He co-wrote a lot of the music on Disney's Tarzan for the younger folks in the crowd. That will be probably how you know him.

2:32

Speaker 1: And also I just learned that Peter Gabriel, none of this is relevant to the talk. You probably figured it out, but Peter Gabriel was the original lead singer of Genesis and Phil took over for him. So why is Phil my co-speaker in spirit? Well, pseudonymization is an incredibly difficult word to say. Try it. How many of you got it right? Yeah. So sudo is close enough, and that was enough reason for me of of like as of any to to do a Phil Collins-inspired data privacy talk. Also I'm pretty confident I'm the only one to have ever attempted this, so we'll see how it goes. Um all right, so if you've never heard of Phil. That's okay. I got you. We've got a

3:18

Speaker 1: Spotify playlist of some of Phil's greatest hits. He is on Twitter. He's Phil Collins feed if you're interested. It starts with Si Studio, which is the song I started this talk off with. It gets kind of sappy towards the middle, like this talk will. I don't know, are there any tissues? If there are, you'll need them. This is heavy stuff, y'all. Um and of course this playlist ends with the air drumming spectacular in the air tonight. So please check it out. Enjoy it. Alright, so let's get to the meat and potatoes. What is this very difficult word to say? Well, it's a data de-identification procedure. Data records are replaced by one or more artificial identifiers called pseudonyms.

4:07

Speaker 1: And the idea behind pseudonyms is that it makes a data data record less identifiable without sacrificing data analysis. and processing. And so why would you do this? Well anything worth protecting is worth protecting well. And it provides you some security through obscurity. So you can secure a data set from identify identification. And it's also kind of required by the law. Not kind of, it is required by the law. I only say kind of because there are these gray areas which I'm not going to get into because I am not a lawyer. So do not ask me legal advice. If you have a question at the end and it smells to me like it's of need of counsel, I will tell you I cannot answer that and that you need a lawyer.

4:54

Speaker 1: So uh a couple more things. As a note, we're not going to get into the mechanics of GDPR. I will reference some articles if it's important. and interesting for you to go read. It's actually not that dense of a regulation. So we will be avoiding things like consent, the difference between collectors and data processors. or how it affects your organization. Again, if you ask me those questions, I am not a lawyer and I will tell you that. But it's important for us to define exactly what it is that we are discussing today and that is specifically personal data. This is also known as personally identifiable information. The gist is that personal data is any identifiable data or PII. Note that GDPR refers to it as just personal data.

5:39

Speaker 1: One of the things about the regulation that I don't like so much is that it does paint in very, very broad strokes. So essentially any information that can be used to identify a user, a person , there's this regulation around. So some examples. Basically it's this is if you can identify someone with it or it can be used to identify someone with or without a secondary data point then yes, it's personally identifiable information. If you are unsure, chances are that it's personally identifiable information. So let's talk about data privacy techniques, right? There's two very popular methods. There is pseudonymization and anonymization.

6:25

Speaker 1: Beyond being very difficult to pronounce the first few times that you practice them, they're the two most common approaches to doing data privacy over PII. Sydonimization we kind of covered a bit already. I want to also point out that according to Article 25 of GDPR Data must be protected by design and by default. So these are important things to consider when you are planning, even in the planning stages of a system. And if you want to understand the requirements underneath the regulation, I recommend you read Articles 25 and 32. I'll note that GDPR only recommends one technique by name, and that is pseudonymization, although they spell it with an S and not a Z.

7:12

Speaker 1: That's something I've learned. Anonymization is a more permanent de-identification procedure. With anonymization, you render the use the user's data unidentifiable. So maybe one of the reasons why the many teams of lawyers that wrote GDPR's regulations avoided using anonymization is that the fact the mere fact of The operation of anonymizing a data set makes it no longer personal or personally identifiable. So it actually doesn't fall underneath the purview of GDPR, which is something I think is really interesting. If you are struggling to understand the differences, I'll have some examples on pseudonymization, but if you're struggling to understand the differences between the two, I've made this drawing. To help articulate the differences between pseudonymization and anonymization.

7:58

Speaker 1: Anonymization is essentially like I think of it as analogy to like Batman's very clever disguise, right? When he puts When he puts the mask on, you can't tell it's Bruce Wayne anymore. Thank you. Um But Superman, not so much. He's got he combs his hair a little bit differently and he puts some glasses on. So at least to all the people in Metropolis That are just not that keen to see that he is the same person, he is pseudonymized. They can't tell it's him. But to us, the readers, there is no anonymization layer going on, right? I really just use this as an excuse to make this incredibly funny slide. I think it worked out. So let's dive deeper into pseudonymization techniques.

8:44

Speaker 1: The one that we are going to go over is a technique called data masking. And so To mask data, characters in a record are shuffled or substituted in words, maybe sub may be substituted or obscured completely. The result is usually a realistic data set that cannot be reverse engineered without the re-identifying information or the or the algorithm to reverse the masking technique. There are a lot of techniques that fall under this broader category. There's also a method known as approximation, which is instead of saving the information the user's PII by itself, you approximate it. So one of the common practices this is used for is for date of births. Sometimes you don't want to save a date of birth. You just want to know how old, or maybe the birth month, or maybe the birth year. So then you have a table with those numbers and you increment those.

9:32

Speaker 1: Once a user subscribes or enters that information in, you don't save that user's date of birth record specifically. Another method, very popular, encryption. And this is something that I expect most people will be familiar with. I do have a question though, is does anybody know if this is required by GDPR? No, it is not. Not as a data de-identification procedure. Encryption is required. for data at rest and in transit, but it is not the recommended nor a requirement under GDPR for how to identify how to de-identify your user data. This is actually kind of interesting because one of the big premises is why I'm doing this talk in pseudonymization is A, it's fun because of the Phil Collins aspect, but two, it's actually better for you as an organization

10:21

Speaker 1: and someone who's serving maybe the role as the data processor and the and the database administrator, pseudonymization gives you a lot of value back but you don't necessarily have to go jump to encrypting that data set, because this will add compute resource or compute resource requirements that you don't necessarily need. This is at least my philosophy. Again, I'm not a lawyer, so. And then the final pattern is tokenization , which is a very common use commonly used by companies like PayPal or Apple Pay or Stripe. They will tokenize a credit card's information and then they use that token to retrieve that information when they need it. They only process those that those data points when they need to. Otherwise it's saved on either the client's

11:07

Speaker 1: the client side, the vendor side as this token representation. This is song two on the playlist if you're following along. Alright. So I'm going to go over a simple implementation example. This is going to set the foundation for how we're going to scale this up in our in in our Django example. Uh so Python already supports a common pattern that allows engineers to replace attributes with a set of methods that can intercept values when they are written and they were and when they are read. Any guesses as to what they are? Not trick question, but they are either getters and setters of the properties. So for the following examples and for the continuing examples through the Django

11:53

Speaker 1: methods that I'm going to show you all, we're going to use this incredibly simple uh masking algorithm. The masking algorithm does it all it does is shift each character uh one ordinal to the right and then when it re-identifies them, it shifts them to the left. It does it in a range so that it can not overrun the ordinal ranges for um ki ASCII characters. So it's it's intelligent from that point, but it's very unintelligent if you use this in production because it's not insanely easy to reverse engineer. I'm also not going to talk about algorithms or best practices for doing masking because We first I shouldn't share it with you. Two, this is being recorded. And why would I like why would I you know implicate all of us by sharing an algorithm that then somebody here may go and use

12:42

Speaker 1: And then that is reverse engineered. Now I am culpable. So it's also you know this is a lot easier for everybody usually to understand. And so if I had a more um complex or sophisticated masking algorithm that would take the bulk of time we have for the talk. So to mask and unmask, we're just going to have two methods, mask and unmask. And then essentially this is how it would work, right? We're shifting my name, Frank Valkar , over every character over one, and that's what the masked version of it would look like. So an implementation of this, if we're just using a basic user class, we have an underscore name property, sorry, an underscore name attribute, and then we have a property method for it called name and a setter on name. And then we just call our mask and unmask methods underneath

13:31

Speaker 1: those two functions. And so it'll look something like this now. So when I instantiate user, I'll set the username as my name. If I print username, it's coming from the property, so it'll return my name. But if I'm looking at the underscore name attribute, it's returning the pseudonymized version. This is important to understand because what's being saved in the object and therefore could be serialized later is the pseudonymized version. It wouldn't be my name. My name is only being re-identified in transit So let's look at a Django example. This is song three on the playlist if you're following along. And so we're going to take the same concepts, I'm going to add a few attributes, but we're going to focus in on the name field. Quick question, how many of these attributes are PII?

14:20

Speaker 1: All of them. Yeah. They're all identifiable. Because together, something like the IP address with one of the other data points makes this the user who's saved identifiable. So we're going to move our shifting algorithm, our masking algorithm into a utils file. The code for this is all available later. I'll share the link with you. And then our mask and unmask methods. Then now here's the user attribute, again focusing in on just the name field. We've done the same process. It's underscore name and we have a getter and a setter applied to it, which will mask and unmask as that data moves in and out of the objects. So the problem is that we're not done, and uh for sake of time I'm gonna speed through the rest of this because I want to get to the second example.

15:05

Speaker 1: Um The models query set doesn't yet support our properties. You cannot filter, you cannot exclude on the identifiable data values, right? You have to know that Frank will be pseudonymized and masked to Gizbol or something like that, right? And so therefore that's not a very intuitive way to interact with your data models. The other thing is that these pseudonyms are now included in all of our user objects everywhere that we're retrieving them It just pollutes the user model, I'm sorry, pollutes the object. It's useless. It's just going to add uh weight to that uh that data object and we don't need it. And then also the Django admin has no idea what to do with this. So first let's start updating the query set. We are going to monkey patch some of the methods on query set so that we can filter and exclude.

15:50

Speaker 1: I'm not going to do all of them. I'm just going to do filter and exclude and I'll show you that we actually get a few more. There's a bit bang for your buck by just monkey patching these. Also, for sake of time, I won't be looking at the source code, but Just to note the reasons why this is the reason why this is here is that when you patch filter or exclude you get filter exclude and get out of the box. So you only have to monkey patch that one function and you can see this in the source code that they all just call filter or exclude. Then we'll insert our mast values and we will super the parent instance of our custom models. query set for everything else. So this is what this will look like in code. We have our masked fields, name. And then we iterate over the masks fields and create a keyword argument that we then pass to our, there's my mouse, our

16:37

Speaker 1: filter or exclude customized method. So now we'll be able to do things like filter on the identifiable name or exclude on the identifiable names. And then the last thing we need to do is override the auth user manager get query set. And you can see how I've done that there for the object. Second thing we have to do is exclude pseudonyms. They pollute the models. So there's actually a method called defer. The gist is that if you don't need a particular field, when you fetch the data, you can tell Django not to retrieve them from the database using defer. So very similar to the last, we'll create a new list. Hey, F strings for the win. We'll iterate all over all the attributes in our model that start with underscore and then we'll add them to our keyword arguments

17:25

Speaker 1: that we pass to defer which is chained. at the end of our monkey patch filter or exclude. And now when we query using filter, we can use the identifiable information. And then also the object that is returned does not have those pseudonyms in in inside of it. I didn't overwrite all. You would have to overwrite all in this method. The last step here is updating the Django admin. So write read is masked and unmasked, but what about Django admin? Well, it doesn't know how to do this. It doesn't know that we want to display the unmasked values in the admin. It doesn't know to mask those values. when you submit the forms in the admin. So we have to start by telling it what fields we want to show.

18:12

Speaker 1: Then we'll begin to define a form that we can swap out in place for the default one Django wants to use. You can see we're overriding the built-in user change form from Django Contrib auth forms and we're creating a form with a new char field On initialization, we get the correct value and check it against the validator for our masked field, which could be important with something like a phone number if you were using phone regex. And when the forms clean method is called, we can get the appropriate value or error out on invalid input. Next, we've got to register this. So we have our base fields of the model, namely username and password, but now we've created a group of subfields called personal data and we've added the name property to it. And last we told Django Admin how we would like to use uh how we would like the users to be displayed in the user list.

18:59

Speaker 1: I will note that this last step could be very important. And you may want to, under the regulations of GDPR, you may want to add some blogging on this because according to Article 30, each processor shall maintain a record of all categories of processing activities carried out on behalf of a controller. And a controller can be someone with access to the Django admin, whereas a processor could be you, the engineer who wrote this process, right? You are obligated and responsible for logging every time someone re-identifies this PII, when it happened, who did it, and sometimes why they did it. what that business process was. And so we are set. We have finally encapsulated some pseudonymization through the entire life cycle of this object and how we manage the data of this object. It's stored in the database as its pseudonymized field, and when we retrieve it, it will be

19:45

Speaker 1: the identified the re-identified values. And best We have access through the Django admin to work with that data as it is as it sits identified, but it's always going to be saved as a pseudonymized field. So the next example is a new and improved method. It's a lot more straightforward than the last one. I show the last one and I build up from it because There's a lot of work that needs to be done on legacy code. And just because the last method was naive doesn't mean that doesn't make it wrong. You may have a system that has very few fields that constitute PII. and you need to create some safety uh and regulation compliance for your clients. That last method isn't bad, it's just that there are better ones if you are starting from the ground up. If you're following along, this is song seven.

20:34

Speaker 1: So we are going to do data masking via custom fields. Using a custom field class, we will automatically mask values on their way in and out of the database. With this approach, we no longer require getters and setters, the custom query set and corresponding user manager, or the bulk of changes we did to the user admin because we were doing at the field level. So we're taking the same user model as before. And here's our customized field. It's called Sodonomized Field. The class constructor and deconstructor methods will accept a field type, so we need to tell it what kind of field we are saving underneath. This will set the appropriate database column. And our deconstruct method has to mirror any argument changes we make in the constructor. This is the only thing I don't like about this method.

21:20

Speaker 1: We will also override the get internal type. Um which specifies the internal type of the field. If you've seen the field, the source code for field, this will look familiar. We are essentially overriding some of it to provide a masking and unmasking method when that data goes in and comes out of the database. And all of that work is done by these two methods. Get prep value is called prior to interacting with the database, and then from db value is called when a value is pulled from the database. So this is the core of this implementation. This is what makes it sing. We'll use get prep value as an opportunity to mask values before they are saved. And we'll mask values for query purposes, which

22:05

Speaker 1: is really cool. Also, we'll unmask our values when they're pulled from the dB and they're before they're converted to a Python object using db value. And so this is what it would look like when we apply it to our user model. We have a field now for name of Sodonomiz field and we tell it what type of field it will be, like what the field is underneath the hood. There's a tuple here that accepts the masking and the unmasking algorithm. So these are not tied to the customized field. You can swap these out You can use different ones for different field types. In fact, that may be one of the ways to improve upon this, is to have something underneath that knows intuitively how to shift something like a phone number, something like a zip code, or something like a name or a date of birth. And you can see that there's still validations that can happen on the phone field.

22:54

Speaker 1: And that's it, y'all. So thank you. I am going to take questions. There is a sample, there's some sample code and a blog post associated to this, and you can find them at those links.

23:11

Speaker 2: I see you have sample code, but have you packaged up uh this So don't pseudonomize field on PyPy PyPI so it can be used by other people.

23:26

Speaker 1: I wanted to So I shouldn't say I. We wanted to. Um this was a collective effort um with the number of engineers at at Cuddle Soft. Um we've chosen not to, mostly because the pseudonymized field um Code isn't that it isn't that verbose and we don't see the value in having something like that easily like injectable and like or available for your code. You can just copy and paste it it. We also think that there's a number of ways to improve upon it, which we haven't gotten around to yet. But it's just not, I don't know, to me it doesn't, it's not significant enough of source code that we need to have a package for it.

24:03

Speaker 3: Hi, thanks for the talk. Um can you explain why you wouldn't use encryption as a pseudonymization method instead of kind of rolling your own pretty much?

24:13

Speaker 1: There's a lot of reasons to use it. So it just depends on the use case. I think pseudonymization doesn't require encryption as its masking, unmasking method. I think you can achieve a lot with a clever sorting or shifting algorithm that you control or that maybe even you seed somehow, right? Encryption, encrypting and deencrypting the the object's attributes adds a lot of compute resources or adds the requirement of needing a lot of compute resources and so sometimes I just don't think it's necessary to add that type of overhead when there are perfectly You know, there are methods that exist that will perfectly handle um the regulatory compliance of it The other reason is that

24:59

Speaker 1: think about it from like a database administrator's perspective. Like sometimes you encrypt something and it fills out a lot more space than you you know, like a phone number would. Um whereas like if you are saving a phone number that's just shifts shifted around using some unique method, the the database space and therefore the representation of that data in the database looks a lot more like the identified data, right? versus some you know long hash

25:24

Speaker 4: um hi uh I was just trying to figure out how to phrase this question. Um so this is a really uh useful example of how to adhere to GDPR I guess that the broader question is if this is a useful way to make data more anonymized irrespective of GDPR criteria. Um do you have general advice for like people building whole apps to be more uh anonymizable

26:02

Speaker 1: Yeah, we use this method um to achieve HIPAA compliance. Just because there's a regulation telling you you should do this doesn't mean that you shouldn't do it if you don't have to adhere to said regulation. Like I said it in one of the earlier slides, like anything worth protecting is worth protecting well. And I think as engineers, our responsibility over time has increased in like how we need to handle users. data especially so I'm a consultant and especially for folks that are consulting for other entities right like if your client doesn't appreciate having some kind of data privacy technique in place to secure users data, there's no reason why you shouldn't have that, right? I don't necessarily think this method adds a lot of complexity over top of what you're trying to achieve.

26:50

Speaker 1: And then at the end of the day That that client could go on three years and have somebody working on the project that doesn't know how to configure an S3 bucket for their database backups, but you did all of their users a solid by just implementing some kind of pseudonymization technique, right?

27:05

Speaker 5: Yeah, so my question is, what is like the specific threat model that this tries to achieve? Um because if your database was dumped, a computer could easily reverse engineer a lot of these because it's not encryption. So I guess the question is what's the threat model that this actually solves?

27:20

Speaker 1: That's a really good question. And the uh if we're just talking about database dumps. I disagree. I don't think that I think it could take a long time to decrypt this if you have a smart masking algorithm. them. This one's not smart. This one's r this one's don't use this. But there's a lot of synonymization techniques that avoid shifting as like the primary basis for moving around that data or sorry, de-identifying that data. And you can mix and match. So tokenizing is not encryption, but yet nobody can reverse engineer a token to get back the valid credit card information to steal those credit card numbers, right? So I just presented masking and the shifting algorithm as a way for us to all easily understand like what was going on with the data in transit.

28:11

Speaker 6: What do you guys do for um I know we had a large um health provider and logging of the audit trails once they're um anonymized.

28:28

Speaker 1: Yeah that's a really good question. Sorry, I don't need No, I

28:34

Speaker 6: Frank will be available in the hallway per

28:37

Speaker 1: That's that is a really good question. So like something in the Django admin like creating a log for that, you know, would solve that problem. On accessing data between systems, we like CloudWatch. And then and then because we can use IAM to create the roles that we need and we can track which roles are accessing the data from which services. And then also IAM and CloudWatch give you a ton of Like transparency into just access of systems or access of services and things like that. And you can control when user passwords need to be reset. You can enforce multi-factor, a lot of cool stuff.

29:21

Speaker 6: Okay, so let's thank Frank one more time for the talk and the nostalgia.

Questions this talk answers

What is pseudonymization, and why use it for personal data?

Pseudonymization replaces identifying values with artificial identifiers so data is less directly identifiable while remaining useful for analysis and processing. It can reduce exposure of personal data and support privacy and regulatory requirements.

Discussed at 3:18

What is the difference between pseudonymization and anonymization?

Anonymization permanently makes a person’s data unidentifiable, while pseudonymization only disguises it: the original identity can still be restored using the re-identifying information or method.

Discussed at 7:12

Is encryption required by GDPR for de-identifying user data?

No. GDPR requires encryption for data at rest and in transit, but it does not require encryption specifically as the method used to de-identify data; pseudonymization is the technique it recommends by name.

Discussed at 9:32

How can I make Django queries work with pseudonymized model fields?

A legacy approach is to customize or monkey-patch the queryset’s `filter` and `exclude` behavior so incoming identifiable values are masked before querying. The same setup can defer the stored pseudonym fields so they do not pollute returned model objects.

Discussed at 15:05

How should access to re-identified data be audited?

The speaker says re-identification should be logged, including who accessed the data, when, and sometimes why. For access between systems, he recommends tools such as AWS CloudWatch and IAM to track services and roles accessing the data.

Discussed at 18:39

What is the simplest way to pseudonymize fields in a Django model?

Use a custom Django field that masks values in `get_prep_value()` before database operations and unmasks them in `from_db_value()` when values are read. This avoids the need for model getters and setters, custom querysets, and extensive admin changes.

Discussed at 20:34

Should I pseudonymize data even when GDPR does not apply?

Yes, the speaker recommends protecting valuable user data regardless of whether a specific regulation requires it. He says the same approach has also been used to help achieve HIPAA compliance and adds relatively little complexity.

Discussed at 26:02

What threat does pseudonymization protect against if someone dumps the database?

It is intended to prevent straightforward identification from the stored values, but the strength depends on the masking technique. A simple shifting algorithm is unsuitable for production; stronger masking or tokenization methods make reverse engineering substantially harder.

Discussed at 27:20

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos from DjangoCon US