It's not a bug, it's a bias

This video features Anna-Livia Gomart at DjangoCon Europe 2018 in Heidelberg, Germany.

It's not a bug, it's a bias
0:24:32
Published May 23, 2018
387 views

https://media.ccc.de/v/hd-61-it-s-not-a-bug-it-s-a-bias

Product makers have a biased view of the world, and this translates to biased algorithms.
How can we take this into account, and create a fairer world though fairer algorithms?

Even though Apple's Siri came out with a built-in response to where to hide a body, it was incapable of pointing a user to an abortion clinic.
How did an Artificial Assistant get iffy about abortion?
And how can I stop my own biases from seeping into the products and services I create?

In this talk, I explore our preconceptions about the nature of algorithms, and how users and makers influence them to model the world according to their perspective.
I conclude the talk by proposing a change of mindset for product designers, to be more aware of our capacity to let our biases infuse our services, and to put in place tools to create more inclusive experiences for our clients.

Anna-Livia Gomart

Summary

Anna-Livia Gomart argues that software is never fully neutral: algorithms encode their creators’ assumptions about what is normal, while user behaviour and training data can introduce further bias. She illustrates this with Siri’s original failure to provide abortion-clinic information, validation rules that reject unusual names, rating systems that amplify negativity bias, and hiring systems trained on historically biased recruitment data. She urges developers to assume bias exists, look for evidence in abandonment and support requests, involve people with different backgrounds, and use checklists to test cases such as low battery, no connectivity, disability, name changes, different languages, and varying levels of English. Bias should be identified proactively rather than waiting for excluded users to report painful experiences.

Key takeaways

  • Algorithms model human views of the world and therefore reflect human assumptions rather than being neutral.
  • Bias can come from program rules, users’ behaviour, or training data, and machine learning can scale existing discrimination.
  • Rating systems may amplify users’ negativity bias and impose disproportionate costs on workers and small businesses.
  • Developers should look for exclusion in support requests, abandonment rates, and other user feedback before problems spread.
  • Teams with varied backgrounds and practical checklists can expose assumptions about names, connectivity, disability, language, and other user circumstances.

Summarised automatically from the transcript.

Chapters

  1. 0:07 Bugs and Biases The talk distinguishes unintended software bugs from systematic biases and introduces the question of whether algorithms can be neutral.
  2. 2:48 Algorithms as Worldviews Algorithms are presented as human models of the world, with examples from games, Facebook, and procedural rhetoric showing how rules encode values.
  3. 6:41 Normalcy and User Diversity The speaker examines how assumptions about names, families, infrastructure, and other forms of “normal” can exclude users.
  4. 8:16 Negativity Bias in Ratings Rating platforms can amplify users’ negativity bias, giving a single bad review disproportionate consequences for workers and small businesses.
  5. 10:36 Biased Machine-Learning Data Training systems on historical hiring decisions can reproduce and scale the biases already present in the underlying data.
  6. 12:12 Detecting Bias Early The speaker recommends looking for signs of alienation in abandonment rates, support requests, and user reviews before problems spread.
  7. 13:45 Diverse Perspectives Random forests provide a metaphor for teams with varied backgrounds and experiences, which can produce better decisions than a uniform group.
  8. 15:17 Bias Checklists A practical checklist helps developers question assumptions about edge cases such as dead batteries, missing service, and language ability.
  9. 16:49 Building Inclusive Technology The talk concludes by locating bias in algorithms, users, and data and urging developers to address it proactively.
  10. 18:10 Questions The speaker discusses biased client data, ways to reveal hidden assumptions, and the possibility of shared counter-bias datasets.

Transcript

3,907 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:07

All

0:07

Speaker 1: right, folks, let's get the party started again. Next up is uh next up is Anna Livia Gomart. She's going to talk about s about bugs or well things that are not actually bugs but biases. Thank you.

0:25

Speaker 2: Thank you. Thank you. Um hello everyone. I'm really happy to be here today to talk to you about a subject that has been on my mind for a few years. And one of the things that um so has been on my mind is this difference between when you're using a product and you're having an awful experience Is it a bug? Like is it something that was unintentional or is it a bias, meaning that the people who created the the application you're using uh actually having this systematic way of thinking of things that are is slightly wrong and just ends up being very alienating for users.

1:14

Speaker 2: So um the example that I think uh summarizes really well is back in 2011. So um at that time we were all listening to Adele 's rolling in the deep, just so you kind of get back into that zone. Um uh Apple introduced Siri and uh how many of you have ever used Siri? So quite a few. Um and I don't know if you remember, but at first one of the things that people used to do a lot is kind of ask Siri really silly questions like uh where can I hide the body? Because You know, we all live that kind of criminal life. Um and actually it had answers which was really, really cool. Um and so people started asking like uh questions about like the meaning of life.

2:01

Speaker 2: Of course it's 42. Um but uh at some point someone asks, well, where can I get an abortion? And the answer of Siri is what And then you can start asking yourself, like it knows where to hide a body, but it doesn't know where to find an abortion cling in a country where it is legal to get an abortion. Um and that's the question, is Siri biased? Or is it a bug? And the answer from uh Apple before they fixed it was it's a bug. And so I'm kind of wondering if algorithms are neutral

2:48

Speaker 2: or can algorithms be biased? So what we're gonna do, we're gonna first kinda look into uh the biases that can how can biases be in algorithms And then we're going to see how biases can actually be in programs, but actually in the user or in the data used to train the program. And finally, One of the big questions I ask myself is you can't really analyze your own biases. It's really hard. So how can you actually prevent your users from being alienated by your own biases that you put in the programs you create? So first of all, are algorithms neutral? Well, the first qu I'm trying to do this. Okay.

3:34

Speaker 2: So what is an algorithm but a model of the way we as human, as individual, make decisions And basically it's all it is. It is our model of the world that we translate in a procedure So if you go from the if you if you start with our model of the world, the way we see the world, our reality is biased because we have unique perspectives and also we're humans, um then you kinda realize that if if the algorithms have our model of the world, then the algorithm become bias. And Dominique Cardon, who is a French sociologist, he wrote something that really um

4:19

Speaker 2: uh resonated with me, he said as soon as we opened the black box of algorithms, we realized that the choices they make for us are questionable and should be discussed because they offer different visions of society. So it's not only that they are biased, is that it is our role as people who know how to read those algorithms to critique them and to have this distance To say, okay, this is something that has been created by human beings, and therefore it is imperfect and biased, and we have to find ways to talk about that. So one of the um interesting frameworks to took talk about that I found is called procedural rhetoric, which sounds very fancy. Um but it's basically it's a concept from the video game industry

5:07

Speaker 2: And in the video game industry there is that um author, he's called Jan Bogost, and he analyzes games And he has a a book called Procedural Um no Persuasive Games. And in his book he says that the rules that create the video games, like all the algorithms behind it, there is actually a rhetoric. in it. So for example, if you take a game like GTA where uh as a gangster you can never and in a poor neighborhood, you can never find fresh vegetables and the only food available to you is fast food. There is like a critique of society behind it. And so as game creators we create w you can create rules of the world, like you make a model of the world.

5:55

Speaker 2: Like I don't know, have you ever played Civilization? It's a game where you like you create an empire. Well if you play civilization, you you're gonna do some things and then you're gonna get better and then you're gonna try some things and your empire is gonna kind of fall or non-ma not um evolve as well. So little by little you're gonna learn the rules of civilization and you're gonna learn the way the game wants you to think And then low by low those rules, you will be uh imprinted by those rules. Well, think about how Facebook redefined friendship. Like Facebook said, you know, this is a friend, and this is not a friend. So I think that we need to have the same uh critical

6:41

Speaker 2: point of view with services that we have with games And actually games is a very good example because they are creating those kind of analytic tools. So how how what does it look like a biased algorithms? Well, first let's think about normalcy. So we all have an idea of what is normal. We all have a model of the world where, for example, um everyone and I'm putting big air quote, everyone has a last name. Or everyone has Parents or everyone lives in a European city where we have 4G everywhere.

7:28

Speaker 2: And the problem with this normalcy Is that it might be not it not might not be normal for everyone. And we need to keep that in mind, that our users are way more diverse than we are in any team. And the I don't know if uh if you ever set saw that, there's a list called uh Things Developers Believe. And it's a GitHub repo and you can find it uh with all the things that developers get wrong. Uh and especially they think that no one has the last name null. Like um So we we we can see that we can put a lot of bad mojo in our algorithms, but we can also create algorithms

8:16

Speaker 2: that prey well not pray that will uh be um blind to our users' own biases. And one of the biases I want to talk about is the negativity bias. And that means that when you read a bad review of something It's gonna affect you more than reading a good review of something. So let's say that you're looking for a restaurant tonight and you're gonna look at the list and then you're gonna have five people saying This is a good restaurant, I had a great time. And then one person will say, This is the most awful experience I've had in my life. And always the bad reviews seems to be very dramatic. So um and and this one review compared to the other five that were good will actually impact you more

9:02

Speaker 2: because we have this negativity bias. So If you think about all the services that use this kind of rating, we're used to rate everything. And not only everything, but now with services such as Uber, we rate kind of people. re-rate a driver, re-rate like this personal inter uh this personal um contact we had with a a person that provided a service And we we start making this, you know, those stars and depending on our mood, we're gonna get, you know, like really um really judgmental or less judgmental. But at the end of the day, especially like for people who are in precarious situation or small businesses, one bad review actually has a huge consequence

9:51

Speaker 2: and has very little consequence on the person actually giving the judgment. And I think that by and there are some um platform who are trying to uh uh mitigate that bias. Um I don't if you st seen that but on Yelp or in um Google Maps now you have uh those local guides or like people who are who give a lot of reviews. So you know it's not just one person who just you know woke up one day and wanted to trash someone. So this is uh one example of how our platform, even though the algorithm inside it is not biased. it will kind of play on the biases of the users. And maybe we have a responsibility to mitigate this

10:36

Speaker 2: because of the alienating impact it can have Finally, let's talk about machine learning because that's something that is getting more and more traction every day. And it's something that is always based on data that you need to use to train your machine learning. And the problem with the data we use, like the data we actually have, is that it's actually there some of it is very biased. Typically, if you want to train an algorithm to do hiring for you. So you you have this company and you say, I don't want humans to hire anymore because they're biased, so I'm going to train a computer to do it. And then you're gonna give it all the records you have of the people you actually hired.

11:24

Speaker 2: But the problem is because you had a bias hiring those people. You're gonna teach it to higher in a very biased way, but on a larger scale. So it doesn't really deal with the problem that we have biased data and the way we get that data is biased So we need to be very careful about um the way we train the algorithms that not only would work in one company But let's say that this company that does the hiring, it gets really popular. Let's say it goes all over the world. Then we have this biased algorithms making decisions for a lot of people. So it scales up. So now that we've had this very grim moment together, what can we do about it?

12:12

Speaker 2: And that's something I've um I'm struggling with because as a software developer I don't want to look at my own code and saying and and seeing biases. Like I want to think that my code is just very neutral and very welcoming and um sometimes it's not. Um so here is the some of the Tips that I we try to use. Some of them are to catch uh try to catch as soon as possible. Go from the um Start with thinking that you have biases in your algorithms. There are somewhere and now your job is to catch them. And If possible, catch it before you have a user's pulling their hair because their last name is like one letter long. And

12:57

Speaker 2: For some reason, at some point in the life of your project, someone said that a last name couldn't be less than two characters. Try to find those experiences. It can be a high abandon rate. It can be found maybe in like the support emails. It can be found maybe on review websites where people are gonna tell about uh talk about their experiences and so you can kinda catch them there and tell like try to as soon as possible um try to deal with their problems so that it doesn't happen to other people. I don't know if uh some of you have uh done a bit of um machine learning, but there is that expression that I found to be the most poetic thing. It's called a random forest.

13:45

Speaker 2: And a random forest is actually a lot of decision trees. And because they're a lot, it's a forest And I think that this idea of like a forest made of like data decision trees is kind of really poetic. But this the idea is that if you have one decision tree, if you train your machine learning to have this one way to take make decisions is going to try to overfit the data, the training data you give it. So if from my understanding what you do is that you create a lot of decision trees Like they make decisions um in a in different ways because you give them very um varying type of data as entry. And then you can have something that is way

14:30

Speaker 2: not trying to overfit the data and actually gives you better results. But the point behind this all is that if your team is a random forest you will have better results than if your team has the same way to make decisions. So different backgrounds, different life experience is actually gonna give you this different perspective. It's not gonna give you all the perspective. But it's going to be better than have just this one way of making decisions. And finally, the thing that I'm trying to do is to check what you find normal. So have a checklist. Have a checklist of uh go to that repo and GitHub of things programmers believe and just take things one after the other and check if

15:17

Speaker 2: Um what happen like for example, have you um ever been taking a plane and you have the boarding pass on your phone and your battery is getting really low? And you get that fear of saying, what happens when I don't have battery anymore? And I I get scared. Like I get like, I don't know what I'm supposed to do once like I have no battery left And I'm thinking, the problem is I don't I I I get this fear because no one told me what to do if like everyone is like, oh yeah, just put it on your phone. It's gonna be great, but what should I do if I don't have batteries anymore? Well now if I make an app, I ask myself what happens if my user needs my app but doesn't have battery anymore? What happens if they don't have service because they're in the subway

16:03

Speaker 2: What happens if and so there's that list and that checklist. And one of the things that uh I realized recently in the project I'm working on I'm working on an open source project, so we have a lot of documentation to kind of be self-serving. And one of the things I realized is that I'm always trying to have this very good English. I'm always trying to to to sound like I really know what I'm talking about. And then I realize that my bias is to think that everyone speaks very good English. Like all developers all around the world, they have this It's beautiful English and the whole point of coming to my to this project is to judge my English level. It is not. So I realized that. But the idea now is

16:49

Speaker 2: can I make my English maybe simpler? or more accessible without oversimplifying what I'm trying to say, but so that people come to the to to our repo and they're actually feel good about reading it because they get all the all the meaning that I'm trying to convey So uh first the algorithms we create as software developers, they're biased. Let's start with that because they just They represent the way we have to represent the world and not one piece of I don't think one piece of software yet can really encompass the whole human experience And that that um that bias it can be in the algorithm, it can be in the users that are going to use your algorithms, or it can be in the data that you use in

17:40

Speaker 2: to make decisions. And so before our users are alienated in some way or another, let's try to be very proactive and let's try to find those biases to make tech more inclusive. Thank you very much. Do we have time for questions?

18:10

Speaker 1: We absolutely have time for questions.

18:12

Speaker 2: Wonderful.

18:13

Speaker 1: You wanna uh

18:14

Speaker 2: if someone has

18:16

Speaker 1: it?

18:19

Speaker 3: Is this already on? Yes it is. Okay. Hi, thank you so much for this talk. This was really useful. Um for thank you. Sorry. I should have done that myself. I'm really short. Um but for Sometimes you you might kind of work on projects where you have a a client or a need um who who feels like they need to represent their data in a way that does not actually represent the full diversity of the human experience. What what are strategies that you have to sort of get clients on board with representing their data in a way that is more truthful?

18:55

Speaker 2: Um that's a very uh I think it's a very good question. I think software development is one of the many areas where this question uh happens. My first question would be Do they have other data? Because if the biased data is all that they have, it's gonna be hard to represent it differently. So how do they gather that data and i are they making a choice in the data to kind of pick and choose? Or just do they might not have like unbiased data. So that's the first the first kind of thing I would uh I would try to to understand. And if they are pick and choosing, um there is actually uh a talk I think to have about how um it can be perceived or how it can be alienating for some people. I've re I've realized in my career

19:41

Speaker 2: that often it's not that there are like some very biased people who hold their bias truth to be evident. Most of the time it's just people who don't think about like what happens if there's a disabled person trying to come in the building. Oh what happens if, you know, there is a same sex couple and uh one of the uh person wants to change their name? Uh I I've had this this case where uh the um a form um where uh you as a man you couldn't change your last name because it just it wasn't used to be done like only women would change their names so there's the option just didn't exist. And just saying like what happens if uh is a a good way to kind of open that Uh but if uh

20:26

Speaker 2: their um their bias is something that is more like something they believe or so like I I have no idea how to deal with that.

20:36

Speaker 1: Thank you. More questions?

20:42

Speaker 4: I'm I'm quite good at spotting biases in other people's work. And um uh sometimes I don't always even react badly when they point it out in mine, but I I know that uh uh understanding that there is a bias in one's work is quite difficult and also difficult to persuade someone else. Not that they should do a certain thing but that there is a problem. What strategies I mean you just mentioned one right now about saying what if. What other strategies will bring this out so people can come to it themselves without having to be taken there by you?

21:20

Speaker 2: Um it's a it's a very good question, I think, because uh one of the strategies I've sa I've seen and what I'm trying to do with my uh some of the tips and some of the strategies I I'm trying to think about. The the problem is most people realize it once they're faced with someone like a client or a a a friend or a user that is actually i took the time to come up to them and tell them their story. Which means you have to wait for what, ten people or more to have a really bad experience for maybe one of them to take the courage and take the time and take the energy to come to you and explain their you know their issue. And most When I I talk about uh

22:05

Speaker 2: biases, uh often this is the answer like, oh, uh I was doing a training and this person came up to me and so I had this epiphany. But I'm trying to uh my question now, and that's I mean I haven't found one answer, but my question now is how can we do that proactively Like without actually going through the whole thing where you have to alienate people to someday have your own realization. And that's where I found it hard also. And Especially I find it hard on my own work. Like I don't want to believe that my work is biased. I want to believe that I'm special. But uh I'm not, and so I'm I'm working hard on it Uh and I hope that we all do because I think that's how we're gonna make a tech a better place to be.

22:58

Speaker 5: Hi, thank you for the talk. Um do you know if there is somewhere uh like uh a database or repository or something of uh counter-biased data? like a list of last names that are usual problems, a list of addresses that are usual problems and things like that.

23:17

Speaker 2: I haven't found it, but you're not the first person to ask me about it, so I think someone here If they have some times should uh s maybe me should actually make it because I think having that training set like should would be a very a great way to start. If someone has it, please put in the Slack because I think we would all uh use it really uh have a a lot of uh of of fun fun. Uh it would be very interesting to use it. But yes, I think there is a need for that Uh definitely. But again, uh any list would not encompass maybe some of the uh issues. Um I I'm thinking for example you can have a list of names but um things like uh You know the languages where you write from uh left to right.

24:04

Speaker 2: And that's kinda hard to have a list of you know like so I'm thinking it's a first step, uh but there's a whole also it and there is also other things to put in place Thank you for your question.

24:19

Speaker 1: Thank you. Are there more questions?

24:22

Speaker 2: Thank you very much.

24:24

Speaker 1: Thank you.

Questions this talk answers

Are algorithms neutral, or can they be biased?

Algorithms model the way humans understand and make decisions, so they inherit the limits and biases of that worldview. They should be treated as human-created, imperfect visions of society rather than neutral systems.

Discussed at 3:34

How do user biases affect rating and review platforms?

People tend to give negative reviews more weight than positive ones, so a single bad review can disproportionately harm a worker or small business. Platforms can partly mitigate this by showing context, such as whether a reviewer is a frequent contributor.

Discussed at 8:16

How can a machine-learning hiring algorithm reproduce discrimination?

If it is trained on a company’s past hiring records, it learns the biases that shaped those decisions and can apply them at a much larger scale. Using more data does not fix the problem when the underlying data is already biased.

Discussed at 10:36

How can developers find and prevent bias in their software?

Start by assuming bias exists, then look for abandonment rates, support requests, and reviews that reveal alienating experiences. Use checklists of supposedly “normal” assumptions, test situations such as no battery or no network access, and address problems before more users encounter them.

Discussed at 12:12

Why do diverse software teams help reduce bias?

People with different backgrounds and life experiences make decisions in different ways, much like the varied decision trees in a random forest. A team with diverse perspectives will not catch every issue, but it is less likely to rely on one narrow model of the world.

Discussed at 14:30

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Anna-Livia Gomart

More videos from DjangoCon Europe