WebRTC with Django, Channels, HTMX, and coturn with Ken Whitesell

This video features Ken Whitesell at DjangoCon US 2024 in Durham, North Carolina, USA.

WebRTC with Django, Channels, HTMX, and coturn with Ken Whitesell
0:40:52
Published December 6, 2024
572 views

Audio/Video conferencing has become standard in many areas for augmenting communications among individuals. Modern browsers facilitate this by including support for Web Real Time Communications (WebRTC).

WebRTC itself is a point-to-point protocol, which means that two browsers using this for a video call are talking directly to each other. But, before that can happen, the those browsers need to know that each other exists and are looking to establish this connection. Then they need to negotiate the parameters for the connection.

Then there are many network-related issues that can affect the ability for those two browsers to connect. Things like firewalls and Network Address Translation (NAT) can affect how each side "sees" the other side, further complicating the situation.

All these issues have known solutions. The WebRTC APIs have matured to the point where they can be considered reasonably stable and reliable. It has become practical to incorporate these solutions in a Django-based website.

This session will discuss one implementation of a Django-based website that facilitates a group video conferencing system, using Channels as the signalling mechanism, HTMX for page-content management, and coturn as the NAT transveral and and gateway server.

This talk was presented at: https://2024.djangocon.us/talks/webrtc-with-django-channels-htmx-and-coturn/

LINKS:
Follow Ken Whitesell 👇
On X: https://x.com/KenWhitesell

Follow DjangoCon US 👇
https://fosstodon.org/@djangocon
https://x.com/djangocon

Follow DEFNA 👇
https://www.defna.org/

Video production by Confreaks
Follow Confreaks 👇
https://confreaks.com
https://x.com/confreaks

Summary

WebRTC video conferencing with Django uses Channels for WebSocket signaling, HTMX to update the page, browser WebRTC APIs for peer-to-peer media, and coturn to handle connections that cannot be made directly. Each participant maintains a separate connection to every other participant, so joining a room requires creating video placeholders and exchanging channel names, ICE candidates, session descriptions, and other signaling messages. STUN helps discover publicly reachable paths, while TURN relays media through a server when NATs or firewalls block direct connections, though this can become expensive and a bottleneck. The speaker also explains how JSON multiplexing and a custom HTMX response extension route different events to page components, how the perfect negotiation pattern resolves simultaneous connection attempts, and why capacity, security, cleanup, scaling, and browser diagnostics matter in production.

Key takeaways

  • Django serves the application, Channels carries WebSocket signaling, HTMX updates the interface, and WebRTC establishes the media connections.
  • A room with many participants creates a separate peer-to-peer connection for each pair, which increases browser, network, and signaling load quickly.
  • STUN discovers possible network paths, while coturn’s TURN relay provides a fallback when firewalls, NAT, proxies, or VPNs prevent direct communication.
  • WebRTC signaling must exchange channel names, connection events, ICE information, and SDP session descriptions through the Channels server.
  • The perfect negotiation pattern assigns deterministic polite and impolite roles so simultaneous offers do not leave peers in conflicting states.
  • A production deployment needs TURN authentication and security, sufficient Channels capacity, room cleanup, scaling across TURN servers, and browser WebRTC diagnostics.

Summarised automatically from the transcript.

Transcript

5,298 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:20

Speaker 1: Good afternoon. Whoa. All right. Hi, I'm Ken Weitzel and it's a real pleasure for me to be with you here this afternoon. I hold the title of Manager Software Development at the Data Science and Advanced Technology Group for WSP USA, based in our Baltimore office. I have the privilege every day of getting to work with some of the most incredible software developers, system admins, data scientists, GIS engineers. and management that I've ever had the pleasure of working with across 45 years. My work there leads me to explore some unusual or out-of-the-ordinary technologies

1:05

Speaker 1: Sidebar discussions, you know, you get into the idea of what if. And this talk is actually a result of one of those sidebar type discussions. We're going to do something a little different here. I have a live demo set up and those of you who are watching remotely can also try this. Caveat, I've never tried it with a group anywhere near this size. This may make my server melt. This may make your laptop melt. I don't know what's gonna happen. We'll find out. So go ahead and give it a try. There's no real security. Basically you put in your full name. Your screen name actually serves as the password. I kind of short-circuited the regular Django

1:51

Speaker 1: security. And that's what we've got. Also, uh to be clear, uh yes, audio is disabled. The last thing I want to do is create feedback among a hundred laptops. Doing it among ten is bad enough. 100 would uh Be real interesting. While people are giving this a shot and seeing what's going to happen, just a little bit of history regarding WebRTC It was first released, it was introduced as a draft standard prototype back in 2011, officially standardized in 2018. In that range of years in between, there was a lot of activity, a lot of prototyping. The APIs were not stable. Finally adopted as a standard in 2021.

2:38

Speaker 1: But what has happened is that there are a lot of blog posts. You go out, you go looking for information to try and figure out how do I do this. And you're going to find a lot of blog posts written in the 2015, 16, 17, 18 timeframe that is obsolete or inaccurate. It is really tough to find good solid information as to how to make this work today I have one recommendation. If you are seriously interested in pursuing WebRTC as a technology By this book. Disclosure: I have no business relationship. I do not know the author. I do not

3:23

Speaker 1: am not in any way affiliated with the pragmatic programmer. I just know that I would not be standing here today giving this talk if this book had not been published. And it's brand new. Metaphorically, the ink's still wet. So that 's a solid source of information. Also, for those of you who may wish to follow along with the code, because obviously I can't dive too deeply into the code while I'm standing up here talking. The Git repo has my code on it and also the slides. There's a PDF of these slides as of Last week. So there are going to be some differences between what those slides are and what you'll see here.

4:11

Speaker 1: No big deal First, I'd like to start out with, well, when you're talking about doing WebRTC, what's really involved? And so I took a picture that some of you who may have seen my tutorials in the past are well familiar with This is just a very basic architectural diagram of what's happening in a Django channel's WebRTC environment. Everything on the right hand side of the diagram is your typical ordinary. You've deployed Django. This is what you're looking at. Nginx sitting in front of Daphne and or you Whiskey. serving your Django project using channels. And oh by the way, a shout out to the Valky

4:57

Speaker 1: people after a conversation with them on Monday. This is my demo installations running on Valky as the storage element. But what I'd really like to bring out about this diagram is The left-hand portion showing that blue box with the WebRTC components. The key Item to take from this part of the diagram is that when you are actually doing video conferencing, those of you that have actually gotten it to work, if it server hasn't crashed already, that are seeing each other. You're establishing a point-to-point connection with each person that you are communicating with. So if you've got ten

5:43

Speaker 1: of the little boxes on your screen showing your connection, with 10 people, that's 10 separate network connections. You have one connection for each of those PCs computers that you are communicating with. So what really are the fundamental parts here? First of all, Django. Hopefully everybody here knows what Django is, what Django does. That's why we're here. serves the web pages and the JavaScript provides the basic authentication. Channels is used for the WebSocket communication signaling. I will get a lot more into detail into that in just a couple of minutes. HTMX. What can I say about HTMX other than what's already been said and that it's fantastic?

6:33

Speaker 1: This is one case where I'm sorry to our wonderful fellow Natalia. We I have a disagreement with your earlier talk. She said there are no evil entities within the Django community. I'm sorry, as far as I'm concerned, Sauron created JavaScript. Okay. I don't do JavaScript unless I absolutely have to. HTMX, thank you. And then finally, the last part of this is COTER. Which is a separate package completely. It's a pre-package you can, you know, apt get or yum install for whatever repo. And you just configure it and it handles some of the networking issues that may not be

7:20

Speaker 1: apparent or obvious that need to be taken care of. For allowing the two computers that are involved in this communications channel to actually intercommunicate. Again, I will be getting into that in a bit more detail a little bit farther So those of you who have gotten on the demo, you see a homepage, standard Django homepage, rather typical template. Those of you looking at the code can see. Nothing special about that. It loads the HTMX, it loads the HTMX WebSocket extension, and then it loads my extension to their extension Which is a hook that they provide within HTMX

8:06

Speaker 1: called the Transform Response Extension. Now, HTMX, as as I perceive it, as I see it, is really fundamentally designed around the idea that you're moving HTML from the server to the clients. My transform response plugin says no. In a channels situation when you have a persistent connection that you're moving data regularly. I'm going to send JSON. Everything that I send between the back end and the front end is JSON. That JSON may contain HTML that needs to be injected into the page, or it may have other data that needs to be sent as well.

8:52

Speaker 1: And I do this to allow me to do what I refer to as multiplexing that WebSocket. It gives me the ability to support more different types of data, more things that I may want to send. back and forth through the channel than just HTML. This comes from uh an architecture of another system that I've worked on. Where the idea is that the page has these pluggable components that can be user selected to be embedded into the page. And so you have your browser window, you have the menu bar, you may check, may select from that menu what part or what modules that you want to see on this particular page.

9:37

Speaker 1: Those modules have some JavaScript in them that register with the Transform Response plugin And so anytime any message comes across the wire for that particular module, it gets handed off to the JavaScript within that module to be handled. Multiple modules can plug into the same events, to the same data stream. So for example, if you have the type of display where you you where you are using Data tables, the JavaScript library data tables, and maybe Chart. js, you could have two of these panels That are hooked into the same data stream, where

10:24

Speaker 1: one of those modules is going to dynamically update your data tables, the other module is going to update your chat JS, all fed by one single data stream coming through the server. Just as a very brief example of what the JSON may look like, I'm showing the example on the screen. It's a key for the apps that are registered with the extension, and potentially some HTML that also comes along that gets handed off directly to HTMX to be injected into the page. Nothing fancy here. So what are these events dealing with RTC

11:10

Speaker 1: that have to come through these JSON responses? There are five events that this setup process to allow this communications to occur that need to be handled. First is the connect event. You get your channel name. Those of you that have worked with channels know that you have a channel name associated with your connection. And it's what it's like an 80-character string of mostly random characters. You get that back. You effectively from the browser side, this is data coming back to the browser from the server. So you from the browser, you don't normally see that channel name. It comes back to you because you're going to need it to send to other people.

11:55

Speaker 1: If I want to establish a connection with you. You need to know how to find me. And in the channels environment, you find me by my channel name. And so channels, the system gives me my own channel name that I can then send to other people. Other is a request to open a connection to one other peer. It's the it's in effect One person has sent their channel name to another person, and that other person is going to try to connect with you. Others is a group setting. There are a whole list of people. Think of a room like a big, like the demo session.

12:41

Speaker 1: where somebody joins a call, there might be 10 other people already on that call, and so you need to connect to this whole list of people. There are some benefit to getting that entire set at once. Signal is a Shorthand way of directly feeding data from one side of the connection to the other. And before they have that connection, before they have that one-to-one connection between the two and the data is being passed through channels. The signal is how you pass send a message from one browser to the other via their channel names. And the last one, the disconnect message, is when somebody has dropped off the call.

13:32

Speaker 1: So that you know from your side that person's gone and you can tear down that session. You don't need that anymore. The transform response plugin is again, it's really some simple JavaScript JavaScript code. Remember, I don't like JavaScript. It's all my JavaScript is simple JavaScript. The transform response gets the message in from HTMX. And it looks through this data structure that is created that has the listing of all the event names and it dispatches that event to the ones listening. The forward function. is then the function that actually performs that dispatch. There's a data structure that keeps track of all the event names, the apps that have registered them, and one

14:23

Speaker 1: or more listeners for that event. And so it just cycles through, it looks through its list of, hey, who do I have that's listening to this event, to this other's event, and can go through and dispatch it to however many functions are registered for that event. The actual act of registering an event With this plugin is simply apps. underscore add, the name of the app that it is registering for, the name of the event, and the function. You're passing it the name of the function that is being registered. So, what we have here, if we have a group call in progress, you've already got some number of people who have already connected.

15:14

Speaker 1: They're all talking to each other. you have these individual point-to-point connections for everybody that is currently on that call When you see that button for joining a call, joining the call within the JavaScript in your browser, it sends a single message through the WebSocket. To the server that says I'm going to join a call, and the call name that I have defined here for this demo is simply called video. Sends the message to channels to join it. Now what actually has to happen is quite a lot because you have this new person who has chosen to join a call.

16:00

Speaker 1: And so you have this situation where all of these other laptops, all these other computers have connections already established. This new person coming into the room. has to establish a connection with everybody in that room. And likewise, everybody who's already in that room needs to turn around and acknowledge and work with that connection. The other way with that new person joining in. So in the channel's consumer on the server side, when a person is joining a room The first thing that channels does is it looks at the room. Who's all in there? It creates a tiny HTML div, and you're going to see that in just another slide or two.

16:49

Speaker 1: And it sends that div of all the other occupants of this person joining the room. The diagram showed five people in the room, six person coming in. So, channels is going to create a div for each of those other five people and send that as one set of divs through the WebSocket back to the person joining the room. It then creates a div for the person joining a room and sends that div to all the other people who are already in there so that both sides now have an HTML element, a div. That is reserved or allocated for that video connection.

17:34

Speaker 1: Channels then sends out another signal to everybody else that is in that room to say somebody new has joined this room and you need to start working on establishing a connection with them. And it sends the connection back to the person joining the room to say, yes, you have connected to the room, and this is information that you may need to forward. when you're trying to work on the connectivity from your side. The divs being created are very simple This is in fact, I believe, the entire div that gets sent out. The key point is the

18:20

Speaker 1: video tag there, which allows the browser to bind a video stream or audio stream as well to that particular element. And the rest of it is just Bookkeeping, really, a div to be able to identify it. Uh the area using the HTMX to add that tag, to add that, add that to that div, that ID, and not replace what's already in there. That's pretty much it. So now we've got These divs, everybody has a spot, a placeholder on their screen

19:06

Speaker 1: for everybody else in the call. Then the real work of what WebRTC comes into play. Both sides want to communicate with each other, but the real world is a messy place. It is never that easy. The issue is that with network connectivity, you are frequently, often, sometimes, operating behind something like a firewall or a NAT box. or a proxy or a VPN that figuring out how to get traffic between two endpoints is never simple.

19:52

Speaker 1: The simplified view, this is part one, and I'm not expecting anybody to actually read this. Okay, this is just, I threw this up there for its visual effect. This is the first 34 steps in the negotiation process that is listed as part one. This is part two, the other thirty-eight steps. And for those of you who are really interested in the guts of what's going on here, the URL is in the title part of the slide, where you can find this diagram and all the accompanying text that goes along with it

20:39

Speaker 1: And if you are so inclined, uh have at it. Enjoy it So what is actually being negotiated here? What is it that needs to be exchanged? The first part is as I started to touch on earlier, is the path as to how this network traffic is going to take From the two endpoints. The protocol is called ICE Interactive Connectivity Establishment The MDN website defines this as a framework for facilitating the connection of two peers, regardless of the network topology. And it is a multi-stage, multi-phase

21:25

Speaker 1: negotiation process that searches for the fastest path between the two computers. And what this means is that based upon who you are communicating with simultaneously, because you have that independent connection between every two computers. You could have a direct UD UDP path between the two systems, or it may have to fall back to TCP. And if necessary, through HTTP or HTTPS, in that order, and because we're running this in a browser, pretty much anybody who's running a browser can expect to be able to make an HTTP request. Doesn't make a whole lot of sense to be running a browser otherwise.

22:12

Speaker 1: Or Indirect through the turn server. Going back to the earlier slide, that's where COTERN is going to come in. And we'll get into more details there And this is where COTERN fits in. Open source package, there's the website for it. It supports two key protocols. Stun and turn. Stun session traversal utilities for NAT and turn traversal using relays under NAT. A way of working around NAT proxies, VPNs, etc. Stun Special Tession Traversal Utilities for NAT

22:58

Speaker 1: This is effectively what's my IP for WebRTC. If you're familiar with the website What's my IP? It's a website that you can go to and it tells you what your public facing address is. Because if you're behind a NAT at home, like for instance, in here, I believe pretty much everybody in here has a 172. 20 address, if I remember correctly That address is private, that's local to within Marriott. You're going out through some type of network proxy to get out to the internet. Somebody else in the hotel Who is on this same subnet domain would be able to connect with you directly

23:45

Speaker 1: because you and I being on the same subnet here Those addresses, those 172 addresses, are going to resolve properly. But for somebody who may be watching us at home, they can't. 172. 20 is meaningless to them. Stun lets you identify, actually, multiple ways of identifying how you can be seen Because it can also handle things like IP4 versus IP6 and other ways of routing traffic. The actual protocol for STUN is a very low traffic utilization. It's a very quick and simple exchange.

24:31

Speaker 1: It's really very little more than just a ping that gets a response that includes what's the address that I saw when you pinged me? Effectively what it is. There are actually public stun servers available. Google runs two You can go out onto the internet and you can do a search for public stun servers and you'll find multiple pages Each containing lists of public stunned servers. I think I did a brief survey and just I stopped counting at 50 public servers that were available. You can configure multiple stun servers to use, but

25:18

Speaker 1: it's not free. It's not like DNS where you can configure five, six DNS servers and you get a response and that's the one you use. Every time that you connect to a stun server that gives you an IP address, that establishes the possibility of a different path for that communications to occur Between you and the other people involved in the call. And so it greatly lengthens, it increases the traffic of the negotiation process and increases the latency and the time that it takes to actually settle down and resolve these communication paths because every single one gets evaluated to find out which one is going to generate the fastest path

26:12

Speaker 1: The other part of this is Turn, traversal using relays around NAT. This is different. This is an actual data distribution hub. All your video data, all your audio data, all your data data that has to go through Turn will get channeled through this one server. So if you've got 10 people that are all connecting to each other and all of those interconnections are needing to use turn. then what you have are 90 connections whose audio video streams are all being channeled directed through one

26:58

Speaker 1: server in the basic case. You can multiplex it, there are ways of scaling it up, but it can become a real bottleneck. You do need to implement, if you're implementing Turn, you do need to implement the security layer. Unless you have truly an unlimited budget for data. Because if you do the math, if you've got 90 audio video streams going through one particular server and you're paying per gigabyte of data transferred ka ching, ka-ching, ka-ching, and not to your benefit. So

27:44

Speaker 1: as a result of this, what you won't find Are any public servers, public turn servers available? Or if anybody puts one up, they generally come down when the first AWS bill comes in. Or when Comcast says, uh, no, you are limited to three terabytes a month and you pass that on day two. So how are they defined? It's a little block of JavaScript that is passed to the RTC APIs within JavaScript. That identifies what your stun and turn servers are. This is the actual set of definitions within the demo that you are using.

28:29

Speaker 1: I've got stun running on my own server uh along with an reference to the one of the Google boxes and the turn server is my server. As a result, those of you who may be playing with this now, this server, believe me, is getting turned down after this talk is over. But the negotiation path, the path for that network connectivity, is only half of the picture Because there's a whole nother part of this that needs to be addressed, and that is what is the actual data that you're sending back and forth.

29:15

Speaker 1: There is no such thing as one format of a video stream. You have the resolution of the data that you that needs to be negotiated. Who can handle what resolution? What codecs are involved Is there any encryption involved? This is part of the data, that AV data that needs to be negotiated. SDP, the session description protocol, is technically it's not a protocol, it's a data format. It's a whole bunch of lines that all Pretty much look like the two sample lines that I gave here. It's a single letter, an equal sign, and stuff. I have no idea what that stuff is, nor do I really care I don't need to know it. That's all handled by

30:02

Speaker 1: the JavaScript RTC APIs. But each negotiation grouping can be a hundred lines or more. You don't need to know what they are. But every grouping of lines has to go through the server. It goes back through your channel server so that it can be distributed out to the people who are trying to be connected with. These all become messages through channels. Those of you who have worked with channels may be aware that channels by default has a 100 message limit on basically the message queue. I'll say. Andrew's probably cringing when I use that phrase.

30:49

Speaker 1: But point is you will need to increase your channel capacity. If you think about a room such as this , Even even 10 people, 10 people trying to each make nine connections, that's 90 connections. If you've got a hundred lines. You're talking in the neighborhood of 10,000 messages that are all being passed through channels in very short order. And when I was talking about my server melting, It's this particular load that I was thinking of and hoping I had things scaled up reasonably well. There's another aspect of this exchange, and the last part of this conversation.

31:35

Speaker 1: What happens when both sides are trying to connect to each other at the same time? Let's say we want to meet. We want to go out for dinner, have drinks, whatever. One person says, let's meet in the lobby. Other person says, okay, what time? How about 6 p. m. ? Great, I'll see you in the lobby at 6 p. m. Good, I'll see you there That's a that's a complete negotiation. That is a conversation that has taken place. All the information has been passed back and forth. and acknowledged that that information has been received. The very last line, good I'll see you there, is the acknowledgement that everything is set. Everybody is in agreement. But you may have been in the situation

32:22

Speaker 1: where both people start talking at the same time. Let's meet at the bar, let's meet in the lobby. Great, I'll meet you in the lobby at 7 p. m. Good. I'll find you in the bar at 6. Cool, I'll be in the bar at 6. Great. I'll see you in the lobby at 7. Both people have heard and understood what the other person has said. No data has been lost in this process. But there's been no agreement on what's going on. Both people are going to walk away from that wondering, what am I supposed to do here? And so when you have these cued messages flying back and forth, this is very much a real possibility.

33:11

Speaker 1: What should happen, they both say, let's meet in the bar, let's meet in the lobby. One person takes a pause. Second person says, I'll find you in the bar at 6 p. m. Other person says, cool, I'll be in the bar at 6. Great. See you then. You're back to a point where there is that agreement, that consistency of message between the two. Both sides know what's going on. There is a protocol that has been developed. They call it the perfect negotiation protocol. The full definition is in that URL at the top of the page. There are two roles that are defined between the peers: polite and

33:56

Speaker 1: impolite. The polite person is the one who is going to pause. The impolite person is the one that is going to assume that they are the one driving the connection process. This assignment between polite and impolite is purely arbitrary. It doesn't matter. pejorative intended one way or the other. It's just a way of desynchronizing these requests such that this agreement can take place. But the assignment of these roles must be deterministic. Everybody must understand what the rule is.

34:45

Speaker 1: So that the two people who are trying to communicate both know this rule. One person knows By definition, that oh, there's a conflict here. I need to be polite, I need to back off and let the other person drive the conversation. In the case of this demo, the new caller, the person joining the call, is the person who is impolite. The people in the room, that group of people in the room, are the ones that are going to be polite when a new person joins. So that sixth person joins the call They are

35:30

Speaker 1: driving effect in effect that negotiation process. If there is a conflict between The people already there trying to start this connection and the new person coming in. It's the new person coming in that has the priority of straightening this out and making it work. And it's a simple fact of logistics. This new person's coming in has got n number of people that they're trying to manage connections with, whereas the people who are already there only have one new person. And so it's a it's a way of allowing that new person coming in to effectively take control of the sequence of events.

36:19

Speaker 1: Once the negotiation is complete, the media streams are assigned to the video tag and You have pictures, live pictures of a camera coming through. You can also use that WebRTC connection to share data other than audio video The book and other resources do talk about data channels that reside along that. So you can do things like direct point-to-point file sharing or like a like a shared whiteboard, something like that The book describes a system where you're using the uh using it for chat And so that wraps it up. At this point you have a live connection

37:06

Speaker 1: and hopefully everything's working here, but I've just touched the surface. It has taken me Three years to get to this point, what it is. A lot of things that I haven't covered. The audio. They're really it's just another data stream, very simple data channels The ability to have multiple rooms in the effect that that has. Implementing COTERn security so that it's actually secure Cleaning up the rooms when people add and drop off of calls, running multiple-turn servers so you don't make a server melt with a larger call. Diagnostics and troubleshooting tools. Because this is browser-based fundamentally, both Firefox and Chrome

37:51

Speaker 1: have very good diagnostic information available to you directly within the browser. That is Chrome. Uh slash slash WebRTC internals shows you a lot of data. Firefox Developer Tools has something very similar. For any information and interest in actually deploying this and standing this up yourself, there are some scattered notes, stream of consciousness thoughts in the README file in the repository. Thank you all. Those of you who know me, that probably the best place to find me is likely going to be at forum. jango. project. com. I'm typically hanging around in there at some point in time or another. And I am here

38:37

Speaker 1: tomorrow and Friday morning at the next break. I'm interested in seeing Elizabeth's Postgres talk. here next. So I'm gonna be in here for the next talk. But anybody that wants to talk to me about this, I'm more than happy to talk to you at the next break this evening. I'm gonna be at the sprints tomorrow and Friday morning

39:07

Speaker 2: So you do have time to take at least one question.

39:11

Speaker 1: Okay.

39:12

Speaker 2: Is there anybody

39:17

Speaker 3: great talk, Ken? Thank you so much. Just one quick question. In the demo, I know it's you the idea of trying to capture somebody falling off the the like losing internet connectivity is something that's way out of scope for your initial demonstration. But from a messaging standpoint, how does that flow? Does it like does the server do like a kick message or how does that work?

39:36

Speaker 1: If it's if it's an intentional disconnect. Okay. If it's unintentional, then At some point in time, your Daphne server is going to recognize that that WebSocket has gone away. And so that does send the disconnect message through to the channel's consumer which then notifies the other members of the room. So the consumer being disconnected sends the signal through the room to the other channels in that room.

40:12

Speaker 2: Anybody else? Awesome. Thank you for a good talk. Thank you And round of applause

Questions this talk answers

What components are used to build WebRTC video conferencing with Django?

Django serves the pages and handles basic authentication, Channels handles WebSocket signaling, HTMX updates the interface, and coturn helps establish connectivity through NATs, firewalls, proxies, and VPNs.

Discussed at 5:43

How does HTMX handle WebRTC signaling data?

The demo uses a custom HTMX transform-response extension to carry JSON over the persistent WebSocket. The JSON can contain HTML for insertion as well as other data, allowing multiple browser components to share and respond to one data stream.

Discussed at 8:06

How does a new participant join a WebRTC room with Django Channels?

Channels finds the existing occupants, sends the newcomer placeholders for each of their video connections, and sends the existing participants a placeholder for the newcomer. It then notifies both sides so they can begin negotiating the individual peer connections.

Discussed at 15:14

How does WebRTC find a network path between two computers?

ICE performs a multi-stage negotiation that evaluates possible paths, preferring direct UDP when possible and falling back to TCP, HTTP/HTTPS, or a TURN relay when network restrictions require it.

Discussed at 20:39

What is the difference between STUN and TURN in WebRTC?

STUN helps a browser discover the public address by which it can be reached, while TURN relays the actual audio, video, and other data through a server when a direct connection cannot be made.

Discussed at 22:12

What does WebRTC negotiate besides network connectivity?

It negotiates the media details, including resolution, supported codecs, and encryption. These capabilities are represented in SDP and are exchanged through the Channels server.

Discussed at 29:15

How does WebRTC prevent both peers from negotiating at the same time?

The perfect negotiation protocol assigns deterministic polite and impolite roles. The polite peer backs off during a collision, while the impolite peer continues driving the negotiation; in the demo, existing room members are polite and the new participant is impolite.

Discussed at 33:11

How does Django Channels notify a WebRTC room when someone loses their connection?

When Daphne detects that a participant’s WebSocket has disappeared, Channels sends a disconnect event to the consumer, which notifies the other members of the room so they can remove that participant’s connection.

Discussed at 39:36

Presenters

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos by Ken Whitesell

More videos from DjangoCon US