Scheming with CSRF: When platforms manage to break things with Katie McLaughlin
Published November 3, 2022
This video features Katie McLaughlin at DjangoCon Europe 2022 in Porto, Portugal.
Keynote: What should you have to worry about by Katie McLaughlin
As a Django developer, there are things you need to worry about. Keeping up to date with Django trends and best practices is one thing, but what about all the other parts of technology that are incessantly advertised towards you? What should you have to worry about?
Katie McLaughlin argues that Django developers should focus on three practical concerns: stable, scalable infrastructure; reliable code and dependency security; and healthy, inclusive communities. Developers should understand the infrastructure layer their applications run on without taking on unnecessary platform operations, design for stateless deployment and persistent external storage, and keep dependencies and operating systems updated while planning for regressions, malicious packages, and supply-chain failures. She also stresses that open-source work extends far beyond commits, so projects should identify, support, and properly credit maintainers, reviewers, documenters, bug reporters, community organizers, and other contributors.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Bom dia amigos. The 1. 18 release of the Go programming language came with it a number of backwards incompatible changes to the language which introduced a number of various issues in Kubernetes orchestration tooling. It caused the blocking of the release of Kubernetes 1. 24 due to the garbage collection scheduling preventing the immediate upgrade to the latest language version. Due to this, the Kubernetes release process already heavily documented was updated to no longer follow main images of Go, but instead conform to the validity of the canary language images before considering the introduction of the main branch. This also had a number of knock-on effects for a number of Kubernetes-dependent projects and infrastructure things. In December 2020, Kubernetes announced the deprecation of Docker
as a container runtime. This was announced as the removal of the Docker shim, initially a temporary workaround to allow Docker and Kubernetes to work together, but was becoming a burden on the Kubernetes maintainers. Earlier in this year, the K-Native project was accepted into the Cloud Native Computing Foundation. Originally developed within Google, the donation of this project ensures that it will continue with open governments and not be beholden on any one corporate entity. As new architectures are adopted and supporting technologies grow, a lot of people are very, very excited about them.
Much like the flood of information I given you in the last couple of slides, depending on what channels you follow, you might be inundated with similar things from time to time. Those who work in DevOps and Systems Administrator roles, excitedly share the latest Kubernetes updates and want to encourage everyone to follow along. Some of this stuff is honestly really interesting. to the right people. That's not you. Probably. You are probably not the intended audience for their enthusiasm You dear audience member are probably a Django developer. Who here is a Django developer? You're here in Portugal, after all, at a conference specifically for Django, so I'm glad that my educated guess was mostly correct.
You might also have skills outside of Django, absolutely. You might already know a bunch about Kubernetes, but I'm talking to the Django developer side of you. And I'm sorry for the mean intro, that was really rude, but I intend to make the rest of this accessible and I am not going to speak at the Keith McGee Breneckey speeds. I am going to slow down now. As a further warning, this keynote contains opinions, conditions apply, often not available in some states, 10% rebate in South Australia. This presentation is going to focus on three areas that I think you should be worrying about. The stability and scalability of your infrastructure.
The reliability and resilientness of your code and the thoughtfulness and thriving of your community To start, stable and scalable infrastructure. This section is going to talk about where you run your code. The state of the web has changed dramatically Since Django was first released into the world. And the way that modern infrastructure works has also changed, but it's also stayed the same in many respects. If you want to run Django, you can absolutely create a server, install Django and Postgres and Nginx onto it and open it up to the world. If that's what you want to do Sure, but there's a lot going on and not going on in that sort of setup.
Where is the server? Is it under your desk in your office? What happens when the power goes out What happens when the network is disconnected? When are you going to stall operating system updates? What happens when you need to update Nginx? What happens when you need to update Django? How do you make updates to your site? How do you apply database migrations? When do you apply database migrations? Who manages your database? Where are your backups? Do you have backups? What happens when your site gets really popular what happens if somebody tries to access your site and make unauthorized changes? How do you tell? What happens when no one goes to your site? How can you tell if it even works? If the answer to any of these who questions are you, are you a Django developer? Or are you a combination Django developer
and database administrator and systems administrator and network administrator and full stack engineer and grossly grossly underpaid Your Django site is unique and special and you've put a lot of work into it that it's to make sure that it suits your needs and your users' needs But your Django site needs are the same needs as many other Django sites. It needs to talk to a database. It needs to have somewhere to host and serve images. It needs to be running somewhere where it's accessible to the internet, but can withstand the internet trying to access it back. Turns out these are the things that many other websites need. Not even just Django sites, not even those written in Python.
But a number of different programming languages and web frameworks all need these things. The things that a Django developer needs are the same things as what a Ruby on Rails developer needs. They need a database. They need somewhere to run their stuff. or a Laravel developer, or most any other ORM desk web framework. The size of the developer base of all these frameworks combined is of course larger than any one Stack. So since there are so many people that need these things, there are now many service providers that can give you these things. I work for one of them. I work for a place called Google Cloud The infrastructure required to deploy any website is now a commodity or a utility for many hosting providers.
They will offer compute options. Compute in this case is anywhere you would run a computation, any way you could run some Python code That might be a virtual machine, it might be a function or a lambda, it could be a container. But these are all ephemeral objects. They don't physically exist, they exist in the cloud, where cloud is not your laptop. But it's also not one specific server, hosting providers won't have somebody in a data center physically clock come in and plug in to a physical computer when you want to host your website. There are still some service providers that allow you to do this
But the scale of the compute you want probably isn't that big anymore. There is no one physical machine hosting your website. The benefit for this model though is you only pay for what you use. Much like if you have plumbing in your house and you have multiple taps, you only get charged for when you take water to have a drink. But if you're not using any water, you don't get charged. Clouds can provide you with the compute that you only get charged for when your site is serving things. and they can handle the peaks in demand because in the great scale of things, your Django app is just a drop in the ocean of the power that they can provide. On top of that base compute, there are also a myriad of other managed service options available that allow you to do things like
Click button get database. You can choose your flavor of database, probably the latest version of Postgres, and in a few minutes you can have a database instance that will be all set up for you where you can adjust your settings for your backup frequency, have a maintenance schedule for you. All these operations you might be used to manually performing, but they're managed for you. After creating your database and users and their password, all you have to do is point your code to this new instance using those credentials and you're off and running. And you can continue to hook things into your compute to expand your platform. You can connect it to source control and have continuous deployment. So when you merge your latest changes, everything is updated for you.
You can hook in managed Redis to handle caching. You can add a CDN, you can plug in event-driven management. You can make all this happen with many samples, tutorials, and guides no matter which platform you choose. Part of my job is to help write these tutorials. If you've tried to run Django on Cloud Run, come speak to me afterwards. I'd really like to chat to make sure it's still working for you all. But it's not just for Django either, but for other frameworks as well. Part of my day job is to help take the worry about these sort of concepts for developers like you. I like my job. Because in all of this, this sort of architecture is a common modern pattern for developers and practitioners to run their code.
But depending on where you run this, even if you were to run this on Google Cloud, you may not even be using Kubernetes. As a developer, Kubernetes is not the tool that you want. Google pay me, and I'm allowed to say this. Kubernetes helps build the tools that you as a Django developer will want to use. For sufficiently complex setups, Kubernetes might help But for sufficiently complex setups you should have a sufficiently complex team who are experts in all that infrastructure stuff so you can focus on being a Django developer. But if it's just you, you should not have to manage that all yourself.
Kubernetes is a platform that hosting providers might use to help their customers deploy and manage their deployments But if you aren't a platform provider yourself, this may not be the tool for you. For example, Google Cloud Run uses Knative mentioned earlier in my big mini intro. But While that can run on Kubernetes, it doesn't have to. But you don't have to care about that. You do not have to worry. One of the things that Kubernetes popularizes is containers. This is not AI generated. This is a real art thing, by the way. Containers are a separate concept that Kubernetes and Docker have popularized. Containers and container images are the latest
in a series of virtualization paradigms that are basically zip files with recipe cards. A a Docker file, if you've seen one of those before, is a series of instructions that help build a container into layers that are all wrapped up, copied up somewhere and then copied out again. Heck, the earlier versions of containers were just tables, just little zip files, and you go zip and you go zop, and then you get your file and it's like, hey, and they have a container. But uh once your container image is built, it's a self-contained bundle of stuff that you can run. depending on your hosting platform, they probably handle it uh like an executable or a jar file.
Just fire it up, check that there's something listening inside. and it'll run it. In your case it'll probably be something like G Unicorn pointing to your Whiskey Django application object. push button get database get website databases websites they're all the same right no they're not they're not all the same it's fine But understanding how containers work isn't the magic, it's how they work in stateless architecture that's the important part. Because those images, those tarbals, once you build them, they're immutable. And you might be serving zero, one, or many of those
uh copies of those images at once and no state persists within. This is why you need to connect to a database and storage and other things because once your containers scale down, they disappear like magic. When I talk about scaling up or scaling down, I'm talking about auto-scaling. If you have a service that's being configured to take so many requests per second You start if you start getting more than that, then your provider can scale up your server so that instead of one instance serving traffic, there's two or more. Think of it like the registration desk outside yesterday morning. If you were here before the opening of the conference, you would have used one of two desks
because there were so many of you. that it helped make things go smoother. Now that we're in the swing of things, there's probably less people that still need to get their name tags. And so there's only one desk in operation. Uh hint it's the one with the people in the orange shirts behind it. That's the one you need to go to. When you use auto-scaling, you can have many copies of your site serving traffic when your site is busy and none when it's quiet. And when you have nothing running, you don't get charged. The tap is off. But just make sure that when you're using this feature you set an upper limit just in case you get really popular and start scaling all the traffic and you end up with a little bit of the bill at the end of the month With containers, whether auto-scaling or not, you do not have
a place where your site is. If you were to try to write temporary files into your container, much like you would if you were still on a virtual machine By the time you want to get access to these temporary files, that container may not exist anymore. Some providers don't even let you write files other than at slash temp to prevent you from having these sorts of issues. So you need to make sure that any data you want to persist is written somewhere that's not in the container itself. How your application works in this stateless architecture, that's one of the things I think you should worry about. Kubernetes is a different part of the stack to Django. And as a Django developer, you're working at a higher layer of the stack. You should know your stuff for your layer. If you want to learn more about Django, we hear this
awesome conference on the next couple of days. But you should at least have some passing knowledge about the interface that you directly interact with. Or to put it a better way that doesn't look like a stock image, you are the lovely custard filling in this pastry You might be aware of the meringue on top, the pretty front-end heck, you might know someone or know how yourself, how to make the wonderful piping and the crispy bits on top. But You need to be aware of the base you're served on in case this case the pastry just to make sure you don't make it soggy. But The plate that your pastry is on, the table that plate sits on, the floor that table is on, and all the layers below that, don't worry about that. Worry about being tasty and working well with your immediate neighboring elements
Best practices change over time and are constantly changing and are in a much better place to what they were ten, fifteen, six months ago. But for right now, this is where we're at, and you should worry about having at least passing familiarity with how your part of the stack works with other parts of the stack. And with that encapsulation out of the way, with you all thinking about morning tea, we can now talk on the next worry. But first If you have slides that say drink, you have to drink. Water. Drink water. This next part of the talk is going to talk about your code itself.
This is the bit your managed service provider is going to presume that you are the expert on because you are You know your site better than anyone else, and your site is going to be made of code, and that code Is going to depend on things. It runs a version of Django, which comes from a Python package as a dependency. And then that is going to have other dependencies. You might rely on other Python packages that may help with developing your application You might also have dependencies outside of Python. Your fancy front-end that meringue on top might rely on the latest JavaScript framework, and that itself has dependencies. When you rely on these packages, you depend on other people's work.
You do not have to create your own package to do that one specific thing. Danny was talking about this in his keynote yesterday. You can rely on the work of others and you do not have to recreate everything yourself. You don't have to work out how to do that weird, complex, mathy thing when you can just pip install a solution. Heck, even using Django means you do not have to worry about a bunch of things. you can let Django handle your calls and CSRF and all that other fun stuff so you don't have to worry. Building on top of anything means that you rely on other people's work and you can focus on what you want to build. But things change. Packages have updates that you should really incorporate into your workflow.
It's more than likely you've already seen this in action. Services like Dependabot, Renovate, and Seek are integrated into a lot of coding platforms nowadays, like GitHub and GitLab and others. It's likely you've seen the automated pull requests asking you to update your package dependencies. It is a very, very, very good idea to be applying updates. Updates to to the packages you use, updates to Python itself, updates to your operating system, to your browser. Installing updates protects you from security issues and protecting yourself from security issues is a very very good idea. Keeping your version of Python in line with the current supported versions is also a very very good idea. No one here is still using Python 3.
6 are they? No one maintains that version anymore. No one here is using Django two point two, are they? No one here clicked away their Windows updates or their Chrome update notification this morning, reminding themselves to do it later, did they? I saw one of your lightning talks with your update window thingy. It was in red. You should have clicked that. Keeping your dependencies up to date is a very, very good idea, but What happens when that goes wrong? What happens when bad updates are pushed out? It might be something as simple as a regression in functionality that wasn't picked up in testing.
If you don't pin your versions, you can sometimes run into these issues unexpectedly. The updates themselves may not be safe There have been issues in the past with updates being made in bad faith, including malicious code. Updates can be removed. Maintainers have the ability to remove their packages from indexes so that their work is just no longer publicly available. Instances where this has happened has been based on many factors from personal political views to lols. But the end result is that that code just isn't available anymore. Or that package that you think you're installing may not be the one that you're actually installing. That one misspelling means you may have ended up downloading a different package than what you intended.
And we're not talking about potentiality here. We're talking about inevitability because there will be a time when this goes wrong And you need to be ready for it. You may have heard of these concepts under a new name, supply chain security. This or at least the concepts around it is one of the things I think that you should be worried about. Who here um has heard the term supply chain security before? Who here knows what that means and how to describe it? Exactly. There are many ways you can think about this concept. You can think of it as a clean room assembly, like NASA making and rocket. You need to make sure that everything going into the clean room
is clean so that when you assemble it there's nothing else that can get in there, no dust, nothing. That object that you end up being that you end up creating, you know what went into it, so you know what comes out of it. You need to make sure that the data going into your system is good to make sure that the product is good. Another way you could think about it is like the supply chain for coffee, everyone's favorite dependency. I miss live audiences so much. When you go to the store to buy coffee, what things do you look for? Do you look for eco-friendly? Do you look for fair trade? Do you look for local providers, specific brands that you enjoy?
Do you care where it comes from? Do you care that the growers are fairly compensated? That the package itself has beans in it that are freshly roasted that that package has been sealed. You wouldn't buy a pack of coffee from the supermarket if the package itself has been ripped apart and resealed with duct tape. So why would you do the same for your package dependencies? What I think you should worry about is thinking about drafting your plan for how you handle these sorts of potential issues. For that one site you created, what's your maintenance plan? When do you plan on applying those security updates? Are you using the latest long-term support or LTS versions of various parts in your stack?
The Django, the Python, the operating system. What is your plan for when those LTS versions expire? How do you update? Does your hosting provider have anything to help with this? They may offer some scanning or some automation. Again, these problems extend beyond Django. So there are a lot of solutions for a lot of different types of developers out there that you might be able to use. There could be a commercial vendor that could support you. You do not have to buy a solution, but this is something you should check. And how you do your dependencies and how they relate to each other is also important. One of the things I think you should check out is deps. dev. a project that I've helped out with that has the cutest fricking logo of all time.
I have stickers. DevStot dev shows dip how dependencies relate to each other in graphs so you can see things like for this one package, what is the common license that all these package dependencies use? Are there any Issues where the latest version of one package relies on a version of another package that hasn't been updated. All this information is also available as part of the Google Cloud Public Data Set program where you can analyze all this in BigQuery and you can have a look at different things in that data set. The depths. dev blog has some really interesting write-ups about the timelines for vulnerability disclosures.
And this little fella is old Captain Napkins and was created by Renee French, the same illustrator that created the Golangopha. There is no related in a Django talk, I swear. For your packages, if you happen to host a package on PyPI, when do you apply updates? How do you plan to prove that your packages are safe? You may want to look up the supply chain levels for software artifacts or salsa Some of these names are tasty. Salsa is a series of levels of compliance for
package integrity. You don't have to try to reach the top level of solve circ compliance today, but consider what you can implement about some of the suggestions for some of the lower levels of compliance. One of the tools that helps with this is SIGSTOR. Which has been getting a little bit of a buzz around Python recently, as the most recent releases of Python are signed not only with GPG signatures, but also with Sigstore. For your CI, your continuous integration, check if your hosting platform has options for automatic dependency update suggestions. robots that can say, oh there's a new version of Django, you should probably install this so that Carlton
will not email you in anger at some point. You should take advantage of these automation options where you can. You may also, depending on the size of your custard cream tart Consider having a good cache of packages. When you run your CI, you're probably going to be installing from PyPI, you're re-downloading them each time, but Consider of having a DevPy or some sort of other cache of dependencies that you can set up and pull from and specifically only add new packages in there when you know that they are safe, when you've checked them, instead of just installing latest from pip. This can also extend to your own system as well.
I know that you ignored your Windows update notification last night. Plan for when you want to click that button. One thing to remember in all this though, in all these analogies about supply chain security, in actual supply chains, there are explicit contracts and trade agreements and money. While there are some improvements to the sustainability of open source projects in recent years, with sponsorships and grants and the like, most of open source is run on volunteer labor. In this article from Ileana, Shay explains that a lot of open source is hobbyists and a lot of people do not plan to be the critical path for a million
billion dollar software build chain. The stacks that we build upon are maintained by people like us, so we need to make sure we have empathy for those maintainers. And understand that not all actors in this play are forces for good, and we still need to be vigilant. Who here has played Untitled Goose Game? Untitled Goose Game for those that are unaware is a game described as it's a lovely morning in the village and you are a horrible goose. This is a game made by the Indie Studio House House, which are based in Melbourne, Australia. G'day And you go around being a chaos goose, stealing baguettes and pushing people into lakes.
Uh this is available on PC, Switch, and a couple of other things. It's a whole bunch of fun. There's a button for honk Honk. While there are many security threat persona, I like to think of an amalgamation of a lot of these as chaos geese. Think about what sort of goose could get into your code and how you can mitigate that. How could someone ruin your day and how can you give yourself time to respond? What can you do to chase these geese away? You are never going to be 100% safe in digital security, but the best thing you can do is try to make things easier for s for yourself and a little bit safer from these chaos geese. Drink break
We're nearly done, it's fine. We've covered a lot so far. There's a lot of work that goes into everything that I've been discussing for the last half hour. There are many groups of general and specializing interests, thought leaders, teachers, learners, and peers. This section is about that community. It takes a lot of labor to make things in open source happen. We've seen just how many things go into your projects that make your community. We're Django developers, right? But the packages that you rely on, the mentors you've learned from, the speakers you've listened to and taken the advice of, the people who inspire you.
They are your community. But who are we exactly? This entire time I've presumed that you are a Django developer. And that prince assumes a lot, I mean, given that you're at a Django conference, you probably identify as a Django developer But that might not be all you care about. You might still be very angry about my awful statements about Kubernetes earlier and Fair. But I cannot presume that Django Developer isn't your only technical identity. It is most likely you have other identities that you associate with. You're a PyCharm user. You're a fan of your local football team. You're a lover of jazz. The people within this community are full of many individuals.
who in a four-dimensional Venn diagram happen to overlap with the identity of Django developer today. And a lot of them are here in this room. But look around you Who isn't here? You may have someone that you were hoping to see this week, but they weren't comfortable travelling given the ongoing pandemic. When I was asked to give this keynote many months ago, I had my own risk assessment to work out if I was okay to even consider flinging myself halfway across the world to be here Not everyone is okay with traveling right now. That's even if they can afford to travel, or even if they know that this event exists. Yes, having student tickets and virtual tickets absolutely help with this
hello from online. But this is just for this particular event, this specific gathering at this point in time. This is a biased representation of the Django community as a whole because this is not the entirety of the Django community, it is just a subset. But what about a smaller group of this? What if we wanted to know everyone who has ever contributed to Django? If you wanted to conduct an exercise to list every person who has contributed to Django the package, you'd need a time machine and a cloning device for all the auditors to be in every hallway, every water cooler, every room, every conference, every company where the ideas and of the creation, development and evolution had ever occurred, and you would still never get all the data.
But wait, we know who every contributor to Django is. We have the full Git and subversion history. That's everyone right there, right? My sweet summer child, it goes so much more deeper than that. The people who turn up on the GitHub subversion history, that's only the code contributions. That is only a subset of what makes Django Django. And Kojo talked about this yesterday and I'm going to continue to explain it to make sure that the message gets across There is so much more that goes into Django than just the people who have commits on the main branch. But sadly, many projects rely on this metric of contributors to the
code in Git only because it's easily obtainable data and they can just say, oh, select all from contributor list, call it a day How many Git stars you have is the indicator of popularity of a project on GitHub after all, right? Mario Bagoli explains it thusly, all metrics of scientific evaluation are bound to be abused. Goodhart's law states that when a feature of the economy is picked as an indicator of the economy then it exorbitantly ceases to function as an indicator because people will start to game it You might have seen this in action in a number of different instances.
If you were graded on the number of bugs that you fix. You're going to do the easier bugs first, right? Or, heaven forbid, you create easy bugs to fix. You game the system. The harder complex bugs that might be the biggest benefit for your users only count as one. You are the gimli to the legoless in this case. Guess there are also New Zealand references in this talk. There's a movie called Lord of the Rings and there's a big elephant and that only counts as one. Pew Pew. Pew. A more concrete example of this instead of twenty-year-old Hollywood movies. Oh god. Um
Who here has heard of Hacktoberfest? Who here remembers when Hacktoberfest was not opt-in only? Yes. Hacktoberfirst for those who aren't familiar and haven't seen the people with the t-shirts running around. It's an annual effort to encourage people to contribute to open source. And before projects were able to opt into Hacktoberfest, projects would be inundated with useless pull requests from people not understanding the assignment, just trying to get their free t-shirt. The amount of project maintainers who dread October, the amount of spam they have to deal with, it was a lot. I was one of them. But the inverse is also true.
We cannot readily rely on an accessible an accessible metric alone to be the one way that we measure things. Because we would exclude so much work. Forge Your Future of with Open Source by VM Brussure describes a list of possible contributions to a hypothetical open source project. And this list may not apply to all projects, but it also doesn't include all the types of contributions that might happen. For example, if you work with hardware, there's going to be a number of tasks like quality control. and manufacture it and that sort of thing. But this sort of list brings up a couple of important points. When you're using metrics based on your social coding platform of choice, how much of this work, and it is
work Is going to be recorded in that platform to even have a chance of being rolled into whatever code metrics you're using. If you're using platform metrics you have to augment them to make sure that all the work that is being done is visible. The people with the skills to make some of these contributions may not have the skills to make others. For example Your marketing person doesn't know what code testing is. Your projects benefit from people with different specializations, which means that the people who contribute to your project are going to have different experiences. and may be a member of your project's community, but not the community that your project itself sits in. For example, the domain expert for your goose
egg-based custard retail site may not know a thing about Django. But they are a part of your community, but they don't identify as a part of the Django community. And in all these different things, only one is code. The rest is still work. Hacktoberfest opens next week with an emphasis on non-coding contributions and have an advisory council now to ensure accessibility and inclusivity for the event. And to track this, to make sure that you can still create four pull requests get t-shirt, you have to make sure that the work you do is tracked in pull requests This helps with data collection for the purposes of that specific event, but for your projects, the ones that already exist, you may not have that luxury.
And yes, I know what you might be thinking, Katie, you've ranted about this before. Why, yes, I have. I have presented thoughts on this sort of topic at many a PyCon and a DjangoCon about all contributors being welcome. That talk was focused on GitHub because at the time it was extremely isolating to contributors being code only. If your code ends up in the main branch of the repo, it counts as a contribution. Any other kind of work doesn't count, doesn't get attributed, doesn't get acknowledged. Since then though, GitHub has improved. For a time, they included a list of contributors to your project's dependencies as contributors to your project. Who remembers when we first saw a frickin' black hole?
Who remembers? Yes? Dr. Katie Baumann and the team from the Event Horizon Telescope were able to make this with Python. As part of GitHub Universe that year, GitHub's big conference doodad, they used this project as an example of the scale of dependencies in open source because over 21,000 Pythonistas contributed to the dependencies that were used in a project to take a picture of a black hole. Black hole so cool! But at the exact same time, Dr. Bowman wasn't given the credit that she deserved because some people on the internet thought she didn't write as much code as the others, so she should not be given the
acknowledgement because code is important. GitHub hasn't been improving since then. The work of Dr. Nicole Forgson and the space framework for understanding developer productivity has helped direct GitHub. in a increasingly more inclusive direction. But a lot of that is still ongoing and a lot of that is about enterprise developer productivity. GitHub is a platform for code and enterprise code, so that's what they are going to focus on. Personally, I hate the term non -coding contributions. It's splitting the work into those who are technical
and everyone else and suggests that anything that isn't code isn't worthy of further delineation because code is the only thing worthy of identification. It's like having technical talks and soft talks. Every talk at this event is technical, even though it may be outside the realm of software engineering. Fight me in the whole way after. Kojo's keynote yesterday already emphasized this, where programming is just a subsection of software engineering, and we need to make sure that we aren't excluding the work of those who use Python to solve problems just because they don't have specific training or experience in software engineering.
As part of my oft repeated hat rack talk, I introduced a tool that would aggregate everyone who has interacted with a GitHub repo and call them the contributors. It would capture the code reviews. the bug reports, all the interactions on a GitHub repo that may not necessarily be code in the main branch. But since those rants, I've been busy I'm part of the University of Vermont Open Source Complex Ecosystems team whose goal is to deepen the understanding of how people, teams and organizations thrive. in technology-rich settings, especially in open source projects and communities. And as part of this, I wrote science Presented at the 2021 Mining Software Repositories Conference, I was a co-author on the paper, Which Contributions
Count? Analysis of attribution and open source. Citing from the abstract of Babby's first academic paper, we found that community-generated systems for contribution acknowledgement make work like idea generation or bug finding more visible, which generates a more extensive picture of collaboration. We also found that models requiring explicit attribution led to more clearly defined boundaries about what is and what is not a contribution. We originally sought to make a taxonomy of contributions, moving away from the code and other. But We found there is no way that we can apply that to every open source project
because every open source project is just so different. We were originally basing the idea on a taxonomy for academic papers known as the credit model, but academic papers are much more structured and formal with a much more defined process. As part of our ongoing work, we've just released a preview of a workshop that some of you helped us with. Cheers. And We have been continuing to develop this as a way to help projects identify the work that happens in their projects. Available at WhoDozThe. dev because I love naming projects. We are offering a multiplayer and single-player version where as a group or on your own you can go through and analyze the work done in your community
to make sure you're correctly identifying existing work and encouraging the work that you are missing and that you desperately need. If you're interested in trying any of these workshops yourself, please get in touch. We've run them a couple of times ourselves, but we're always looking for ways to improve them to be used by others. You should worry about where you let people hang their hats. You need to make sure that people can show that they have done work and that you should proudly display those people who make your project possible Make sure that you're acknowledging the contributions of others and make sure that they are attributed so that those people can point back and say what they have done.
Especially when many of these contributions are voluntary and unpaid. Being able to show proof of work is helpful helpful in many aspects. Not just being a good community member, mm whichever community that is So that's what I think you should worry about. You should worry about how your Django app works in modern infrastructure stacks. You should worry about your plan to prevent issues with your dependencies and provide clarity when you are the dependency. And you should worry about the work that goes into your projects and make sure that your contributors are properly acknowledged for their efforts. And for those who have been paying attention, yes, I have been directly working on each of these three things for a number of years now. And that's my bias and that's what I think you should worry about
because hopefully one day if you're presented with issues like this you can be like an Australian and say no worries Happy job
Focus on the stability and scalability of where your application runs, including hosting, databases, backups, updates, deployments, migrations, security, and monitoring. Managed cloud services can take much of this operational work off your plate.
Discussed at 3:11Usually not. Kubernetes is often part of the platform that hosting providers use, but unless you are running a sufficiently complex platform yourself, you should work at the Django layer and leave that infrastructure to an expert team or provider.
Discussed at 9:27Container images are immutable and instances may scale down and disappear, so state does not persist inside the container. Persistent data must be stored in an external database, object storage, or another service rather than in the container's filesystem.
Discussed at 11:46Applying updates to dependencies, Python, the operating system, and other software protects against security issues and keeps the stack on supported versions. Version pinning and testing are also important because updates can introduce regressions or malicious code.
Discussed at 18:06It means making sure that the packages and other inputs used to build and run your application are the ones you intended and have not been tampered with, removed, or replaced. The speaker recommends planning maintenance and security updates, checking dependency relationships, using automation carefully, and considering trusted package caches or signing tools.
Discussed at 20:31Code commits are only one kind of contribution; reviews, bug reports, ideas, documentation, support, design, marketing, and other work also matter. Project metrics should be supplemented so this work is visible, and contributors should be explicitly acknowledged and given proof of what they did.
Discussed at 35:15Projects should identify the work happening in their communities, make it visible, and attribute it to the people who performed it. This is especially important for voluntary, unpaid work, because contributors should be able to show evidence of their contributions.
Discussed at 43:18Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025