A Data Harmonization Engine on top of Django - The Good, the Bad, and the Ugly

This video features Jonathan Ströbele at Django Day Copenhagen 2024 in Copenhagen, Denmark.

A Data Harmonization Engine on top of Django - The Good, the Bad, and the Ugly
0:27:13
Published October 13, 2024
113 views

Django Day Copenhagen 2024 talk descriptions: https://2024.djangoday.dk/

Summary

Jonathan Ströbele presents Data Hub, an open-source, self-hostable Django system that harmonizes weather, land-use, census, geospatial, and other data across administrative and temporal boundaries for tropical-medicine research. He explains how Django provides structure, authentication, the admin interface, stability, and an extensible way for users to add custom apps and processing code, while also describing difficulties around environment configuration, reusable projects, dependency management, dynamically loading user code, APIs, and front-end tooling. His central argument is that Django was a good choice despite gaps in use-case-oriented documentation and the need to build some surrounding tooling himself.

Key takeaways

  • Data Hub maps varied sources such as rasters, vector files, CSVs, spreadsheets, and APIs onto geographic and temporal administrative units.
  • Researchers can view metadata and maps, while developers define source-specific download and processing routines in Python.
  • The project combines database-backed data-layer metadata with dynamically loaded Python classes, using error indicators to flag broken or missing routines.
  • For exports, straightforward Django views and pandas are sufficient for read-only JSON, CSV, and Excel responses without adopting Django REST Framework.
  • Django’s structure, admin, ecosystem, stability, and support for customization outweighed challenges involving settings, dependencies, documentation, and front-end asset management.

Summarised automatically from the transcript.

Transcript

4,521 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:05

Speaker 1: All right, welcome on stage, Jonathan. Thank you so much for joining us here today. I think we should start by giving a big round of applause. And the stage is yours.

0:19

Speaker 2: All right, uh thank you so much. Um yeah, today I'm going to uh present to you our And um before I start, just a few words about myself. Uh hi, I'm Jonathan as you can. already heard I'm based uh from Hamburg Germany and I have a background in PHP and full stack development worked there in an agency for many years but I also uh still up to day uh work with PHP and Laraville framework for example And last year I finished my computer science degree in at the Harvey Hamburg. And there I started also to work a lot with Python already for so coming from the data science perspective. And now at the moment I'm working at the Bernhard Nocht

1:06

Speaker 2: Institute for Tropical Medicine. This institute has a focus on research in tropical diseases. And there I am in the Data Snack project and we are two people project uh where I am the only developer in the team. Um so and you might be wondering how did I as a computer scientist end up in a research institute full of tropical uh X? experts in tropical diseases. So what are we building and why? First and all, context data like weather or land usage, census data can provide crucial insights for health professionals. But those information usually come in a variety of different formats. For example, we have geo rasters, we have vector files, we have plain CSV files, or maybe Excel files, or maybe even

1:52

Speaker 2: APIs and stuff like that. So we somehow need to yeah handle all those different data inputs. And additionally, health professionals usually work on the adm administrative units of countries. For example, in this case, the country of Tanzania is divided into regions and the regions are subdivided further into districts. And usually the healthcare system also provides information and data based on those administrative units. So we need to somehow map those different variety of the raw data inputs onto those administrative levels. For example, we are counting facilities in a region. or the proportion of forest inside a regional district. So that we get some nice heat maps, for example, which the healthcare professional can use to

2:39

Speaker 2: base his further enrich his decision making on. So and that's actually what we are building. It's a data fusion engine we called the data hub. And it's an data harmonization engine written in Python so the user can Define data processing and actual Python code. The system is um yeah open source and can um yeah and uh the since the data is based on Python code, the actual data processing routines are reproducible because well you You need to just execute the code and then can go from the raw data to the processed data again. Yeah, and this system we built uh aims to provide um To harmonize the data on a temporal and geographic extent.

3:25

Speaker 2: So, and you might be wondering, okay, that's nice, but why Django? So we not only want to process the data, we also to want to show the researcher the process data. So for that we used a web application based on Django. And also we want to enrich the information shown to the researcher by showing documentation and metadata about the data sources we aggregate in our system. And also I didn't want to mix any programming languages. The data processing is also already fixed to the Python ecosystem due to the large environment of data science and stuff like that. And then I also wanted to create a web front end for it in Python. And so I chose Django. But I before that I also already did a false prototype and for this I used Flask.

4:11

Speaker 2: which was nice at the first week and after that it was quite smessy because Flask doesn't provide any well that not not that much structure. Also it doesn't provide, for example, user management out of box. and you had to pretty much do everything yourself and it was quite messy and so also that was a further point for using Django where you have a provided structure. So and uh that's actually the data hub and I want to give you a quick uh live demo clicking around in it So that's the data hub system, it's the live system. For example, we can click around in it, see the different uh regional um shapes and extents for an country, in this case it's Ghana, and then we can click on the data layers menu item and see all the different data sources which are integrated at the moment.

4:57

Speaker 2: And for example, we can get in, search for forest. And then find the forest land cover data layer and see the metadata of this data source where it's coming from, and then can quickly visualize the data, for example, on the district level Inside Ghana, and now we see in the south part there is lots of forest, the blue parts, and in the north of Ghana there is less forest. So, um genau exactly. The data hub is open source, it's self-hostable, so you can put it in your own infrastructure, it's updatable and aims to be a platform for reproducible data. So, and now comes the ugly part. I started to build the Django application for it.

5:43

Speaker 2: And the first thing when using Django was like I'm have a background in uh yeah lachable framework for example and if I create and setup I have an environment file which uh the environment specific settings like database credentials for example And this is coming from the 12-factor app principle where you separate, for example, the environment configuration from your actual code And in Django I had just the settings PI. And I was like, okay, I need to commit it to version control, but I also have the environment-specific settings in it. How should I handle that? And where should my environment-specific configuration go? And I think the documentation didn't guide me, or the tutorial in Django didn't guide me in anything how I should solve this conflict. And then after a bit of search, I found the

6:29

Speaker 2: Django Environment package, which now allows me to create an environment file outside of version control and then inside the settings BI load this package and yeah uh put in um the different variables defined in environment file and reuse that and so can decouple my settings pi file from yeah my actual code The next big part was the Django app, which probably you're all familiar with. For me as a newcomer, uh this was quite a big roadblock because um Well, the documentation tells you you have the project with the app. The project is the combination of different apps with a configuration, and the app is a singular module containing a specific functionality. And well, I wanted to provide my users with a redistributable out of the box system.

7:15

Speaker 2: So I wanted to redistribute Anjango project and not just a single app because my user shouldn't be bothered with configuring this thing the application. So in the documentation I only only found how to write reusable apps, but there was no word on how to, yeah, how do I provide a reusable project for new use for other users? So and that was uh yeah quite uh difficult for me to grasp and that's actually how I ended up structuring my app or my project into apps. First of all, I have the data layers app, which contains everything regarding to data sources and data processing. The next app is my Shapes app, which contains logic regarding the different shape files and yeah ge geog um geo the the processing for the geometries of the areas, districts and shapes.

8:02

Speaker 2: And then in the last I have an app which is called app, because I wasn't really sure how to solve the problem of having global stuff, template stuff, user authentication. Where should all the stuff go? Because then everything is well has a singular modular scope. Where should yeah global stuff go? And also, for example, a junky documentation um tells you and advises you to override the user model in your new project. But it doesn't tell you where this overwritten user model should actually go. And I was completely confused so and also I put it into my app which is called app and for me the system works quite well. But also I think it's not really in the sense of how Django apps were um yeah envisioned to be used because I not actually structure single uh singular self-containing parts so no

8:49

Speaker 2: more of like uh meant the mental model um of the different parts of my application and all those three apps have well depend depend on each other and I can't really share any of it. So uh it's maybe not that conventional, I'm not sure. Nevertheless, I got that out of out of the way and it worked, and now I was faced with the issue of dependency management. Um and I mean you know all Python stuff, it's a little bit complex. We have requirements, TXT, we have pip, we have pip tools, PowerTree, PDM, lots of different tools. And I was completely confused on how to organize the dependencies of my project. And I was very reminded of the XKCD comic regarding Python environment and generals. And well, I was at a loss. So and always I do what I always do in such a case, I start looking at different projects.

9:35

Speaker 2: And so I started looking at all those different open source projects, from which I know they are based on Django. And well uh after looking at them I was more confused than uh before. Because every one of those tools does something completely different regarding the tendency management and deployment. So that was uh quite confusing. Nevertheless, since I mentioned so every Django project does its own thing regarding that part. There's no guidance in the documentation, or at least I didn't find it. I had no idea how to differentiate because Between my first level dependencies and my yeah, the depth whole dependency tree. So my first idea was, okay, I just use the requirements, TXT, and hope for the best. Nevertheless, in this year I was

10:20

Speaker 2: in a bit of luck because the UV project from Astral kicked off and now I have found some guidance. I just use UV. I have my Pi project Tomulf in my project route define my first level dependencies in there and then use UV to generate the requirements. txt file which I then can yeah re uh roos use in my Docker setup as with a classic pip install and then have on ready to use uh yeah container uh for my whole project which I can dist r redistribute. So so far we have the Django project configured, we have structured it, we have redistribution in in place. So where's the data harmonization? Well Let's enter the bad part. So now was the the thing in the Django model. I have my data layer object which encapsulates

11:07

Speaker 2: the description, the metadata often data source, which the user can edit via the Django admin and yeah, then um is able to change everything. And on the other hand, I have my data layer classes with the actual methods for data processing, which are based on source code and should be created by the developer on the um on the file system so he can put And I somehow I needed to find a way to merge the database part. with the actual user contributed file system part with actual code from the user space I wanted to execute in my system, which is uh a little bit unconventional, I think. So, and that's what I actually came up with. So I have the data layer model, which uh encapsulates the data source.

11:56

Speaker 2: It has an a key which is a unique identifier written in snake case And then I have the actual file on the file system, which matches this ID, also the file name matches the Snake case file name. And inside that file I have the same name but in WordCap as all Python classes are should be in WordCap, I guess. And then I dynamically have this getClass method inside my model, which then just looks if the file is present and loads it and tries to initiate the class inside the file. Which of course this can fail. If the file is missing or the class has some bug. So it's like a bit um yeah, not that stable So how to solve that? Well, put on trial accept clause on it and then just catch the error if the module can't be loaded dynamically.

12:44

Speaker 2: And then but if it's all good to go, I then have the downloading and processing routines accessible from the user-contributed source code and can execute them inside the system. And then the actual system. And that's actually how it looks like in the front end for the user. So you will see here the green icons, there the source file could be loaded and executed. And where it's great there was some sort of exception. We didn't know what exception exactly, but at least we know there's a problem we as user need to solve. So the next difficult part was the researchers, well now we have the yeah the the data process in our system. But the researchers are usually doing their own thing, for example, in Jupyter Notebooks, or they rely heavily on R, or maybe they want to just look at the data in

13:35

Speaker 2: Excel, for example. And we somehow need to free the data from the data hub and make them accessible to the user. We are download or APE access, for example. And so Django plus IPI, the first Google result were of course Django REST framework. And I looked at it and I was like, okay, wait a second. I need serializers, I have abstractions, I have lots of docu to read and grasp. And actually all I want to do is provide my user more or less with a simple CSV file he can use because I don't have any write um yeah operations going. I just need to read operations. So and then I looked at a documentation from another library, from the Pandas library, which I already use heavily for data processing. And well

14:21

Speaker 2: Pandas provides us with a two CSV, two Excel and two JSON methods based on our actual data frame which uh holds the table of data. So an anti crazy idea. Well, for example, if that's my IPE request, I can just identify in my view the data layer I'm requesting. I load the data. And then pandas data frame. Well and then when I want to t return it to the user. Well, I just use the pandas function to uh create a JSON based on the data frame and return that with a JSON response object from uh Django back to the user. And the same also works, for example, if the user requests a ZV or Excel file as download, I can also again you reuse the pandas

15:08

Speaker 2: library to Create those file types and I just need to create a virtual bytes. io object in memory for creating a virtual file. Right, tell pandas to write the C S V or Excel file into that buffer And then with the file response object from Django, I return this virtual file to the user and it actually works. Yeah, so and the last part was front end. So in my application or our application I don't have any state inside the users um yeah inside the user space. It's just viewing at stuff. Um except for charts and maps, which is dynamic, but also it don't has any state. So I didn't want to use any React or single-page application stuff. And I just went with classic Django templates and a little bit of JavaScript to enhance it.

15:55

Speaker 2: Or to be honest, a lot of JavaScript to enhance it. I have Bootstrap for styling and interaction, for example, uh pop-ups on models I have data tables for dynamically sorting and filtering tables. I have leaflet which we heard before for the uh dynamic map parts. I have Protly. js for time trend charts and stuff like that, and a little bit of D3, and also I use Swelty to create custom elements. which is quite a lot. So I need some some way to handle all that uh yeah dependency stuff. And uh then I remembered my PHP and Laravel background and Laravel actually does it really nice because because when you install the framework, this all works out of the box. And then I started to re uh refactor that from Laravel into my Django project. I use weed for bundling

16:40

Speaker 2: all my assets and have a custom template helper for each asset, for example a javascript or CSS file. Which also supports hot reloading inside the development mode. And in production mode, I just return with the same helper the bundled JS or CSS file. And then well, it kind of works and it's really nice. So, but nevertheless, now I told you about uh I'm not happy with Django, the documentation doesn't tell me what I need to know, and um I need to write my own front-end tooling and stuff like that. Why am I still using it? So let's enter the good part. Because actually I'm really quite happy with using my decision to use Django in my project I mean I have ext extensive documentation.

17:25

Speaker 2: The only point is I would wish it would be more use case driven more use case driven and not only that much I think it's very descriptive in some points and doesn't tell you how to actually combine all the stuff Also I have the provided structure, everything, it's a Django project, everything has its own place, routes, views, commands, everyone immediately knows where to look for things. Also, uh the longevity, I mean it's there for almost 20 years, probably will be there for the next 20 years. Um it's has it's helps really great stability. I mean we have security and release process and maybe it's a little bit of slow, but it also means for at least until now in the last year all my minor updates went went through fluently, there was no hiccups or anything like that. And also we can and I can build on the community and

18:11

Speaker 2: the ecosystem provided by Django for information, apps and plugins. for example. And also a big part is the Django Admin interface for us came for free, so we didn't had to build anything for the user to actually edit of the any of the information he wants to, yeah, for the metadata and stuff like that. And one part is the last uh part I think which is really really great. This, for example, is the file structure for a reusable project based on our system. We pull in the dependency, the base image, we are Docker Docker Compose YAML, we can configure it over the environment file. Then the user can create its customization, its custom data classes within custom Python-based data processing routines inside the folder and then one really big part is

18:57

Speaker 2: the user can also just create a new Django app. with all the structures it provides and it allows the user to override the existing templates in the base system and it also allows the user to create new data input forms for example or just custom pages and aggregate the existing data in the system. All by still allowing us to update and uh provide updates for the system, even through the extensive customization the user might have done in this instance of our system. So yeah, thank you so much. That's how we or I unchained Django for our use case. Um if you have any questions or want to read reach out you can reach me on following systems. Also we have a live demo available at demo at datasnack. org and uh yeah we have a project page with some more documentation

19:44

Speaker 2: And information about the system and the whole system is publicly available on GitHub. You can check it out if you want. Yeah, thanks.

19:59

Speaker 1: Thank you. We have time for questions. So if the this uh question yes uh

20:11

Speaker 3: hi thank you so much for your talk uh I'm from Africa but so I I was uh actually excited to see that you used Ghana as your use case. Is there a reason why you chose Ghana?

20:22

Speaker 2: Uh actually there is the Bernhard North Institute for Tropical Medicine has a partner institution inside Ghana in Kumasi And because of that reason we have existing projects based in Ghana and for that we have the use case of providing data aggregation for those projects inside Ghana.

20:44

Speaker 1: Yeah.

20:46

Speaker 4: You mentioned a lot of different formats for input data, uh PDF, uh all kinds of wonderful things. Um can you talk a bit about how you manage

20:57

Speaker 2: Yeah, that was uh the easy part because well it's also a little bit of cheating because uh we provide the user with the hooks in the data layer class for downloading the raw data and processing it And of course we provide the user with some guidance on how to download the data and where the where they will be stored. And also, for example, for DutyF we provide and aim to also provide for different data sources to provide to to provide uh boilerplate code and best practices for handling those data types. But in the end, the user can just define his complete data processing inside this Python function. And so he can do anything Python will allow him to do. And so yeah, okay, to come back to your question, for some data types we provide ready-to-use templates. Which you can put your custom data in

21:44

Speaker 2: and if that's for example for uh geoTIF processing or for vector file counting stuff when you have location data for example. And for everything else, in the future we try to build more and more examples and um yeah provide guidance on how to input different data types. But in the end, if your data type is not yet an example, have not yet an example or is provided, you can just write Python code and do it yourself.

22:18

Speaker 5: Thanks for the talk. Uh I saw that you serialized some raw data to the users. So while you are doing this have you experienced any timeouts because of the size of the row data?

22:31

Speaker 2: Um you mean the quick demo time with the heat map start chart Oh

22:35

Speaker 5: uh I remember a CSV exporter like you create some kind of virtual file on the fli uh uh on the fly.

22:42

Speaker 2: Ah yeah. Uh well We have no timeouts until now. But of course if the data ingested into the system is really, really big, then probably there will be some timeouts in the future. I think at the moment I have just um put in in in engine x the the timeout value to infinity. So that's not had yet become uh become a problem until now, but I'm expecting it will be somehow coming up in the future and we probably need to find out how to handle that.

23:23

Speaker 1: Yeah, any more? Yeah, there is a question in the front. Thanks for a wonderful talk.

23:34

Speaker 4: What unexpected feedback have you got from your users?

23:41

Speaker 2: Uh they all liked it until now uh um um But no I um to be honest, that's actually uh when we present to our peers in the institute or to other people the project, they are often quite like Whoa we need we we need to have this. There is nothing else we can use to provide that and that's information in the data aggregation. And so that there yeah, that is um yeah so welcomed um in those projects that was uh irritating to me because for me from the computer science stuff I mean it's a big fancy Excel table with a nice and with nice visual visualization on top of it. So that was also for me as a technical person that it

24:26

Speaker 2: was quite confusing sometimes.

24:34

Speaker 1: Another question, and this time at the other end of the room, Emile. Hi, thanks for your talk. Um

24:44

Speaker 6: I like how you're now the third person today to mention leaflet, first of all. Um secondly, um you mentioned that you used D3 for data visualization and um I wondered what specific parts you've been using because uh in my team we've been using that kind of and we found it really challenging. for multiple reasons. But yeah, I was just curious what you've used.

25:08

Speaker 2: I didn't know the part actually, which is but for example, in this case to create this map. I load the GeoJSON from my database from my API to provide the GeoJSON and then I load the raw data from my same API and then on the client with D3 create the color scale based on the actual data and then use uh in a in the in a loop to to fill all the geometries from Leaflet and there I use the the three function to get the actual color which is needed for um the correct representation representation.

25:47

Speaker 3: Uh yeah, also a question about visualization uh is like

25:51

Speaker 7: so you also you have the geographical one, but you also have other vi visualizations, right? Or is it just is it just geographical?

25:59

Speaker 2: It's also a time trend. So for example this year this is based on uh plotly. js where you can see for this data source is available for five years. And we can then uh extrapolate the time trend change over uh the available dates of the data source and there we use this uh yeah line shard stuff.

26:18

Speaker 7: Which brings me to the next The next part of my question, have you at any point considered to use a data visualization uh ready-made data visualization app instead like Streamlit or P

26:35

Speaker 2: Um you mean for the client stuff?

26:38

Speaker 7: For the for the you for the data visualization you Use a user interface.

26:54

Speaker 1: Thank you for all the good questions. And um I think that's it. Thank you so so much for the great talk. Um I wanna say that um we

Questions this talk answers

What is the Data Hub data harmonization engine, and why was it built?

The Data Hub is an open-source Python engine that combines data from formats such as rasters, vector files, CSVs, Excel, and APIs, then harmonizes it across geographic and temporal boundaries. It maps those sources to administrative areas so health researchers can analyze and visualize them consistently.

Discussed at 1:06

Why use Django for a data-processing web application instead of Flask?

Django provides a stronger project structure, built-in user management, and an admin interface for editing metadata, while keeping the web front end in Python alongside the data-processing code. The speaker found Flask’s lack of structure and built-in features made the initial prototype messy.

Discussed at 3:25

How do you manage environment-specific settings in Django?

The project keeps environment-specific values such as database credentials in an environment file outside version control. The `django-environ` package loads those values from the Django settings file.

Discussed at 6:29

How should a reusable Django project be structured?

The speaker separates the system into apps for data layers, geographic shapes, and global application concerns such as templates and authentication. He notes that this arrangement is practical for the project, even though the apps are tightly coupled and may not match Django’s ideal model of independently reusable apps.

Discussed at 7:15

How do you manage Python dependencies in a Django project?

The project declares its direct dependencies in `pyproject.toml`, uses `uv` to resolve them and generate `requirements.txt`, and then installs those requirements in the Docker setup with pip.

Discussed at 10:20

How can Django combine database metadata with user-written Python processing code?

Each database data-layer record has a unique snake-case key that corresponds to a file and a similarly named Python class on the filesystem. The model dynamically loads and instantiates that class, catching errors when the file or class cannot be loaded, so the system can run user-contributed download and processing routines.

Discussed at 11:56

How do you export processed Django data as CSV, Excel, or JSON?

The application loads the requested data into a Pandas DataFrame and uses Pandas’ export methods to create JSON, CSV, or Excel output. CSV and Excel files are written into an in-memory buffer and returned with Django’s file response.

Discussed at 14:21

How do you build the frontend without using React or a single-page application?

The application uses classic Django templates with JavaScript enhancements, Bootstrap, DataTables, Leaflet, Plotly, D3, and Svelte custom elements. Vite bundles the assets and a custom template helper serves development assets with hot reloading or production bundles.

Discussed at 15:08

What are the main advantages of using Django for the Data Hub?

The speaker values Django’s documentation, conventions, long-term stability, security and release process, ecosystem, and built-in admin. Its structure also lets users customize the system with new apps, templates, forms, and pages while still receiving updates to the base system.

Discussed at 17:25

Why did you choose Ghana as the Data Hub example?

The Bernhard Nocht Institute has a partner institution in Kumasi, Ghana, and existing projects there needed aggregated data. Ghana therefore provided an established use case for the system.

Discussed at 20:22

How does the Data Hub handle different input data formats?

Users receive hooks in a data-layer class for downloading and processing raw data, and the project provides templates and guidance for formats such as GeoTIFFs and vector files. For unsupported formats, users can write arbitrary Python processing code themselves.

Discussed at 20:57

Does exporting large raw datasets cause request timeouts?

The speaker had not encountered timeouts yet, partly because the Nginx timeout was set to infinity. He expects very large datasets to create a problem eventually and says the project will need a better solution.

Discussed at 22:42

How does the Data Hub use D3 and Leaflet for map visualization?

The client loads GeoJSON and the corresponding data from the API, uses D3 to create a color scale, and then applies the calculated colors to Leaflet’s geographic features.

Discussed at 25:08

What visualizations does the Data Hub provide besides maps?

It also provides time-trend charts, built with Plotly.js, showing how a data source changes over the years for which data is available.

Discussed at 25:59

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos from Django Day Copenhagen