Working with (and around) external data sources to create www.kotuia.org.nz

This video is from Wagtail CMS 2024 .

Working with (and around) external data sources to create www.kotuia.org.nz
0:21:48
Published July 11, 2024
160 views

Have you ever had to build a Wagtail CMS where most of the content comes from an external source? What about two? What about two that really should be linked together, but aren’t? In this talk, we’ll explore ways to combine multiple external data sources and augment them to fit in seamlessly with a Wagtail CMS.

Using the case study of www.kotuia.org.nz, a centralized online whare taonga and museum collections aggregator for Aotearoa New Zealand, we’ll explain how we managed to combine external sources with managed content to produce a cohesive user experience.

We’ll explore:

  1. combining external sources with managed content to produce a cohesive experience
  2. combining multiple external sources that are linked, when you can’t make a copy of the data for licensing issues
  3. augmenting external data sources
  4. managing crowdsourced content with variable quality and data structures, and display it in a consistent way

We will also touch on:
5) what compromises to make when your external data source doesn’t have content in the language you need
2) content caching: how to respect data sovereignty while serving content efficiently.

šŸ“¹ Related Videos To Watch Next:

ā–¶ Improving weather warnings and climate information dissemination in Africa https://www.youtube.com/watch?v=rtuN92AbRWA
ā–¶ Quick video tour of Wagtail 6.0 https://www.youtube.com/watch?v=_Vg_lPMipcQ

Wagtail future proofs your CMS system, as it’s open source, continuously updated and built on Python, one of the most popular global programming languages, used widely in machine learning and big data. So you’re always ahead of the curve when it comes to CMS platforms

Wagtail is the #1 choice for accessibility, is scalable and most importantly, secure.

šŸ‘‰ Get started with a FREE Wagtail CMS TRIAL: https://wagtail.org/get-started and see how easy it is to build a website that works for you.

šŸ“Š Read why Google, NASA, and the British NHS, are powering their digital estates with Wagtail: https://wagtail.org/about-wagtail/

šŸŽ„ More Wagtail Videos: https://www.youtube.com/watch?v=cne2kxemMAQ&list=PLfwZ-fob20cPvSQ_v1hkjto8BAPN21tLJ

šŸ“£ Follow us on social:

#WagtailCMS

Summary

Kōtuia brings together museum objects and stories from across Aotearoa, using Wagtail for editorial content, eHive for museum information, and DigitalNZ for collection items. Because the sources have mismatched identifiers and fields, uneven coverage, no reliable language filters, and restrictions on storing data, the team used a mix of API-driven pages, CMS-managed pages, and limited synchronization. The speaker argues that projects working with external data should identify what is reliable, accept practical compromises, and prioritize the solution that best serves users—even when it is not the most technically ideal.

Key takeaways

  • Kōtuia uses eHive for museum profiles and DigitalNZ for collection items, while Wagtail manages stories and other editorial pages.
  • Different identifiers, inconsistent fields, incomplete coverage, API outages, and data-use restrictions shaped the site’s architecture.
  • Hybrid CMS pages and limited synchronization helped make museum listings more usable and offered a place to add controlled or translated content.
  • With no dependable way to filter external data by language, the site puts bilingual material mainly in CMS-managed content and uses Māori for elements it controls on API pages.
  • The team chose to show reliable common fields and a larger set of items rather than wait for complete, uniformly high-quality data.
  • The speaker’s central lesson is to understand which data can be trusted and favor a workable user-focused solution over an idealized one.

Summarised automatically from the transcript.

Transcript

3,267 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:00

Hi everyone, my name is Shisela Delavisha and I work at Springload Topipitanga. We are a digital agency based in Wellington, New Zealand, and we believe in creating human-centered digital experiences that make a positive impact. In this video I'm excited to share with you how we combined multiple external data sources and work with them to fit seamlessly with the Wagtail CMS build as we go through the journey and challenges we faced while building COLTUA. But first, let me tell you a bit about ourselves. Springload was founded in 2002 and was the first digital agency to become a B Corp in New Zealand. Over the years, we've had the pleasure of working with a wide range of high-profile clients, including large banks in New Zealand and Australia, the New Zealand Red Cross, and various government agencies, and many successful corporations in New Zealand as well.

0:52

So our journey with Wagtail began several years ago and it has become our CMS of choice for dozens of projects. We've seen firsthand how Wagtail's flexibility and user-friendly interface can empower organizations. To manage their content effectively. Our team has expensive experience in customizing and extending Wagtail to meet the unique needs of our clients, ensuring that their digital platforms are both powerful and easy to use. So, what is Kotuya? Kotuya Ngakete is a museum sector sharing site developed by National Services Tepaerangi, a team at the Museum of New Zealand Tepa Patongarewa. The site connects museum objects and taunga

1:39

or treasures across Aotearoa, New Zealand with people and places through the use of stories and collections. Kotya brings together over 1. 5 million digitized tanga and artworks from over 80 museums, galleries and faretanga, which loosely translates to treasure houses. Through Kotya, you could, for example, connect data between a museum in the far north and a museum in the far south and tell the story of how their items share a common origin or significance. Kotya lives in a very special and unique context, that is New Zealand. For those of you who are not familiar with the country, New Zealand has a very unique relationship with Tangatamaori, the indigenous population of the country. Which gives it a very unique set of requirements when it comes to bilingual presence on the website.

2:28

We've actually done a short talk about that recently, so if you're interested in the bilingual approaches in Oteadoa then follow this link and you will be able to watch that talk as well. But mainly what you need to know is that there is a strong focus on Māori language revitalization And one of our preferences is to do first languages first. So native language should always come first. And as a national organization, Kotuya has a strong interest in showing as much multicultural, bicultural and bilingual content as possible. This is going to be important later on on two accounts. One, the Maori treasures, including language, need to be honored properly in itself.

3:14

And two, that language is also data and will impact how we interact with the APIs. So let's talk a bit about the external sources. In its essence, Klotuya is a museum aggregator. It pulls and displays information about museums and information about the museum objects. So what do you do when most of the content on your site comes from an external source? What about two? What about two that really should be linked together but aren't? First, we need to look at a problem beyond the technical solution. Kotuya needed to display information about museums, and it needed to display information about museum objects

3:59

This content had to be searchable, listable and as complete and consistent as possible. And we talked as well about creating stories and connecting Tonga to each other and the land. Then, as we mentioned earlier, all the content needed to be as bilingual as possible. If you were creating a system from scratch, you would probably look at having your own data models You would want to have your museum model that has all of its information, and you would want to have your museum item model. You would want to also have an authentication method to give access for each museum to manage their own data. That is not the case here. To understand the data constraints and architectures of the site, first we need to look at its history.

4:47

Kotuya is actually the rebuild of a platform that used to be called NZ Museums. The platform NZ Museums actually had the data structure that we explained before. It had museums and it had museum items, all managed through a platform called eHive, which each participating museum had access to. However, the platform didn't account for one key thing. Most museums were already managing their collections through other systems, and neither NZ Museums nor eHive were integrated with them. To revisit, Kotya is showing over 80 museums and it's showing over 1. 5 million objects. You can imagine how much effort it would be for each museum to duplicate their data entry.

5:35

NZ Museums was facing the issue that many museums Some of the largest in New Zealand even were only using e-hive to put their venue information, but not their catalogue. As a national museum aggregator, it was sorely lacking information. One of the goals for Kotuya was to increase the amount of objects that was being shown and leverage the existing data that was distributed across the country. Here is where DigitalNSET comes in. DigitalNZ is a search site for all things New Zealand It fetches information from libraries, museums, galleries, government departments, the media and community groups. So, DigitalNZ

6:21

as a platform was already scraping all of the collection management systems that each individual museum was using. And as it was exposing that information through an API, shaping the data to be mostly uniform, that was easy to consume. However, it did not contain information about the museums themselves, only about their items So, we have museum information on eHive and we have museum items on DigitalNZ. Now, the way that these two relate is more of a logical relationship We have the museums on EHive which have their names and their IDs, and we have museum items on DigitalNZ, which are tagged with the source.

7:08

There are a few things that we needed to note. Number one is that the ideas of the museums on eHav did not match the IDs of the museums on Digital and Z. Number two, that not all of the museums that are on eHive gave them permission to have their data scrapped and put into Digital and Z. Number three is that not all of the objects that were present on e-hive were present on DigitalNZ. This clearly indicated that whatever solution we chose, there was going to be some data loss and some compromises. However, DigitalNZ had the most up-to-date data. because it directly connected to the main platforms that these museums maintained. This led us to the mixed approach of using eHive

7:56

for fetching information about the museums. and DigitalNZ for fetching information about museum items. Now this solution obviously doesn't come without its challenges. You might remember a couple of slides back that we mentioned that on Kotuya we wanted to show as much bilingual content as possible. Now there were two factors that affected this goal. Firstly, that managing bilingual content It's really hard, especially with so many different museums and content contributors. Remember that we mentioned that the Maori language is a national treasure Unfortunately, at this stage there are not enough fluent speakers and content writers nationally to support all of the museums in translating their catalogue.

8:44

Getting bilingual content of a quality high enough to honour the language is a very difficult task. Secondly, neither the Digital NZ API nor the eHive API had a language filter. Now, beyond the human and operational factor of producing bilingual content, it was not really possible to request content in a specific language, making an API-based bilingual approach impossible. The second challenge is that since the information is effectively crowdsourced, each museum is cataloguing their objects with their own information and their own criteria. For starters, they might not be recording the same information, but even if they did, they might be calling it different ways.

9:31

So for example, you might have one museum that indicates that nature is a category, and you might have another museum that indicates that nature is actually a tag. And so that information is getting scraped and put into Digital NZ and served using the unified format. However, it might not be relating the same data to the same fields. The third challenge is that Digital NZ collects information not only from museums but also from other sources that should not appear on Kotuya. This content needed to be excluded from API calls in an efficient and robust way.

10:16

Finally, we're looking at system stability and availability The data protection constraints that we were subject to meant that we needed to use data from DNZ and EHive API without storing the data locally. While this option is straightforward to implement, it doesn't scale very efficiently. The robustness of the system depends on multiple data sources. So if one data source is down, people are shown incomplete information on the website, based on the information they're looking for, affecting user experience. For example, e have being down means people can see objects from museums and stories about the museum, but not information about the museum itself.

11:04

So we're getting into what we do with that information and with those constraints and how we actually manage what we can and accept what we cannot. So we did a couple of things to actually weave this data into the CMS. Remember that as the CMS, we of course want the experience to be cohesive. But we have Wagtail CMS for handling site copy and CMS's default search for the site content search Digital Z for me museum item this includes searching, filtering items, listing, fetching item details. We have eHive for museum profiles also including searching and filtering museums, listing them, fetching museum details

11:52

As a bonus, well, I mean this allowed us to reduce retraining of existing contributors as museums were already updating their information through e-hive. So there's that and let's not forget the variety of languages. So getting down into the technical details We've taken a few approaches to make sure that this was as consistent as possible and that all felt like part of the same site. First of all, we have two different types of pages. We have pages that are API based, so that'll be the item pages, the museum listing, and the museum pages mostly. And then we have purely CMS manage pages. That would be the home page,

12:38

stories, which allow us to weave the items together and talk about the relations. And the collections, which are ways of grouping items that do not belong to the same museum necessarily So what we've done is a few things. First of all, we set up API-based item pages. So for that we have Set up a route pattern that's handled by a standard request function, which makes an API call to the DitanNZ to fetch the details related to the item ID found in the route. and then passes it as context to the rendered page template. We have also set up hybrid way of having pages, which is the case for organizations

13:23

and museums. These exist in the CMS but also fetch most of their data through their API. The reason why we have the museums as pages as opposed to as purely APIs based is to improve performance of the visit page. The eHav API has rate limiting and low performance when listing, as it returns the full information about each museum. This was making the visit page very unusable. When synchronizing museums were only storing the data necessary for the map view to operate. And we run the synchronization task monthly so that we can stay up to date with changes. An approach like this, this mixed page hybrid approach, could also be used for augmenting API data.

14:12

So you could keep API data synchronized and then have additional fields in the page model. This could be useful for adding translations, for example. We have also created stream field components that connect to the APIs. So, for example, you will see that in many of the content pages you can introduce either a museum component or collection item component. In those we indicate the ID and that makes a fetch from the API and embeds the content in the page. And finally we have the search solution So this solution is quite interesting. Obviously as a user you would want to search through the largest set of information. So that is the collection items.

14:57

You might also want to be looking for a museum or you might also be wanting to search for some content through the site. So when it comes to the search, you will see that we have separated it in three sections. This makes the experience logical. for the user, but under the hood it's really using three different search endpoints. Both eHive and DigitalNZ provide their own endpoints that can be used for querying And then we are using the standard out-of-the-box search for Wagtail Content itself. We have unified the user experience by giving the users the possibility of searching across all three knowledge bases or selecting which one in particular they want to look for.

15:47

As with any project, there were compromises that we needed to do for the rebuild, particularly around data and architecture Here is a comparison between the old NZ Museum site and the new Cotuya site. In green, you will see the fields that are presently being shown on the new site, and in yellow, those that have not made it across If you remember, at the beginning of the talk we talked about NZ Museums and EHAV having data structures representing museum items and a uniform UI that gave all contributors the same fields That uniformity is not something that we have on Digital NZ. Again, because different museums use different collection management systems that have their own data structures. And the Digital and Z

16:33

team did their best to map the data and make it all fit into one common data structure, but you can see that there are quite a few fields that are not reliable. So what we did for that is we highlighted what information actually was across all of the museums. We identified which fields were consistent and were bound to have information quite often. and we prioritize putting those into the final UI. The final solution might not show all of the information that the different museums have available. But that is fine. If you want to see more information, you can go to the museum's original listing, and if you want to dig deeper into that particular museum object, you can This is still meeting the final goal of actually showing the data together, joining it

17:20

and being able to tell stories about it. Beyond the technical constraint, there were also other limitations around the data, such as licensing and the lack of permission to store and modify collection item data. This impacted more than anything the search. As we mentioned, we are using three different search endpoints. Those endpoints provide different capabilities, which make it really hard to have a consistent search experience Ideally, we would not be doing that at all. In an ideal world, we would have created a copy of the data in the CMS and synchronized it with the different data sources We would have also built a robust search engine that indexed all of the data to enable more advanced search functionalities.

18:08

This approach will be revisited when we work on finding ways to improve the search experience later on. Another compromise that we've needed to do was in the language management of the site So obviously, if we couldn't trust that the collection items were in one language or the other, or that the museum information was in one language or the other, we needed to assume that all of the external API content was in English. Now, back to the bicultural approach in New Zealand, we wanted to introduce as much bilingual content as we could. And the compromise that we needed to do was to actually put as much bilingual content as possible in the CMS components only. So, we have things like fully bilingual content pages, components that have hybrid content that is in both languages, and the cultural elements such as a MIHI or a welcome.

19:05

For items and museum pages, we are using Tereo Mauri and the elements that we do control, which in our case limits itself to the static titles Finally, another compromise that we needed to do was to accept which data we could show and which data we couldn't So we are choosing to show a limited amount of data per item and go for more quantity as opposed to higher quality This is not normally an approach that you might want to choose, but in this case it was really important because we gained so much in amount of new collection items that it really makes the site richer. And it really allows the team to tell stories that are deeper and more deeply connected to the country.

19:52

Now this project was a journey and it left us with a good number of learnings that we'd like to share with you. Obviously, as you will have seen, this is a story of compromises. It is not a story of the best solution ever. As a story of compromises, one of the things we have learned is primarily that your data might not be perfect. And that is fine. What is important is for you to identify what data you can rely on. This goes not only for a problem like this one, but for any data problem at all. You need to make sure that you have reliable content that follows a certain structure, that is clean and follows the same terms. We have also learned that even if you don't have all of the control over the data, you can still work around it, particularly through the use of external elements or through augmentation.

20:45

We learned that sometimes you need to make some compromises in favor one dataset, and perhaps not the most optimal technical approach, over another in order to give users a better experience. Sometimes the best approach might be mixing many approaches. This goes both for the bilingual elements and the data integrations, especially when so many diverse data sources are involved. That might not make me your codebase the most maintainable out there, but your users will benefit from it Finally, find what drives the success of your project and prioritize accordingly. A working solution is better than the best solution. And

21:31

that's it for me. Thank you very much for watching. If you have any questions or suggestions, even let's connect. You can find me on LinkedIn or you can go to the Springlow website Or you can find any of our awesome div and operas on GitHub. We hope this was useful for all of you and thanks for watching

Questions this talk answers

How do you combine data from eHive and DigitalNZ for a museum website?

Kōtuia uses eHive for museum profiles and DigitalNZ for collection items, since eHive has the museum information while DigitalNZ provides broader, more current item data. Because the sources have mismatched IDs and incomplete overlap, the solution accepts some gaps and compromises.

Discussed at 7:56

How can Wagtail work with external APIs while keeping the site manageable?

The talk describes API-driven item pages, hybrid museum pages that store only the data needed for performance, and StreamField components that fetch and embed API content. CMS-managed pages handle editorial content such as stories and collections.

Discussed at 12:38

How does Kōtuia let users search across Wagtail, eHive, and DigitalNZ?

It presents search in three sections, each backed by the relevant endpoint: Wagtail’s built-in search for site content, eHive for museums, and DigitalNZ for collection items. Users can search all three or choose a particular source.

Discussed at 14:57

How do you handle bilingual content when external APIs don't support language filtering?

The site assumes external API content is in English and puts as much bilingual material as possible in CMS-managed pages and components. For API-driven item and museum pages, te reo Māori is used for the elements the site controls, such as static titles.

Discussed at 18:08

What should you do when external data is inconsistent or incomplete?

Identify which fields are reliable and consistent, and prioritize those rather than trying to show everything. The speaker also recommends working around limited control through augmentation and choosing compromises based on what best serves users and the project’s goals.

Discussed at 19:52

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos from Wagtail CMS