Running a multi-site newsroom in Wagtail

This video features Ryan Verner at Wagtail Space US 2018 in Philadelphia, Pennsylvania, USA.

0:41:50
Published June 21, 2018
740 views

Summary

Ryan Verner describes consolidating several Australian financial-news sites in one Wagtail instance, using shared page models, templates, validation, structured metadata, custom image tools, compliance previews and multi-site URL helpers. Wagtail improved consistency, deployment, testing and structured content compared with the former WordPress setup, but required custom work for legacy URLs, permissions, AMP rendering, image renditions and performance. He argues that newsroom workflows are much more complex than draft-and-publish: journalists need comments, readable revision history, tracked changes, offline drafts and configurable approval states, while many still prefer Word or other writing tools. The presentation ends while comparing Wagtail StreamField editing with a rich-text workflow that journalists found easier for copying and pasting formatted documents and images.

Key takeaways

  • Several brands can share Wagtail models, templates and logic while retaining site-specific fields, URLs and publishing rules.
  • Structured article metadata supports compliance, reporting, search aggregation and social sharing more reliably than the former WordPress setup.
  • Custom validation on publish, external compliance previews and editor-only publishing helped enforce financial-news requirements.
  • Performance problems from image renditions and rich-text content were reduced by prefetching related data and avoiding unnecessary queries.
  • Newsrooms need more than draft and publish: configurable approvals, comments, readable revisions, tracked changes and offline drafts are important.
  • Journalists often prefer Word or similar editors, so StreamField structure must be balanced against familiar copy-and-paste workflows.

Summarised automatically from the transcript.

Chapters

  1. 0:00 Introduction and Newsroom Context Ryan Verner introduces his background, Stocks Digital, and the challenges of running a small financial news operation.
  2. 2:24 Multi-Site Newsroom Structure An overview of the publications, article metadata, promotional features, and shared Wagtail instance.
  3. 5:29 Why Wagtail The reasons for replacing inconsistent WordPress sites with an explicit, Django-based CMS.
  4. 7:49 WordPress Migration and Site Architecture The talk covers content migration, data cleanup, URL compatibility, and the inheritance model used across sites.
  5. 10:50 Production and Publishing Features Wagtail supports frequent releases, testing, alternative renderings, structured metadata, and compliance previews.
  6. 13:53 Custom Editorial Tools Custom image embedding and editing tools improve how journalists search, upload, caption, link, and resize images.
  7. 16:58 Wagtail Implementation Challenges The speaker discusses rich-text cleanup, URL handling, deletion restrictions, and AMP rendering problems.
  8. 19:17 Multi-Site and Performance Optimization Site metadata abstractions, absolute URLs, dynamic templates, image renditions, and SQL query reductions are explained.
  9. 23:59 Newsroom Workflows Complex editorial, photography, legal, client, and publication workflows expose gaps beyond Wagtail’s draft-and-publish model.
  10. 28:35 Revisions, Comments, and Change Tracking The talk examines Wagtail’s revision support and the commenting, track-changes, offline drafting, and configurable workflow features journalists expect.
  11. 31:43 Newsroom CMS Alternatives and StreamField Demo Existing proprietary and large-publisher CMS approaches are compared before demonstrating the newsroom’s Word-based workflow and rich-text editing approach.

Transcript

7,797 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:00

Speaker 1: Uh a newsroom in worker.

0:03

Speaker 2: Hello. Is this working? Yes. Excellent. Hi, so despite getting two days ago, I'm still really crazy jet lagged, so this is gonna be fun. Um and that's not working. Awesome. Yeah, beautiful. So it's unlucky any of you know me. Um you may have heard of uh Next Day Video. So I run the Australian side of Next Day Video. We do lots of uh conference recordings so I record a lot of the PyCon Australia, PyCon New Zealand, all of that sort of region Python conferences. I've also been a previous organiser of PyCon Australia. I currently work for a company called Stocks Digital. Stock Digital is a Australian company that writes about

0:48

Speaker 2: the Australian stock markets, the kind of companies, activities, stuff like that. The equivalent of it's ASX, it's equivalent of your Nasdaq, basically I previously worked for a company called Gizmag, which we named New Atlas. You may not have heard of us, but you probably read an article. So Gizmag and New Atlas write about science, tech Yeah, improvements to humanity. Very occasionally an article will go kind of nuts. So think sort of n gadget, gizmodo, it's the same kind of site. That was a Django-based CMS. That got about 20 million NICs per month. But that was written from scratch. That was my previous job about a year ago. This is a very short talk about building a new site in WebTowl. Um everyone's experience is different. Um the company I work for uh writes about the Australian stock market, which is a very different sort of newsroom than your traditional sort of newsroom that might be writing about stock market.

1:38

Speaker 2: something different. So some of the experiences we've had are going to be the same, so going to be wildly different. We're very, very niche. Yep. Um and new sites are kind of interesting because they seem simplistic, but once you get into them, there's lots of parts that aren't obvious that you need to think about when you're uh developing and working on a new site. And obviously the challenges of a small newsroom is going to be different than a big one. So Stocks Digital is fairly small. I think there's about eight or nine journos total, some on contract, which is very, very different than a newsroom with you know like dozens or hundreds of journalists. So again I experiences be a bit different than a larger neutral room, but there's certainly things in common. So this talk is really about our experiences using WagTowl for a multi-use site instance. use instance

2:24

Speaker 2: um and the challenges we have moving forward things that webcloud currently doesn't do that i would love for it to do We have a number of publications. I'd imagine no one here is a speculative investor interested in the Australian stock market, so we won't talk about the brands, we'll talk about the actual tech. This is the end result of one of our sites. So if Finfeed uh is one of our sites, that's all Wagtail. So a newsroom, a news site typically has something like this. And we've got three different brands or three different sub brands, um sorry, six different sub brands. You typically have things like a section, a subscribe panel, very important for new sites to get subscribers, you want to promote articles, stuff like that. Articles will often have a hearer image, so each article has a

3:10

Speaker 2: required image that basically displays in lots of places. So things like social media, you know, your main page. If I can alter um That's the site here. So these change effectively and that's all dynamic. based on your image uploaded into the WagTowl. And that's all dynamic. Other things a new site will have is a way to pin images at the top of the site. So these are basically defined by the editor as in the actual articles they want to promote primarily. And that links to other things like our newsletters and stuff like that.

3:58

Speaker 2: Sir. Um, that's not a cool page Um so an article will have a bunch of metadata, some which is obvious, some which isn't. So you'll typically have a headline, an author, a date, a section or category. Sometimes you'll have multiple categories You have a here image, which I spoke about before, so that one image is used everywhere on the main page in Open Graph tags or social media, various other places where news is aggregated. Special flags, things like editors pick. uh trending stuff like that are typically flags on an article. Body, you've got things like embeds, images, stuff like that. You've got a published date, which might be different than the last modified date on the article You've got a company. So a lot of the stuff we write about is about a company, so we want to keep that metadata for reporting reasons.

4:43

Speaker 2: Most traffic to a news site is typically either promoted or through aggregated through something like Google News. So you actually want to capture people's attention. Typically they won't go to your front page and then go to a sub uh an article. They'll generally land directly in an article. They'll read it, they'll close it off. So you generally do things to try to keep them on your site. So this is a um little widget that basically grabs uh there's a little bit of logic in the back end that grabs another wagtail. um article within a certain criterion on the site and that changes in every load. Sure. Um so we have a bunch of different brands. Um each brand basically lives in the same way instance and we've done things like uh shall go to my next slide, um use inheritance. So we have a uh uh article uh model which is inherited from page and we do things like common templates, common logic.

5:29

Speaker 2: Uh basically means we can make changes to one side of reflex and everything. I'll talk about some of the challenges in doing that. Next investors doesn't show a little folder because that's basically just a standard Wagtail site. That's sorry, standard Django site using UralConf. So nothing particularly special there. Combined traffic to these is about 350,000 unique visitors a month, so not that big, but enough that there's certainly performance challenges we've had to overcome. Why Wagtail? I joined the company about a year ago. They currently had all their sites running on WordPress and every site was set up differently. um ultimately be consistent but there's lots of plugins like gravity forms and other bits and pieces that just trying to get structured data out of it is a complete nightmare Yeah. Um I wanted a Django-based CMS, slightly biased, been a Python developer

6:17

Speaker 2: for like 15 years. Uh used a Django-based CMS in my last job You know, I wanted I wanted to be able to write a CMS that was explicit. Um and I'd used WagTab previously in a previous job, not for a new site, it was for more like an educational site, but uh I certainly understood its uh pros and cons. Um We are writing about the financial sector, we actually have legal compliance requirements. be obviously kind of hairy if you do the wrong thing there. So we want to enforce that certain thing that an article goes through a certain workflow. We know that things being published or legally correct, there's no incorrect information, stuff like that. Doing that in WordPress is kind of interesting. Obviously writing around explicit CMS in something like Wagtail, we can enforce this stuff. The editors will define a style guide.

7:02

Speaker 2: We can start enforcing some of that stuff to the CMS as well. So all sorts of more blue sky sort of thinking but stuff we can do. And we also need a report and some stuff too. So we have some internal reporting tools based around what we're doing to kind of give the editorial team some focus. Again, trying to get structured data out of WordPress is a bit of a nightmare. Doing it in Wagtail is much easier. And my grand plan is to bring journalists into the CMS earlier. So typically when a journalist coming from look at some more of a recruitmentality, write a piece, they're thinking about the actual piece itself, not thinking about metadata that's important for promotion. Things like the small exercise that will appear in search engines, the images used, you know, things like tags, things are really, really important when you're promoting your content that if you're writing an article outside of a C mesh you're not necessarily thinking about

7:49

Speaker 2: Um so bringing a journalist in earlier into Wagtail hopefully would uh assist with that. And uh having bought CMSs from previous new sites, um I know that works Building the site. So we took about a month and a bit to build the first site, which was a lot quicker than I expected, which was nice. We looked at a few things. I originally took all the images from WordPress. I uh dumped out the database and it was in interesting and every instance was different. So we took the other um mentality, but actually WordPress has a REST API that gives you kind of the rendered content. So basically using that iterating through each article, doing some data synodization, doing lots of data data. data sanitization, uh, because it was all over the place. Um, and doing things like taking the images and actually creating a

8:34

Speaker 2: wait till image and storing it all properly. So the data's really nice and consistent. And a lot of our sites might have information about the stock code of the company, you know, how they're performing, stuff like that. So actually using Beautiful Suit to actually pass it out and keep it a structured data inside Wagtail is really nice. So we can do smarter things with our data rather than just being content you know in a blog basically. Um we did some funny things, so every site was set up differently. So like you We had a section slug and an article slug, some had a date, some didn't have a date, some sites had subsections, some didn't. So the default logic of Wagtail was typically to follow your structure you define in your UI and all generate the URL like that. We ended up having to override URL path conditionally. And as we added each site, I think the first site took about a month and a half, the next site was about a month, the next site was about a month.

9:20

Speaker 2: And most of that wasn't actually build because we're using the same templates and just building on top of it. It was uh dealing with kind of trying to do it elegantly in a way that it could actually um have multi-site um multi-site functionality way you could actually make changes to the base site and have everything reflected. We really have complicated this. So the set URL path ended up being this crazy thing I'll show I'll show you guys later. We're a vertical page at the time, that would have made our lives much, much easier. Cool. So after lots of attempts we end up with this basically. So we hear it from Paige, an article, I it's red and green intentionally, red code, green code. Article has a whole bunch of logic in there. Things like

10:05

Speaker 2: site-specific fields, custom validations. We have things like a hearer image and some metadata required on our sites to ensure that we meet compliance and the article looks nice. Journalists complained when they were doing drafts inside the editor that they had to put this data in their NAT images before they could actually save a draft. So we do things like actually define um arbitrary validation um on save, sorry, on publish rather than save. And putting it up there is really nice. Everything else gets it for free effectively. We also do that so we can lock down where the pages appear. So using parent page types and subpage types, we can actually say which site a article has to appear on each so we don't end up with weird articles and weird spots. And having it in here, it means that we know, for example, we're using an API to actually talk to an Excel reporting suite.

10:50

Speaker 2: We know there's always going to be certain fields always available. Yeah. Um in production, so um any of you using Webtime production didn't number this is a surprise, but uh coming from WebPress this is a huge difference. We can the ability to ship multiple changes per day. So previously the team I think will do a release every two weeks and hope would break. We just release stuff you know as it gets approved for the PRs, which is really nice. We've got a full CL of the test suite. So again tests are unheard of previously. We have full test coverage now, which is well full-ish. as it always is. Um using Docker and ECS. So everything's containerized, which makes things nice and easy. We're also doing alternative page renderings. So things like AMP, Apple News, stuff like that all comes out of Wagtowell. Um and we also do things like J JSON LDJSON open graph

11:36

Speaker 2: tags. So you guys aren't familiar with this, um what's important when you're promoting articles is um Sorry, I should sorry, the other site, so that's Finfeed. Let me show you this before I forgot. Um on the site. Like so images, data. Actually this one doesn't have one, so go figure. Interesting article I picked as well. Uh that one. Yeah, so an image. So FinFeed articles are very different. So typically they'll basically be generally unformatted text, which Makes this easier and harder in some ways, but I'll run through that later again. We have links, all the usual

12:22

Speaker 2: kind of stuff, images, headings, stuff like that. Catalyst hunter, again, that was the first site we converted over. Very, very simple. Heading There's a metadata up here which will normalize that when you converted it over. So the company named their stock code, you know, that sort of speaks about the kind of risk involved in the stock. These are really, really important. So on our assets we have these things that get either pushed in. So we've rewritten the rich text uh template tag to actually arbitrarily insert these into certain parts of the content. We don't like pop-ups because they're horrible, but obviously making um getting people's information is really really important in terms of um promoting content. And this arbitrarily inserts it based on some rules. and the other side as well. So again, very simple stuff on the front page

13:08

Speaker 2: and dip cross-promote sites. Um to show you the OEG tags which I haven't put up that I had. Ah, here we go. Great. So these are really really important. So you've got things like um open graph tags, so things like Twitter, Facebook, et cetera, use these. And you basically define things like when an article is published, you know, the URL to it, the actual URL. the image to use, uh the width and the height, stuff like that. So if you want your stuff to appear really, really nice on social media when people share your link, you need this sort of stuff in it. What else we have also the LDJSON. So this is a spec that Google released that basically gives you, it basically gives Google better ability to index your content and other sites that support This

13:53

Speaker 2: as well. There's a huge spec on it. It's massive and ridiculous and not many places implement it fully. But we've found this has been really key to actually get our stuff aggregated in more places. Again, this is sort of stuff that's not very obvious, but you really want it for a new site. We've added some custom functionality. So we want either we generally want X in a review by a compliance team. So we've got a little thing there we've built that basically generates a UUID for each article based on a revision and we basically send that link to an external post at LASM to review the actual link without actually having to log in, which is really nice.

14:42

Speaker 2: Which is something we haven't heard before. They used to email Word documents, which is very different than what you see in the actual site. And legal compliance wants to see how it actually looked on the actual page itself. So this has been really nice to do. We implemented our own draft tower image element. So we migrated the draft tower probably about a month ago. One of the things we discovered was right now draftar doesn't give you the ability to actually have an image with a link. I believe it's due to atomic blocks, but I think that might be just not implemented yet. We also had some complaints from staff that adding images was kind of clunky do I didn't agree, but I took their comments since we end up building this basic building in about two or three days. So if I can show you, it's effectively a custom um image embed that lets you

15:27

Speaker 2: um search existing images um and I played on with them with very minimal clicks basically. There's one else. So if you drop in an image, replaces the um the standard image embed, an image, this is all React. with a single API added. You can search existing images, you can upload a new one, you can also paste an image as well, which is I'll show you guys later. Select, give it a caption, which is required, give it a link optionally, and resize it. So this is something they

16:13

Speaker 2: um writers were saying that quite often they'll uh author content. And uh they didn't know how wide or small an image would appear. So we built this thing that basically gives them the ability to see how wide or narrow an image is. Which again really simple stuff, but it just helps the journalist actually author content in a way that's really palatable when it's published And again, this is really convincing of drivetales the way forward, because this took about two or three days to build total. And that was with no experience previously, which was nice Satan your own one Cool. Um and now that the staff know we can add new custom functionality fairly quickly, the issue I'm now dealing with is kind of keeping back the requests

16:58

Speaker 2: very different than before, which is nice. Um yeah, run through that. Um bumpy bits. So hello. js was a bit of an issue initially. It wasn't bad, but we tend to find, and I'll run through this later, but quite often our journalists will actually paste in um content from Externally sourced word documents, things clients sent them, you know, various kind of sources that kind of have really bad formatting. Um and they pasted into hello. js and used to do weird things like wrap everything in a h5 tag, which you couldn't remove You could actually lock down the rich text editor to be like, these are the only the elements we want to do. So it only gives us an H2, a H3, a link, stuff like that. But it would actually let you paste in stuff that you couldn't remove. So we end up basically on the save method for articles used to do lots of data standardization. Just we had to. It's the only way to get data in there effectively. With Driftile, we've written a lot of it out.

17:45

Speaker 2: Like it's just better out of the box, which is all Awesome. Um the custom URL logic, so again I think that was a snuffer and error in, but we end up doing lots of custom URL handling logic based on, you know, we had different URLs from the start of the sites we converted back from. Again, we didn't know about writable page, and that probably would have made a big difference. Things like uh yeah, so we want to prevent articles from being deleted by either editors or writers. We want to be able to unpublish a page, not delete them. And that was something I would expect to be pulled it out of the box. We couldn't find out how to do it. That doesn't mean we it you can't do that. We just couldn't figure out how to do it. So we end up basically doing a bunch of CSS hacks and a bunch of permission hacks. The primary problem with that was I think you could disable it, but you the permissions weren't granted you enough. So we wanted to do it based on specific groups and I think it was either a global or a wealth of recall correctly.

18:31

Speaker 2: And um things like AMPR with alternative rendering. So AMP is a is a, I'm not going to say subset of HTML, but it's an alternate version of HTML that Google and use, for example, will use to promote your news piece and it's done in a way that um is uh that's not true it's also used normally as well. So when a mobile phone used an amp page, an amp page basically is a subset of HTML that basically is designed to be rendered very, very rapidly. So lots of things are not allowed. Um and uh with my previous job we stored stuff in Markdown or similar. Um so we just basically had a different sort of uh transformation. What we do with AMP is basically take the uh HTML that's stored inside the um um inside the database and basically munge it, which is not ideal. And occasionally someone in manages to put a tag in that's not supported and it all blows up.

19:17

Speaker 2: So it'd be awesome to fix that. The other thing we had a challenge with is multi-site. So this is something I've dealt with a few times now. So Domain's typically environment specific, so if you set up Wagtow sites, it's very based around the host name in the port. And if you move it to a local dev environment or somewhere else, it'll go It also doesn't give a canonical university accessible slug. Actually, canonical is incorrect. I should have used uh what have I used? Uh semantic, sorry. So basically you've got a PK or domain that port and a human readable name. They can all change. So if you've got bits in your code that are conditional based on that and you change your site name or the domain or something like that, they all break. Conditional views, conditional logic inside of views can get kind of hairy as well.

20:05

Speaker 2: Because again, you're checking things like either a PK, which is not very readable, what does that mean? You're checking for like a human readable name, you're checking for a domain name. Again, it's not a really easy way to do it And it's a limit helpers as well. So one of the biggest challenges you get is get an absolute URL on Django gives a relative URL. So you're doing multi-site, like what site is this for? What's a domain name? You end up with all these really sort of hacky kind of solutions there. So we've got a thing called site data, which basically gives you the ability to define a semantic multi-site setup. So you use a label and everything comes off that. I'm not going to go to that too much, uh, because the talk's not about that, but it it basically provides a whole bunch of um uh convenience methods, things like reverse get absolute URL just work and give you the actual full URL. Um we define it in settings.

20:51

Speaker 2: py. The reason why we do that is um you generally aren't changing your site data very often, having an RM hit every single time. is kind of bad because you don't put an HTTP request and then there's ways to solve that but there's caches and then you've got more problems. So um this gives you a way to basically get arbitrary metadata about a site into a settings. py which is initialized at Django runtime And then you can basically access that. So site data comes up as request. site data that gives you a lot of this data is basically as attributes or properties and also provides a bunch of convenience methods as well. So in your temp in your um in your templates you can do things like uh request. site data dot social links dot Facebook, which is way more readable than having conditional logic in there. Um also have a have a have a template tag called URLFQDN

21:39

Speaker 2: which basically wraps around URL but takes two um attributes takes either site data or URL conf. So it basically lets you create a full URL for the entire domain Which again is much, much nicer than the normal ways you do it. You can also do things like SQL use, class-based views or pages without having to define one for every single subsite as well. So there's a mix-in, well it might mix inherited, but yeah, you can basically define a dynamic template name and things just work. I should release this at some point. But yeah, this is a a pattern that works really really well for us. Other challenge I've had too was um I don't know if else has hit this, but there's a challenge we've certainly had is renditions. Um lots and lots of SQL

22:25

Speaker 2: requests. So we had um I think this is an article with like something like 30 or 40 images in it. And of something like 80 SQL queries because it gets oh no sorry, this is at 55 SQL queries, so 15 images inside the rich text field. So I did 15 images with 15 image renditions. Um Django has some nice tools to do that, like Select related and prefetch related. We do some stuff which again there might be easier solutions for this. I'd love to hear from people. But we basically uh we hijack the rabble surf pattern, implement our own get children for render, which basically does things like select related and pre-trick related We also store the PK of the images on the actual page or article model directly as well, which again just means it's much, much more efficient. Again, um

23:10

Speaker 2: there might be better ways of doing this. Maybe I misunderstood something, but we were just finding that every single HTTP request to uh our Wagnetile sites was resulting in lots and lots of uh SQL requests. And even though we can wrap do things like lazy views and template fragment caching, you add a cache and you've got more problems. So I tend to avoid caching unless I absolutely have to And databases actually fast surprisingly. So after those optimizations, we get 22 queries in about a quarter of the time, which again can optimize further, but that was just fewer fewer through a few tricks which I just spoke about. Something a lot of people don't look at, but Django debug toolbar is your friend. Looking at SQL queries is a really useful thing to do.

23:59

Speaker 2: Right, so things we haven't solved yet. Um So I wanted the journalists to use the CMS directly. We made a really nice looking, I'll show you a picture. We made a really nice looking um editor. Again, they'll use the WordPress or other systems that let you do absolutely anything but formatting. which is a double-edged sword, because you can do anything you want, but you don't want them to do anything they want. So you know, we had a lockdown, so only H2, H3s are allowed. Um, you know, i i really easy to use editor, effectively. And we gave it to them and they're like, nope. Which was really surprising. Um these are some features that WagCloud doesn't currently have. Um and we um I thought this was a solved problem because at New Atlas we actually built a custom CMS editor and all the writers used it. And I thought, fantastic, everyone's using the CMS early, they're writing pieces in it, they're getting it

24:46

Speaker 2: reviewed. This is a solved problem. It turns out it's not a solved problem. Lots of journalists write in things like Microsoft Word and copy and paste. And no matter how much you try to get them out of Microsoft Word, they love Microsoft Word or some other tool. And a run through as to why that is. It's a little surprising to me. So um so a simple editorial workflow um which is probably what most people think would happen was the writer submits an allocal onto the CMS Editor reviews it, might make some changes, and then they publish it. That's probably the most simplistic uh workflow, um, which usually doesn't reflect reality. Um every workflow, uh sorry, every newsroom is different and they often got insanely complicated.

25:31

Speaker 2: workflows. There's another one. So you've got a writer submits an article, you have a photographer who submits images separate to the writer. An editor reviews that a chief editor will double check it and then they'll publish it. This obviously doesn't include things like data going back to the writer as well. So again, this is a very, very simplistic kind of example of a workflow. Another one, so NIDA might assign a story in a GDA to a writer, and for example, we have a legal compliance So before anything is published, it has to go through a legal compliance. If they knock it back, it goes all the way back to the site again. Another one, you know, I can go on and on and on. And again, this is simplistic. Um some news and workflows are crazy complicated with lots and lots and lots of steps. Doing this on Wagtail is kind of difficult because you've kind of got the idea of a draft

26:18

Speaker 2: and then a publish. All these other states are not currently supported Yeah, there's another one. So this might be an abattorial piece, for example. So it's something effectively it's like a written piece that's uh promoted. Um so at the top of the legal compliance for example, it was a a financial article, you also might want clients to approve it as well. And again if a client rejects it, it might go all the way back to the start again. So crazy sort of number of steps there. These are the features that write and editors want in their in what they write into. They want the ability to have revisions, so go back and actually view all the changes to an article. They want to be able to comment in an article. So an editor might go through and be like, can you change this? You know, can you reword this? Stuff like that. They want to be able to track their own changes. Now that's what this one was enlightening for me.

27:04

Speaker 2: It's always so track changes a bit like the way Google Docs doesn't. Like I want to track the changes that someone else has made to my document. Lots of journalists will basically go through, write a bunch of paragraphs, go back, change it, and they actually want the ability inside their um their editor to actually do this and see their own changes as they're working on it. They want an offline safe draft, so if they're writing drafting to a CMS and they get knocked offline, they don't want all their articles, all their work to go disappear. And you want to be able to have configurable workflow rewards. So um Like I explained, having a simple workflow, adding more steps is not sufficient. Every newsroom is different, so you want to be able to have different approval steps based on the requirements of the newsroom. room and uh with us for example we've got six different publications um and each one of those has a slightly different compliance um and review framework

27:49

Speaker 2: so even for us even on the same instance we'll have different rules. They don't necessarily need or want real-time collaborative editing. That's a cool feature, but it's kind of hard to do and a lot of journalists don't actually want that they want clean versions of every revision. And a newsroom schedule which is part of our newsroom workflow is out of scope of workflow. Wagtow. So where I see Wagtow comes into it is basically putting content in or putting content together. Working out when that content goes in is sort of out of scope. What does Wi-Fi do currently? So we have basic revisions. So we've got that. So different revisions of a single piece, you then click. compare with previous revision and you get that. So that was me, uh which is the really cool thing about this though is it actually gives you um

28:35

Speaker 2: it tells you exactly what changed in any single uh field on the model. But that's not very readable though. That was me, I think, changing two words and leaving one paragraph. And it's just that's hard to read. Um Google Docs, for example. Um I'll take it off. It's going off. Um Google Docs are something like this. So um you can basically click on every single revision. I mean a lot of people probably use Google Docs if you're familiar with this, but you can click on every single revision there and see exactly who changed and what changed. Something like this for Wagtail would be awesome. Again, the states are fairly simplistic. So you've got drafts scheduled and live only. Um which again works for um lots of content strategies, but uh for a newsroom it's not sufficient.

29:22

Speaker 2: Um what we can do at the moment, which we do for some of our uh for some of our sites is we actually lock down publish on the editor's group only. So we haven't enforced kind of like the the writers will put it into the CMS. They can't actually make it go live or publish it. The writer sorry the editor actually has to look at it. But again feedback is kind of hard. Feedback is done either in person or by email. It's very decoupl process. There's some parts here from WagTile core models. py. I'm not a core developer, so I don't know, but from what I can see that's kind of the understanding of the current state of the page at the moment. There are ways we can build this that are kind of hacky and keep states separately, and I don't really want to go down that path, so I suspect doing this well will require some changes in core. But I'd love to talk to other uh developers about the best way of writing

30:10

Speaker 2: Commenting is another feature. So being able to comment on a piece, or not necessarily making changes, but being able to comment on a piece and say, hey, you know Let's look at changing this and have a discussion around it. That's something else that writers use a lot. The challenge with Wagtail though is it's like A Google Doc or a Word document is kind of like a big rich text field. When Excel has other fields and you want to be able to comment or add revisions to lots of different fields as well. So there's a UI challenge here in how we do this But again, this is something that I believe a lot of people would gain benefit from. Tracking changes. So that's one of our documents. You can see how many changes there are. There are crazy amount. Being able to review that and see exactly what's happening is useful.

30:56

Speaker 2: Also, track changes are something that riders will do even on their own piece like I expect. So quite often I might write a piece out, that paragraph doesn't seem right, I want to go back and rewrite it again, they'll go back and rewrite it again. Oh, you know, I actually like that. So they should want to go through kind of piece by piece and actually approve or or deny their own changes even. Also you might notice there's images inside the document. So um this this workflow of writers using Word and then copy and pasting into the CMS is something that I don't like, but it's something that happens a lot. and lots and lots of newsrooms. There are other tools out there I've found or used in the past, but they're very, very decoupled. So basically the kind of these cloud suites or proprietary Suites that will aid a journalistic team through authoring content, approving content, and then putting to the CMS

31:43

Speaker 2: as a distinct separate step, either done through an API or something like that. Um what open source solutions already exist? Not many. Um There are some proprietary cloud tools that basically will help again a team pull together content. Um but it basically means no one's using WagTaz Excellent Admin. And I know we certainly use it because we can do sorts of Awesome stuff and uh using Wagtail not using its admin to me seems like a real cop-out. Um it's also unlikely to support custom embeds, so we've built that custom image embed for example. And the way that other the way these suites work is they generally will pre-render the HTML and stick it into the the the database of the CMS. Um so it's very very unlikely that one of the powerful things about Wagtail is a bit of

32:29

Speaker 2: customized stuff. Uh you would lose all that, which is not good. Um bigger newsrooms often build their own CSM workflow. So this seemed to me like it should be a solved problem, and it actually kind of is. Um Some examples, I mean these are from 2014 so they're four years old. Um but the Guardian bought their own uh CMS and it's got a lot of this stuff there like track changes, commenting, stuff like that I found many, many examples of this. It seems like lots of large newsrooms, Buddha and CMOS is kind of like their competitive edge. It's their way of authoring better content effectively. And even though parts of it are are open sourced, I'm not seeing any full solutions. Another example is New York Times. Again, this is from 2014, so I would imagine this is very different now. But they even say things here like our editors prefer will demand word style track changes of the text editor.

33:17

Speaker 2: So, you know, very common theme here, right? Um again, they've open source I think called Text Editor. ICE, but the GitHub repo seems to be about two or three uh years out of date. Um again, maybe I'm not looking in the right spot. Um and obviously one of the solutions here would be to replace uh Drive Tail. Well 33 minutes. I'm gonna speed this up, sorry. Didn't realize how um Yeah, um one of the ways we're gonna publish push is building editor effectively and What are an editor and basically replace draft tile, but again draft tile is excellent. It seems like the sort of thing that'd be awesome to add. Yep. And as I mentioned, even the best CMSs. So I've spoken to journalist friends who work out some of these large publications

34:02

Speaker 2: And even though they provide a really awesome interface, they seem to like use Word or their own favorite Word processor anyway. And they'll copy paste it into the CMS when um you know w when it's gone through viewed effectively. Something features writers expected as a result. Something I wanted to show you guys as well is actually how well DraftTower does this. Actually, in fact, I'll skip this. Sorry, I didn't realize I have 33 minutes, so my apologies. Yeah, so our current work tower workflow is basically the writing team drive comments in a web processor. They use uh track changes, comments a document management system to track your revisions. When the piece is actually done for approval, then it goes into the CMS and there's a dedicated publishing team that do this.

34:49

Speaker 2: Which is again very normal for small small small newsrooms. Something we do which is kind of naughty, which I wanted to demonstrate, is we uh but we're using a rich textile for the body. Which is an oh no, because uh stream fields are awesome, and it's one of the reasons I was actually drawn to Waytown in the first place. Um But we tried stream fields and the journalists didn't like it, and I actually wanted to demonstrate as to why that worked. what's. Stream fields are great. It does say using them for news stories. Slightly controversial, but I at the moment don't necessarily agree. But again maybe I'm doing something wrong, so please educate me if that's the case. Um, if I go here

35:36

Speaker 2: Cool. So where are we? So if I go to pages, for example, let's go to here. Let's go to here. So we've got the article, we've got the metadata about there. Um Oh the body here. So if I go back here again, where are we? So here is the document we've had earlier. Word decides it wants to move its over. Sorry.

36:25

Speaker 2: Well, let me move it over. How do we use computers? Here we go. Sweet. So this is basically a Word document. The Word document has things like images, etc. etc. etc. Previously had lots of issues with this, but um and you can see as well Um I can't even use what anyway. Copy, paste It's come through brilliantly. So H2 tag, stuff there, very, very minimal uh munging of data. Draft Hardio dot draft Uh the old one, sorry.

37:10

Speaker 2: What was the audit the auditor called again? Hello, that's the one. Yes, Hello didn't do as good job as this, which is fantastic. Images are fairly easy. So we've actually got a feature now that basically lets them grab the image. Copy. There they go. Upload, you can actually paste the image in Can upload it.

37:55

Speaker 2: We can also copy and paste that. So if they want to go through and change stuff, really easy to go through and copy and paste this, including the image. Very easy to use. If I want to change move around headers, it's great to use. Again, maybe I'm doing this wrong, but the equivalent using Stream fills. So heading, paragraph, image, embed. Let's put a paragraph in. So you then have that plus add a heading, copy and paste this in. Things like images obviously you can't copy and paste like I did before. So um

38:41

Speaker 2: even though Tune fields are awesome and I use them in a lot of situations, we tend to find for the body of an article they're much more cubicome to use. Um this might be something we're gonna prove in a future update. I don't know. Sorry, I'll try to finish up. Um yeah, that's it. Thank you. Oh sorry, any questions? Sorry, I'm conscious I've gone away over time. I apologize. Thank you.

39:23

Speaker 3: Very important question. Is the image um plug-in for GraphTale open source quite chair?

39:30

Speaker 2: Uh no, but we might open source it, so yeah. But I I guess the powerful thing about that is it took us like two or three days to build it. Like it was awesome how fast it was to actually build that thing and have it work.

39:41

Speaker 4: The multi-site thing, is that also something custom that you built?

39:44

Speaker 2: It is. I'm I'm gonna open source that. Yeah, it solved a lot of problems for us, which is really good. So at the moment it's a manual thing, but we're building it back in for it. So yeah, yes, manual thing. So every day that a journalist go through and basically look at the statistics and actually update it. But we have a we have a separate application now that aggregates data from Google Analytics, so we're just going to create that for an API and actually pull data out dynamically every hour.

40:16

Speaker 5: It's not the intention to kind of tell that

40:20

Speaker 2: Yes, correct. Yes. Yeah, I guess the big goal for us, the overarching goal of using Wagtail is to kind of automate human things. So a journalistic team or editorial team will do lots of things by hand. We're just trying to get rid of that all and just have it all just happen by itself. And we've done a lot already. And this has only been we started developing this in I think September, October last year. So there's a lot that's got done in that period of time. Any other questions?

40:45

Speaker 6: capabilities in terms of um thinking of content um in terms of SEO. Uh

40:51

Speaker 2: how do you mean sorry?

40:52

Speaker 6: Of of search engine optimization ways Um the keywords that pop up.

40:57

Speaker 2: Uh sorry, I'm not understanding the question. For the URLs?

41:01

Speaker 6: Yes.

41:01

Speaker 2: You're talking about the metadata where we're populating on the page or

41:04

Speaker 6: um were you populating um uh I I thought I saw a slide at some point that said that But um the URL was populated through the

41:13

Speaker 2: Oh yeah, we we use we use the article slug for that. But a lot of the a lot of the validations on our safe, for example, actually like have checks for sanity and stuff like that. Also good actually feels like excerpt, which are used in things like the OEG, the Open Graph data. you know, the the meta description stuff like that. So we'll I guess we're kind of one of the goals of this we're trying to encourage the journalists to kind of think about this sort of stuff first rather than kind of an afterthought. But a lot of that's just done through basically rules on the uh on the model. So Well good, thank you. Thank you very much, Ryan.

Questions this talk answers

Why choose Wagtail instead of WordPress for a multi-site newsroom?

Wagtail provides an explicit Django-based CMS with structured data, enforceable workflows and validation, easier reporting, and more consistent behavior than the company’s separately configured WordPress sites. It also supports frequent releases, testing, and alternate renderings such as AMP and Apple News.

Discussed at 5:29

How do you migrate multiple WordPress news sites into Wagtail?

The migration used WordPress’s REST API to retrieve rendered article content, sanitized inconsistent data, converted images into Wagtail images, and extracted details such as stock codes into structured fields. Existing URL differences were handled with conditional URL-path logic.

Discussed at 7:49

How can Wagtail support multiple brands in one Django project?

The sites share inherited page models, templates, and common logic, while site-specific fields and validation remain available on the shared article model. A site-data abstraction gives each site a stable semantic label and provides helpers for domains, URLs, metadata, templates, and conditional behavior.

Discussed at 19:17

How do you reduce Wagtail’s database queries for articles with many images?

The speaker uses related-object prefetching and custom rendering logic, including storing image primary keys on the article model, to avoid repeatedly querying image renditions. In the example, this reduced the request from roughly 55 queries to 22 and cut the response time to about a quarter.

Discussed at 22:25

What newsroom workflow features does Wagtail currently lack?

Wagtail’s built-in states are mainly draft, scheduled, and live, which is too simple for newsrooms needing legal review, editorial approval, client approval, comments, track changes, and configurable multi-step workflows. The current workaround is to restrict publishing to editors and handle feedback in person or by email.

Discussed at 24:38

Why use a rich-text field instead of StreamField for news articles?

The newsroom found that journalists preferred writing in Word and copying content into the CMS, including images and formatting. A rich-text editor made that workflow much easier, while StreamField required separately adding blocks and did not support the same straightforward copy-and-paste experience.

Discussed at 34:49

Presenters

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos from Wagtail Space US