Automated spell-checking in Django projects

This video features Jakob Schnell at DjangoCon Europe 2018 in Heidelberg, Germany.

Automated spell-checking in Django projects
0:32:04
Published May 24, 2018
545 views

https://media.ccc.de/v/hd-63-automated-spell-checking-in-django-projects

I'm aiming to show how avoid spelling errors by showing ways to implement automated spell-checking.

Nearly all Django applications have two main textual bodies that users come in touch with: First, any text in the application and its translation, second the documentation. Since both are usually written by humans, they will contain spelling errors.
This is considered harmful and can from time to time hinder the user trying to understand what to do.
Therefore, an automated spell-checking tool should be a part of any CI-cyle.

For spell-checking documentation, I will give a short demonstration on how to use the "sphinxcontrib-spelling"-tool written by Doug Hellmann, the problems we had and how we overcame them.
For spell-checking text in the application and its translations that are usually found in .po-files, I have implemented a small tool called "potypo" (name and development in progress).
I will present this tool and show challenges and problems on the way to implementing automated spell checking for .po-files.

Jakob Schnell

Summary

Jakob Schnell explains how to add automated spell-checking to Django projects. Sphinx documentation can be checked with sphinxcontrib-spelling and integrated into CI, while application text is harder to find in scattered Python, HTML, and CSS files; Django’s gettext translation workflow collects these strings into PO files. His PoTypo tool combines the Babel/PO translation format with PyEnchant to check source and translated strings, with configuration for languages, custom word lists, ignored languages, phrases, edge-case words, URLs, HTML, and Python format strings. He presents it as a young, still-evolving project and notes that it does not solve spell-checking arbitrary untranslated strings or identifiers.

Key takeaways

  • Sphinx documentation can be checked with sphinxcontrib-spelling and run as part of continuous integration.
  • Django translations make UI text easier to spell-check because gettext collects strings in PO files.
  • PoTypo combines PO-file processing with PyEnchant and supports multiple languages and custom dictionaries.
  • Projects need word lists or exclusions for names, technical terms, URLs, mixed alphanumeric words, and phrases.
  • The approach does not automatically find and check arbitrary strings scattered through untranslated Python, HTML, and CSS code.
  • PyEnchant provides a common Python interface over different underlying spell-checking engines and dictionaries.

Summarised automatically from the transcript.

Chapters

  1. 0:00 Introduction and Motivation Jakob Schnell introduces the problem of typos in Django projects and explains why automated spell-checking is useful.
  2. 4:05 Documentation Spell-Checking The talk covers Sphinx-based spell-checking for project documentation and integrating it with continuous integration.
  3. 7:54 Spell-Checking Application Text The speaker explains why checking text embedded in Python, HTML, and CSS is harder, and how translations provide a practical solution.
  4. 10:15 PoTypo Overview PoTypo is introduced as a tool combining gettext translation files with PyEnchant spell-checking.
  5. 11:02 Gettext and PO Files The talk explains gettext, PO-file structure, message extraction, translations, and contextual comments.
  6. 14:06 Spell-Checking Engines and PyEnchant The speaker surveys ISpell, Aspell, MySpell, Hunspell, and the PyEnchant abstraction over different spell-checking engines.
  7. 18:07 PoTypo Setup and Configuration This chapter covers installation, language dictionaries, setup.cfg options, locale directories, and CI failure handling.
  8. 22:47 Word Lists and Edge Cases The talk discusses custom word lists, punctuation-heavy terms, mixed alphanumeric words, phrases, HTML, URLs, and custom filters.
  9. 26:44 Project Status and Resources Jakob describes PoTypo's current status, invites contributions, and shares contact information and repository details.
  10. 28:33 Questions The speaker answers questions about Sphinx documentation, translated strings, docstrings, and unusual class or function names.

Transcript

5,198 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:06

Speaker 1: Our next speaker will be Jacob Schnell and he will be talking about automated spell checking.

0:38

Speaker 2: All right.

0:39

Speaker 1: Okay, please welcome Jakob.

0:47

Speaker 2: Perfect. Hello everybody, my name is Jakob. Some of you might know me by my nickname Kubi. I'm a maths and computer science student currently on an Erasmus year in Milan in Italy, but uh originally studying at the University of Heidelberg. This is my first talk at any uh DjangoCon or PyCon or whatsoever. Um so I thought uh as an introduction it might be fair uh to Just uh have a few words about why I am here. And it all starts uh back in October last year. I was sitting in a database class I think and uh compared to Katie's talk from earlier it was uh boring at least to me and so I decided to do something uh worthy of my time

1:33

Speaker 2: And a friend of me, uh Rafael, that has just introduced me, has a project called PreTix that you might know that you all bought your ticket from. And so I decided to start hacking uh on pretex. And I fired up the uh the program and after a few seconds I discovered a typos somewhere in the code and I was like wow great cool I can just search for this wrong word and I can fix it and I can submit a PR and it will be great and I will be doing something useful and I will be doing good stuff with my time. Um and so I did that and that worked fairly well. And then I figured maybe there are more typos than just this one. And So I fired up a spell checker, I think a spell. Um and spoiler alert, there were more typos, and so I fixed them all, or all of them I could find.

2:23

Speaker 2: And uh I asked whether maybe Automated spell checking would be something that PreTix would like or that the uh Django world could use. And uh Rafael's answer was yes of course, please do so. And now you might be wondering, well, but I am typing my text and I am typing uh good and perfectly and here are some examples of uh wrong words. Uh that that that were fixed, uh that are to be found in uh or that were found in the pretext project. Um And I think since all the text that we write is written by us, by humans, we can all agree on the fact that it will contain errors And a friend of mine even goes as far as saying whenever you publish anything, be it a book or a thesis or an article or whatsoever, once it is published, once you handed it in, you can open your thesis, your book at whatever page and the first thing you will find.

3:19

Speaker 2: is a typo and this will never change. Oh yeah, the audience agrees with me. Cool. And um for most of these typos it's not too bad because most of us are just so used to reading all the time um that they will just skip those typos. Um but if you're not an English native speaker or if you have trouble with the language uh altogether, then you might stumble uh across some of these typos and uh every now and then There might be typos that uh wouldn't even be uh well let's say code of conduct compliant. Uh for example I encountered a typo in the word account where the second C and the O were switched. Um and that's why it's not on the slide. And uh yes, so obviously

4:05

Speaker 2: we need uh spell checking because we ca we can't find those errors all by ourselves. Um some of them will slip past us and Computers are just better at doing so. In general, there are uh two places where uh we need spell checking in a Django project Uh one first is uh the documentation and the second is the code itself or the uh user interface that we provide And uh the code is the much more complicated part, and I'll come to that uh in about five minutes, I think. Um so I first uh want to talk about uh documentation. For documentation, spell checking is rather easy. That is because a um All of your documentation is usually in one place.

4:52

Speaker 2: Right? You have one or uh one folder called docs and it contains all of your all of your documentation. It's usually large text files that are easily um checkable. Usually you only provide documentation in one language, so you only need one spell checker for that language. And In the Python world documentation is sorry usually done uh with uh Sphinx, with a Sphinx uh project, package, whatever. Um And for Sphinx, uh fortunately uh spell checking already does exist. Uh and it's uh a package called SphinxContript. spelling And the usage is uh it was written by Dac Alman, the usage is quite simple. Uh you fire up strings minus build um with a directive. Oh nice, I can look at here, not there.

5:37

Speaker 2: Um You fire it up via things minus build minus b spelling, then you put all your other other options that you need, and then you specify in what folder your your output should should be and usually that will be something like underscore build slash spelling. The requirements are fairly easy. You can just install SphinxControl-spelling via pip and it needs PyEnchant, which I will talk about in a few minutes. Um those those are readily available. You can just add those two lines to your uh requirements. txt file or whatever. Um if you are using Travis, you can easily implement uh the Sphinx build thingy um by just adding uh the enchant package that is needed for Pi enchant uh

6:23

Speaker 2: in your uh Travis setup and you can generally easily integrate it into your um CI by just checking whether after the build is done um the file output. txt in the directory underscore build slash spelling does exist And if it does, then you have some errors and you should go fix them. And if it does not, then everything is fine and you can just move on. So this is uh this is fairly straightforward. This is nice. It works just fine. It has been working for PreTix for uh I don't know let's say four or five months now and it's great Uh it needs a bit of configuration of course. Uh first of all your uh Sphinx config in conf. py uh needs to uh have sphinx config. spelling as an extension

7:08

Speaker 2: Then you should specify a spelling language, uh which I guess will be En underscore US for most of our applications. Um since your application will probably contain some words that are valid for your application but not for the English language in general. For example in the Pretex project, this is pretexts as a word itself. It will obviously uh come up in the documentation, but it is not a word to be found in an English uh dictionary, I think. Um You you you need a so-called word list that just lists all these words that are not found anywhere in uh in a dictionary but in your documentation and that are correctly spelled. And you can just specify a word list via the spelling underscore word underscore list

7:54

Speaker 2: underscore file name directive. And if you want, you can ask the sphinxcontript. spelling to show suggestions for things that you misspelled, which makes it even easier for you to fix all those spellings. And this is easy and nice and uh works great. For code, everything gets a little bit more complicated. Um This is because in general wherever you interact the the second part where you interact with your users apart from documentation or manuals or handbooks or whatsoever is your user interface either um Usually in in some kind of a graphical user interface or text user interface.

8:40

Speaker 2: And actually most GUIs contain a lot of text. And if you want to fix all the errors that are there, then you would have to parse uh Python code, you would have to parse HTML code, you'd have to pars uh CSS probably and extract all the strings and then it's unclear what strings should be translated and what ching strings shouldn't be translated. And uh to be fair I have no clue on how you would do this. Um but there is one thing that makes it very very easy for us as a developer to indeed check all those strings and that thing is translation. Once your project scales beyond a certain point, you probably want to translate it in another language. For example, for myself, I am German, so uh English is not my uh my mother tongue

9:29

Speaker 2: So I would probably develop my application in English since it should be accessible to more than just the German and German-speaking population. But I would probably want uh to have a German translation because it is easier for me and for my probably for my use cases to work um with the product. It is easier for me to sell it to other people and um so I want that translation. Even if you are an English native speaker then you might still want translation because you want uh to sell your product to other people. You might want to translate it into Spanish, which is a which is a language that many many people speak. And once we have translation, it makes it quite easy for us. To automate our spell checking. Um because translation gathers all those strings and there are uh there

10:15

Speaker 2: is uh a system that does this for us that I will present in a minute. Um And since we now have all of our strings gathered in one place, we can check them for spelling errors rather easily. And basically what uh the project I want to present to you today is, is uh it combines the PO library for translation and PyNchant for spell checking into one program called uh Po typo. And of course it's not just those two components, but it's also a lot of sweat and tears and uh work and glue code wrapped around uh stuff. To bring this all together. So let's talk about the two components uh that go together. The first uh is GetText. GetText is a uh localization and internationalization uh

11:02

Speaker 2: system That is, I think, the standard for uh translation, for internationalization, for localization of things that handle text. And it is based on so-called uh. po files. This is an example of uh such a uh file It has some metadata in the top like the creation date, the date of the last revision, um whoever was the last translator, and most importantly what language we are translating into in the language tag That is written up there and then some more metadata and the top. Uh and then it has a a bunch of entries that all basically look like the last three lines here. Um the first is where that where does uh this text

11:47

Speaker 2: appear in my project, for example in this sample PO file. po file, the first message would appear in the file junglecon slash talks slash intro. pi dot pie in line one. And uh every uh entry has a message ID that is the string that is uh there in your code, which would be hello DjangoCon, I am so happy to be here. I truly am. Um and the second would be the message uh string, which is the translation of that code into the language that is given above And the usage of this system in general is uh fairly simple as well. Um you can just uh import pget text and ugText uh from Django. utils. translation, not T9N, but otherwise it wouldn't fit.

12:33

Speaker 2: Um And you would for example have the the two statements um print underscore bra uh open brackets my name is name closing brackets And the underscore is then the call to you get text, and you would have, for example, a call like print pget text sorbox sizes 20. And if you now let uh getText do its thing on the files and gather all those message IDs for you, then this would uh render as uh something like the following. po file Um you see we have the file jjangocon slash talks slash poetypo. p dot py. And in the third line we have the my name is name in curly brackets um message ID and uh the comment that we gave before this, which is leave name as is the code will handle it, uh gets automatically drawn into uh the.

13:21

Speaker 2: po file Um it adds that this is uh the Python brace format, which is a uh thing that tells uh getText that this is uh special or the thing in curly brackets is special. It shouldn't handle this. Uh and you can add a message string or a translation for this. In German this would be Ich heiße name. Um and what the pgetText thing does is uh the first string is a context for your application. Um so usually uh a translator wouldn't would n I don't I don't think you would be able to handle the word venti, but once you add the context, oh it's about Starbucks sizes, um then of course the translation or the the meaning is large and it translates to growth in German. Um so this is fairly easy uh or if fairly fairly simple to to handle.

14:06

Speaker 2: We can use the uh PO library to extract uh the message IDs and the message strings and then we can check them separately for errors. And since we have a base language given uh we know what our message ID ID's language is and since we have a language tag in our. In both of these languages quite easily. The second part that goes into PoTypo is a pie enchant. And if you're working with uh spell checking then you will probably have encountered some of those words. And I will try to give a short overview of on uh what they are. Um and it all starts with iSpel, which is a spell checker that is really old.

14:52

Speaker 2: It was written in 1971, if I'm not mistaken. Um Originally for the English language it works quite well. Even today it is sort of the de facto standard, but along came UTF-8 and uh with that well ice bell wasn't really able to handle that, so They uh so a spell was created. And to this day I think ice spell is the best spell checker for the English language When the OpenOffice Office project came along, they wrote uh they implemented their own uh spell checker, MySpell, as part of their Abby Word word processor. Um and that has worked that has replaced ASPL as a spell checker in OpenOffice. And later uh Hunt Spell was developed originally for the Hungarian language.

15:39

Speaker 2: um and it has by now replaced my spell as a spell checker. And uh from my point of view, A spell is the spell checker that you want if you uh spell check English text, and a hand spell is the spell checker that you want to use. whenever you check other European languages and for non European languages I'm very sorry but I have no idea. I don't encounter them frequently. Um In general this uh the the uh uh those this mass of spell checkers uh produces kind of a problem, um, because we would have to implement them all or we would have to handle them all and they all work. Sort of the same because they all take text and then they tell you what is wrong in the text, but they also all work sort of differently. And that is why uh

16:25

Speaker 2: the enchant project exists. Um the lib enchant is a library that wraps all of those uh spell checkers uh and uh provides you as a developer an API. uh with which you can easily check all the things and it handles all the specifications of the spell checkers for you and even if they don't implement some functionality but others do uh enchant will try to emulate the the functionality that other spellcheckers have for you and it is very convenient um as developer to just have an uh one one framework that you need to know and you don't need to uh Interface with a spell or an hand spell or uh Finnish spell checkers or H spell for the Hebrew language or whatnot.

17:10

Speaker 2: There is a list on Wikipedia. Go look it up. It's uh huge And um there exists a Py enchant um written by Ryan Kelly, which is uh currently unfortunately unmaintained, um which wraps uh the Libanchant and gives you Python bindings. So it's quite nice to interface uh with py enchant. Alright, these are uh the two main components that go into uh Po typo. And I will now talk about uh how you would use uh PoTypo in your project and uh how you would set this up, but first I will have a short drink with some water somewhere here

18:07

Speaker 2: Ah mais Alright. The usage of uh Po typo is uh fairly simple. Uh you just fire up a Po typo in your directory where your um Setup. cfg lives. I will talk about this in a second. Um the requirements are uh fairly simple. Uh Po typo itself is installable via pip install PoTypo and um You of course need uh Py and Chant and maybe the PO library as a um as requirements for Po typo, uh but if you use translation you have the pot uh the PO library installed anyways You will then have to

18:52

Speaker 2: have packages for every language that you want to uh check your spelling in. So for a German and English project, you would install something like uh a spell minus en for English dictionaries for a spell, and you would install something like my spell minus d minus d e. four dictionaries for the German language or actually you would install Hanspel but I haven't checked whether Hanspel is as good as my spell it should be but I don't know So again it is fairly easy, uh it is fairly easy to use. And the configuration is um compliant with uh the setup So you have one one part where you configure Po typo and your setup. cfg where you also configure Flake 8

19:37

Speaker 2: and other Django and Python thingies And you first of course specify a default language, which is the language that everything will be translated from. So it's basically the language of your message IDs in your. po file Then you specify where your local uh where your. po files live. Usually they will live in a uh locals directory like Django-project slash locale or somewhere. Um And uh Po typo handles uh the finding of all those uh dot Po files by itself. Um Currently it just assumes that you follow uh the structure that is given uh in the left that I will explain in a second, um

20:25

Speaker 2: but it may add some magic to find your. po files automatically. Then of course you can specify languages for which you do not want uh PoTypo to to fail or to report any errors. For example, if you are just in the process of translating your application into Danish Then you might of course want uh PoTypo to report what uh errors there are in the translation into Danish, but you do not want these errors to uh to break your continuous integration process or to break your Travis build. And this is why the no-fail directive exists. And then of course, as before, we need a word list for words that are present in our application that are not present in an English dictionary.

21:11

Speaker 2: And you have basically two ways of handling these, of implementing this. The first is you can put them all in a wordless directory. If you do so, you should specify the uh WL underscore dir variable and you should give the directory that uh your word lists live in. And if you have that, then they should be named language tag. txt, which is the example on the right here You would have uh on in your Django project you would have a folder word list and then then for every language you would have a file d e. txt en. txt and so on and so forth If you don't like this, um then you can also put them in various uh places in the uh locales d directory And usually in your locals directory you have one directory for every language that you're translating into, one directory for Danish, one

22:00

Speaker 2: directory uh for German, and so on and so forth. And in those directories you have directory lc underscore messages and in that directory live your. po files. And you can just put your wordlist. txt, in this case called wordlist. txt, in this uh LC messages folder or in the folder above, in the folder for German or for Danish or for whatsoever. And for your base language, for your default language, uh you would just put a file called wordlist. txt into your into your locale folder. Those are the two uh ways that wordless are currently handled. Uh if you have any other ideas or if you think wow this is uh Not as nice as I wanted, I want this another way.

22:47

Speaker 2: Um please come talk to me and we can figure something out. But for now this uh works quite well Now there are a few exceptions to simply uh using word lists or to just uh spell checking. Um and the first are what I call edge case words. And these are words that uh contain punctuation um or that contain a mix of uh numbers and letters. Um this is because of the way that a pie enchant works It basically uh splits your text by a white space and then strips punctuation from your text. And so the word add-ons would be split into a part adds and a part ons. And uh well add is an English word, onz

23:34

Speaker 2: is not, or not really, and it would report an error for that. Uh or for example for uh the For translate. pretex. eu it would uh report an error because it would split it uh into translate pretext and eu And it probably won't be able to handle those. And so all those words that contain punctuation, all the words that contain a mix of uh numbers and letters, like for example two hundred and fourteenth, which is a part of an address I think, um you can specify them. And uh then it will just skip those words across all of your text. And uh you would not want to add them to your word list because again they are not uh the the parts the single parts are not really words MT940 is not really an English word so you would not want it into your in your English word list

24:25

Speaker 2: Um but it's correct, it should be spelled that way, and so we just tell uh ProTypo please skip this word. And the next thing that is um Somewhat different is what I called phrases. And a phrase is uh something that is present in your text, although it doesn't technically belong. For example, ticketing powered by is a text that might be used in a German application. And German people probably would understand it. Because a ticket in uh English is a ticket in German, so ticketing works and powered by well by and the German word by are quite quite similar and also powered would be understandable. So you can uh You can use the phrase ticketing powered by in a German uh in a German text.

25:11

Speaker 2: But neither the word ticketing nor the word powered nor the word by are German words Uh so you'd have to you would have to put them in your word list if you do not want them uh to pop up as errors, which is bad because by might be a misspelling of the German word bi, which is written B-E-I. Um because you are typing English and German parallel. And so you want this to be an error, but you do not want the phrase ticketing powered by by itself to be an error Um so this is also handled by uh phrases and those are um The two the two main things that you have to filter out when you are checking for your code. If you encounter any other things where neither HK's words nor phrases are enough for you to

25:58

Speaker 2: spell check your application in the way that you want to. Um I'm very happy to to see more edge cases and to find a way to work around them. And then last but not least uh you might uh have some HTML uh code in your in your code somewhere that will be uh chunked out by an HTML chunker, which is provided by the enchant project. And then you have some filters that filter uh the Python braced format from before or that filter URLs or that filter HTML from uh from your code and those are fairly easy to use you might uh write your own it's it's quite nice. And this is uh the complete feature set of PoTypo to this point. Um The project is

26:44

Speaker 2: uh is not that old. Um it's quite new. It's uh still working project uh working progress. Um but it is uh already used in one application which is uh pre-tix uh once Raphael is done organizing this conference and gets around merging pull requests um which I totally cannot blame him for. If you have any wishes on features that you think you might use if you want to use this project, you are very welcome to open issues and we can discuss about it. If you think, oh wow, this is something that I can use then please do so. Please come to me, please talk to me. I will help very gladly help you set all of this up. Um as I said, this is still work in progress. There might be things that change. This is

27:31

Speaker 2: my contact information, also where you can find the slides. and the GitHub repository. And before I finish, I want to thank uh three entities. First and foremost, I want to thank uh Rafael for his uh many many times that he has provided help uh for me whilst going through the process of writing a Python project and publishing it and so on and so forth. He has been uh super super helpful and is in general an awesome person. Um second, I would like to thank Matthias Vogelgesang for the uh Metropolis beamer theme that I have used to create this presentation. And third, uh I want to thank you. As my audience uh for your time and for your attention. Uh if you have any feedback, please go to uh the pre -talk system, click on this talk to the schedule, click on this talk, and give me some feedback.

28:17

Speaker 2: And you have questions, please go. Thank you.

28:33

Speaker 3: Hello. Um I have one question regarding the documentation. Do you have actual ex example of what kind of documentation we can um validate and run uh the library to the spell uh to run the spell checking

28:52

Speaker 2: Um I don't know if I got your question right, but uh uh any any kind of documentation basically that is uh that is using Sphinx to document its stuff, you can use the SphinxContript. spelling uh project to to spell check your documentation. Does this answer your question?

29:10

Speaker 3: Yeah, kind of. Okay. And another one is you showed actually how to um spell check the strings representation that we show to the user. Uh can we use this on the sp

29:27

Speaker 2: Um as I said, uh if you do not have a translation, so if you're All of the strings that should be spell checked are scattered throughout your project in Python files and HTML files and CSS files. Um then I uh don't know on how to how to spell check those. If you want to uh use uh to to translate your project into any other languages, then you would probably use uh the get text system and in that case you have a.po file and you can use PoTypo to spell check. But in the other case, unfortunately I have no idea on how to do this

29:57

Speaker 3: Okay, thank you.

29:58

Speaker 2: Welcome.

29:59

Speaker 1: Next question from the microphone on the back.

30:09

Speaker 4: Frustrations for my reviewers. I think I will have plenty of need for that. But one of the things I'm a bit unclear is doc strings. Um so if you use it for automated documentation of your code. Um clearly all your class names, your function names, all this stuff um is not spelled according to regular syntax rules. Would I have to maintain all these things in the edge case um list or is there a different procedure to handle Doc strings and function names and class names and the weird spelling.

30:42

Speaker 2: Um well class names and uh function names are spelled correctly by default I would say. Um because you you really can't cannot cannot I don't know spell spell them wrong.

30:56

Speaker 4: But for example one one of the projects I had, I had the misfortune to name the controller with one L, which is quite annoying. And Of course this was something that was not just a mistake I made when naming the class, but all in all the documentation that I wrote referring to this very class. So obviously once spotted it's easy to rectify, but It would so how would I how would I approach this particular problem?

31:20

Speaker 2: Okay. Um the ThingsContra. spelling uh project does not uh check it only checks for uh errors in the pros that you write to document your process and not in function names or anything else. If it in c if you use the word controller for example with only one L in pros to refer to a controller then it will of course tell you this is wrong, but it doesn't check uh function names or uh I don't know methods or whatever.

31:47

Speaker 4: Thank you.

31:47

Speaker 2: You're welcome.

31:49

Speaker 1: Wonderful. If we don't have any further questions, thank you very much, Jakob, for the talk.

Questions this talk answers

How can I automatically spell-check Sphinx documentation in a Django project?

Use SphinxContrib Spelling with Sphinx’s spelling builder, install it with PyEnchant, and add the spelling build to CI. Configure the spelling language and a word list for valid project-specific terms.

Discussed at 4:52

How can I spell-check text in a Django application's user interface?

The practical approach is to use Django’s translation system to collect interface strings into gettext PO files, then spell-check those message IDs and translations. PoTypo combines PO-file processing with PyEnchant to automate this.

Discussed at 9:29

What is PoTypo and what does it combine?

PoTypo is a tool for checking spelling in translated Django projects. It combines the PO library for gettext translations with PyEnchant, which provides Python access to multiple spell-checking engines.

Discussed at 10:15

How do I install and configure PoTypo in a Django project?

Install PoTypo with pip, provide PyEnchant, the PO library, and dictionaries for each language, then configure the default language and PO-file location in setup.cfg. You can also mark incomplete translations as no-fail and configure project-specific word lists.

Discussed at 18:07

How does PoTypo handle project-specific words, URLs, HTML, and other strings that spell-checkers misread?

Use word lists for valid vocabulary, exclusions for words containing punctuation or mixed letters and numbers, and phrase entries for multiword text. PoTypo also provides filters for Python brace formatting, URLs, and HTML, and can use an HTML chunker.

Discussed at 22:27

Does SphinxContrib Spelling check Python function names, class names, and docstrings?

It checks the prose written in documentation, including prose in docstrings when that prose is processed as documentation, but it does not check function names, methods, or class names themselves.

Discussed at 31:20

Presenters

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos from DjangoCon Europe