Closing session
Published June 13, 2025
This video features Anssi Kääriäinen at DjangoCon Europe 2018 in Heidelberg, Germany.
https://media.ccc.de/v/hd-110-banking-with-django-how-to-not-lose-your-customer-s-money
Topic: Banking with Django - how to not lose your customer's money
Abstract: How Holvi decided to pick Django as part of their core infrastructure. I'll go through both the business reasons for using Django for banking (5-10 minutes), and technical details of how we do reliable distributed software which keeps our and customer's money safe (10-20 minutes).
Anssi Kääriäinen
Anssi Kääriäinen explains how Holvi uses Django to provide business banking services and why payment systems must be designed so that failures do not leave customers’ money missing. He describes a payment flow that separates customer-facing features, core accounting, compliance approvals, and communication with external banks, whose batch-oriented interfaces contrast with Holvi’s internal real-time messaging. The central recommendation is a reliable inbox-outbox messaging pattern: record business changes and outgoing messages in the same database transaction, send messages only after commit, retry failed deliveries, and deduplicate them with unique constraints on the receiving side. Reconciliation between inboxes, outboxes, and account balances exposes missing messages quickly. He also stresses testing, code review, monitoring, rapid incident response, fixing root causes, and choosing simpler manual processes at small scale while adding stronger reliability before payment volume makes failures expensive.
Summarised automatically from the transcript.
Automatically transcribed, so expect mistakes in names and technical terms.
Speaker 1: Our next speaker is Anzi Karian, and I hope I pronounced her last name correctly. He's uh one of uh the he's a part of the Django team and wrote a huge part of the R M. And now is he's going to talk to us today about how we lose all our customers' money uh sorry, how we not lose our customers' money.
Speaker 2: So So hello, my name is Hansikar and I'm here to talk about using Django in a banking setup. And of course about making reliable banking with Django. First a bit about me. So I work at Holy as a back-end lead. I've been working over there for a couple of years. I enjoy it a lot. I'm also has be have been doing some ORM work so uh from 2012 I have been Django Core contributor, worked on the ORM a lot until A couple of years ago. Uh now I have a couple of small kids, so that's taking all of my time, but I'm hoping to get back to Django
Speaker 2: at some point. So first a quick introduction to what Halloween does So we offer business banking services to microentrepreneurs in Finland, Germany, Austria and to e-residents of Estonia. The e-residents is actually a pretty fun setup. So anywhere in the webilla kan johta Estonian embassi , registra as an e-estonian. And then you can set up banking you can set up a business like you were a Uhril Estonian uh uh origin and that way you can without ever visiting Estonia you can actually set up a business in Europe
Speaker 2: Holvi has a payments account, so it's very much like a banking uh bank account, except for legal reasons it's kind of different. for practical purposes it's uh it offers you SEPA payments, it offers you MasterCard, it offers you uh online payments which you can use with the whole way online shop. On top of the bank account we have uh uh we have some business tools. So we have invoicing, you can create invoices, send them to somebody, collect money so you get the separate transfer in. We automatically match the payment to the invoice, we do the bookkeeping for you. We have an online shop so you can uh sell stuff online. It's it's for small shops, but it's really quick to set up so it takes a couple of minutes to start selling something online.
Speaker 2: Again, you get the money immediately when somebody pays uh buys something from you, you get the money immediately on your account and we do the bookkeeping all of that for you Finally we have expense handling so when you are traveling you can go to somewhere with taxi, use your mastercard, you get a push notification when you paid. You can do the bookkeeping easily on the go. You take the pictures to receipt, you do the categorizations, the VAT stuff. So you get the bookkeeping preparation done real time. So that is what hold with us. It's an interesting offering. Try it out if you are in one of the markets or well if you want to be an E Estonian
Speaker 2: Our tech stack, nothing more special going on in there, but I'll go through it in any case so you understand what we are using on the backend side. So we have Django of course, then we have Django REST framework using PostgreSQL. Then we have Celery for asynchronous tasks for doing these scheduled tasks using Redis behind that. We have Angular on the front end side. I don't know that much about it , but it works really well for us. We have the native mobile apps. We have iOS, Android. And finally we are running all of this on Amazon so we use S3, we use the EC2, we use
Speaker 2: or a lot of services, RDS for example from Amazon and it's it's working really well for us. Django for us has been a good choice The ecosystem is really nice. For pretty much ever anything you want to do, you find something already uh existing. So that's a big plus when using Django. It's it's an overstatement to say that it's easy to find developers, but there are a lot of developers using Django. At the moment in Finland, of course, you know, anybody Who has seen a computer is immediately hired. So it's kind of not that easy to find the developers, but at least there are a lot of developers who know the Django. And finally it's reliable. It's uh
Speaker 2: proven software, it works really well for backing setup. We are using Django for 99% of the code we have. And it it works really well. It's reliable, it's uh has all of the basic things you need for building something for banking So my subtopic for today is that how not to lose your customers' money. There are two options. The first option which we try to use is reliable payments. You don't lose the payments. You always get them go through the system and everybody's happy. But in case There's an error omen. Just lay on message, you
Speaker 2: on bug ominessä. It not that you lay kustemminessä muin, you actually lose your own muin. You you need to use a lot of time for each error case to sort it out that hey is this really an error case? Uh what's happening in the system, where is the money? And in the worst case you for example can double execute a payment, you send thousand euros two times to somewhere and it's often impossible to get the money back. So there it went you lost your own money. But you actually don't ever lose your kusters' money because. That's the way banks work. You you take the losses for yourself instead of anything going for the customers So I'm quickly going to introduce the play payment flow.
Speaker 2: So just to have the mindset for how how this thing works in Halloween. This is a bit simplified, but it it has all the major pieces. in there. So it starts with customer verifying payment, typing in the IBAN and stuff like that. Then they receive a push notification with a sekurytoken. So we do two factor autistic with that. They type it in, click verify, the payment goes out. Then actually starts the uh kind of the payment processing inside Holy. So first we make a message, we we record this in the the the kind of the front end side. or we have the customer facing side of the uh backend and then we send a message to core accounting the next layer
Speaker 2: which handles actually the all of the money we want to keep that At least conceptual conceptually separate from the customer-facing features for security, for performance, for You can change the customer facing features quickly, but you want to keep the core accounting such that you don't change it too often Then for core accounting we send a message to approvals. This is something that was kind of a surprise for me that actually the approvals, how do you verify that this payment is okay? That's uh that's causing a lot of the issues when you are uh trying to build a banking setup. So you have a lot of compliance rules from the law. For example that you don't send money to uh
Speaker 2: Terrorist you don't send uh money for money laundering purposes and then you of course need to handle your own risk. So you check that Uh is everything okay with the payment? Uh from kind of is is somebody trying to empty the account illegally Then okay somehow we approve the payment by some machine learning approach or manually we get another message for to core accounting that okay this money is now okay to move out of side of Holby Then you get the message to Bank Gateway. Bank Gateway is a system we use for interacting with the external banking world. The external banking world in many cases is based on technologies from the 80s. It's very much batch based. So in the bank gateway we take all the payments
Speaker 2: jail jail jail batch of tästä, say an XML same outside our system. But inside Holy everything is message space, everything is real time but when it gets to the external banking then it turns to a into a batch. Into the other direction it's pretty much the same setup. So now you get the batch from external banking, then you get the payment message to core accounting from your bank gateway. You do the approvals, a couple of messaging go going over in there. Then you get the message to the business tool. So the customer facing back end where we do the all of the bookkeeping and uh stuff like that and finally you try to automate the bookkeeping And or
Speaker 2: depending on the case, you send a notification to the customer, hey, now you received 100 Euros of money So the reliability in this case. For a startup, you might have 100 payments a day, it's low amount. At this point you probably have a kind of a simple system still. So you just use a couple of messages. You don't do the approvals by having a separate system somewhere, but you do uh it Maybe by just using a Django admin and checking that hei it it seems okay. Your reliability might not be that high, three ninths. So one in thousand uh message of these two messages is lost. This means that you get one case a week.
Speaker 2: It's fine. That's easily doable by hand, so no problem in that case. Actually an interesting story about Holy. Beginning very beginning of hallway the system worked such that when when the user types the uh payment out, they filled the form, clicked okay, pay this. What happened in the back end was that somebody copied the data from Django admin to an external bank by hand. And it worked for the varia beginning of Holy. So kind of getting back to the sofistikation vi hade in the uh morning by Daniel. So If you are a startup you can do something very unsophisticated and it works really well
Speaker 2: for for the beginning so you can prove the concept and then start improving it. Now all this now in the mid-sized These numbers are not real numbers from Holvi. We have different numbers, mutta this kind of reflect mitä in Holvi. So 10,000 payments a day, five messages per payment. So now as you saw the flows, the messaging amounts are now a bit Higher, we have the approval system, we have the back end or the customer facing backend, we have the core accounting backend talking with each other Uh the reliability has gone up a bit, so you lose one message out of ten thousand. This means that you get on average
Speaker 2: five cases a day. And if it takes you a couple of hours, if you haven't automated resolving the payments, if it takes a couple of hours per payment, failed payment, you need now a dedicated person, maybe even two dedicated persons to solve out these payments. and it's really not cost-effective. So at this point you want to start to have reliability and we are now building the reliability for the messaging. Finally, on the enterprise level, just as an example, if you get million payments a day, 10 messages per payment, so now you're making it more microservice style. You have improved your reliability, you have invested a lot on your architecture, kind of the hardware And now you have
Speaker 2: 5-9 reliability, which is really good, but you still get 100 cases a day, which means that you have a department of solving these cases. So when you scale up, you want to build more reliability into your system. So the thing that I see very often is that people build these systems, microservices or otherwise systems where you have messaging, but they don't build the reliably into the messaging. That's fine if you don't have that many messages or if it's fine to lose some messages. So if you for example to push notification to the kustemman, it's fine to lose one in ten thousand, it's not a problem at all. But the payment messages each time we lose it, pretty much each time we will get a call from the customer that hey now my money is missing, do something.
Speaker 2: And we get to do the work or we get to pay the money. So we want to build the messaging reliably. So the idea we are using is that we in the origin system, for example in the approval system, when we record that hei okei this payment is now approved. At the same time we record a message to the local database inside the same transaction. So if the approval gets committed to the database, also the data message gets committed to the database. Then on-commit, using Django 's on-commit, you can find more about it in the documentation. We send the message. So only after it has been committed to the database we send the message to the next
Speaker 2: system. The idea over here is that if you send the message before it's committed to the system you can have the hardest to debug fail case. For example, if you have the customer verifying a payment, you first record that inside the transaction to the database. Then you send a message about that and then for some reason the transaction doesn't get committed. You have a message over in the core accounting that okay customer send 100 Euros of money out of the account. But in the origin system you see nothing about that. You might have something in the logs but you have this fantom message you don't know see anything in the database
Speaker 2: And it's really hard to debug these cases. So use on commit if you are doing messaging because In the other direction it's quite reliable, but when you get the error case it's uh really hard to debug. Okay now on commit is not guaranteed to run or you can have some errors. Actually we are using on commit we are actually in Holy we are creating an asynchronous task which we then execute because on commit you don't want to run anything heavin But you can kind of do lightweight HTTP requests, for example, in the on-coming. But it might fail. You might have a Network error. The other system might be down, it might be overloaded. So if you get a batch of payments in, thousand
Speaker 2: payments in jail fire as fast as you can to the other system, it will get overloaded. And then you need to retry. So the idea is that only on commit you send a message from the local database, then you retry if it doesn't go first through. Finally on the receiving side when you are retrying you need to de-duplicate so you get this idempotentia setup where basically the semantics are exactly ones. This way we have reliable mess messaging. This works. I have also been talking with m some consultants from other companies and they are using for banking and they are using something very much like this. It's a proven system. It's
Speaker 2: some some might say that it's a maybe a bit heavyweight, but it's it's really reliable Okay, now you can abstract this model. You can have this inbox-outbox model. So in the origin system you have the outbox where you record the messages. Basically you have an ID, you have. The payload you have a couple of timestamps when the message was created, when was it processed or sent out from the system Then you do the on-comment send and so on. On receiving side all of the messages from the outbox of the other system you store them in an inbox. And you should have a unique constraint on the inbox so that the origins ID you don't process the same message
Speaker 2: multiple times. So doing the deduplication by having a PostgreSQL unique constraint Or well database unique constraints. It doesn't need to be about scratch well. Finally, you do reconciliation checks with the inbox and outbox. You check, okay, in the last hour how many messages did I. Send from system A milka message do I have in the outbox. And in the resiving system you sait check milka message there on the in-box. If they don't, you can react fast Also you can easily, if if some message is lost somewhere, even with this setup, you can
Speaker 2: from the altbox you can just have an admin view and click some message that say, okay, resend this. Okay, how do you actually transport the messages? For reliability, it turns out it's not important. You can use HTTP with REST. So now the simple setup is that you have one origin system, you have one receiver system, and you send the messages just using the standard requests library and then rest framework on the other side. It works pretty well. Of course you can have these cases where you uh overload the other system but it's fine you have this retrying logic so it will be eventually sent over to the other system
Speaker 2: But if you want to do for example this pop sub type of thing , it's really nice to use something in between that does the uh pub sub so you from the origin you still have the outbox from the outbox on commit you send to Kafka Kafka does its thing. It can do things like check which receivers are allowed to see which messages. you kind of get the rate limiting because the receivers can play the lob as fast as they can but if they fall behind it's fine they can just keep on doing their stuff until they catch up And of course the receivers over here they can be different systems.
Speaker 2: And so so you get true pubs up so you can do stuff like okay either to payment in at the same time send a push notification to customer at the same time try to do some machine learning use some machine learning based approach on there kategorisation so we can check that okay how to push this message to bookkeeping and so on. So you can you can do the true pubsub with Kafka. So the transport using Kafka It's important or it's very useful when you are trying to get these benefits, but it's not actually that important for reliability. And finally, of course, you have the outbox inbox module over here still, so you can do the reconciliation of okay, have these messages
Speaker 2: actually been sent So this was mostly about reliability considerations uh of messaging. So now you have a messaging setup that works really well. It turns out as you very likely know that the most likely reason for error is that somebody makes a uh programming mistakes. So what we are doing is that we have these uh of course we are we are using pretty standard uh practices for developing our software so we have this desk drive driven development uh approach we 've right tests we have continuous
Speaker 2: integrations we do reviews in GitHub so everything that goes in it's actually from compliance point of view also important that you review the uh code so everything needs to be checked by somebody else. We react quickly to failures. So This is actually pretty important from the customer's point of view that if a payment is delayed for a couple of hours, it's it's fine if we react to it. If the customer needs to call us You have a problem in kind of this customer trust in Holy. And of course you get to do a lot more work because then you need to do the customer communications that okay yes this payment was
Speaker 2: was lost we we fixed it by doing something and now it's fine. So it costs you a lot of money. Always fix the original reason. So if you find that hei something strange is happening, try to find out what's the original reason. Do do not just uh check that okay I made the payment now go through I just clicked somewhere that okay retry this. But try to find out what's what's wrong in the system because when you are scaling up you want to get uh more and more reliability otherwise it's going to be really really painful if you go for the enterprise level And we are using monitoring and reconciliation. I said a bit about the reconciliation of the in-box-outbox modul ,
Speaker 2: but there is also reconciliation you can do on kind of a business logical level. one system that hurts money du har muit out. Just checking the ira system. Dö ammounts match. If you do this konstant jul. quickly see the cases where you have these uh error cases and you can fix them quickly without needing to wait for customers to Kalla. In muli käsiä itse asiassa. Jos kustomer , itse take time to reaktiä. where it has taken more than a week than when you get the initial call that hey something is broken, we we lose one payment in ten thousand say. Something is broken and then you
Speaker 2: let it run for a while and you lose more payments and then it gets really painful to fix the data. The uh kind of the data is already in the accounting of the company and you need to somehow then fix past data which is something you you really don't want to change the account statement of your uh bank account or the payment account. If you need to do that it's going to cost you a lot in customer trust For monitoring we are also using Datadoc, so on the infrastructure layer that works really well for us. So we get uh and century also so we get a bit uh nice alerts when we have these error. Usually you actually want to check this error rate
Speaker 2: instead of individual errors when you have a bit more volume because it's it's kind of uh assumed that some of the cases are going to fail once and then you retry and then they go through Okay, that was pretty much it. Uh One thing I could say a bit more about is this sophistication. So what I said in the startup slide is that when you have a small system You don't need to have this reliability yet built in. But if you aim to scale or if you are now scaling, you will face this problem at some point if you have these messages. If you're working with money, you can't uh lose if you are working with messages you can lose.
Speaker 2: So in that case I really uh would recommend that you use this this kind of setup where you have absolute reliability on the messaging setup So that was it. Thank you. I'd like to thank the organisers. This has been a really nice example of how to set up a kind of inklusive professional. well organized event while still keeping it relaxed.
Speaker 1: We got some time for questions before we go into lunch. Um
Speaker 3: so about the inbox-outbox model, uh can you talk a little bit about that in the sense of Um are there two completely database clusters also? Are the system separated for security reasons? Because you could as well just have the inbox and outbox model in one database or even in the same table to do the reconciliation
Speaker 2: Uh if you have these two different systems of course then you have need to have them in two different uh databases also if you are using two different databases. I'd we are actually doing we have kind of a large monolith codebase right now, but we are doing messaging from the same codebase back to the same codebase. So uh even in that setup It's kää to have this inbox different modul, outbox different modul. . The outbox records a bit different data. So in the outbox you record that okay the in-box message ID on this. I received it at this time. The kind of in-box create time
Speaker 2: was this. In the outbox you record basically just the data and when you sent it. When you created it. So then you can reconciliate on the create time of the outbox and inbox. And even if you are using it ins inside one. one system, one code base, I'd go for two different database models just for the reason of uh If you want at same point split it up, then it's easy to do. Also when you are a small kompany , I'd not go for mikrosystems but have one large codebase. But have these messaging approaches even if you are using a bigger uh bigger code base so that when you scale you can split it up if you need to
Speaker 1: do more questions This one? To the microphone.
Speaker 4: Uh hey. I was wondering whether you considered instead of this inbox and outbox model to use only append only data structures. and keeping pointers and basically following following the logs from one and copying to the other one. I mean I I'm asking b this because I'm involved with a project. We do hundred million m messages per day and it's not okay to lose a single one. And it's not really hard to do if you emulate the patterns of how Kafka does it internally, also in in your own databases.
Speaker 2: You could depending on the case you could you could use Kafka directly. So so instead of recording anything in the database you first record the message and then you process it in the database. So you go kind of moving from Kafka to transport for messages you move for Kafka as the data storage of the messages. So that's one way. Then you use Kafka for these pointers and stuff like that. You could also do that in the database. I'm I haven't considered how to do that so not not sure how how it would work. The idea over here is that you would use the uh outbox module and send it to Kafka so if you want this hundred m 100 million messages per day so you get it
Speaker 2: reliably to Kafka ja then you can use Kafka for the kind of pointers and replays and stuff like that.
Speaker 4: Yeah. Mm my question was because the outbox introduces a mutable intermediate storage, it introduces also the chance to lose a message.
Speaker 2: Uh it's up and only the outbox. So so or well you will they delete after 30 days. So you just insert new messages over there And the only update you do is that okay now this has been pushed to somewhere. But that's the only update.
Speaker 1: Okay, are there more questions? Going once, going twice. Thank you, Ansi.
Record the business change and an outbox message in the same database transaction, then send the message only after commit. Retry failed deliveries and deduplicate them on the receiving side so messages are eventually processed without losing or double-processing payments.
Discussed at 13:37Using `on_commit()` ensures a message is sent only after the transaction that created its underlying state has committed. Sending before commit can leave the receiving system with a message for a payment that does not exist in the originating database.
Discussed at 13:37The origin system stores outgoing messages in an outbox, while the receiving system stores them in an inbox with a unique constraint on the originating message ID. Reconciliation checks compare the two, and an admin can resend messages that appear to be missing.
Discussed at 16:40No. HTTP and REST can be reliable when combined with an outbox, retries, and deduplication. Kafka is useful for pub/sub, rate handling, multiple consumers, and replay-style workflows, but it is not required for the core reliability guarantees.
Discussed at 18:11Use monitoring and reconciliation both for message delivery and for business-level balances, checking that amounts match between systems. React quickly to discrepancies, investigate the original cause, and monitor error rates rather than treating every isolated retryable error as a crisis.
Discussed at 21:48Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025
Published June 13, 2025