Northbank: replacing a settlement core while it was still running
Reconciliation ran as a nightly batch. Disputes took four days to clear, the batch had started overrunning into business hours, and the three people who understood it were within six years of retirement.
The situation
Northbank processes settlement for a mid-sized card and payments book. The core was COBOL batch on a mainframe, written in stages between 1994 and 2006, with roughly fourteen years of live history. It worked. That's the thing people forget about legacy systems — it had been correct every night for two decades.
The problem was the edges. Dispute resolution required manual reconciliation across three systems and took four working days. Volume growth had pushed the batch window from four hours to nearly nine, which meant a bad night bled into the trading day. And the institutional knowledge sat with three engineers, one of whom had already announced his retirement date.
The two false starts
We wrote both of these proposals, and we were wrong twice before we were right.
Attempt one was a lift-and-shift of the batch logic into Java, preserving the batch model. It was cheap and quick and would have bought about three years. We built a prototype and killed it ourselves at week seven, because the dispute-resolution problem — the actual business pain — was structural to batch processing. We'd have delivered a faster version of the wrong thing.
Attempt two was a full event-sourced rebuild with a big-bang cutover over a long weekend. Northbank's risk committee rejected it, correctly. We had assumed a level of confidence in the migration that we could not evidence, and the rollback plan was, on honest inspection, a hope.
Attempt three worked: event-sourced ledger, built alongside the mainframe, dual-run against live volume for eleven weeks with automated comparison of every single position. Cut over account segment by account segment. It cost more than either of the first two proposals and took four months longer.
What we built
An append-only event ledger in Java on Kafka, with projections for balance, position and dispute state. Every business event is immutable and replayable, so a dispute becomes a query over history rather than a reconstruction exercise across three systems. Fourteen years of mainframe history was migrated in and replayed to verify the projections reproduced known-correct closing positions for every month-end in that period.
The comparison harness was the real deliverable. For eleven weeks, every transaction went through both systems and a reconciler flagged any divergence within sixty seconds. We found nineteen genuine discrepancies. Eighteen were our bugs. One was a mainframe rounding behaviour on a rare currency pair that had been quietly wrong since 2011, which Northbank's finance team described as the most expensive free gift they'd ever received.
What went wrong
The dual-run cost roughly eleven weeks of double infrastructure and two full-time engineers watching a reconciliation dashboard. Northbank's programme office pushed to shorten it to four weeks at the halfway point, and we pushed back hard. Three of the nineteen discrepancies surfaced after week eight — in month-end processing, which only happens, by definition, monthly.
We also underestimated training. We budgeted two weeks for operations handover and needed five, because the operations team had spent twenty years developing intuitions about the batch that simply didn't transfer. That was our miss, and we absorbed the cost.
Where it stands
Live since March 2022. Settlement completes in under six hours, disputes clear same-day, and the system has had no unplanned outage over eleven minutes. Northbank's own team has run it since month four after launch; four of their eight platform engineers were hired through our recruitment practice and trained on the codebase before go-live.