Postdata AIAI
WorkServicesResearchAbout
WorkServicesResearchAbout
Postdata AIAI
WorkServicesResearchAboutContact

© 2026 Postdata AI. All rights reserved.

PostdataAI Pty Ltd·ABN 97 688 913 340·ACN 688 913 340·NSW 2010, Australia

  1. Research
  2. /LLM-based credit scoring from bank statements

LLM-based credit scoring from bank statements

credit scoringLLMfintechdocument extraction

How we turn bank statement PDFs into an explainable Financial Health Score: DotsOCR extraction, a five-stage pipeline, and a strict rule about where the LLM may speak.

Most credit models are trained on labels: loans that were repaid and loans that were not. In a young market there are not enough of those labels to train on, and the applicants who most need credit are exactly the ones with no history. Mission Mobile, a South African fintech, asked us a direct question: can large language models support modern credit underwriting where labelled data is scarce?

The answer we shipped in six weeks is a two-part system. The first part turns bank statement PDFs into clean transactions. The second turns those transactions into a Financial Health Score (FHS): twelve quantitative indicators and a short, evidence-backed credit report that an analyst can audit line by line. This article explains how the pipeline is built, where the LLM is allowed to speak, and where it is not.

Why bank statements

A bank statement is the richest document most applicants can produce. It shows income timing, spending after payday, savings discipline, recurring obligations and the small behavioural signals that a bureau score flattens into one number. The catch is that every bank formats it differently, and most of the useful structure is in tables.

Earlier attempts at the problem combined rules with OCR. They struggled with layout variability, especially borderless tables, and produced superficial scores with no behavioural explanation. The client needed three things at once: support for many banks, a score with a "why", and costs that stay sane at scale.

Step 1: clean transactions out of PDFs

We iterated through three generations of extraction. PDF-to-Markdown tools were fast but lost table structure. MistralOCR handled scans better but still broke on borderless layouts. The final architecture is built on DotsOCR with a bank classifier in front of it and bank-specific formatters behind it, written with BeautifulSoup and Pandas.

The classifier decides which bank produced the statement. The matching formatter knows that bank's quirks: where the balance column sits, how multi-line descriptions wrap, which rows are summaries rather than transactions. Adding a new bank means adding a classifier rule and a formatter to a registry, not retraining anything. The service runs in the cloud with both synchronous and asynchronous ingestion, so a single statement returns quickly and a batch does not block.

Step 2: the five-stage Financial Health Score

The scoring pipeline ships as an installable Python package with typed configs, schemas, logging and tests. It has five stages, and the division of labour between code and model is the whole design.

  1. Processing and features. Regex categorisation of transactions, engineered features and monthly aggregates. Deterministic, testable, cheap.
  2. LLM enrichment. Batched, asynchronous calls to gpt-4.1-mini enrich individual transactions with context that rules cannot infer, such as what kind of merchant a cryptic description refers to.
  3. Insights. Grounded prompting: the model must cite the transactions it relies on and the base rates it compares against. An insight without evidence is rejected.
  4. Indicator calculation. Twelve core Prosperity and Stability indicators, computed in code from the features and the enriched data.
  5. Score adjustment. A holistic LLM review that may move indicators by bounded deltas, each with a plain-language justification.

The rule that runs through all five stages: deterministic math lives in code, and the LLM handles enrichment and interpretation. The model never invents a number. It explains numbers that the code produced, and when it proposes a change, it has to say why in words a credit analyst can check.

Where the LLM is allowed to speak

Three risks shape an LLM-based credit report: hallucination, cost and fairness.

Hallucination is controlled by grounding. Every insight must point at transaction evidence and a base rate, so a claim like "spending rises sharply after salary" is backed by the rows that show it. Cost is controlled by scope: the model sees batches of transactions for enrichment and a compact summary for review, not whole statements over and over. Fairness is the reason the score is built from behaviour rather than from labels: with few repayment outcomes to learn from, a model that scores post-salary spending and savings discipline is more defensible than one that learns proxies from a small, skewed history.

The report an analyst can audit

The output is not a single number. It is twelve indicators with concise, evidence-backed summaries, plus the deltas from the final review and their justifications. An analyst can read why an applicant scored as they did, trace each statement back to the transactions, and disagree with a specific line rather than with a black box. That is what makes the result usable in underwriting and explainable to a regulator.

In practice the behavioural indicators, such as post-salary spending and savings discipline, make the scoring more competitive than traditional models for applicants with thin files, which is the population the client cares about.

What we would do next

The roadmap we agreed with Mission Mobile includes OpenRouter failover so no single model provider is a point of failure, confidence intervals on each indicator, continued benchmarking of models against the same statements, and retrieval-augmented generation for merchant metadata and local economic signals, so that "a supermarket in Soweto" and "a supermarket in Sandton" are read in context.

The case study has the team, timeline and stack: LLM-based credit scoring for Mission Mobile. If you are building a credit report tool on top of LLMs and want the deterministic parts to stay deterministic, talk to us.

Work with us