AI in Finance

Developing an LLM-Based Methodology for Automated Data Remediation in Credit Risk Modelling: A Case Study at a Large Dutch Bank

Xander, Bon (2026).  Developing an LLM-Based Methodology for Automated Data Remediation in Credit Risk Modelling: A Case Study at a Large Dutch Bank

Abstract

This research investigates how a modular Large Language Model (LLM)-based methodology can be developed and evaluated to support automated data remediation in such a regulated environment, demonstrated through a case study at a large Dutch bank. Using a Design Science Research Methodology approach, the study develops a governance-aware, evidence-grounded methodology consisting of a two-stage document classification pipeline and a retrieval-augmented generation pipeline for structured information extraction. The methodology is evaluated through a bounded proof of concept focused on valuation-related remediation tasks using an expert-validated proxy dataset. The final two-stage classifier identi fied all relevant valuation reports without false positives. The strongest extraction configuration, based on attribute-specific retrieval and stronger retrieval representations, achieved 11 out of 12 market values correct, 10 out of 12 valuation dates correct, an attribute-level accuracy of 87.5%, and 9 out of 12 reports fully correct. The findings show that performance depended primarily on retrieval quality and evidence ranking rather than on model scaling alone. An indicative operational assessment suggests a task-level effort reduction of 64.7% and a broader process-level throughput reduction of 26.4%, reducing average case handling time from 14.0 to 10.3 working days per case under continued human validation. Overall, the thesis contributes a modular, evidence-grounded, and governance-aware methodology for responsible LLM-based document processing in regulated financial environments.

Supervisors

UT Supervisors

  • Dr. Wouter van Heeswijk
  • Dr. Marcos Machado

ING Supervisors

  • No

Thesis Repository