AI in Finance

Machine Learning for Predicting Energy Performance Certificates: Evidence from a Dutch Residential Mortgage Portfolio

Bjorn B.J.T. Oude AVenhuis (2026). Machine Learning for Predicting Energy Performance Certificates: Evidence from a Dutch Residential Mortgage Portfolio

Abstract

This research examines whether missing or invalid energy performance certificate labels for residential properties can be estimated using only publicly available data. This question is highly relevant for financial institutions, as incomplete EPC coverage limits the quality of ESG reporting, portfolio monitoring, and transition risk assessment. To address this problem, open data at the building, neighbourhood, and grid level were combined to predict EPC labels without relying on private customer data or detailed building-envelope variables. Three models were compared: an ordered logistic regression as an interpretable baseline, XGBoost as a non-linear ensemble model, and a multilayer perceptron as a neural-network approach. The results show that XGBoost performs best overall, with its advantage being most pronounced for exact class prediction, while differences in ordinal agreement between models are more modest. The ordered logistic regression remains useful as a transparent reference model and produces relatively few extreme errors, but does not perform best on average performance metrics. The most important predictor is construction year, while secondary determinants are model-dependent and primarily consist of neighbourhood-level and energy-related proxy variables. Additional validation supports the practical usefulness of the approach. XGBoost passes a label-permutation sanity check, significantly outperforms naive benchmark models, and maintains a useful level of performance when applied to the ING mortgage portfolio. At the same time, model performance remains moderate rather than highly precise: lower EPC classes are particularly difficult to distinguish, and predictions should therefore not be interpreted as substitutes for certified labels at the individual property level. The main managerial implication is therefore not that missing EPC labels can be “automatically resolved,” but that a bank can use open data to perform a defensible and scalable indicative enrichment of its portfolio. Recommended applications include portfolio analysis, ESG reporting, transition-risk screening, and prioritization of additional data collection, particularly for older and more mature loans where label coverage is missing or outdated. Governance should explicitly monitor systematic overprediction of favourable EPC labels, as this could create greenwashing risk in ESG reporting. For more recent or post-2021 certified dwellings, it is generally preferable to retain an existing certified or expired label rather than generate a new model-based estimate. In practical use, it is essential to account for temporal inconsistencies in the data, potential noise in historical labels, and the fact that secondary model drivers are less stable than the dominant influence of construction year. The approach is specifically tailored to the Dutch data environment and is not directly transferable to countries with different EPC systems or open-data availability

Supervisors

UT Supervisors

  • Dr. Laura Spierdijk
  • Dr. Marcos Machado

ING Supervisors

  • Ruben Siemerink

Thesis Repository