AI in Finance

Confidence scores for Large Language Models

Joost, DEKKER(2026). Confidence scores for Large Language Models

Abstract

This research compares four confidence-estimation methods, P(True), P(IK), Semantic Entropy, and a trained corrector model, together with the raw token probability on multiple choice, across three QA datasets that span closed-book (TriviaQA), short-context multiple choice (CosmosQA), and grounded long-document (NarrativeQA) question answering. Answers are generated by Gemini 2.5 Flash Lite with reasoning enabled and disabled, methods are selected on a development split, and all results are reported on a held out split, including pairwise combinations of methods fitted on a separate split. Three findings stand out. First, the best method is highly dataset-dependent: Semantic En tropy leads on closed-book QA but degrades on grounded long-document QA, where question-level signals hold up best and overall discrimination drops from 0.94 to 0.71 AUC-ROC. Second, reasoning improves accuracy but saturates token-derived confidence scores towards 0 and 1; running verification calls without thinking or scoring by sampling frequency restores a graded signal. Third, combining two confidence scores improves dis crimination on every dataset, contradicting results form literature that stated that most exiting methods where not complementary.

Supervisors

UT Supervisors

  • Dr. Marcos Machado
  • Dr. Wouter van Heeswijk

ING Supervisors

  • Nilay Karahan Afacan
  • Dr. Max Baak


Thesis Repository