2026 Research Programme

Testing LLM reliability in corporate credit decisioning.

Capital Benchmark has built an executable research platform to test whether general-purpose LLMs can reason through structured credit policy, challenging edge cases and final decision aggregation with the reliability required in a bank-controlled process.

Independent research · Executable benchmarks · Capability transfer
What we found

Model choice and implementation architecture matter.

The research separates fluent memo writing from the harder operational question: whether a model applies the supplied policy correctly and reaches the right final decision.

45–100%

Wide dispersion in decision accuracy

The four tested configurations produced materially different final-decision accuracy on the same controlled task.

93.7%

Rule accuracy can hide decision failure

Across the 640-memo V2 benchmark, granular policy accuracy materially exceeded aggregate decision accuracy.

100%

Scale result for the selected configuration

GPT-5 retained perfect observed decision accuracy across 1,000 additional public-company obligors.

Business model

Research agenda in, capability transferred out.

Capital Benchmark's operating model is designed to keep research open, independent and actionable. We use a public testing environment to convert a research agenda into a working reference implementation before transferring code, methods and test assets into a bank-controlled environment.

1
Research agendaWe define the questions that matter to credit-risk teams: policy edge cases, model-provider comparisons, architecture choices and operational controls.
2
Public / open testing environmentWe test sample obligor data, policy manuals and LLM configurations in a non-client environment so results can be challenged and discussed without requiring confidential bank data.
3
Capability transfer to banksOnce the reference approach is validated, we transfer code and test data into bank-specific pilots so internal teams can adapt the framework to their own policies, ratings and systems.
Capital Benchmark business model showing research agenda flowing into a public testing environment and then capability transfer to banks.
Suggested asset path: static/business_model_exhibit.png. Save the generated exhibit to your Flask static folder and the landing page will render it automatically.
Research report

LLMs in Credit Decisioning

The report sets out the research goals, data model, policy framework, challenger cases, deterministic ground-truth construction, model configurations, scoring methodology, results, limitations and next research programme.

  • V1 and V2 experimental progression
  • CP-01–CP-11 policy framework and edge-case design
  • Controlled four-model comparison
  • 1,000-obligor scale validation
  • Implications for bank credit operating models
Free research download

Get the full report

We will email the report and use your details to manage research updates and portal-access invitations.
Method

Evidence first, implementation second.

We use a structured test harness rather than anecdotal prompts. The benchmark combines a common obligor dataset, generated facility requests, an explicit policy manual and a deterministic ground-truth engine.

Obligor data & facility context
LLM memo generation
Structured scoring and comparison

The result is an evidence base that can support research, dialogue with credit teams and later institution-specific pilots.