Wide dispersion in decision accuracy
The four tested configurations produced materially different final-decision accuracy on the same controlled task.
Capital Benchmark has built an executable research platform to test whether general-purpose LLMs can reason through structured credit policy, challenging edge cases and final decision aggregation with the reliability required in a bank-controlled process.
The research separates fluent memo writing from the harder operational question: whether a model applies the supplied policy correctly and reaches the right final decision.
The four tested configurations produced materially different final-decision accuracy on the same controlled task.
Across the 640-memo V2 benchmark, granular policy accuracy materially exceeded aggregate decision accuracy.
GPT-5 retained perfect observed decision accuracy across 1,000 additional public-company obligors.
Capital Benchmark's operating model is designed to keep research open, independent and actionable. We use a public testing environment to convert a research agenda into a working reference implementation before transferring code, methods and test assets into a bank-controlled environment.

static/business_model_exhibit.png. Save the generated exhibit to your Flask static folder and the landing page will render it automatically.The report sets out the research goals, data model, policy framework, challenger cases, deterministic ground-truth construction, model configurations, scoring methodology, results, limitations and next research programme.
We use a structured test harness rather than anecdotal prompts. The benchmark combines a common obligor dataset, generated facility requests, an explicit policy manual and a deterministic ground-truth engine.
The result is an evidence base that can support research, dialogue with credit teams and later institution-specific pilots.