How to Build an LLM Evaluation Suite for a Bank or Insurer
A practical build guide to LLM evaluation for regulated institutions: golden datasets from real work, calibrated judges, CI release gates, production sampling, and the documentation validators expect.

A bank or insurer that ships an LLM feature without a written, repeatable evaluation suite is one model update away from an incident it cannot explain to its regulator. The fix is not a bigger model or a better prompt. It is an evaluation harness: a versioned set of test cases drawn from real work, scorers that measure the failure modes you actually care about, CI gates that block releases when scores drop, and production sampling that keeps measuring after launch. This guide walks through how to build one that a model validator or supervisor can read and accept.
Why regulators expect a written evaluation practice
The supervisory logic is older than generative AI. The Federal Reserve and OCC's SR 11-7 guidance on model risk management defines a model broadly, as any quantitative method or system that applies theories, techniques and assumptions to turn input data into estimates, and requires banking organizations to manage the adverse consequences of decisions based on models that are incorrect or misused. That means robust development, effective validation, and sound governance. An LLM that drafts credit memos, classifies complaints, or summarizes claims fits that definition comfortably, and validators are increasingly treating it that way. You can read the letter itself at the Federal Reserve's SR 11-7 page.
NIST's Generative AI Profile of the AI Risk Management Framework (NIST AI 600-1, July 2024) makes the measurement duty explicit. Its MEASURE 1.1 action says approaches and metrics for AI risks should be selected starting with the most significant risks, and that risks which cannot be measured must be documented as such. That second clause matters: "we cannot measure hallucination rates yet" is an acceptable position only if it is written down with a rationale and a plan. The profile also formalizes red-teaming as controlled exercises to surface adverse behavior and stress-test safeguards. The full text is at NIST AI 600-1.
The cost of skipping this is no longer hypothetical. IBM's 2025 Cost of a Data Breach Report, cited in Galileo's evaluation framework guide, found that 13% of organizations reported a breach involving their AI models or applications, and 97% of those lacked proper AI access controls. For institutions under the EU AI Act's high-risk regime, which includes credit scoring, the most serious violations carry penalties of up to 35 million euros or 7% of global turnover.
What an evaluation suite actually is
Strip away vendor language and an eval suite has four parts:
- Golden datasets. Versioned test cases built from real work: inputs, the context the system saw, and expected outputs or grading criteria. These live in your repository, not in a slide deck.
- Scorers. A mix of deterministic code checks, reference-based metrics, LLM-as-a-judge graders, and human review. No single method covers every failure mode, a point the Arize LLM evaluation guide makes well.
- Release gates. The suite runs in CI on every prompt, model, retrieval, or orchestration change. A score drop beyond an agreed threshold blocks the release, the same way a failing unit test does.
- Production sampling. A slice of live traffic gets scored continuously, so drift, new input distributions, and silent regressions surface before customers or examiners do.
The deliverable your validator cares about is the combination: the datasets, the scorer definitions, the thresholds with their rationale, the run history, and the sign-off record.
Step one: build golden datasets from real work
Generic benchmarks tell you almost nothing about whether your system handles your policy documents, your product names, your edge cases. The first artifact to build is a golden dataset per use case, assembled from production traces, historical tickets, or documents your subject-matter experts select.
Practical guidance:
- Size. Start with 100 to 300 cases per use case. Below 100, a single bad case moves your pass rate by a full point and release gates become noisy. Above a few hundred, runs get slow and expensive enough that teams stop running them.
- Stratify deliberately. Cover the happy path, known hard categories, adversarial inputs (prompt injection attempts, out-of-scope requests), and the long tail of formats your documents actually arrive in. A dataset that is 95% easy cases will report a score your risk team mistakes for safety.
- Include what the system saw. For RAG features, store the retrieved chunks alongside the input and output. Retrieval failures and generation failures need different fixes, and you can only tell them apart if you kept the context.
- Version it. The dataset is code. Changes to it get reviewed, because a quietly weakened test set is a quietly weakened control.
OpenAI's open-source Evals framework captures the right philosophy: it supports private evals built on your own data that represent the patterns in your workflow without exposing that data publicly, and OpenAI's own team describes building high-quality evals as one of the highest-impact activities when shipping LLM systems. You do not need to use that specific framework, but your datasets should follow the same principle.
Step two: choose metrics that map to failure modes
Pick metrics by asking "how does this feature fail expensively?" and working backward. For the common banking and insurance patterns:
- RAG over policy or product documents. The open-source Ragas library provides the standard set: faithfulness (does the answer stick to the retrieved context), response relevancy, context precision and context recall (did retrieval fetch the right chunks). Faithfulness is the one your compliance team will ask about first, because an unfaithful answer over a lending policy is a misstatement to a customer.
- Extraction tasks. Covenants, dates, amounts, names. Use deterministic checks: exact match, numeric tolerance, schema validation. These are cheap, perfectly reproducible, and the easiest to defend to a validator.
- Agentic workflows with tool calls. Ragas and similar libraries offer tool call accuracy, tool call F1, agent goal accuracy, and topic adherence. Measure the trajectory, not just the final answer: an agent that reaches the right conclusion through the wrong tool call is a control failure, not a success.
- Tone, completeness, and style. This is where LLM-as-a-judge earns its cost. Define a rubric, grade on a small scale, and calibrate the judge against human labels before you trust it.
On calibration: an LLM judge is itself a model, and under SR 11-7 logic it inherits model risk. Before using one in a release gate, label 150 to 250 outputs by hand, run the judge on them, and measure agreement. If the judge agrees with your reviewers less than about 85% of the time on a critical metric, tighten the rubric, give it fewer grade options, or keep humans in that loop. Document the agreement study; it is the validation evidence for your validation tool.
A worked example: a credit memo summarizer
Make this concrete. A regional bank builds a RAG assistant that answers credit officers' questions over policy documents and draft memos. The evaluation build looks like this:
- Dataset: 180 goldens, versioned in git. 60 policy questions with cited passages, 60 covenant and figure extraction cases with exact expected values, 60 memo summarization cases with rubric-graded expected outputs. 25 of the 180 are adversarial: injected instructions inside documents, requests for customer data the officer is not entitled to, out-of-scope legal advice prompts.
- Scorers: exact match and numeric tolerance on extraction; faithfulness and context recall from Ragas on the RAG answers; an LLM-as-a-judge rubric for summarization quality, calibrated against 200 human-labeled outputs with 88% agreement, documented in a one-page study.
- Thresholds: faithfulness at or above 0.95, context recall at or above 0.90, extraction exact match at 100% on critical financial fields, and zero passes allowed on the 25 adversarial cases. Each threshold has a written rationale.
- Run cost: one full run is roughly 180 cases times about 3,000 tokens of generation, around half a million tokens, plus roughly 1.5 million judge tokens across four LLM-graded metrics. At current API prices that is single-digit dollars per run, cheap enough to run on every merge.
- Human review: the initial 200-label calibration took about 33 analyst-hours at ten minutes per label. Ongoing, reviewers sample 20 production outputs a week, roughly 3 hours, and disagreements feed back into the golden set.
The numbers that matter to the validator are not the scores themselves but the structure: stratified dataset, calibrated judge, thresholds with rationale, and a run history showing the gates actually fired and blocked two regressions before release.
Step three: wire evals into CI as release gates
An eval suite that runs when someone remembers to run it is a demo, not a control. The gate belongs in the same pipeline as your unit tests.
promptfoo is the lowest-friction starting point: an open-source CLI that evaluates prompts, models, and RAGs against your test cases, red-teams for security and compliance risks, and integrates with CI including a GitHub Action. A minimal config looks like this:
prompts: ["prompts/credit_memo_v3.txt"]
providers: ["openai:gpt-4o", "anthropic:claude-sonnet-4-5"]
tests:
- vars:
question: "What is the maximum LTV for owner-occupied mortgages?"
context: "{{retrieved_chunks}}"
assert:
- type: llm-rubric
value: "Answer states 80% and cites the policy section. No invented figures."
- type: not-contains-any
value: ["85%", "90%"]
If your team is Python-first, DeepEval gives you pytest-style tests: an LLMTestCase with input and actual output, research-backed metrics like GEval for custom criteria, assert_test in CI, and tracing that scores agent trajectories span by span so you can tell whether the retriever, the planner, or the generator caused a failure. It also supports online evals over production traces, which gets you step four partly for free.
Two rules make gates stick. First, gate on deltas against the baseline, not just absolute scores, so a model upgrade that drops faithfulness two points gets caught even if the absolute number still looks fine. Second, make the gate hard to bypass: overrides require a named approver and a written reason, and both land in the audit log.
Step four: keep evaluating in production
Pre-release evals answer "is this change safe to ship." They do not answer "is the system still safe," because inputs drift, source documents change, and upstream model versions shift under you. The production layer is simpler than teams expect:
- Sample, do not sieve. Score a fixed slice of live traffic, 1 to 5% for most volumes, with the same scorers you used offline. Full-traffic scoring is usually wasted money.
- Watch deltas on the same goldens. Re-run the golden set against production weekly. If production scores diverge from CI scores, your staging environment is not representative and that itself is a finding.
- Close the loop. Every production failure that human review catches becomes a new golden. The dataset should grow fastest in the first quarter after launch.
- Log for audit. Inputs, retrieved context, outputs, scores, and the model and prompt versions that produced them, retained to your normal records schedule. When an examiner asks "what did the system say to this customer and why," this log is the answer.
The documentation a validator will ask for
Budget for the paper trail from day one. Assembling it retroactively is where projects stall:
- Use-case inventory: every LLM feature, its owner, its risk tiering, and which decisions it influences.
- Dataset documentation: how goldens were sourced, stratified, and reviewed; who signed off; change history.
- Metric and threshold rationale: why each metric maps to a real failure mode, why each threshold is set where it is, and what the residual risk is below it. NIST AI 600-1's requirement to document what you cannot measure belongs here.
- Judge calibration studies: agreement rates between LLM judges and human reviewers, refreshed when judges or rubrics change.
- Run history and gate overrides: every CI run, every blocked release, every override with approver and reason.
- Red-team results: scope, findings, remediations, and retest evidence for each use case tiered as high risk.
- Change management: how model version upgrades, prompt changes, and retrieval changes flow through the gates, and who may approve them.
Common mistakes
- Benchmark shopping. Reporting MMLU-style public benchmark scores as evidence your policy Q&A bot works. Public benchmarks measure the base model; your risk lives in your data, your prompts, and your retrieval.
- Easy goldens. Datasets built from documentation examples rather than messy production reality. The suite passes, the system fails.
- Uncalibrated judges. Trusting an LLM-as-a-judge score to three decimal places without ever measuring its agreement with humans.
- Gating on vibes. Thresholds chosen so the current system passes, rather than from the failure cost. If every threshold was set after looking at the scores, they are decoration.
- Evaluating only the final answer. In agentic systems the wrong trajectory with the right answer is still an incident waiting for a slightly different input.
- Treating the suite as done. Model providers ship new versions, policy documents get revised, attackers iterate. An eval suite with a last-commit date eight months ago is evidence of a control gap, not a control.
When a full eval suite is the wrong investment
Honesty helps here. If you are still in discovery, running a two-week pilot on non-sensitive internal documents with no customer exposure, a 30-case spreadsheet and three reviewers is proportionate; build the harness when the pilot earns a production decision. If the use case is low-risk and fully internal, say drafting meeting notes, the full SR 11-7-style apparatus is overkill and a lightweight version of steps one and three is enough. And if your volume is so small that a human can review 100% of outputs, human review is the control; automate evaluation only when review stops scaling. The full build is justified when outputs reach customers, influence regulated decisions like credit or claims, or run at volumes where sampling is the only review possible.
Frequently asked questions
How many test cases do we need before go-live?
For a single use case, 100 to 300 well-stratified goldens is a defensible starting point. Below 100, pass rates are too noisy to gate on. More important than the count is coverage: hard categories, adversarial inputs, and the formats your real documents arrive in.
Is LLM-as-a-judge acceptable to a regulator?
Yes, if you treat the judge as a model and validate it like one. That means a documented calibration study against human labels, agreement thresholds, periodic recalibration, and human review retained for the highest-risk decisions. An uncalibrated judge is an unvalidated model grading another unvalidated model.
What is the difference between evals and red-teaming?
Evals measure whether the system does its job well; red-teaming tries to make it fail on purpose, through prompt injection, data exfiltration attempts, and abuse scenarios. NIST AI 600-1 treats red-teaming as a distinct control, and for anything customer-facing or high-risk you need both.
Can we just buy an observability platform instead of building this?
Platforms like the ones behind DeepEval, Ragas, and Arize give you scorers, tracing, and dashboards, and they are worth using. What no vendor can give you is the golden dataset from your documents, the thresholds tied to your failure costs, and the sign-off trail. Those are the parts a validator actually reads, and they are yours to build regardless of tooling.
How much does it cost to run an eval suite?
Less than the meeting where you decide to build it. A few hundred cases per run typically costs single-digit dollars in tokens, and CI integration is an afternoon of work with open-source tools like promptfoo or DeepEval. The real cost is the analyst time to build and label the initial golden dataset, typically a few person-weeks per use case.
Do open-weight models change the evaluation picture?
They help. Running an open-weight model such as Llama, Qwen, Mistral, or Gemma inside your own perimeter means your golden datasets and production logs never leave your environment, and the model version cannot change under you without a deliberate, gate-checked upgrade. You trade provider-managed scaling for version control and data residency, which is usually the right trade in a regulated setting. Note that Claude models are not available for customer fine-tuning; where fine-tuning is needed, open-weight models are the path.
Where to go from here
The build order that works: one high-value use case, one golden dataset built from real documents, deterministic and RAG metrics first, a calibrated judge second, CI gates third, production sampling fourth, and the documentation assembled as you go rather than after. If you want a second pair of hands, our enterprise AI practice for banks and insurers designs and builds evaluation harnesses, private model deployments, and the governance documentation around them, and our Claude for financial services work covers the same ground for funds and advisors. For the broader engineering picture, see our guides to enterprise LLM integration and function calling in production.
About AI Agents Plus Editorial
AI automation expert and thought leader in business transformation through artificial intelligence.



