AI Vendor Due Diligence for Banks and Insurers: A Working Guide
How to vet an AI vendor so the file survives a bank exam: the regulatory floor, supply chain map, questions that surface risk, contract clauses, a scoring model and the monitoring plan.

When an AI vendor fails inside a bank, the regulator writes the letter to the bank, not the vendor. The US Interagency Guidance on Third-Party Relationships says it directly: using third parties does not diminish or remove a banking organization's responsibility to operate safely, soundly, and in compliance with the law. Standard software due diligence will not protect you here, because an AI system is not standard software. This guide gives you the full working process: the regulatory floor you will be measured against, the supply chain you need to map, the questions that surface real risk, the contract clauses that matter, a scoring model you can defend in an exam, and the monitoring plan for after signature.
Why AI vendor due diligence is different from software due diligence
Traditional vendor review assumes you are buying a deterministic product. You test it, it passes, and it keeps behaving the same way until someone ships a change you can read in release notes. Four properties of AI systems break that assumption.
The artifact is a model, not code. The vendor's application is a thin layer over a foundation model the vendor often does not control. Your due diligence on the application tells you little about the layer that generates the output.
The behavior is probabilistic. A conventional system either works or throws an error. An LLM returns fluent, wrong answers with the same confidence as right ones. Classical model errors, like misspecification, are traceable through statistical tests; as the Moody's model risk team points out, generative AI breaks that traceability, which is why your validation team cannot rely on its old playbook.
The model changes under your feet. Providers fine-tune, retrain, and deprecate model versions on their own schedule. The system you approved in January may not be the system running in June. The Alan Turing Institute's work on GenAI model risk in financial services stresses that these systems are vendor-dependent and continuously adaptive, which strains traditional model risk management built for static models.
Your data leaves the perimeter in a new way. Prompts, retrieved documents, and outputs flow to infrastructure you do not control. That raises data protection questions that a standard security questionnaire barely touches, especially around cross-border transfer.
None of this means you should avoid AI vendors. It means the due diligence file needs extra sections, and a few old sections need sharper questions.
The regulatory floor you will be measured against
Before you design a questionnaire, know what your examiner will compare it to. Four anchors cover most jurisdictions our readers operate in.
The third-party risk management life cycle
The Interagency Guidance on Third-Party Relationships (OCC Bulletin 2023-17 and its Federal Reserve and FDIC equivalents), finalized in June 2023, organizes third-party risk into five stages:
- Planning: assess the risk of the activity before choosing a vendor.
- Due diligence and third-party selection: gather evidence, evaluate, document why this vendor.
- Contract negotiation: lock in the protections the evidence says you need.
- Ongoing monitoring: verify performance and controls for the life of the relationship.
- Termination: exit without losing data or continuity.
The guidance explicitly rejects a safe harbor for smaller institutions. You may tailor depth to risk, but you may not skip stages. An AI vendor that touches customer data, credit decisions, claims, or client communication sits squarely inside this framework.
DORA for anything touching the EU
The Digital Operational Resilience Act has applied since 17 January 2025. The provision that bites hardest in vendor management is the register of information: financial entities must maintain and update, at entity and consolidated level, a register of all contractual arrangements for ICT services from third-party providers, and produce it to the competent authority on request. An AI vendor is an ICT third-party provider. If your register cannot answer "which vendors run models on which of our data, where" you have a gap, and DORA supervisors can ask for the register at any time.
Kenya's Data Protection Act for East African operations
The Data Protection Act 2019 sets conditions that shape AI vendor selection directly:
- Section 25(h): personal data must not be transferred outside Kenya unless there is proof of adequate data protection safeguards or the data subject's consent.
- Section 31(1): a data protection impact assessment is mandatory before any processing likely to result in high risk to rights and freedoms. An LLM processing customer files qualifies in most cases.
- Sections 48 and 49: set the conditions and required safeguards for transfers out of Kenya.
- Section 41: requires data protection by design and by default, which is an argument you can use internally to demand privacy controls from the vendor rather than bolt them on afterwards.
The Data Protection (General) Regulations 2021 remove any doubt about whether AI triggers the DPIA obligation: regulation 49 lists automated decision-making with legal or similarly significant effect, large-scale processing of personal data, and the innovative use of new technological solutions as processing that requires one. Note also section 31(5): the DPIA report goes to the Data Commissioner sixty days before processing starts, so build that lead time into the procurement plan.
If an AI vendor cannot document where inference happens and which subprocessors touch the data, you cannot complete a section 31 assessment honestly.
NIST AI RMF as the organizing frame
The NIST AI Risk Management Framework gives you four functions: GOVERN, MAP, MEASURE, MANAGE. GOVERN is cross-cutting; the other three are broken into categories and subcategories you can turn into due diligence artifacts. MAP becomes your supply chain inventory. MEASURE becomes the evaluation evidence you demand from the vendor and produce yourself. MANAGE becomes your monitoring and incident plan. None of this is legally binding, but it is the vocabulary US examiners and many group risk functions now speak.
Map the AI supply chain before you send a questionnaire
Most AI vendor failures are fourth-party failures. Before any questionnaire, draw the chain for the specific product you are buying:
- Foundation model provider: Anthropic, OpenAI, Google, Meta weights, and so on. Whose model generates the output?
- Model host: where does inference run? The provider's own cloud, AWS, Azure, GCP, or the vendor's tenancy?
- Application vendor: the company you are actually contracting with, whose workflow, RAG pipeline, and guardrails wrap the model.
- Data processors: vector databases, logging and observability tools, annotation subcontractors.
- Integrators: anyone with API access to your data flows.
For each layer you need the same five answers: what data reaches it, where it is processed, how long it is retained, whether it is used for training, and what happens to it at termination. A vendor that cannot produce this map for its own product has not done the work, and that is itself a finding.
Build the due diligence file: a working procedure
Here is the procedure we use when assembling an AI vendor file for a bank or insurer. Budget two to four weeks for a material vendor.
Step 1: Classify the use case. What decision or output does the system influence, and what is the worst plausible error? A drafting assistant for internal memos and a claims triage model are different risk classes and deserve different depth.
Step 2: Collect the corporate evidence. Financial statements, funding runway, ownership structure, key person dependencies. AI startups fail and get acquired at high rates; continuity risk is real.
Step 3: Collect the security evidence. SOC 2 Type II report or ISO 27001 certificate, penetration test summaries, encryption posture, and incident history. For reference, OpenAI's enterprise documentation describes a completed SOC 2 audit with AES-256 at rest and TLS 1.2 or better in transit; that is the baseline shape of an answer, not the ceiling.
Step 4: Collect the model evidence. Which model versions the product uses, how the vendor evaluates before and after model upgrades, known failure modes, and guardrail architecture. Ask for their evaluation harness results on tasks like yours, then run your own evaluation on your own documents before signature. Our guide to building an LLM evaluation suite for a bank or insurer covers how to build that test set.
Step 5: Collect the data evidence. The data flow map from the previous section, retention schedules, training-usage terms, and subprocessor list. Check that retention is configurable and that deletion is contractual, not aspirational.
Step 6: Complete the privacy assessment. For Kenyan operations, the section 31 DPIA. For EU operations, the GDPR DPIA and a DORA register entry. For US operations, document how the arrangement fits the interagency guidance life cycle.
Step 7: Score, document the decision, and set monitoring triggers. The scoring model below. The file is not done when the contract is signed; it is done when the monitoring plan has owners and dates.
The questions that actually surface risk
Generic questionnaires get generic answers. These are the questions whose answers discriminate between vendors, grouped so you can lift them straight into your template.
Model behavior and change control
- Which exact model versions power the product today, and which are planned for the next two quarters?
- How much notice do you give before a model upgrade, and can we pin a version?
- What is your regression evaluation process before you roll a new model version to customers?
- Can we run our own evaluation set against your system before go-live and after each material model change?
Data handling
- Is any customer data used to train or fine-tune models, ours or anyone else's? Anthropic's commercial terms state that Anthropic may not train models on customer content from its services, and OpenAI states it does not train on business data by default. If your vendor's answer is weaker than the market leaders' published position, that is a negotiating fact.
- Where does inference physically run, and which countries does our data cross? This answer feeds the Kenya DPA section 48 and 49 analysis directly.
- What is retained in logs, for how long, and can we shorten or disable retention?
- Which subprocessors see our data, and how are we notified when that list changes?
Operational reality
- What happens to our data, configurations, and fine-tunes if you are acquired or shut down?
- What are your RTO and RPO commitments, and have they been tested?
- Who at your company owns our account when the founding team moves on?
Red flags that justify walking away
- The vendor cannot name the model version it serves you.
- Training on customer data is opt-out rather than opt-out-by-default, or buried in a changeable policy page rather than the contract.
- No subprocessor list exists, or it includes entities in jurisdictions your regulator will not accept.
- The vendor refuses pre-contract evaluation access. A confident vendor lets you test.
Contract clauses that matter for AI vendors
Standard SaaS paper misses the AI-specific risks. Push these into the agreement or the DPA.
No training on your data, in the contract. Not in a privacy policy the vendor can edit. The clause should cover inputs, outputs, and retrieved context, and survive model provider changes.
Model change notification and pinning. Minimum notice period before model version changes, the right to pin a version for a defined window, and the right to re-evaluate before accepting a change.
Evaluation and audit rights. The right to run your evaluation suite against the production system on a schedule and after material changes, plus audit or questionnaire rights proportionate to risk.
Data location and subprocessor control. Committed processing regions, a contractual subprocessor list, advance notice of changes, and the right to object. This is what makes a DORA register entry and a Kenya DPA transfer analysis possible.
Incident notification. Defined security incident and material model-failure notification timelines, measured in hours, not "without undue delay" alone.
Exit and portability. Export of your data, prompts, fine-tune artifacts, and evaluation results in usable formats, with certified deletion afterwards. Termination is stage five of the interagency life cycle; plan it at stage three.
Output ownership and IP indemnity. You should own outputs, as both Anthropic's and OpenAI's commercial terms already concede, and the vendor should indemnify against third-party IP claims on outputs within stated limits.
A scoring model you can defend in an exam
Examiners do not expect perfection; they expect a documented, repeatable decision. A weighted scorecard gives you that. Illustrative weights for a material vendor:
- Data handling and privacy: 25%. Training usage, retention, data location, subprocessor control.
- Security: 20%. Certifications, penetration testing, incident history.
- Model governance: 20%. Evaluation evidence, change control, guardrails.
- Contract terms: 15%. The clauses above, as actually agreed.
- Financial and operational resilience: 10%. Runway, continuity, concentration risk.
- Exit and portability: 10%. Reversibility of the decision.
Worked example with illustrative numbers. You score Vendor A, a well-funded AI document review platform, against Vendor B, a cheaper alternative:
- Vendor A: data handling 90, security 85, model governance 80, contract 75, resilience 85, exit 70. Weighted score: (90x0.25) + (85x0.20) + (80x0.20) + (75x0.15) + (85x0.10) + (70x0.10) = 22.5 + 17 + 16 + 11.25 + 8.5 + 7 = 82.25.
- Vendor B: data handling 55, security 70, model governance 60, contract 50, resilience 65, exit 45. Weighted score: 13.75 + 14 + 12 + 7.5 + 6.5 + 4.5 = 58.25.
Set a floor before you start: say 70 overall and no category below 50. Vendor B fails on both counts regardless of price, and the file shows why. Adjust the weights to your risk appetite, but fix them before vendor demos begin, or the demo will fix them for you.
Ongoing monitoring after signature
Stage four of the life cycle is where most programs go quiet. For AI vendors, monitoring has an extra dimension because the product changes without a release you asked for. Supervisors keep extending classic model risk discipline to GenAI rather than exempting it; our guide to SR 26-2 and GenAI model risk covers what that means for your model inventory and validation standards.
- Re-run your evaluation suite on a fixed cadence, monthly for material systems, and after every notified model change. The Turing Institute's guidance is blunt: GenAI systems require continuous monitoring of performance, with input-level data checks linked to RAG-specific checks where retrieval is involved.
- Track model versions actually served against versions approved. Ask the vendor for a version log if they do not publish one.
- Review the subprocessor list quarterly against your contract.
- Watch the vendor's corporate health: funding news, layoffs, leadership exits, acquisition rumors.
- Keep the register current. Your DORA register of information, or its equivalent in your jurisdiction, should reflect reality at all times, not at exam time.
- Rehearse the exit at least annually for critical vendors. An exit plan nobody has tested is a document, not a plan.
Common mistakes
- Reviewing the application and ignoring the model. The foundation model provider is in your supply chain whether or not you contract with it. Read its terms too.
- Accepting policy-page promises. Anything that matters belongs in the contract. Policy pages change unilaterally.
- Scoring after the demo. Decide weights and floors first, or you will rationalize the vendor you liked.
- Treating the DPIA as a form. Under Kenya DPA section 31 it is a pre-processing obligation with real content; a thin DPIA is worse than none because it documents that you did not look.
- No evaluation before signature. If you have not run your own documents through the system, you have approved a demo.
- Forgetting termination. Data return, deletion certificates, and continuity plans belong in the contract, not in a post-incident scramble.
When full due diligence is the wrong use of time
Tier your effort or the program collapses under its own weight. A low-risk tool, no customer data, no decision impact, reversible in a day, needs a light review: security posture, data handling basics, contract essentials. Save the full file for material arrangements. The interagency guidance explicitly allows tailoring depth to risk; it does not allow skipping stages. The mistake is applying one depth to everything, which either buries low-risk procurement in process or lets high-risk vendors through on a light checklist.
Frequently asked questions
Is AI vendor due diligence legally required for banks?
In substance, yes. The US interagency third-party guidance applies to any business arrangement, and AI vendors qualify. DORA requires EU financial entities to manage ICT third-party risk and maintain a register of information. Kenya's Data Protection Act forces a DPIA for high-risk processing and restricts cross-border transfer. The combination makes a documented process unavoidable.
Can we rely on the vendor's SOC 2 report alone?
No. SOC 2 covers security controls at the vendor, not model behavior, training-data usage, or fourth-party risk from the foundation model provider. Treat it as one input among several, and check which systems and trust service criteria it actually covers.
What if the vendor uses a foundation model from Anthropic or OpenAI?
Then your supply chain includes that provider. Read the provider's commercial terms, since they govern training usage, retention, and output ownership, and verify the vendor passes the relevant protections through to you in the contract rather than relying on the provider's defaults changing.
How long should AI vendor due diligence take?
Two to four weeks for a material vendor is realistic if the vendor responds promptly. Low-risk tools can clear in days. If a vendor cannot produce basic artifacts like a subprocessor list or evaluation evidence within a week, that response time is itself a finding.
Do we need to redo due diligence when the vendor swaps model versions?
Not fully, but a material model change should trigger your regression evaluation, a check that contract protections still apply to the new model and provider, and a register update. Your contract should obligate the vendor to notify you before this happens.
What if no vendor passes our scorecard floor?
That happens, and it is useful information. Options include negotiating the gaps into the contract, narrowing the use case so less data flows to the vendor, redacting PII before data reaches the model, or moving to a private deployment where the model runs in your VPC or on your hardware. Our guide to private LLM deployment for banks compares those options, and the RAG or fine-tuning decision guide covers the build-side trade-offs.
Where to go from here
Vendor due diligence is one file inside a larger governance pack: model inventory, validation standards, evaluation suites, monitoring plans, and the registers your supervisor expects. We build that documentation and the evaluation infrastructure behind it for banks, insurers, and funds through our enterprise AI practice, and we implement Anthropic's finance agents for fund teams through Claude for Financial Services. If your next exam or procurement round needs a defensible AI vendor file, talk to an engineer.
About AI Agents Plus Editorial
AI automation expert and thought leader in business transformation through artificial intelligence.



