LLM Monitoring and Audit Logging for Banks and Insurers
Production LLM features in banks need traces, scores and audit-grade logs, not just pre-launch tests. What regulators require, what to capture, and what it costs.

A bank that tested its LLM before go-live and then stopped watching it has done half the job. Regulators, auditors and your own incident reviews will ask a harder question: what has the system been doing since launch, and can you prove it? The answer is a production monitoring and audit-logging layer built on traces, and this guide shows what to capture, which rules it satisfies, what it costs, and where teams get it wrong.
This is written for technology and risk leads at banks and insurers running LLM features in production: a credit-memo assistant, a claims summariser, a KYC copilot. The same discipline applies whether you call a managed API or run open-weight models in your own VPC.
Why pre-deployment testing is not enough
Pre-deployment evaluation tells you how a system behaved on a fixed test set, on a fixed day, against a fixed model version. Production is none of those things. A 2026 position paper from practitioners validating GenAI systems in financial institutions puts it plainly: benchmark performance alone is not evidence of readiness, and financial LLM systems need system-level validation across data, retrieval, agent behaviour and governance, continuously, not once (Benchmarks Are Not Validation, arXiv 2607.28840).
Three forces degrade a good launch:
- Model drift you do not control. Providers update model versions. A prompt that scored well in March can behave differently in October on the same input.
- Data drift you cannot see. Customer queries shift with products, seasons and news. Your test set freezes the past.
- Prompt and config change. Someone edits a system prompt, bumps a temperature, swaps a retrieval index. Without versioning, you cannot correlate a quality drop to the change that caused it.
The fix is not more pre-launch testing. It is treating the LLM feature like the payment system next to it: instrumented, logged, alertable, auditable. If you have not yet built the pre-launch side, start with building an LLM evaluation suite for a bank or insurer, then come back here for what happens after go-live.
What regulators actually ask for
The EU AI Act is the clearest written standard, and it matters even for non-EU institutions because it is becoming the reference text auditors reach for. Three articles do the heavy lifting:
- Article 12, record-keeping. High-risk AI systems must "technically allow for the automatic recording of events (logs) over the lifetime of the system", with logging appropriate to the system's intended purpose (EUR-Lex text via artificialintelligenceact.eu). Credit scoring of natural persons and life or health insurance risk pricing sit in the high-risk Annex III categories, so a large share of bank and insurer LLM use cases land here.
- Article 26, deployer duties. If you deploy a high-risk system, you must keep the logs it automatically generates for a period appropriate to its purpose, at least six months unless other law says otherwise (Article 26). Deployer, not just provider: using a vendor model does not move this duty off your desk.
- Article 72, post-market monitoring. Providers must run a documented post-market monitoring system that collects and reviews experience from real use (Article 72). If you fine-tune or substantially modify a vendor model, Article 25 can reclassify you as the provider, inheriting this obligation.
The Digital Omnibus package pushed the Annex III high-risk obligations from August 2026 to 2 December 2027 for stand-alone systems and August 2028 for AI embedded in regulated products, enacted through Regulation (EU) 2026/1744, in force since 27 July 2026 (timeline analysis). That is breathing room, not an exemption. Building a logging and monitoring layer takes months, and retrofitting it onto a live system costs far more than shipping it with the system.
For Kenyan and East African institutions, the EU act applies directly only where you serve EU customers, but the pattern travels: data protection authorities, central bank ICT risk guidance and internal audit all converge on the same ask. Show what the system did, when, on whose instruction, with which model and prompt version.
The three signals that matter
LLM observability tooling converges on three signal types. Langfuse's observability documentation frames them well, and the framing holds across vendors:
Traces
A trace is the full execution path of one request: the exact prompt sent, model parameters, retrieved documents, tool calls, intermediate steps, the final response, latency and token counts. Traces answer "what exactly happened for customer case 4815 at 14:32?" This is the audit log. For an agentic workflow with multiple model calls and tool invocations, the trace is often the only way to reconstruct why the system took an action.
Metrics
Aggregated over traces: latency percentiles, error rates, token usage and cost per feature, per team, per customer segment. Metrics answer "is the system healthy and what is it costing us?" Token-level cost attribution matters more than most teams expect, because an LLM feature's unit cost can drift 2x without any visible failure.
Scores
Quality signals attached to traces: automated checks (did the response cite a source, did it refuse an out-of-scope request, did PII leak), sampled human review, and LLM-as-judge ratings. Scores answer "is the output still good?" For retrieval-grounded features, grounding and citation checks earn the most, because retrieval failures are exactly what static benchmarks capture worst. Which checks you need follows from the architecture choice; our RAG or fine-tuning decision guide for banks and insurers covers that fork. The arXiv paper above is right to insist that a judge model is a structured evaluator, not ground truth: calibrate judges against human review and re-calibrate when models change.
Traces without scores tell you the system ran, not that it worked. Scores without traces tell you quality dropped, not why. You need all three.
Standardise on OpenTelemetry GenAI conventions
Instrument once, keep your options open. OpenTelemetry now maintains dedicated semantic conventions for generative AI, and the main observability vendors read them. The current spec defines inference spans named {gen_ai.operation.name} {gen_ai.request.model} with attributes that map neatly onto what an auditor asks for:
- gen_ai.request.model and gen_ai.response.model: which model you asked for and which actually answered. These differ more often than you would think during provider version rollouts.
- gen_ai.usage.input_tokens and gen_ai.usage.output_tokens: the cost record.
- gen_ai.prompt.name and gen_ai.prompt.version: which prompt template produced this output.
- gen_ai.conversation.id: ties multi-turn sessions together so a complaint about "the assistant" resolves to specific turns.
- gen_ai.request.temperature, max_tokens and friends: the configuration in force at request time.
One honest caveat: most GenAI attributes are still marked at "development" maturity, and the conventions have already moved repositories twice while stabilising. That is an argument for putting an instrumentation layer between your code and any single vendor's SDK, not against adopting the conventions. Teams already running LLM agent telemetry on OTel can extend the same pipeline to GenAI spans rather than standing up a parallel stack.
PII in logs is your biggest liability
Here is the uncomfortable part: a faithful audit trail of an LLM system is a database of everything customers typed and everything the model said back. OWASP's 2025 Top 10 for LLM Applications ranks sensitive information disclosure as the number-two risk, and logs and traces are exactly where that sensitive data accumulates (OWASP LLM02:2025).
Practical controls:
- Redact before you log. Run PII detection on prompts and responses before they hit the trace store, and log the redacted version plus a flag that redaction occurred. We covered detection approaches in PII redaction before the LLM; the same layer belongs in front of your observability sink.
- Separate payload from metadata. Keep token counts, latency, model version and scores in your monitoring store with long retention. Keep raw prompt and response payloads, if you keep them at all, in a separate, access-controlled store with shorter retention and tighter audit.
- Set retention deliberately. The AI Act's six-month floor for high-risk deployer logs is a minimum, not a target. GDPR and local data protection law push the other way: do not keep personal data forever because it is cheap. Write the retention schedule down, per data class, and enforce it in the tooling.
- Restrict and audit access. Trace viewers show customer data. Role-based access and an access log on the logging system itself are not optional extras.
Tooling: self-hosted versus SaaS
The tooling decision for a bank is mostly a data-residency decision.
Langfuse is the default open-source answer: MIT-licensed, self-hostable with Docker or Kubernetes on your own infrastructure, running the same codebase as its cloud service, with some add-on enterprise features behind a licence key (Langfuse self-hosting, GitHub repo). Self-hosting keeps every prompt and trace inside your perimeter, which is usually the deciding factor for a regulated institution.
If you can use SaaS, Langfuse Cloud pricing as of October 2026:
- Hobby: free, 50k observability units per month, 30-day data access.
- Core: $29 per month, 100k units included, $8 per additional 100k, 90-day retention.
- Pro: $199 per month, three-year retention, SOC 2 and ISO 27001 reports.
- Enterprise: from $2,499 per month, audit logs, SLAs, premium support.
LangSmith, Arize Phoenix, Helicone and others occupy the same space with different trade-offs. Whichever you pick, insist on OTel-compatible ingestion so the instrumentation survives a tooling change. The same logic applies to the rest of a bank's stack: we deploy and manage open-source systems for clients under our managed open-source stack service precisely because self-hosting shifts control back to the institution.
Worked example: a credit-memo assistant
Make it concrete. A regional bank ships an internal assistant that drafts credit memos for 60 relationship managers.
- Volume: 60 users x 50 requests per day x 21 working days = about 63,000 requests per month. Each request is one trace with a handful of spans.
- Cost of observability (SaaS): Langfuse bills in observability units, and every trace, span and score counts, so 63k requests at roughly six units each is about 380k units per month. Core includes 100k; the remaining 280k bills at $8 per 100k, about $22. Call it $50 per month all-in, rising to roughly $130 at 5x growth as graduated overage rates step down. The observability bill is noise next to the model API bill, which is itself visible because every trace carries token counts.
- Cost of observability (self-hosted): a Docker host or small Kubernetes namespace sized for Langfuse, plus object storage for payloads. For this volume, a single well-specced VM handles it; the real cost is the engineer time to run upgrades and backups.
- Retention: high-risk deployment in scope of the AI Act means at least six months of logs (Article 26), so Hobby's 30 days and Core's 90 days fail on their own. Either self-host with your own retention policy, pay for Pro's three-year retention, or export traces to your own archive monthly.
- Alerting that earns its keep: error rate above 2% for 15 minutes; p95 latency above the memo workflow's tolerance; daily refusal-rate jumps; a quality score sampled by a judge model dropping below threshold two days running; and per-user volume spikes that suggest misuse or an integration loop.
That last category is underrated. A broken integration that retries in a loop shows up as a cost metric anomaly days before anyone files a ticket.
Common mistakes
- Logging everything raw, forever. You built a PII honeypot. Redact, split payload from metadata, set retention.
- No prompt versioning on traces. Quality drops after a prompt edit and you cannot prove the correlation. Log gen_ai.prompt.version on every span.
- Metrics without scores. Dashboards show green latency while the model quietly hallucinates a new policy. Sample and score production traffic continuously.
- Alerts nobody owns. An alert channel without a named owner and a runbook is decoration.
- Treating the judge model as truth. Calibrate automated scores against human review on a schedule, and after every model change.
- Forgetting the access log. The audit system itself needs audit: who viewed traces, who exported data, who changed alert thresholds.
Go-live checklist
Before an LLM feature reaches production in a regulated institution:
- Every model call emits an OTel GenAI span with model, prompt version, tokens and conversation ID.
- PII redaction runs before any payload reaches the trace store.
- Retention schedule documented per data class, meeting Article 26's six-month floor where applicable, and enforced by the tooling.
- Quality scoring on a sampled share of live traffic, calibrated against human review.
- Alert thresholds for errors, latency, cost anomalies and quality scores, each with a named owner.
- Trace access is role-based and itself logged.
- A model or prompt change procedure that deploys through shadow or canary mode before full rollout.
- A written record that ties this monitoring layer to your model risk and AI governance documentation.
If you are still deciding where the model itself should run, our guide to private LLM deployment for banks covers the VPC versus on-premise versus managed API trade-off.
Frequently asked questions
Is LLM monitoring a legal requirement for banks? For high-risk use cases under the EU AI Act, yes in substance: Article 12 requires automatic event logging and Article 26 requires deployers to keep those logs at least six months. Credit scoring and life or health insurance pricing are high-risk categories. The high-risk obligations apply from 2 December 2027 after the Digital Omnibus delay.
How long should we keep LLM logs? At least six months for AI Act high-risk deployments, but write the schedule per data class. Keep metadata (tokens, latency, scores, versions) longer than raw prompts and responses, and let data protection law cap how long personal data sits in any store.
Can we just use our existing APM tools? For latency and error rates, yes. APM tools do not natively understand prompts, token usage, model versions or quality scores, which are the signals an LLM audit actually asks for. Use LLM-aware tooling for traces and scores, and pipe metrics into your existing dashboards via OpenTelemetry.
What does LLM monitoring cost? Tooling is the small line: Langfuse Cloud starts free and costs $29 per month at moderate volume; self-hosting Langfuse costs a VM plus engineer time. The larger costs are instrumenting the application properly and staffing the review of quality scores and alerts.
Do we need monitoring if a vendor runs the model? Yes. Article 26 puts log-keeping duties on the deployer, and your incident and complaint processes need records regardless of who hosts the weights. Vendor dashboards rarely give you trace-level export, per-prompt versioning or the retention control you need.
Where to go from here
Monitoring is where AI governance stops being a document and becomes a running system. If your institution is moving LLM features toward production and needs the evaluation, logging and governance layer built to withstand audit, that is the core of our enterprise AI work for banks and insurers. We will tell you plainly what to build, what to buy, and what to skip.
About AI Agents Plus Editorial
AI automation expert and thought leader in business transformation through artificial intelligence.



