LLM Claims Triage for Insurers: Accuracy, Human Review, Cost
Real numbers on LLM claims triage for insurers: what fine-tuned models score against ground truth, how the human review baseline actually behaves, per-claim API costs, and the governance supervisors expect.

Claims triage is the highest-volume, most text-heavy decision an insurer makes, and most of the evidence sits in unstructured documents that nobody reads twice. Large language models can read every page of every claim and produce consistent structured triage, but the accuracy numbers behind the vendor decks are thinner and more interesting than the marketing suggests. This post lays out what the published research actually shows, how to design the human review layer, what it costs per claim at current API prices, and what insurance supervisors expect you to have in place.
What claims triage means and where LLMs fit
Triage is the routing decision made in the first hours of a claim: how severe is this, how complex, how urgent, and who should handle it. Get it wrong in one direction and a straightforward windshield claim waits in a senior adjuster queue for a week. Get it wrong in the other direction and a claim with litigation risk or early fraud signals sits in a fast-track lane until the evidence goes cold.
The inputs are mostly text. Medical records, adjuster notes, phone transcripts, claimant statements, repair estimates, and settlement documents carry the real signal, and over 80% of potentially useful business information exists in unstructured form, as a Casualty Actuarial Society-funded research project presented to the NAIC puts it. Manual extraction from that text is slow, inconsistent between reviewers, and impossible to scale.
LLMs are good at exactly this shape of problem: read a large, messy document set and emit structured fields. Adoption data says insurers have noticed. In EIOPA's February 2026 survey of 347 insurance undertakings across 25 countries, nearly two-thirds were already actively using generative AI, and 64% of the reported use cases were back-end productivity work such as data extraction from invoices, audio recordings, and medical reports. Claims triage sits squarely in that category.
The useful way to think about an LLM in triage is not as a decision-maker. It is a reader that never gets tired, applies the same rubric to every claim, and produces structured output a routing engine can act on. The routing rules and every adverse decision stay with you.
What the accuracy research actually shows
The most honest published result comes from a February 2026 paper by researchers at the University of Illinois and PCMI, a warranty administrator processing millions of claims. They built a locally deployed language model component that generates structured corrective-action recommendations from unstructured claim narratives, fine-tuned with LoRA on historical warranty claims and evaluated with a framework combining automated semantic similarity and human review.
The finding that matters: domain-specific fine-tuning substantially outperformed both commercial general-purpose models and prompt-based approaches, with roughly 80% of evaluated cases producing near-identical matches to the ground-truth corrective actions (arXiv:2602.16836). Two lessons sit inside that number. First, fine-tuning on your own claims history is worth real accuracy points over a generic model with a clever prompt. Second, even the best configuration in a controlled study was wrong or materially different in about one case in five. Triage output needs a review layer, not blind trust.
Vendor numbers should be read with the same skepticism you apply to any internal benchmark. EXL, for example, claims its insurance-specific LLM achieves 30% better accuracy than top general-purpose models at 30% lower cost, with practitioner efficiency gains of 30% rising to 75%. Those figures come from the vendor's own internal studies published in a sponsored article. They are directionally consistent with the independent research, fine-tuned domain models beat generic ones, but treat the specific percentages as marketing until you can reproduce them on your own claims.
If you are weighing the fine-tuning route against retrieval-augmented prompting on a general model, that trade-off has real cost and governance consequences either way. We wrote a separate decision guide on RAG versus fine-tuning for banks and insurers that covers when each approach earns its complexity.
The uncomfortable baseline: humans disagree with each other
Here is the finding that reframes the whole accuracy debate. In the CAS-funded NAIC project, two independent clinician reviewers scored LLM outputs on a structured Likert rubric, with a third reviewer adjudicating disagreements. The two trained human experts reached exact agreement only 51.5% of the time. Their overall quadratic weighted kappa was 0.53, which statisticians classify as moderate agreement, and on 19% of items their scores diverged by two points or more on a five-point scale. Agreement was lowest precisely where judgment mattered most: the "understanding and reasoning" dimension had a kappa of just 0.34.
This is not a criticism of the reviewers. It is what decades of underwriting and claims research have always found: expert judgment on complex narratives is noisy. The implication for triage is profound. If you benchmark an LLM against a single adjuster's disposition, you are measuring it against a noisy target and calling the noise "errors." The LLM's actual selling point is different: it applies the same examination to every claim, every time, with no fatigue effects at 4pm on a Friday. Consistency is the product. The NAIC deck makes this point explicitly: the key value is reliability and repeatability, the same process for every claim, feeding structured output to downstream systems.
Practically, this changes how you evaluate. Build your golden dataset from adjudicated outcomes, claims where the eventual result is known, not from one adjuster's opinion. And measure what matters downstream: did routing accuracy improve cycle time and leakage, not did the model mimic a particular reviewer. Our guide to building an LLM evaluation suite for a bank or insurer walks through golden datasets, scoring rubrics, and regression testing in detail.
Designing the human review layer
The review layer is where most triage projects succeed or fail. A design that holds up with regulators and in production looks like this.
Step 1: Define a structured output schema. Each extracted variable should carry three things: a categorical value from a fixed list, a confidence score between 0 and 1, and a mandatory rationale field that cites specific document evidence. The NAIC/CAS prototype used a 36-variable actuarial taxonomy spanning reserving signals like injury severity and recovery trajectory, ratemaking signals like causation type and pre-existing conditions, and claims-management signals like litigation risk, settlement likelihood, and intervention priority. You do not need all 36 on day one. Pick the 10 to 15 fields that drive your routing rules.
Step 2: Build a golden dataset. Pull 200 to 500 historical claims where the final outcome is known: settled value, litigation, fraud referral, cycle time. Have senior adjusters label the triage fields the model should have produced. This dataset is your permanent regression test.
Step 3: Measure before you deploy. Score the model against the golden dataset field by field, and check that its confidence scores are calibrated, meaning a claim scored 0.9 confidence should actually be right about 90% of the time. Re-run this suite every time the model, prompt, or schema changes.
Step 4: Set routing thresholds. High confidence plus low severity goes straight through. Low confidence or any high-severity flag goes to a senior adjuster with the extracted fields and rationales already assembled, which is where the real time saving lands: the adjuster reviews a structured brief instead of reading 200 pages.
Step 5: Run in shadow mode. For four to eight weeks, let the model triage in parallel with your existing process and compare its routing against what adjusters actually did. Divergences are your calibration data.
Step 6: Sample in production, forever. Review a random 5 to 10% of auto-approved claims plus 100% of claims the model flagged as high-severity. Track the disagreement rate monthly. If it drifts, investigate before expanding automation.
One hard rule: the model recommends, humans decide anything adverse. Denials, fraud referrals, and reserve-setting above thresholds stay with named people. This is not just good practice; it is the compliance position, as the next section covers.
Building this layer against a live claims system is plumbing work: schema design, extraction pipelines, integrations into your claims platform, and evaluation harnesses. That is the scope of our Claude implementation service, which covers MCP servers, skills, and workflow integration for exactly this kind of deployment.
Cost per claim: a worked example
Token costs for triage are almost embarrassingly small. Take a typical first-notice-of-loss pack: adjuster notes plus medical report excerpts, roughly 20 dense pages, which is about 15,000 input tokens. The model emits 500 tokens of structured JSON. At current Anthropic API list prices:
- Claude Haiku 4.5 at $1 per million input tokens and $5 per million output: $0.015 input plus $0.0025 output, about $0.018 per claim.
- Claude Sonnet 5 at $2 in and $10 out: about $0.035 per claim.
At 100,000 claims a year, the API bill is roughly $1,750 on Haiku or $3,500 on Sonnet. Prompt caching, where repeated context like your rubric and schema is cached at 10% of the input price, and batch processing for overnight runs push that lower still.
Compare the manual baseline. If triage takes an adjuster 25 minutes at a loaded cost of $50 an hour, that is about $21 per claim, or $2.1 million a year at the same volume. Even if the LLM only eliminates the reading time and humans still make every routing decision, you are buying back most of those 25 minutes for under four cents of compute. The economics are not close. The dominant cost in the new model is the human review of flagged claims, which is exactly where you want your money going.
The fine-tuned open-weight route changes the arithmetic, not the conclusion. A LoRA-tuned model in the 7B to 14B range, the class the PCMI study used, runs on a single mid-range GPU. You trade API pennies for infrastructure and MLOps effort, which makes sense at high volume, under strict data-residency constraints, or when the accuracy gap on your golden dataset justifies it. Anthropic's Claude models are not available for customer fine-tuning, so the fine-tune path means open-weight models such as Qwen, Mistral, Llama, or Gemma. Many insurers run both: a hosted frontier model for the long-document reading, and a small fine-tuned model for the narrow classification steps.
What supervisors expect
If you operate in or sell into the EU, the governance bar is set. EIOPA's August 2025 opinion on AI governance and risk management does not create new rules; it clarifies that existing insurance law already applies to AI systems, under Solvency II Article 41, the Insurance Distribution Directive, and DORA. The expected governance areas map directly onto a triage deployment: fairness and ethics, data governance, documentation and record keeping, transparency and explainability, human oversight, and accuracy, robustness and cybersecurity.
Three specifics for claims triage:
- AI Act scope. Under Regulation (EU) 2024/1689, AI used for risk assessment and pricing in life and health insurance is high-risk. Claims triage generally is not in that Annex III category, but do not take comfort too fast: the EIOPA opinion's governance expectations apply regardless, and GDPR Article 22 restricts solely automated decisions with legal or similarly significant effect, which covers automated claim denials.
- You own the vendor's model. EIOPA is explicit that the undertaking remains responsible for AI systems built or supplied by third parties. Expect to need contractual assurances, audit rights, and due diligence evidence from any AI vendor, including foundation model providers.
- Write it down. Documented explainability, a human oversight register, and log retention are the artifacts supervisors ask for. The good news: if you built the review layer in the previous section, you already generate most of them.
Adoption of formal governance is moving fast. The EIOPA GenAI survey found 49% of undertakings now have a dedicated AI policy, up from about a quarter in 2023, and hallucination was the top-cited risk, ahead of cybersecurity and data protection. Outside the EU, the trajectory is similar: in Kenya, for instance, the Data Protection Act already constrains how insurers use personal data in AI systems, which we covered in our guide to the Kenya Data Protection Act and AI in banking and insurance.
Common mistakes
- Benchmarking against one adjuster. The NAIC data shows two experts agree exactly half the time. Evaluate against adjudicated outcomes, not individual reviewers.
- Automating adverse decisions. Denials and fraud referrals made solely by a model are a regulatory problem under GDPR Article 22 and a reputational one everywhere. Keep a named human in that loop.
- Skipping the rationale field. A triage label without cited evidence cannot be QA'd, cannot be explained to a policyholder, and will not survive a supervisor visit. Force the model to show its work.
- Believing vendor accuracy claims. Internal studies like EXL's 30% figures are hypotheses. Reproduce them on your own golden dataset before putting them in a board paper.
- One giant prompt. Triage is a pipeline: extract, validate, score, route. Monolithic prompts are impossible to debug and impossible to regression-test.
- No production sampling. Models drift, claim mixes shift, document formats change. If you are not reviewing a random sample of auto-approved claims every month, you do not know your live accuracy.
When LLM triage is the wrong choice
Be honest with yourself on these before starting.
- Low volume. Under roughly 10,000 claims a year, integration and evaluation effort usually outweighs the saving. A well-run manual process wins.
- Already-deterministic claims. Parametric products and simple rule-eligible claims do not need a language model. A rules engine is cheaper, faster, and fully auditable.
- No usable history. If past claims lack recorded outcomes, you cannot build a golden dataset, and deploying without one is guessing. Fix the data first.
- The goal is removing humans from denials. If the business case requires fully automated adverse decisions, the project is pointed at the one thing regulators and courts least accept. Reposition it as adjuster augmentation.
Frequently asked questions
What accuracy can an LLM reach on claims triage? In the strongest published study, a fine-tuned model matched ground-truth corrective actions in about 80% of evaluated warranty claims, beating both general-purpose commercial models and prompt-based setups. Your number will depend on the task definition, document quality, and how much domain data you fine-tune on, so measure on your own historical claims before committing.
Is claims triage high-risk under the EU AI Act? Generally no. The Act's high-risk insurance category covers risk assessment and pricing in life and health insurance. But EIOPA's governance expectations under Solvency II, IDD, and DORA apply to claims AI regardless, and GDPR Article 22 restricts solely automated decisions with significant effect, which includes automated denials.
How much does LLM claims triage cost per claim? At current Claude API prices, roughly $0.02 to $0.04 per claim in token costs for a 20-page document pack. Human review of flagged claims is the dominant cost. Compare that with $15 to $25 per claim in adjuster time for fully manual triage.
Should we fine-tune a model or prompt a general one? Start with prompting plus retrieval on a frontier model to validate the workflow and build your golden dataset. Fine-tune an open-weight model when you have the data, the volume, and either an accuracy gap or a data-residency requirement that justifies the MLOps overhead. The research is clear that fine-tuning buys accuracy, but only once you can measure it.
Can an LLM deny a claim automatically? Technically yes, and you should not let it. Adverse decisions need a named human decision-maker, both for GDPR Article 22 exposure in the EU and for basic defensibility in disputes. Use the model to assemble evidence and draft recommendations.
What do we need in place before starting? Three things: a taxonomy of the structured fields triage should produce, 200 to 500 historical claims with known outcomes for evaluation, and agreement on which decisions stay human. Everything else is engineering.
Where to go from here
Claims triage is one of the cleanest LLM use cases in insurance: the documents exist, the baseline is measurably noisy, and the review layer is designable. The insurers getting it right are the ones treating accuracy as something you measure on your own claims, not something a vendor slides across the table. AI Agents Plus works with banks and insurers on exactly this: private LLM deployment, fine-tuning open-weight models, evaluation suites, and the governance documentation supervisors expect. If claims triage is one piece of a wider rollout, our enterprise LLM integration guide covers the architecture, security review, and sequencing around it. If you want an engineer to look at your claims pipeline, see our enterprise AI work for banks and insurers.
About AI Agents Plus Editorial
AI automation expert and thought leader in business transformation through artificial intelligence.



