How to Fine-Tune Llama on Bank Policy Documents in Your VPC
A practical pipeline for fine-tuning an open-weight Llama model on internal bank policy documents inside your own perimeter: dataset building, QLoRA training, GPU costs, evaluation, and vLLM serving.

A bank that wants an LLM fluent in its own credit policy, operations manuals, and product terms has two realistic options: retrieval over those documents, or fine-tuning an open-weight model on them. When the work demands consistent phrasing, a fixed house style, and answers that follow internal procedure step by step, fine-tuning wins. This guide walks through the whole pipeline: turning policy documents into a training dataset, training with QLoRA on rented or in-VPC GPUs, evaluating the result, and serving it with vLLM, all without a single policy page leaving infrastructure you control.
Everything below uses open tooling: Axolotl or TRL for training, PEFT for the QLoRA method, vLLM for serving. The compute bill for a serious first run is measured in tens to hundreds of dollars, not thousands. The expensive part is your people's time, and this post shows where that time actually goes.
Decide fine-tuning is what you need before touching a GPU
Fine-tuning changes how a model behaves. It does not reliably teach a model new facts. That distinction decides whether this project makes sense for you.
Fine-tuning is the right tool when you need:
- Consistent procedure following: draft a complaint response that always follows your eight-step internal process, in your register, citing the right policy section.
- House style and terminology: your institution says "facility" not "loan", uses specific regulatory references, and writes in a specific tone.
- Structured output discipline: produce extraction JSON that matches your core banking fields every time, with no prompt engineering per call.
- Behaviour at latency: a small fine-tuned model that does in 200 tokens what a frontier model needs a 4,000-token system prompt to approximate.
Fine-tuning is the wrong tool when the underlying facts change monthly (interest rate schedules, fee sheets), when you need answers traceable to a specific document version, or when the corpus is small and static. Those are retrieval problems. We wrote a full decision framework in RAG or Fine-Tuning for Banks and Insurers: A Decision Guide, and the short version is: most banks end up with both, a fine-tuned model for behaviour plus RAG for current facts. This post assumes you have made that call and fine-tuning is on the critical path.
Choose the model and read the license first
You cannot fine-tune Claude. Anthropic does not offer customer fine-tuning, so private fine-tuning means open-weight models: Meta's Llama family, Alibaba's Qwen, Mistral, or Google's Gemma. That is not a compromise. For narrow, well-scoped behavioural tasks, a fine-tuned 8B to 70B open model routinely matches a much larger general model on that task, and it runs on hardware you own.
Two license points matter for a bank, using the Llama 3.1 Community License as the example:
- Attribution: anything you distribute or make available that contains Llama materials must prominently display "Built with Llama", and a model you fine-tune and redistribute must have a name starting with "Llama". Internal use inside one institution is far lighter than redistribution, but legal should read the license before procurement signs anything.
- Scale clause: organisations with more than 700 million monthly active users need a separate license from Meta. No bank reading this is near that line, but your legal team will ask, and now you have the answer.
If your policy team prefers a permissive Apache 2.0 license with no custom terms, Qwen and Mistral are the usual picks. We compared the current open-weight field in Meta Llama 4 Scout: The First Enterprise-Ready Open Source LLM. For the worked example below the choice barely matters: the tooling and the costs are the same for any Llama-family 8B or 70B instruct model.
If what you actually need is a top-tier general model with financial-services tooling around it rather than a private fine-tune, our Claude for Financial Services work covers that route. Be honest with yourself about which problem you have.
Turn policy documents into a training dataset
This stage is 60 to 70 percent of the total effort. Budget for it. The pipeline that works:
- Export the corpus to clean text. Convert policy PDFs, procedures manuals, product term sheets, and internal circulars to plain text or Markdown. Preserve section numbers and headings. A policy that says "per section 4.2" is useless training data if section 4.2 lost its number in conversion.
- Chunk by logical section, not by token count. One policy section, one chunk. Keep the document name, version, and effective date attached to every chunk as metadata.
- Draft instruction pairs. Each pair is a question a staff member would actually ask, plus the ideal answer grounded in one or two chunks, ending with the citation. Source the questions from real places: front-desk FAQs, internal helpdesk tickets, onboarding quizzes, examiner findings. These are gold because they are the distribution the model will actually face.
- Review every pair with a subject-matter expert. A credit officer reads the credit pairs. A compliance officer reads the compliance pairs. This is the step everyone wants to skip and the step that decides whether the model is trusted. Do not train on unreviewed synthetic data for a regulated use case.
- Deduplicate and hold out an eval set. Remove near-duplicates, then set aside 10 percent of pairs before training. That holdout is your evaluation set. It must never touch training, or your eval numbers will lie.
- Format for the trainer. TRL's SFTTrainer expects a conversational format, one JSON object per example with a
messagesarray of system, user, and assistant turns, as described in the TRL SFTTrainer documentation. Axolotl accepts the same shape.
On dataset size: quality dominates. The QLoRA paper's Guanaco models were trained on roughly 9,000 examples and reached 99.3 percent of ChatGPT's performance on the Vicuna benchmark, per Dettmers et al., arXiv:2305.14314. For a bank policy assistant, 1,000 to 10,000 expert-reviewed pairs is a realistic and effective range. Ten thousand pairs where a credit officer checked each one beats a hundred thousand machine-generated pairs nobody read.
One governance note: log the provenance of every pair (source document, version, reviewer, date). Your model risk team and your regulator will eventually ask how the model was trained, and "here is the manifest" is a much better answer than "we scraped SharePoint".
Train with QLoRA on GPUs you control
Why QLoRA
Full fine-tuning of a large model means updating billions of parameters, which means multiple high-memory GPUs and a serious MLOps setup. Two ideas from the literature remove that requirement:
- LoRA freezes the base model and trains small rank-decomposition matrices injected into each layer. Against full fine-tuning of GPT-3 175B, the LoRA paper (arXiv:2106.09685) reports 10,000 times fewer trainable parameters and three times lower GPU memory, with no added latency at inference because the adapter merges back into the base weights.
- QLoRA goes further: the frozen base model is quantized to 4-bit NF4 precision with double quantization, and optimizer states are paged to CPU memory when the GPU fills up. The paper fine-tuned a 65B parameter model on a single 48GB GPU. The Hugging Face PEFT quantization guide documents the same recipe, so it is a standard path, not exotic research code.
Practical consequence: an 8B model trains on one 40GB GPU. A 70B model fits on one 48GB GPU in principle, and trains comfortably on a single 80GB GPU or a small multi-GPU node. This is why the whole exercise costs hundreds of dollars, not a data centre.
Tooling
Use Axolotl: one YAML file describes the whole run, it wraps TRL and PEFT, and it supports FSDP and DeepSpeed if you later scale to multi-node. A minimal QLoRA config for an 8B instruct model looks like this:
base_model: meta-llama/Llama-3.1-8B-Instruct
adapter: qlora
load_in_4bit: true
lora_r: 32
lora_alpha: 64
lora_target_linear: true
sequence_len: 4096
micro_batch_size: 4
gradient_accumulation_steps: 4
num_epochs: 3
learning_rate: 2e-4
datasets:
- path: ./data/policy_sft_train.jsonl
type: chat_template
val_set_size: 0.02
output_dir: ./out/llama-policy-v1
Unsloth is a reasonable alternative that claims roughly 2x faster training with 70 percent less VRAM on supported models, per the Unsloth repository. Either way, the output is the same artifact: a LoRA adapter of a few hundred megabytes, not a full copy of the model.
Keeping it inside your perimeter
Nothing in this stack phones home. The model weights download once from Hugging Face (or from your own artifact registry if you mirror them), the dataset never leaves the training machine, and the training run is a local process. For a bank, the deployment shapes that satisfy "data does not leave our control":
- Your own cloud account: an AWS p5 instance in your VPC, in your region, with your security groups and no public ingress. The AWS P5 page lists p5.4xlarge with one 80GB H100 and p5.48xlarge with eight H100s and 640GB of HBM3. Equivalent shapes exist on Azure and GCP.
- On-premise or co-located hardware: one workstation-class GPU server handles 8B training; a small 4 to 8 GPU node handles 70B.
- A GPU cloud for the training run only: if your policy allows training on de-identified or synthetic-only data off-premises, a specialist GPU cloud is far cheaper than hyperscaler on-demand rates. This is a policy decision, not a technical one. If any real customer data is in the training set, keep it in your own account.
We compared the broader trade-offs among these shapes in Private LLM Deployment for Banks: VPC vs On-Premise vs Managed API.
What the compute actually costs
Using Lambda's published on-demand GPU rates as the benchmark (H100 SXM 80GB at $3.99 per GPU per hour, A100 SXM 80GB at $2.79, A100 SXM 40GB at $1.99):
- 8B model, 6,000 pairs, 3 epochs: a few hours on a single A100 40GB at $1.99 per hour. Call it $5 to $15 of compute per training run.
- 70B model, same dataset: QLoRA fits this on one 48GB GPU per the paper, but a full 8x A100 80GB node at about $22.32 per hour finishes in roughly a day. Call it $250 to $550 per run.
- Iteration budget: expect 5 to 15 training runs before the eval set and the reviewers are both happy. Even if most of those runs are at the 70B price, total GPU spend for the project stays in the low thousands of dollars.
Hyperscaler on-demand rates for the same GPUs are typically two to three times higher, and reserved or spot capacity cuts them again. Either way the conclusion holds: compute is a rounding error. Engineer time for dataset construction, SME time for review, and evaluation are the real budget lines, usually one to two engineer-months for a first production-quality run.
Evaluate before anyone relies on it
A loss curve going down is not evidence the model is ready. A bank-grade evaluation has three layers:
- Held-out set scoring. Run the 10 percent holdout through both the base model and the fine-tuned model. The fine-tuned model should win clearly on your task. If it wins by a little, train longer or fix the data.
- Regression checks. Fine-tuning can damage general ability. Test the model on general instruction-following and on deliberately out-of-scope questions. The failure mode to hunt for is a model that answers every question with policy-sounding text, including questions no policy covers. Confident nonsense is the worst outcome in a bank.
- Human review of a live sample. Have the same SMEs grade 100 to 200 fresh model outputs against a rubric: correct procedure, correct citation, correct tone, correct refusal when out of scope. Record the scores. This rubric and these scores are exactly what a model risk framework wants to see.
Version everything: base model, adapter, dataset manifest, eval results. When a policy changes and you retrain, you need to prove the new model is at least as good as the old one before it replaces it. We cover the broader discipline in LLM Fine-Tuning Best Practices: Complete Guide for 2026.
Serve it with vLLM, still inside the VPC
The trained adapter merges into the base weights, producing a standard model directory. vLLM serves it as an OpenAI-compatible API:
vllm serve ./out/llama-policy-v1-merged \
--host 0.0.0.0 --port 8000 \
--max-model-len 8192
vLLM gives you continuous batching and PagedAttention for throughput, FP8 and INT4 quantization for smaller footprints, and multi-LoRA serving, which is genuinely useful here: one base model can serve several adapters (credit policy, KYC procedures, complaints handling) on the same GPU, each adapter a few hundred megabytes. Point your applications at http://<internal-host>:8000/v1 with an OpenAI client library and nothing in the request path leaves your network.
Size the serving hardware to the model: an 8B model serves well on a single 24GB to 48GB GPU; a quantized 70B needs one 80GB GPU or a pair of smaller ones. Steady-state serving cost for a mid-size bank workload is typically one GPU running business hours, or around the clock if it backs customer-facing channels.
Common mistakes
- Training on unreviewed synthetic data. Generating 50,000 pairs with a frontier model and skipping expert review produces a model that sounds authoritative and is wrong in ways nobody catalogued. In a regulated environment that is worse than no model.
- No held-out eval set, or a contaminated one. If test pairs leaked into training, your scores are fiction.
- Fine-tuning for facts that change. Rate sheets and fee schedules belong in retrieval. Baking them into weights means retraining every month and still being stale.
- Skipping the regression check. The model gets better at policy questions and quietly worse at everything else, including saying "I do not know".
- Forgetting the license attribution. "Built with Llama" is a legal requirement for distributed products, not a nicety.
- Treating the adapter as the only artifact. Without the dataset manifest and eval results, you cannot retrain, audit, or defend the model.
When fine-tuning is the wrong choice
Be suspicious of your own project. Fine-tuning is the wrong call when the corpus is under a few hundred good examples (prompt engineering will get you there), when the answers must cite the current version of a document (that is RAG), when the task is broad general assistance (a frontier model behind a managed API is cheaper and better), and when nobody in the organisation owns model risk (fix governance first, then train). The honest default for most banks is a hybrid: RAG for facts, a fine-tuned small model for behaviour, and a frontier API for the long tail.
Frequently asked questions
How much data do I need to fine-tune Llama on bank policies?
Plan for 1,000 to 10,000 expert-reviewed instruction pairs. The QLoRA paper reached near-ChatGPT quality on a general benchmark with about 9,000 examples, so volume matters far less than review quality. Below a few hundred pairs, use prompting instead.
How much does it cost to fine-tune a 70B model?
GPU compute for one QLoRA run on a 70B model is roughly $250 to $550 at published on-demand GPU rates, and a full project with a dozen iterations stays in the low thousands of dollars of compute. The dominant cost is one to two engineer-months of dataset building, review, and evaluation.
Does my data leave the VPC during fine-tuning?
No, if you run the training yourself. The open stack (Axolotl, TRL, PEFT) runs entirely locally: weights download once, training data stays on the machine, nothing is sent to a third party. The exception is managed fine-tuning services, which do process your data on their infrastructure; avoid them for regulated corpora.
Can I fine-tune Claude on our documents instead?
No. Anthropic does not offer customer fine-tuning of Claude models. Private fine-tuning means open-weight models such as Llama, Qwen, Mistral, or Gemma. If you want Claude specifically, the routes are prompting, retrieval, and agent tooling rather than weight-level training.
How often should we retrain the model?
Retrain when behaviour needs to change, not on a calendar. Triggers: a policy rewrite that changes procedures, eval scores drifting down on new question types, or a better base model release. Facts that change frequently belong in a retrieval layer so the model itself rarely needs retraining.
Is QLoRA as good as full fine-tuning?
For task adaptation at this scale, the evidence says yes. The LoRA paper reports on-par or better quality versus full fine-tuning with orders of magnitude fewer trainable parameters, and QLoRA preserves that quality while cutting memory by quantizing the frozen base. Full fine-tuning still has its place for deep domain pretraining, which is a different and much more expensive project.
Where to go from here
The pipeline above is deliberately boring: standard tools, published methods, costs you can verify line by line. What makes it hard inside a bank is not the training run; it is the dataset discipline, the evaluation evidence, and the governance documentation that lets risk and compliance sign off. Our enterprise AI work for banks and insurers covers exactly that: private fine-tuning of open-weight models, evaluation suites, and the model risk paperwork, all deployed inside your perimeter. If you want a senior engineer to walk your team through a first run on your own documents, that is the place to start.
About AI Agents Plus Editorial
AI automation expert and thought leader in business transformation through artificial intelligence.


