Private LLM Deployment for Banks: VPC vs On-Premise vs Managed API
Managed API, VPC isolation or on-premise GPUs: what each private LLM deployment model costs a bank at real volume, what Kenyan and regional regulators require, and how to choose with current pricing from Anthropic, Lambda and Together AI.

A bank that wants to use large language models has four realistic places to run them: a managed API, a managed API with geographic controls, a private VPC deployment of an open-weight model, or hardware it owns outright. For most bank workloads the managed API is both compliant and the cheapest option by a wide margin, and private infrastructure only wins when sustained token volume is very high or the law forces data to stay inside a perimeter. This guide prices all four options with current vendor numbers and maps them to what regulators in Kenya and the wider region actually require.
The four deployment models, in order of control
Managed API. You call Claude or another frontier model over HTTPS. The provider runs everything. Anthropic's current pricing is $2 per million input tokens and $10 per million output tokens for Claude Sonnet 5, a price that was announced as introductory but has now been made standard, and $1/$5 for Claude Haiku 4.5, according to the official pricing documentation. Data leaves your network and is processed on the provider's infrastructure under their commercial terms.
Managed API with residency controls. The same API, but with inference pinned to a geography. Anthropic supports US-only inference on Claude 4.6 and later models through an inference_geo parameter at a 1.1x price multiplier on every token category, and Azure offers a US Data Zone deployment type with the same multiplier. Claude is also available through AWS and Google Cloud at standard token rates on global endpoints, with a premium of roughly ten percent on regional ones, which lets a bank keep procurement and network routing inside an existing cloud relationship. If this is the route you take, our enterprise LLM integration guide covers the procurement, identity, and network patterns in detail.
Serverless open-model API. Providers like Together AI serve open-weight models per token: Llama 3.3 70B currently costs $1.04 per million input and $1.04 per million output tokens. You get no data residency guarantees beyond the provider's terms, but the unit economics of open weights are dramatically better than frontier models.
Private deployment, VPC or on-premise. You run an open-weight model (Llama, Qwen, Mistral, Gemma) on GPUs you rent or own, inside a network perimeter you control. Meta publishes a private cloud deployment guide for Llama covering the standard patterns: private subnets with no internet gateway, private endpoints, network ACLs, security groups, and VPC flow logs for audit, plus cross-region replication for disaster recovery and residency. This is the model our enterprise AI practice implements for banks and insurers when the compliance analysis demands it.
What the law actually requires in Kenya and the region
The regulatory answer is more permissive than most bank risk teams assume, with one hard exception.
Kenya operates what a recent analysis of African data sovereignty regimes calls a control-and-accountability model. Cross-border transfers are legal. Per AWS's guidance on the Kenya Data Protection Act 2019, personal data can leave Kenya under any one of three conditions:
- The controller or processor gives the Data Commissioner proof of appropriate safeguards, which includes transfers to jurisdictions with commensurate data protection laws.
- The transfer is necessary for a contract, a matter of public interest, a legal claim, vital interests, or a legitimate interest not overridden by the data subject's interests.
- The data subject consents.
The hard exception: personal data classed as strategic to the interests of the state must be processed through a server and data centre in Kenya, and at least one serving copy must be stored in a Kenyan data centre. On top of the DPA, financial systems have been designated Critical Information Infrastructure, which pushes core payment rails toward local processing, and the Office of the Data Protection Commissioner actively audits and fines rather than waiting for complaints. The Central Bank of Kenya's migration of the KEPSS high-value payment system to ISO 20022 in late 2024 means every payment now carries richer structured data, which raises the stakes on how that data is handled.
The practical reading for a bank's AI program:
- Internal productivity workloads (drafting, summarisation, coding assistance) on non-strategic data can lawfully use a managed API, provided the safeguards documentation exists.
- Customer-document processing sits in the same category for most document types, but expect the ODPC and your internal risk committee to want the transfer-impact analysis in writing.
- Anything touching core payment rails or state-strategic data belongs in-country, which today means a private deployment in a local data centre or a local cloud presence.
Neighbours differ. Nigeria mandates domestic switching of payment transactions. The WAEMU bloc requires prior authorisation from the BCEAO. South African banks are adopting the local AWS and Azure regions in Cape Town and Johannesburg to cut both legal risk and latency. If you operate across borders, design for the strictest regime you touch.
What each option costs: a worked example
Take a mid-size bank workload: 30,000 customer documents a day (statements, KYC files, claim forms, loan applications), averaging 2,000 input tokens and 500 output tokens each. That is 60 million input and 15 million output tokens daily, or roughly 1.8 billion input and 450 million output tokens a month.
Managed frontier API, Claude Sonnet 5: 1,800 million input tokens at $2 plus 450 million output tokens at $10 comes to $8,100 a month. Two levers cut this substantially: the Batch API halves Sonnet 5 to $1/$5 for work that tolerates latency, and prompt caching bills repeated context at 0.1x the input rate, which matters enormously when the same policy documents or templates prefix every request.
Managed frontier API, Claude Haiku 4.5: the same volume costs $4,050 a month at $1/$5. Anthropic's own worked example prices 10,000 support-ticket conversations at about $37 on Haiku 4.5, which is consistent with these rates.
Serverless open model, Llama 3.3 70B on Together AI: 2.25 billion tokens at a flat $1.04 per million comes to $2,340 a month. Quality on extraction and classification tasks is close to frontier models; quality on nuanced reasoning and drafting is not.
Private VPC, self-hosted: a 2x H100 node rents for $3.19 per GPU-hour on demand at Lambda's published rates, about $4,660 a month before you pay anyone to run it. Add a second node for high availability and you are near $9,300, plus MLOps staffing, plus the evaluation and monitoring stack.
The counterintuitive result: at this workload, self-hosting is the most expensive option per useful token. The reason is utilisation. Fifteen million output tokens a day averages about 174 tokens per second, and even a 3x peak is under 600. A single 2x H100 node serving a 70B model can sustain thousands of output tokens per second under continuous batching, so the GPUs sit mostly idle and you pay for idle silicon. A 2026 inference-engine benchmark derives a cost of roughly $0.08 to $0.18 per million output tokens for self-hosted vLLM or SGLang on H100s, but the authors are explicit that this assumes saturated batch throughput. Their formula is the right one: divide the hourly GPU cost by measured tokens per second, and be honest about your utilisation.
Self-hosting starts to win on pure cost somewhere in the hundreds of millions of tokens per day of sustained load, or immediately when regulation forbids the alternatives. Everything between those two poles is a judgement call about staffing and risk appetite.
GPU sizing and the serving stack
If the compliance analysis pushes you to private deployment, three engineering decisions follow.
Model size. A 70B-class model in FP8 quantisation fits comfortably across 2x H100 80GB GPUs; an 8B-class model fits on one. For document extraction, classification, and RAG over internal policy, an 8B to 32B model is often enough and costs a fraction to serve. Reserve 70B for workloads where you can measure the quality difference. Our comparison of RAG versus fine-tuning for banks and insurers covers when smaller tuned models beat bigger general ones, and our review of Llama 4 Scout for enterprise use looks at the current open-weight frontier.
Serving engine. Three open-source servers dominate. On H100s with an 8B model, the benchmark cited above measured aggregate throughput of roughly 16,200 tokens per second for SGLang, 12,500 for vLLM, and 9,800 for Hugging Face TGI, with GPU utilisation of 85 to 92 percent for the first two against 68 to 74 percent for TGI. At 70B the gap narrows to a few percent because the workload becomes compute-bound. The decision rules that fall out of the data:
- Default to vLLM. Largest community, broadest model support, and prefix caching since v0.6 closed most of SGLang's historical advantage.
- Choose SGLang for agent-style workloads where many requests share long prefixes; its RadixAttention reuse cuts time-to-first-token sharply on prefix hits.
- Do not start new deployments on TGI. It remains feature-complete on paper — continuous batching, tensor parallelism, an OpenAI-compatible Messages API, Prometheus metrics, per its GitHub repository — but the project is in maintenance mode and the upstream repo is archived.
Resilience. Meta's deployment guide is blunt about the trade: private deployment gives you control over model versions, infrastructure, and security configuration, and in exchange you own capacity planning, performance optimisation, and operational maintenance. Budget for a second region or at least a second availability zone, replicated model artifacts, and a tested failover — our production AI deployment strategies guide walks through the multi-region patterns. A single GPU node is a single point of failure with a blast radius your regulator will ask about.
Common mistakes
- Buying GPUs before measuring tokens. Instrument one month of actual usage through an API first. Most banks discover their volume is one to two orders of magnitude below the self-hosting break-even.
- Assuming an API is automatically non-compliant. Under Kenya's DPA, cross-border processing is lawful with documented safeguards, necessity, or consent. Write the transfer assessment before you write off the cheapest option.
- Quoting saturated-throughput costs for an idle fleet. The $0.10 per million token figure only exists when the GPUs are busy. Divide your real monthly token count into your real monthly GPU bill.
- Skipping the evaluation harness. A private model you cannot measure is a governance incident waiting to happen. Build the LLM evaluation suite before go-live, not after the first audit finding.
- Ignoring prompt caching and batch pricing on APIs. These two levers routinely cut managed API bills by half or more, which moves the break-even point for self-hosting even further out.
- Forgetting that Claude cannot be fine-tuned. If your strategy requires a model trained on your documents, that path runs through open-weight models in your own environment, not through Anthropic. Anthropic's enterprise track is about deployment, tooling, and agents, which is what our Claude for financial services work implements.
When private deployment is the wrong choice
Be suspicious of the on-premise instinct when any of these hold: your measured volume is under a few hundred million tokens a day; your workloads are spiky rather than steady; you need frontier-model reasoning quality that open weights have not reached; you do not have, and cannot hire, at least two engineers who have run GPU inference in production; or the driver is a general feeling that data is safer at home rather than a specific legal clause. In all five cases a managed API with the residency and caching levers applied will be cheaper, better, and easier to defend in an audit, because the audit question is never "where does the GPU sit" but "show me the controls and the assessment."
Decision checklist
- Classify the data: does any workload touch state-strategic data or designated critical infrastructure? If yes, that workload goes in-country.
- Measure real token volume for one month through an API before sizing anything.
- Document the cross-border transfer basis under DPA section 48 and 49 for every remaining workload.
- Price all four models at your measured volume, including batch and caching discounts on the API side and realistic utilisation on the GPU side.
- Choose the smallest model that passes your evaluation suite, then size GPUs to it.
- If you self-host: vLLM by default, SGLang for prefix-heavy agents, two-node minimum, replicated artifacts, tested failover.
- Write the governance documentation (data protection policy, retention schedule, model risk assessment) as a deliverable of the project, not an afterthought.
Frequently asked questions
Can a Kenyan bank legally use a foreign LLM API?
Yes, for most workloads. The Data Protection Act 2019 permits cross-border transfers with documented appropriate safeguards, on necessity grounds such as contract performance, or with consent. The exceptions are personal data classed as strategic to the state, which must be processed and stored in Kenya, and systems designated critical information infrastructure, which face localisation pressure.
How many GPUs does it take to run a 70B model privately?
A 70B model in FP8 quantisation fits on 2x H100 80GB GPUs for inference, which rents for about $6.38 per hour on demand. An 8B model runs on a single GPU. Production deployments need at least two nodes for availability, so double whatever the model alone requires.
Can we fine-tune Claude on our own documents?
No. Anthropic does not offer customer fine-tuning of Claude models. Fine-tuning on proprietary data is done with open-weight models such as Llama, Qwen, Mistral, or Gemma, typically inside your own VPC so the training data never leaves your perimeter.
Is vLLM or SGLang better for a bank deployment?
vLLM is the safer default: largest community, broadest hardware and model support, and prefix caching built in. SGLang earns its place when your workloads are agent-style loops that repeatedly share long prompt prefixes, where its RadixAttention meaningfully cuts first-token latency. Avoid starting new projects on TGI, which is in maintenance mode.
What does self-hosting an LLM really cost per million tokens?
Benchmarks derive roughly $0.08 to $0.18 per million output tokens on H100s, but only at saturated throughput. The honest calculation is your monthly GPU bill divided by your actual monthly tokens. At moderate bank workloads that figure is often ten to fifty times higher than the saturated number, which is why utilisation is the whole game.
Does data residency pricing exist on managed APIs?
Yes. Anthropic charges a 1.1x multiplier on all token categories for US-only inference on Claude 4.6 and later models, and Azure's US Data Zone deployment type applies the same multiplier. That premium is almost always cheaper than running your own GPUs.
Where to go from here
Start with measurement, not hardware: run your candidate workloads through a managed API for a month, record the tokens, and let the compliance matrix above tell you which deployment model each workload actually needs. If the analysis lands on private deployment, our enterprise AI team designs and runs exactly this stack for banks and insurers: fine-tuned open-weight models, VPC-isolated serving, evaluation suites, and the governance documentation your regulator will ask for.
Related reading
About AI Agents Plus Editorial
AI automation expert and thought leader in business transformation through artificial intelligence.



